From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs
Explorar
Noticias de IA
30329 elementos — filtrados, clasificados y sin duplicados
A Graph-Native Bitemporal Memory Store for Conversational AI Agents
MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning
The Capability Paradox: How Smarter Auditors Make Multi-Agent Systems Less Secure
HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Mod…
MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval
Position: Evaluation Scores Are Perishable Knowledge Claims
Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Def…
GoGoTB: Agentic RTL Verification with Specification-Grounded Coverage Closure
CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games
Misalignment Has a Personality: A Big Five Account of Emergent Misalignment
SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution
GuidedRAG: Semantic Steering of Retrieval-Augmented Generation
Model-Driven Requirements Configuration with Three-Valued Uncertainty Scoring
How does downsampling affect needle electromyography signals? A generalisable workflow fo…
FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing
When benchmark inferences do not compose: Projectibility in AI evaluation
Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems
GuideSkill: Evolving Executable LLM Agent Skills for Guideline-Grounded Clinical Reasoning
Optimizing Sensor Placement for Hydrogen Leak Detection in Enclosed Infrastructure: A Com…
ClinLens: Towards Long-Horizon Coding Agents for Longitudinal Multimodal Clinical Data Sc…
HealthSLM-Bench: Benchmarking Small Language Models for Mobile and Wearable Healthcare Mo…
Decision-oriented joint optimization of evidence fusion based on event-conditioned credib…
Large-Scale ChatBot Validation Through Customer Digital Twin Simulations
DLAM: Distributional Latent Actions with Temporal Constraints
ReCo: Reweighting GRPO Against Distributional Concentration
Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility
Linguistic Monoculture in LLM-Assisted Language Use
AI as Friction for Reflection Support in Ideation
Crossing-Free Probabilistic K-Line Forecasts Without Retraining