Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watche…
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
From State to Action: OODA-Tool for Reliable Multi-Turn Tool Use
FedV-KGQA: Multi-Hop Question Answering over Vertically Partitioned Knowledge Graphs
Identifying Latent Declarative Representations of Code for Assisting Repository Migration
Method, Mind, and Morality: How People Make Sense of Artificial Intelligence
ExPhy: A Benchmark for Explicit Physical Property Learning in Multi-Object Trajectory For…
When Less Is More: An Empirical Study of Minimal Responses in Counseling Dialogues and th…
ChorusTIC: Training-Free Multivariate Time Series Classification via Chorus In-Context Le…
Restoring Without Forgetting: Continual Learning Across Image Degradations
Ensemble of Convolutional Neural Networks for StrokePrediction: Towards Improved Diagnost…
From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Caus…
PARTAB: Partition-Aware Reasoning with Structured Evidence for Scalable Table Understandi…
StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Ba…
Beyond Executable Models: The Pufibara Agent Harness and the Modelica Agent Workflow Benc…
Rebuild Dossier: Mechanically-Enforced Specs for Agentic App Rebuilds, and What Model-Tie…
Macro-Operator Generation and Predicate Selection for TAMP Operator Learning
CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving
Robust Motion Generation using Part-level Reliable Data from Videos
Quasar: A Programming Language Specialized for LLM Code Actions
PatientHub: A Unified Framework for Patient Simulation
Progressively Learning Heterogeneous Skills in a Unified Latent Space
REFINE: A Multi-Agent LLM Approach for Evidence-Guided Code Refactoring
When May an Agent Stop? Evidence-Carrying Termination for Tool-Using LLMs
The Limits of Automatic Evaluation of Creativity in Large Language Models
ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect General…
Revelation Control
Strictly Causal Streaming Video Anomaly Detection with a Theoretically-Grounded State-Spa…
Score-Based Ideal Observer Approximation via Denoising Score Matching for Signal-Known-Ex…
Can a Dynamic Internal Field Govern a Transformer's Cognition? Certifiability, not Superi…
MolEmb: Multimodal Large Language Models Can Be Strong Molecular Embedding Models