When Audio Separation Hurts Zero-Shot ASR: Evaluating SAM-Audio with Whisper on Bengali a…
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
PC-Diffuser: Path-Consistent Capsule CBF Safety Filtering for Diffusion-Based Trajectory …
Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies
Designing Safety-Constrained LLM Systems for Public Health Information Access
Compaction as Epistemic Failure: How Agentic LLM Tools Fabricate Confirmed Results from K…
Interventional Grounding Audits: Black-Box Premise-Dependency Tests for LLM Chain-of-Thou…
Probabilistic Extension of Neuro-Symbolic AGI Robots based on Belnap's Typed Intensional …
Oracle Agent Memory as an Enterprise Memory Substrate for Long-Horizon AI Agents
Networked Intelligence: Active Shared Context Graphs for Human-AI Team Science
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
Theory-Level Autoformalization: From Isolated Statements to Unified Formal Knowledge Bases
How Far Can Root Cause Analysis Go on Real-World Telemetry Data?
Multi-Agent Collaborative Reasoning with Tool-Augmented Evidence for Urban Region Profili…
AI advice suppresses people's willingness to say "I don't know", even when the advice is …
Automatic Ordinary Differential Equations Discovery For Biological Systems Using Large La…
STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Orac…
Explaining Reinforcement Learning Agents via Inductive Logic Programming
When Bots Join the Team: Bot Adoption and the Institutional Fabric of Open-Source Softwar…
Beyond Backbone Backpropagation: A Decoupled Strategy for Efficient Transfer Learning
CAVA: Canonical Action Verification and Attestation for Runtime Governance of Agentic AI …
Experience Memory Graph: One-Shot Error Correction for Agents
AIMO Interpretability Challenge
Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0
AI-accelerated End-to-End Framework for Rapid Professional Upskilling
Deep Interaction: An Efficient Human-AI Interaction Method for Large Reasoning Models
FixItFlow: Automated Troubleshooting Guide Generation from Cloud Incidents
Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychi…
LessonBench-V1: A Benchmark Dataset for Evaluating AI Lesson Generation Agents
Federated Explainable Artificial Intelligence: Roles, Architectures, Evaluation, and Open…
Uncertainty-Aware Sequential Decision Rules for Event-Triggered LLM Invocation in Streami…