Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
Analysis of Prompt Engineering for Drug Toxicity Prediction
SimSkill: A Lifelong Learning AI Agent for Autonomous Mastery of Traffic Simulation
Counterfactual Routing Using Integer Programming with Constraint Generation
Rethinking World Models for Safety-Critical Embodied Systems
GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis
HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Re…
The Attention Triangle in Audio-Video Models
KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents
NeoRed: A Knowledge-Logic-Alignment Multimodal Large Language Model for Neonatal Respirat…
Feature Reconfiguration With Visual Prior for Medical Lesion Segmentation
DNative-Twin: Decision Graphs and Digital Twins for Reconstructable Agentic Decisions
CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning
AutoGraphForge: Towards Automated Graph Theory Discovery
PPO-STGNN: A Proximal Policy Optimization Approach with Spatio-Temporal Graph Neural Netw…
What Matters for Aggressive Decoding-Time KV Eviction? Temporal Aggregation and Ranking P…
Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GU…
Interface-Induced Trajectory Censoring
Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty
A Prompt-Engineering Approach to Develop Scalable, Flexible, and Real-Time Hybrid Micro-L…
Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation
DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice …
GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving
Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory
MasterControl Seventeen Every Time
</think> Doesn't Stop Reasoning: Analysis of Spurious CoT Termination
A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Ha…
Masked Autoregressive Speech Enhancement with Continuous Neural Audio Codec Representatio…
Common-Witness Certificates and Sharp Feature Bounds for Counterfactual Image Auditing
FailBench: How Reliable are VLMs at Judging Robot Task Success?