Explorar

Noticias de IA

37834 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Research & Papers
DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation
arXiv cs.AI Research & Papers
Self-Reports Are Not Verification: Environment-Grounded Auditing of LLM Operators in Evol…
arXiv cs.AI Research & Papers
A Closed-Loop Evaluation of Capability Loss and Recovery in Compressed Driving Policies
arXiv cs.AI Research & Papers
SOVER: Formal Certification of Optimization Reformulations via LLM-Assisted SMT Verificat…
arXiv cs.AI Research & Papers
Topic Matching in the Wild: Benchmark and Lessons from Real-World ASR Transcripts
arXiv cs.AI Research & Papers
DiffHDR: Re-Exposing LDR Videos with Video Diffusion Models
arXiv cs.AI Research & Papers
Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs
arXiv cs.AI Research & Papers
Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement
arXiv cs.AI Research & Papers
Dual Process Motion Planning
arXiv cs.AI Research & Papers
MAS-ProVe: Understanding the Process Verification of Multi-Agent Systems
arXiv cs.AI Research & Papers
LifeAgentBench: Benchmarking LLMs for Long-Horizon, Cross-Dimensional Lifestyle Health Re…
arXiv cs.AI Research & Papers
HEAL: Hindsight Entropy-Assisted Learning for Reasoning Distillation
arXiv cs.AI Research & Papers
HugAgent: A Human Simulation Benchmark for Individual-Level Reasoning
arXiv cs.AI Research & Papers
Multi-Agent LLM Orchestration Achieves Deterministic, High-Quality Decision Support for I…
arXiv cs.AI Research & Papers
CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harne…
arXiv cs.AI Research & Papers
Mechanism Design for Alignment and Control
arXiv cs.AI Research & Papers
Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation
arXiv cs.AI Research & Papers
ViPlan: A Benchmark for Visual Planning with Symbolic Predicates and Vision-Language Mode…
arXiv cs.AI Research & Papers
GeoGR^2:Zero-Shot Geospatial Inference via Geostatistically-Guided Iterative Refinement w…
arXiv cs.AI Research & Papers
Making Prospective Memory SLM-Shaped: Typed Intention Stores for Small-Model Agents
arXiv cs.AI Research & Papers
Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs
arXiv cs.AI Research & Papers
Feedback-Assisted Trust Propagation over Document Relation Graphs for Retrieval-Augmented…
arXiv cs.AI Research & Papers
Consistency Without Alignment: Item-Sensitive Language Models Indistinguishable From Rand…
arXiv cs.AI Research & Papers
Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning
arXiv cs.AI Research & Papers
Wave Function Backpropagation with Explicit Temporal-Interval Dynamics
arXiv cs.AI Research & Papers
Retrieved but not ranked: surface-form bias in structural retrieval, from mathematics to …
arXiv cs.AI Research & Papers
BS: Take the Hint - Interactive Multitracer PET/CT Lesion Segmentation with a Scribble-Co…
arXiv cs.AI Research & Papers
RecalibrateGPT: AI Fatigue Resilient Conversational Interfaces
arXiv cs.AI Research & Papers
Denoising Diffusion Generative Models Secretly Calculate Attentions
arXiv cs.AI Research & Papers
When the Algorithm Becomes the Brand Crisis: A Sociotechnical Theory of Distributed Respo…