DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
Self-Reports Are Not Verification: Environment-Grounded Auditing of LLM Operators in Evol…
A Closed-Loop Evaluation of Capability Loss and Recovery in Compressed Driving Policies
SOVER: Formal Certification of Optimization Reformulations via LLM-Assisted SMT Verificat…
Topic Matching in the Wild: Benchmark and Lessons from Real-World ASR Transcripts
DiffHDR: Re-Exposing LDR Videos with Video Diffusion Models
Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs
Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement
Dual Process Motion Planning
MAS-ProVe: Understanding the Process Verification of Multi-Agent Systems
LifeAgentBench: Benchmarking LLMs for Long-Horizon, Cross-Dimensional Lifestyle Health Re…
HEAL: Hindsight Entropy-Assisted Learning for Reasoning Distillation
HugAgent: A Human Simulation Benchmark for Individual-Level Reasoning
Multi-Agent LLM Orchestration Achieves Deterministic, High-Quality Decision Support for I…
CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harne…
Mechanism Design for Alignment and Control
Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation
ViPlan: A Benchmark for Visual Planning with Symbolic Predicates and Vision-Language Mode…
GeoGR^2:Zero-Shot Geospatial Inference via Geostatistically-Guided Iterative Refinement w…
Making Prospective Memory SLM-Shaped: Typed Intention Stores for Small-Model Agents
Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs
Feedback-Assisted Trust Propagation over Document Relation Graphs for Retrieval-Augmented…
Consistency Without Alignment: Item-Sensitive Language Models Indistinguishable From Rand…
Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning
Wave Function Backpropagation with Explicit Temporal-Interval Dynamics
Retrieved but not ranked: surface-form bias in structural retrieval, from mathematics to …
BS: Take the Hint - Interactive Multitracer PET/CT Lesion Segmentation with a Scribble-Co…
RecalibrateGPT: AI Fatigue Resilient Conversational Interfaces
Denoising Diffusion Generative Models Secretly Calculate Attentions
When the Algorithm Becomes the Brand Crisis: A Sociotechnical Theory of Distributed Respo…