COMAP: Co-Evolving World Models and Agent Policies for LLM Agents
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
S3TS: Stochastic Scenario-Structured Tree Search for Advanced Planning Under Uncertainty
Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectori…
Community-Aware Assessment of Social Textual Engagement and Resonance: A Human-Centric Pe…
Does Compression Preserve Uncertainty? A Unified Benchmark for Quantized and Sparse LLMs …
Beyond Access: Guided LLM Scaffolding for Independent Learning in Undergraduate Statistics
Token Predictors Are Not Planners: Building Physically Grounded Causal Reasoners
Structure-Guided Adaptive Propagation for Protein-Protein Interaction Site Prediction
TrafficRAG: A Multimodal RAG Framework for Traffic Accident Liability Determination
Characterization of Multi-Model Agentic AI Systems on General Tasks via Trace-Driven Simu…
Joint Agent Memory and Exploration Learning via Novelty Signals
Transferring Information Across Interventions in Causal Bayesian Optimization
Early Diagnosis of Wasted Computation in Multi-Agent LLM Systems via Failure-Aware Observ…
Recognize Your Orchestrator: An Entropy Dynamics Perspective for LLM Multi-Agent Systems
ANDES: Agent Native Data Evolving Synthesis Tool for Autonomous Instruction Alignment
Application of Algorithms in Energy-Efficient Design Platforms for Green Building
Expected Value Alignment for Generative Reward Modeling in Formal Mathematics Verification
Diagnosing LLM Arbitration Behavior over Pre-evidence Epistemic States in RAG-based Fact-…
TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Plan…
Property Prediction of Stacked Bilayer Materials: A Multimodal Learning Approach
Relational Intervention During Functional Collapse in Large Language Models: A Lexical-St…
Mitigating Hallucinations in Large Language Models Via Decoder Layer Skipping
TukaBench: A Culturally Grounded Jailbreak Benchmark for African Languages
Latent Reward Steering: An Adaptive Inference-Time Framework that Implicitly Promotes Cog…
TRACE: Trajectory Risk-Aware Compression for Long-Horizon Agent Safety
TAPS: Target-Aware Prefix Tree Selection for Diffusion-Drafted Speculative Decoding
Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight
VESTA: Visual Exploration with Statistical Tool Agents
Evaluating Bivariate Causal Statements Based on Mutual Compatibility
Capability Self-Assessment: Teaching LLMs to Know Their Limits