TROVE: Adaptive Agent Skill Orchestration via Trace-Grounded Route Validation and Editing
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
Do LLMs Exhibit Coherent Knowledge Structures in Mathematical Reasoning? A Perspective fr…
Quality-diversity in dissimilarity spaces
Why We Care About Understanding: Competence through Predictive Compression
A Tree-based RAG Framework for Evidence-Intensive QA via Adaptive Planning and Topology-A…
Artificial Intelligence in Equity and Crypto Markets: Progress, Profitability Evidence, a…
Language models judge war differently when tested for alignment
Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing
Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference
From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Di…
CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical…
The History Is the Detector: Executing CVE Patch History, End-to-End
From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use Agents
AutoLR: Automating the Path from Research to Launch Review in Industrial Recommender Syst…
Diffusion TV: Experiencing Diffusion Models through Tangible, Embodied Interaction
Long Horizon Transformer Quantile Fault Prediction for Multi Site Industrial Predictive M…
ElderBench: Benchmarking Autonomous Mobile Agents for Older Adults
CHAMP: Cross-domain Hybrid Architecture for Matchmaking and Prediction in Online Multi-Pl…
PetQA: Benchmarking Veterinary Knowledge and Clinical Reasoning
Whose record is this? Diagnosing and authorizing record use in personalized multimodal mo…
Hierarchical Possession-Aware Graph Pointer Network for Pass Receiver Selection
Diffusion Language Models for Mobile Edge Agentic AI: Foundations, Applications, and Chal…
Aplaud: Adaptive Personalized Low-Rank Decomposition for User-Specific LLM
Shadow Queries for Private Retrieval in Vector Databases
When Financial Fine-tuning Fails: A Three-Level Detectability Analysis of Numerical Hallu…
Predicting Spatiotemporal Mobile Sensing-Based PM2.5 Concentrations Using Low-Rank Adapte…
Model Retirement Creates Reproducibility Risk in Biomedical AI Publications
FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflictin…
Commonsense Reasoning in Computer Vision: Foundations, Recent Advancements, and Future Di…
Harness-agnostic detection and immunization of reward hacking in self-evolving language m…