HYDRA: Hyperbolic Dynamic Representation Architecture for Kolmogorov-Arnold Networks
Explorar
Noticias de IA
21010 elementos — filtrados, clasificados y sin duplicados
The Sleeping Agent: What Gist-Based Context Compression Loses and Why
Dynamics Models for Offline Hyperparameter Selection in Real-World RL
Let it Cook: Learning to Wait in Sequential Decision Making
Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Tex…
VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies
Backdoor Decontamination Dynamics in LLM Agents
EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under Non-Ideal Ente…
Consolidator: Learning Persistent Routed Memory Across Context Boundaries
Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration
Benchmarking LLM Judges for Mobile Agent Evaluation
CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Networ…
From Safety Documentation to Safety Knowledge Support: An Evidence-Grounded LLM Framework…
Every pooling rule has its world: matching probability combination rules to situations an…
Federated Learning for Distributed CNC Tool Wear Prediction
Learning from Online User Feedback for Shopping Agents
Variable Selection in the Context of AI Fairness
RecSys Factory: Bounding LLM Agent Autonomy to Decision Points in the Industrial Recommen…
VICBench: A Multi-Language Benchmark for Code Vulnerability Detection
High-dimensional Multi-objective Bayesian Optimization with Learned Variable Interactions
EvoGraph-Mem: Failure-Aware Editable Graph Memory for Long-Term Language Agents
Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic …
Social Meaning in Large Language Models: Structure, Magnitude, and Pragmatic Prompting
TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation
One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL
Adaptive Hybrid Particle Swarm Optimization with Gradient Descent
User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware Autoscaling
Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge
Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs
No One to Blame: A Framework of Constitutive AI Unaccountability