Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isome…
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents
Recursive Multi-Agent Systems
Algorithm Selection with Zero Domain Knowledge via Text Embeddings
Open, Reliable, and Collective: A Community-Driven Framework for Tool-Using AI Agents
GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforceme…
Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs
ECG-LDC: A Hardware-Efficient Low-Dimensional Computing Framework for ECG Arrhythmia Clas…
Reverse Engineering Compliance: A Dual-Graph Verification Framework for Auditing Legacy I…
When Sensing Varies with Contexts: Context Probing for Tactile Few-Shot Class-Incremental…
Where Experts Disagree, Models Fail: Detecting Implicit Legal Citations in French Court D…
Do Implicit Personalization and Explicit Styles Conflict? PsPLUG: A Lightweight Plug-in f…
InqEduAgent: Adaptive AI Learning Partners with Gaussian Process Augmentation
Playful AI in Professional Email: A Field Experiment on Tone and Recipient Engagement
Reproducing human biases in route choice using large language models: Toward scalable beh…
LLM-Driven Collaborative Model for Untangling Commits via Explicit and Implicit Dependenc…
Measuring AI Ability to Complete Long Software Tasks
Auditing the Risk Claims of Distributional Reinforcement Learning
Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks
A Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous Manipulation
Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias
Evidence-Backed Video Question Answering
Transformer-Guided Swarm Intelligence for Frugal Neural Architecture Search
MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents
Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Aud…
StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description
Time-Lag-Aware Deep Reinforcement Learning for Flexible Job-Shop Scheduling in PPVC Modul…
An Explainable Agentic System for Detection of Conversational Scams with Summary-Based Me…
Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming
Interaction Scaling: Grounding the Third Axis of Test-Time Compute