PASE: Leveraging the Phonological Prior of WavLM for Low-Hallucination Generative Speech …
Explorar
Noticias de IA
21272 elementos — filtrados, clasificados y sin duplicados
In-Context Collapse in Vision-Language Models and How to Mitigate it?
Less Traffic, Better Outcomes: Competition-Aware Request Dispatch in Real-Time Ad Exchang…
Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning
Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents
HyperAgent: Planning and Acting over Tool-Schema Hypergraphs for Tool-Use LLM Agents
Policy Fragmentation or Institutional Alignment? Institutional Governance of AI in Univer…
Large language models for partial differential equation workflows
CUADebug: Diagnosing and Repairing Computer-Use Agent Failures
When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO
Towards Improving Sequential Decision-Making in LLM Agents via Experience Memory
Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks
LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models
BAP-SQL: Budget-Aware Observation Planning for Agentic Text-to-SQL
Dr. AGENTONOMICS: A Didactic Experiment of AGENTONOMICS
Enactive Artificial Intelligence: A Decision-Centric Architecture for Complex Systems
Instruction Stacking Collapse: A Benchmark and the Capability-Dependent Value of Prompt C…
AI World Cup 2026: Benchmarking Large Language Models for End-to-End Football Tournament …
When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Rea…
UrbanAgent: A Tool-Augmented Agent for Cross-System Urban Tasks
Hypercubes, Hyperplanes, and Constraint-Induced Complexity Collapse in Atomic Concept Lea…
Neurosymbolic Reasoning with Incremental Knowledge for Sample Efficient Hierarchical Rein…
GraphCliff: Short-Long Range Gating for Modeling Critical Activity Changes Caused by Subt…
TACT: Taxonomy-Aligned Post-Training for Pedagogically Adaptive English Tutoring
WCM: World-Cognition Model for Generalizable Human-Robot Interaction
Agentic Reinforcement Learning with Self-Distilled Reward Shaping
A Unified 2D Framework for DeepLesion Detection, Segmentation and Short Report Generation
Robust Counterfactual Policy Optimisation via Nondeterministic Causal Models
When Outputs Disperse, Does Epistemic Revision Follow? A Black-Box Diagnostic for Machine…
Mixed-Initiative Human-Robot Teaming under Suboptimality with Online Bayesian Adaptation