XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding
Explorar
Noticias de IA
29670 elementos — filtrados, clasificados y sin duplicados
WM-Cov: Test Adequacy for Interactive World-Model-Style Autonomous Driving Simulation
SLMs as Multi-Agent Routers: A Progressive SFT and Reinforcement Learning Approach
Google's AI & Economy ATLAS v1.0: Mapping Gemini Usage in the Economy
Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories
Trustworthiness Costs of Domain Adaptation in Small Language Models:A Cross-Architecture …
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning
Automated ECG Interval Measurement and Wave Delineation Using Fast Fourier Convolution Re…
Homebot: A Personal AI Agent for Conversational Home Assistance and Automation
Neuro-Symbolic Participation Governance for Verifiable AI Agents in Open Digital Twin Eco…
CRAFTS: Collaborative Role-Adaptive Fine-Tuning of LLM Agents for Chemical Process Simula…
A False Average: Chain-of-Thought Monitors Collapse Where They Are the Only Defense
Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for …
Unleashing the Potential of Large Language Models: A Blueprint for Real-Time, Enterprise-…
MemArbiter: Decision-Time Memory Arbitration for Long-Horizon LLM Agents
The Impact of LLM-Assistants on Software Developer Productivity: A Systematic Review and …
Less Is More: Tuning Configurable Systems with Imperfect Fidelity
Slides2MindMap: Reconstructing Cognitively Efficient Knowledge Hierarchies from Lecture S…
Inference-Time Policy Alignment for Fair Reinforcement Learning
Symbolic Attack Chain Generation from Atomic Red Team Techniques: An Empirical Study of P…
Beyond Solution-Centric Search: Adaptive Inquiry and Knowledge Revision for Autonomous ML…
DGA$_2$D: Directed Graph-Guided Automated Algorithm Design with Large Language Models
Large language models improve physician accuracy but lead to false reliance
Cost-Based Semantics for Querying Inconsistent Weighted Knowledge Bases
Don't Offer What Can't Be Done: Deterministic Executability Gating for LLM Skill Selectio…
REIMU: Efficient Heterogeneous Hierarchical Reasoning for SSL-Based Speech Deepfake Detec…
Interpretability-Guided Soft Pruning of Attention Heads in Vision Transformers
xPress: Parallel Refinement for Diffusion Drafters in Speculative Decoding
Logographic Character Visual Pretraining via Semantic-based Contrastive Learning
Perspectives on Tsallis Statistics for Artificial Intelligence