LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
Argumentation for Common Ground: Finding Zones of Possible Agreement between Individuals …
Eigenanalysis framework for autoregressive neural emulators of multi-scale chaotic dynami…
ReForge: Keeping ABR Algorithms Never Finished with Verified Large Language Model Edits
Translating finite-domain integer constraint models to CP/SMT/ILP/PB/SAT solvers with CPM…
Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterizat…
ATLAS: Scaffold-Free Algorithm Synthesis by LLMs via Embedding-Guided Quality-Diversity S…
When Context Misleads: Intent-Guided Decoding for Robust Retrieval-Augmented Generation
Constitutive Priors for Machine Intelligence: A Legitimacy Theory of the Artificial Physi…
Mental Model Management: An Operator-Based Framework for LLM Memory
Broken Symmetry in LLM Refusal: Answer Release Is More Local Than Refusal Restoration
Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning
FLOPs vs Real Work: The Importance of Replication in AI Efficiency Assessment
Large Language Models Show Metacognitive Sensitivity in Medical Reasoning
Position: AI Lock-In Is in Progress, and We Must Be Prepared
Position: Certified Correctness in Neural Constraint Reasoning Requires Symbolic Integrat…
Anatomy of a Quantized Agent: VRAM Stability and Forecasting in Code-Synthesis Agentic Wo…
Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI Detect…
TokenSTFormer: A Tokenized Spatial-temporal Attention Model for Holistic Motion Analysis …
LongDocBench: Benchmarking TOC Hierarchy and Contextual Relationship Recovery in Long Doc…
From Errors to Proofs: Minimal-Core-Guided Repair for Neuro-Symbolic Constraint Solving
MUPA$^{2}$E: Multimodal Unified Perception with Asymmetric Attention for Emotion Assessme…
Understanding and Stabilizing Deep Q-Learning via Controlled Bootstrapping and Regulated …
CardiacMamba: Fair and Robust RGB-RF Fusion for Remote Heart Rate Estimation via State Sp…
Platform Adaptation Under Governance Interventions: Actor Best-Response Modeling and an E…
Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking
Semantic Uncertainty-Guided Orchestration in Hierarchical Multi-Agent Systems
JarvisBench: Always-on Intelligence Between Humans and Agents
Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems
Beyond Correctness: Toward Automated Novelty Verification with Lean 4