Browse

AI News

25879 items — filtered, classified, deduplicated

arXiv cs.AI Research & Papers
Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge
arXiv cs.AI Research & Papers
Is More Privileged Information Better? From Solution Traces to Problem-Solving Structure …
arXiv cs.AI Research & Papers
ProWorld: Progress-Aware Hyperbolic World Models for Long-Horizon Visual Goal Reaching
arXiv cs.AI Research & Papers
Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories
arXiv cs.AI Research & Papers
When Memory Updates but Behavior Does Not: Repairing Implicit Stale Dependencies in Perso…
arXiv cs.AI Research & Papers
MemSIF: From Structured Interactions to Dual-Track Fact Memory for LLM Agents
arXiv cs.AI Research & Papers
EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers
arXiv cs.AI Research & Papers
TrajWiki: Source-Grounded Memory Trajectories for Long-Horizon Dialogue Agents
arXiv cs.AI Research & Papers
PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent
arXiv cs.AI Research & Papers
Passing Coarse Marginal Checks Can Be Cheap: Persona Mixtures and Imprecise Treatment-Res…
arXiv cs.AI Research & Papers
Can You Trust the Confidence? ConfBench for Vision-Language Models on Document Extraction
arXiv cs.AI Research & Papers
Diagnosing Search Behavior and Failure Modes in Long-Horizon Search Agents
arXiv cs.AI Research & Papers
MemArbiter: Decision-Time Memory Arbitration for Long-Horizon LLM Agents
arXiv cs.AI Research & Papers
MEGRAG: Multi-Granular Evidence Graphs for Answer-Aware Multi-Hop RAG
arXiv cs.AI Research & Papers
Cooperative Coevolution for Resource-Constrained Agentic LLM Post-Training
arXiv cs.AI Research & Papers
RubricReviewer: From Direct Critique to Objective and Comprehensive Rubric-Driven Peer Re…
arXiv cs.AI Research & Papers
Auditing Discovery Claims: A Two-Sided Criterion for Agentic Science, with the Negative S…
arXiv cs.AI Research & Papers
AgentMemBench: A Systematic Benchmark for Evaluating Long-Term Memory Management Strategi…
arXiv cs.AI Research & Papers
AgentSLABench: Evaluating and Benchmarking Agentic Systems Under Resource Constraints
arXiv cs.AI Research & Papers
The Scaling Paradox in Human-AI Collaboration
arXiv cs.AI Research & Papers
From AI Technical Debt to Agentic Technical Debt: A Systematic Mapping of Root Causes and…
arXiv cs.AI Research & Papers
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets
arXiv cs.AI Research & Papers
Toward Fine-Grained Forgetting:Attribute Unlearning for Multimodal Large Language Models
arXiv cs.AI Research & Papers
FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Cro…
arXiv cs.AI Research & Papers
Assuming You Knew: Fixing an Epistemic Semantics for Flow Policies Using Agentic AI
arXiv cs.AI Research & Papers
Cross-Fitted Residual Utility for Primary-Preserving Cognitive Decision Correction in Aut…
arXiv cs.AI Research & Papers
AI-Based Thesis Assessment: An Empirical Study of Human Evaluation Priorities and Their I…
arXiv cs.AI Research & Papers
Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation
arXiv cs.AI Research & Papers
Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process…
arXiv cs.AI Research & Papers
Tracing the Cascade: A Topology-Aware Evaluation Framework for Scientific Agent Hallucina…