Trajectories That Segment Themselves: Agent-Declared Boundaries as a Training Unit
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
MechGeo: Autoformalizing and Proving Euclidean Geometry in Lean 4
SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation
Self-Certification of Representation Adequacy: Sequential Certification at Minimum Task L…
Trustworthy AI in Digital Health: A Comprehensive Review of Robustness and Explainability
PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with…
From Profiling to Synthesis: Benchmarking Implicit Behavioral Alignment in Personalized L…
Magnet: Detecting Cross-Session AI Misuse Through Capability Accumulation
Optimizing Minimax Regret in Uncertain MDPs with Small Sets of Policies
Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision
MonitrLLM: A Community-Centered Evaluation Infrastructure for Large Language Models
SkillTrace: Traversing a Query-Skill Graph for Composable LLM Agents
Mamba with Hierarchical Memory: Solving Representation Bottleneck in Long Sequence Modeli…
Instruction-Conditioned Exploration with Asymmetric Reinforcement Learning and Self-Disti…
HPFA: Hypergraph-Based Paired Failure Attribution for LLM Reasoning
Before Reasoning Fails: Pre-Evidence Procedural Failures in Agentic RAG
HALT: Verification-Aware Stopping for Retrieval-Augmented Search Agents
A Contractualist Argumentation Framework for Moral Decision-Making
From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Tas…
DAPD: Dual-Anchored Policy Distillation
LaCache: Robust Semantic Caching for LLM Serving
Beyond Single-Use Tokens: Durable Authorization State for Replay-Resistant LLM Agent Acti…
ReasonCast: Towards Explainable Time Series Forecasting with Reasoning
Physics-Informed Neural Networks for Complex Eigenfrequency Identification and Mode Struc…
Rewriting or Reweighting? A Geometric Account in Language Models
FRAMES: Guarded and Dual-Objective Skill Evolution for Agents in Policy-Governed Enterpri…
Post-Training on Office Work Improves Software Engineering: A Behavioral Account of Cross…
V-Mem: Modality-Routed Retrieval for Long-Term Multimodal Agentic Memory
Where Reasoning Diverges: Localized Multi-Agent Debate