Cautious Context Steering for Language Model Personalization
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
The Ignition Index: Measuring Global Workspace Dynamics in Language Models
LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs
TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Tr…
Recursive Synthesis for Long-Horizon Terminal Tasks
SkillTrace: Multi-Trace Provenance Auditing for LLM-Agent Skill Reuse
PaDoc: Layout-Grounded Parallel Decoding for Document Parsing
C$^3$PO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal M…
Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agentic Operations
Search2Skill: Skill Distillation Beyond Knowledge Boundaries Via Rubric-Based Reinforceme…
ECG-LENS: Lead-Aware Clinical Context Enriched ECG Report Generation and Evaluation
Challenges in Evaluating Explanation Methods for Static and Evolving Data
Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay
AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents
When Self-Evolution Backfires: Pre-Commit Gating against Skill Contamination in LLM Agents
When History Lies: Evaluating and Improving Tool Use under Misleading Multi-Turn Histories
Correct Answers from Sound Reasoning: Verifiable Process Supervision for Language Models
Skill Neologisms: Towards Skill-based Continual Learning
Supervised Learning Has a Geometric Blind Spot
A note on conditional PAC-efficient reasoning in large language model routing
Trajectory-guided discharge stratification for heart failure using short-context electron…
Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap
Hardware Keystores for AI Agent Signing Workflows: A Zero-Trust MCP Enforcement Architect…
Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New S…
Counterfactual Analysis via Large Language Models
Post-Hoc Trajectory-Risk Certification for Modular LLM-Based Security Agents
DoctorAgents: an agentic framework to iteratively refine AutoML pipeline for small clinic…
Adaptive Arena-based Contestable Argumentative Network-of-Experts for Open-Ended Care Pla…
SCP-NL2TL: Selective Conformal Prediction with Semantic Verification for Natural Language…
Integrating Implicit and Explicit Relational Biases through Graph-Based Multiple Instance…