Useful Memories Become Faulty When Continuously Updated by LLMs
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
Safe to Resume? Breaking Execution Continuity of Agent Execution via Rollback
What Does an Evaluation License? A Commit-Bound Census of Claim Replay in Inspect Evals
AGRICAM: A Track-Mounted Crop Pollination Monitoring Robot
COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models
M2K: Making the Model-Kernel Interface Explicit for Reliable CUDA Kernel Verification
FAMPWQ: Fisher Information-based Adaptive Mixed Precision Weight Quantization for Effecti…
Learning Composable Chains-of-Thought
Retrieval-Augmented LLM Agents: Learning to Learn from Experience
WebXSkill: Skill Learning for Autonomous Web Agents
A Multi-Agent Human-LLM Collaborative Framework for Closed-Loop Scientific Literature Sum…
PeopleSearchBench: Evaluating AI-Powered People Search Platforms with Criteria-Grounded V…
Uncertainty Makes It Stable: Curiosity-Driven Quantized Mixture-of-Experts
Correctness Forensics for Batch Speculative Decoding: Diagnosing the Ragged Tensor Problem
Disappearing Ink: Obfuscation Breaks N-gram Code Watermarks in Theory and Practice
MobileDreamer: Generative Sketch World Model for GUI Agent
Not Your Typical Sycophant: The Elusive Nature of Sycophancy in Large Language Models
Proof2Hybrid: Automatic Mathematical Benchmark Synthesis for Proof-Centric Problems
MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
GUI-PRA: Process Reward Agent for GUI Tasks
ReGraP-LLaVA: Reasoning enabled Graph-based Personalized Large Language and Vision Assist…
Personas Differ from Native-Language Generation: Language Pathways Shape LLM Interpersona…
Stride-k Subsampling: Train-Free Audio Token Reduction for Whisper
A Composition-Aware Pretraining Framework for Geospatial Foundation Models
SimGuide: Typed Multi-Context User Representations for Preference-Conditioned Agent Plann…
Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space
Fully Distributed GNE Algorithms for Multi-Robot Placement without Consensus on Multiplie…
Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems
On the Prospects of Dynamic LLM Conversations in Software Development
A Mental Model Based Framework of Trust