Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Pe…
Explorar
Noticias de IA
22318 elementos — filtrados, clasificados y sin duplicados
GRASP: GRanularity-Aware Search Policy for Agentic RAG
Co4ICF: Co-evolving Physics-Informed Surrogate and RL-based Pulse Optimizer for Inertial …
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory
Measure the Sim-to-Real Gap: Designing an Affordable Real-World Benchmark Platform for Re…
SPARK: Susceptibility-Guided Profiling and Steering of Latent Reasoning States in Large L…
Can Agentic Trading Systems Pay for Their Own Intelligence?
Behavioural Signatures of Risk-Sensitive Decision-Making in Large Language Models
When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Base…
KGCQual: An Interpretable Framework for Evaluating the Knowledge Graph Construction Quali…
GRATE: Temporal Extensions for Inductive KG Foundation Models via Gated Rotary Attention
UNIT: Unleash Large Language Models Potential for Graph Continual Learning
IdeaTrail: Full-Process Agent Trajectories for Scientific Ideation
Dynamic Agent Skills: A Lifecycle Survey and Taxonomy of Evolving Skill Libraries
Looped State-Space Language Models with Adaptive Exit-State Selection
MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation
From ambiguous utterances to governed reuse classes: canonicalization, quotient invarianc…
AgentAbstain: Do LLM Agents Know When Not to Act?
A Symbolic Neural CPU for Quantization-Simulated Writeback and Interpretable Program Exec…
TopoExplore: Topological Discrimination for Archive-Based Exploration
Exploring Agentic Workflows for Generating High Quality Math Visual Aids
Agentic Context Learning with Self-Discovered Specification
PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input…
EvoCUA-1.5: Online Reinforcement Learning for Multi-turn Computer-Use Agents
Verification of Adaptive Agentic Controllers through Finite Rule Revision
Norm Enforcement for AI Agents: Robustly Shaping Behavior in Multi-Agent Systems
BatteryLake: Agentic, Physics-Grounded Curation of Heterogeneous Battery Aging Data and B…
Task-Conditioned Synthetic Data Generation for Improving Machine Learning Performance in …
A Dynamic Scene Interaction Reasoning Framework for Scene-level Lane-Change Intention and…
Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Bench…