EA-Graph: Artifact-Anchored Verification Memory for Coding Agents under Upstream Drift
Explorar
Noticias de IA
21757 elementos — filtrados, clasificados y sin duplicados
Strategic Evaluation of Planning Strategies for LLM Agents in Cyber-Physical Systems
Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary
Out-Of-The-Loop Multi-Fidelity Bayesian Optimization
Provable Limits and Certified Deferral for Verbalized Uncertainty in Small Language Models
Training-Free Hashing-Based Attention via Binary Principal Components
Is Monitoring Enough? Strategic Agent Selection For Stealthy Attack in Multi-Agent Discus…
Can Post-Training Transform LLMs into Causal Reasoners?
VQ-VAD: Vector-quantized Motion Representation Learning for Human-centric Video Anomaly D…
TRNet: Topography-Guided Frequency Rectification and Structure-Aware Decoding for Multimo…
Behavioral Skill Reconstruction: Reconstructing Hidden Functionality from LLM Agent Skills
Patients-like-me: A Variational LM--GNN Framework for Explainable Clinical Prediction
MultiPathFormer: Towards a Foundation Model for Multipath Wireless Propagation
LiNC: Lightweight Noise Correction via Per-Sample Trust and Gaussian Mixture Modeling
Corrigibility Transformation: Constructing Goals That Accept Updates
DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
FBID: Adaptive Personalized Federated Learning for Robust Out-of-Distribution Attack Dete…
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Tra…
Spatiotemporal Graph Transformer for Traffic Intelligence in Edge Computing
Interpretable Fuzzy Inference for UAV Target Tracking Using Bounding-Box Geometry
Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings
Beyond Linear Dynamics: Neural Bilinear Dynamical Models for Time Series Forecasting
A Model Merging Approach for Continual MLLM Unlearning
TourSynbio-Search: A Large Language Model Driven Agent Framework for Unified Search Metho…
FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents
Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports
Design Choices That Matter: A Functional ANOVA Analysis for Remote Sensing Multi-Label Cl…
The Luna Bound Propagator for Formal Analysis of Neural Networks
Modality Agreement- and Conflict-Aware Prototype Hypergraph Learning for Multimodal Inten…
Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools