ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
Explorar
Noticias de IA
30724 elementos — filtrados, clasificados y sin duplicados
From Distances to Trajectories: Real-Time Signed Distance Function Mapping and Distance-A…
Supra Cognitive Modes: A Routed Architecture for Agent Memory
OpenRTAG: A Comprehensive Benchmark for Robust Text-Attributed Graph Learning under Data …
Decode-Time Grammars: Constrained LLM Generation over a Refinement Order of Grammar Fragm…
CODENS: Transforming Code Changes into Living, Accessible, and Queryable Documentation
CPGRec+: A Balance-oriented Framework for Personalized Video Game Recommendations
Hazard or Anomaly? Evaluating VLMs for Understanding Dangers and Discrepancies
Sequential Learner Modeling Using Multi-Relational Graph Convolutional Networks
MIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical …
Hyperdimensional Probe: Decoding LLM Representations via Vector Symbolic Architectures
ImplicitRDP: An End-to-End Visual-Force Diffusion Policy with Structural Slow-Fast Learni…
LinguistAgent Technical Report: A Reflective Multi-Model Platform for Automated Linguisti…
CompilerKV: Risk-Adaptive KV Compression via Offline Experience Compilation
Fly0: Persistent Metric Anchoring for Zero-Shot Aerial Vision-Language Navigation
Give Them an Inch and They Will Take a Mile:Understanding and Measuring Caller Identity C…
When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines
Large Language Models Explore by Latent Distilling
Global Automation Atlas
Tunable MAGMAX: Preference-Aware Model Merging for Continual Learning
LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes
Team RAS in 11th ABAW Competition: Multimodal Ambivalence Recognition Approach
Robust Reasoning Benchmark
Why Do Vision Language Models Struggle To Recognize Human Emotions?
OntoBook: Ontology-Grounded Synthetic Textbooks for Medical Encoder Pretraining
NaviAIS: A Scenario-Level Vessel Trajectory Prediction Dataset withVectorized Lane Priors…
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
SkillSight: Seeing Through Shared Descriptions for Accurate Skill Retrieval
DWM: Separating World Effects from Actions in Latent World Models
Lifting Embodied World Models for Planning and Control