Superficial Beliefs in LLM Decision-Making
Explorar
Noticias de IA
22116 elementos — filtrados, clasificados y sin duplicados
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World…
What Fits (Into Few Tokens) Doesn't Overfit: Compression and Generalization in ML Researc…
Monte Carlo Pass Search: Using Trajectory Generation for 3D Counterfactual Pass Evaluatio…
ReasonAlloc: Hierarchical Decoding-Time KV Cache Budget Allocation for Reasoning Models
Failure Modes of Deep Multi-Agent RL in Asynchronous Pricing: Reproducible Triggers, Trac…
Transformer Based Model for Spatiotemporal Feature Learning in EEG Emotion Recognition
++nnU-Net: Scaling nnU-Net with Prefix-Based Data Augmentation
From Volume to Value: Preference-Aligned Memory Construction for On-Device RAG
SHAPO: Sharpness-Aware Policy Optimization for Safe Exploration
Linguistically Augmented Audio Speech Data (LinguAS)
Unifying Data, Memory, and Compute Efficiency in LLM training: A Survey
Multi-Level Analyzation of Imbalance to Resolve Non-IID-Ness in Federated Learning
Co-GLANCE: Uncertainty-Aware Active Perception for Heterogeneous Robot Teaming
What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agen…
Towards Robust Arabic Speech Emotion Recognition with Deep Learning
The Confident Liar: Diagnosing Multi-Agent Debate with Log-Probabilities and LLM-as-Judge
Content-Induced Spatial-Spectral Aggregation Network for Change Detection in Remote Sensi…
Building Change Detection in Earthquake: A Multi-Scale Interaction Network and A Change D…
Democratising Camera Trap AI: An Open-Source Model for Detecting UK Mammals
Diffusion Forcing Planner: History-Annealed Planning with Time-Dependent Guidance for Aut…
Test-time Adversarial Takeover: A Real-time Hijacking Interface against Robotic Diffusion…
Agentic Hybrid RAG for Evidence-Grounded Muon Collider Analysis
CLP: Collocation-Length Prediction for Zero-Loss Adaptive Multi-Token Inference
PreAct-Bench: Benchmarking Predictive Monitoring in LLMs
SkillResolve-Bench: Measuring and Resolving Same-Capability Ambiguity in Agent Skill Retr…
Harnessing the Collective Intelligence of AI Agents in the Wild for New Discoveries
AuRA: Internalizing Audio Understanding into LLMs as LoRA
Recoverable but Not Stationary:Local Linear Structures in Weights and Activations
Hierarchical Policies from Verbal and Egocentric Human Signals for Natural Human-Robot In…