BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economic…
Explorar
Noticias de IA
29343 elementos — filtrados, clasificados y sin duplicados
A Study of the Scale Invariant Signal to Distortion Ratio in Speech Separation with Noisy…
VGGSounder: Audio-Visual Evaluations for Foundation Models
Time Series Forecasting as Reasoning: A Slow-Thinking Approach with Reinforced LLMs
Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics
MesaNet: Sequence Modeling by Locally Optimal Test-Time Training
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Abi…
Speculative Thinking: Enhancing Small-Model Reasoning with Large Model Guidance at Infere…
LaVIDE: Language-Prompted Satellite Change Detection via Map-Image Alignment
SSSD: Simply-Scalable Speculative Decoding
Transformer-Based Autonomous Driving Models and Deployment-Oriented Compression: A Survey
Test-time reward-guided alignment of language models by importance sampling on pre-logit …
Bad Seeing or Bad Thinking? Rewarding Perception for Multimodal Reasoning
Belief-Aware VLM Model for Human-like Reasoning
KITE: Kernelized and Information Theoretic Exemplars for In-Context Learning
Bilevel Autoresearch: Meta-Autoresearching Itself
Interfaze: The Future of AI is built on Task-Specific Small Models
Reasoning or Fluency? Dissecting Probabilistic Confidence in Best-of-N Selection
A Unified Geometric Space for Topological Alignment Between Transformer-Based Models and …
BRAINCELL-AID: An Agentic AI Created Brain Cell Type Resource for Community Annotation
Aligning Deep Implicit Preferences by Learning to Reason Defensively
Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification?
Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning
Streaming Communication in Multi-Agent Reasoning
GeM-NR: Geometry-Aware Multi-View Editing for Nonrigid Scene Changes
Who Needs Labels? Adapting Vision Foundation Models With the Metadata You Already Have
UniCAD: A Unified Benchmark and Universal Model for Multi-Modal Multi-Task CAD
Self-Reflective APIs: Structure Beats Verbosity for AI Agent Recovery
DAR: Deontic Reasoning with Agentic Harnesses
From Agent Traces to Trust: Evidence Tracing and Execution Provenance in LLM Agents