Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Informati…
Explorar
Noticias de IA
21271 elementos — filtrados, clasificados y sin duplicados
RW-TTT: Batched Serving for Request-Owned Test-Time Training State
MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Aud…
The Missing Piece in Pre-trained Model Evaluation: Reward-Guided Decoding Unlocks Task-Or…
ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains
KSAFE-MM: A Multimodal Safety Benchmark via Localized Contextualization for Korean Cultur…
Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training
AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scena…
MEMO: A Modular Framework for Training a Dedicated Memory Model on New Knowledge Without …
An Evolutionary Approach for Designing Stable and Highly Expressible Low-Immunogenicity T…
KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evalua…
Continual Learning in Modern Hopfield Networks with an Application to Diffusion Models
ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning
Skill-as-Pseudocode: Refactoring Skill Libraries to Pseudocode for LLM Agents
Is Backpropagation Optimal? When Synthetic Gradients Improve Sample Efficiency
GeneralThinker: Domain-General Reasoning through Likelihood-Guided Answer-Conditioned Opt…
When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?
BrickAnything: Geometry-Conditioned Buildable Brick Generation with Structure-Aware Token…
Inference-Time Search Using Side Information for Diffusion-Based Image Reconstruction
HiSpec: Hierarchical Speculative Decoding for LLMs
EconCausal: A Context-Aware Economic Reasoning Benchmark for Large Language Models
Mechanistic Interpretability of Antibody Language Models Using SAEs
Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling
L2Rec: Towards Dual-View Understanding of LLMs for Personalized Recommendation
SEAL: Self-Evolving Agentic Learning for Conversational Question Answering over Knowledge…
Towards Generalization-Oriented Models for Vehicle Routing Problems with Mixture-of-Exper…
SeDT: Sentence-Transformer Decision-Transformer Conditioning for Multi-Turn Conversation …
HTMLCure: Turning Browser Experience into State Guided Repair for Interactive HTML
The Kalman Evolve: Closing the Gap in Kalman Filtering via Interpretable Algorithm Discov…
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models