EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Pre…
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
Silent Failures in Multimodal Agentic Search:A Diagnostic Taxonomy and Cross-Judge Evalua…
Cross-Subject Semantic Decoding with Shared-Space Alignment for Generalized Neural Repres…
Simulating Eutopia: Revisiting Long-term Fairness with Outcomes, Performativity, and Dyna…
Scale-Aware Learning of Chaotic Dynamics on Unstructured Meshes via Binned Spectral Losses
Economic Evaluations of Language Models
Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Expla…
Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight La…
PRIME-SVR: Physics-infoRmed Implicit Multi-Echo Slice-to-Volume Reconstruction for Fetal …
NEXUS: Structured Runtime Safety for Tool-Using LLM Agents
OpenEvoShield: Dual Non-Stationary Continual Defense for Open-World Multi-Agent System At…
An Intelligent-Cloud Edge Multimodal Interaction System for Robots
Knowledge-Centric Self-Improvement
ITPEval: Benchmarking Formal Translation Across Interactive Theorem Provers
Geometry-Guided Constraint Learning for LLM Safety Classification
Formal Foundations for Known Good Reliable Die Screening in Chiplet-Based AI Systems-on-C…
SenWorld: A Digital-Twin Simulation for Generating Context-Rich Evaluation Data
Sidewalk Moments: Are Richer Representations Always More Human-Aligned? Evidence from Cit…
Anti-Goal Reasoning: Rethinking the Theory of Goal Reasoning in Non-Axiomatic Logic
DINO-VPT: Hierarchical Visual Prompt Tuning for Joint Physical-Digital Face Anti-Spoofing
Engine-Native Editable 3D World Reconstruction with Objects and Lighting
Webly Supervised Multi-Label Recognition: Evaluation Benchmark and Dual-Branch Multi-Labe…
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?
Probabilistic Residual Learning for Online Recommendations
Code Monitor Red Teaming for Public-Test-Passing Code
Offline RL with Hierarchical Action Chunking
REFACT: Adaptive Fact Restatement for Compact and Faithful Chain-of-Thought Reasoning
New Complexity-Theoretic Frontiers of Tractability for Neural Network Training
Profiling Lightweight Large Language Models