Personalized Auto-Research: Towards a True AI Co-Scientist
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
Frontier AI Forecasting Has a Measurement Problem: An Audit of Progress Evidence
LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays …
Small Models Scout Bottleneck Order for Large-Model Data Control
Skill Blocks: How Should an Agent Load Its Skill? A Caching-Correct Comparison of Pre-loa…
Trust Is Not Enough: Influence Calibration for On-Policy Self-Distillation in Agentic RL
T-LLM Compiler: Trusted LLM-based Code Optimization and Verification Framework
Does a Tool Result Carry More Authority Than Plain Text? Three Prospective Studies of Fal…
S2-MoE: Enabling Efficient Self-Speculative Decoding for Mixture-of-Experts on Edge Devic…
Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form
Enhancing Science Classroom Discourse Analysis through Joint Multi-Task Learning for Reas…
TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation
Andy: A Mathematical Agent for Rigorous Proof and Autonomous Research
CEDAR-GRPO: Process-Aware Reinforcement Learning for General Abductive Reasoning in LLMs
RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
GraphLoom: Reliability-Calibrated Graph Evidence Routing for Multimodal KG-RAG
Funnel of Thoughts: Efficient Test-Time Scaling via Early Voting and Rollout Pruning
Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents
Beyond Thresholds: A Quality-Aware Decision Intelligence Framework for Cold Chain IoT Sys…
StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Ha…
Agentic Data Cleaning Without a Clean Reference: An Experimental Study of Capabilities an…
Advanced modelling and data analytics in aviation
Second-Order Policy Effects as State Transitions: A Source-Linked Benchmark for Policy Si…
Anatomy of a Quantized Agent: VRAM Stability and Forecasting in Code-Synthesis Agentic Wo…
Flow Motion Policy: Manipulator Motion Planning with Flow Matching Models
I-CALM: Incentivizing Confidence-Aware Abstention for LLM Selective Answering
Beyond Pass@k: Measuring Reliability and Security of Agentic Code Generation
VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models
ReForge: Keeping ABR Algorithms Never Finished with Verified Large Language Model Edits
Translating finite-domain integer constraint models to CP/SMT/ILP/PB/SAT solvers with CPM…