Procedural Refinement by LLM-driven Algorithmic Debugging for ARC-AGI-2
Explorar
Noticias de IA
21270 elementos — filtrados, clasificados y sin duplicados
RCTs & Human Uplift Studies: Methodological Challenges and Practical Solutions for Fronti…
Is Human Annotation Necessary? Iterative MBR Distillation for Error Span Detection in Mac…
MARS: Margin and Semantic-Aware Data Augmentation for Reward Modeling
You Can Ground Earlier than See: An Effective and Efficient Pipeline for Temporal Sentenc…
Distilling Game Code World Model Generation into Lightweight Large Language Models
Kolmogorov-Arnold Fourier Networks
Methodology for Creating a Clinically Verified Dermoscopic Image Dataset
JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment
Guess the Unified Model: How Much Can We Recover from Generated Images?
Eureka: Intelligent Feature Engineering for Enterprise AI Cloud Resource Demand Prediction
Second Guess: Detecting Uncertainty Through Abstention and Answer Stability in Small Lang…
SomaliBench Eval: Measuring English-to-Somali Refusal Gaps in Open-Weight Language Models
Meta-Agent: From Task Descriptions to Verified Multi-Agent Systems
Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation
TopoAlign: Topology-Aware Visual Representation Alignment
NeurIPS: Neuro-anatomical Inductive Priors for Sphere-based Brain Decoding
Extreme Region Policy Distillation
Towards the Connection between Activation Sparsity and Flat Minima
Self-supervised Hierarchical Visual Reasoning with World Model
Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs
BoxLitE: A Faithful Knowledge Base Embedding Based on Convex Optimization
SoK: DARPA's AI Cyber Challenge (AIxCC): Competition Design, Architectures, and Lessons L…
Operationalizing Reconstructive Authority: Runtime Construction, Dependency Resolution, a…
BODHI: Precise OS Kernel Specification Inference
Reducing Credit Assignment Variance via Counterfactual Reasoning Paths
LLMs Uncertainty Quantification via Adaptive Conformal Semantic Entropy
AutoSG: LLM-Driven Solver Generation Solely from Task Prompts for Expensive Optimization
Solving Combinatorial Counting Problems with Weighted First-Order Model Counting
AI as Equalizer or Amplifier? Task Complexity as the Moderating Factor for Human Expertis…