Woodpecker Distillation: Weak Models Diagnose Reasoning Bugs in Strong Models
Explorar
Noticias de IA
21271 elementos — filtrados, clasificados y sin duplicados
Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors
Evaluating Investment Logic in Large Language Models: A Real-World Benchmark Towards Pers…
iARCS: Iterative Agentic RL for Controllable 3D Scene Generation
SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents
Schema-Guided Hierarchical Information Extraction and Semantic Evaluation Using Generativ…
EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement …
Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering
MicroEvo: Knowledge-Guided LLM Sampling for Efficient Microarchitecture Design Space Expl…
Signal or Spurious Cue? A Randomized Audit of Survey-Country Metadata in LLM Social Infer…
TS-RAG: Retrieval Augmented Generation for Time Series Forecasting
RA-CAD: Learning Post-Execution Critique for State-Aware Text-to-CAD Generation
Skill Neologisms: Towards Skill-based Continual Learning
CogVis: Must Open-Vocabulary Change Detection Perceive the Scene Anew for Every Query?
DoctorAgents: an agentic framework to iteratively refine AutoML pipeline for small clinic…
Adaptive Arena-based Contestable Argumentative Network-of-Experts for Open-Ended Care Pla…
Abstract Event Causal Rules: Induction and Application
Coherence-Oriented Dream Scene Visualisation
Integrating Implicit and Explicit Relational Biases through Graph-Based Multiple Instance…
Mind the Gaps: Mixture-of-Minds for Human Simulation
Reducing belief in conspiracy theories as they unfold using large language models
When Do Prompt-Side Agent Playbooks Transfer? Accuracy, Cost, and Runtime Shift in Agent …
Poli-Bias: Understanding and Measuring Large Language Model Biases in International Polit…
ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models for Multimodal Knowledge …
Automatic Detection of Deaths from Social Networking Sites
Marginal Matching Does Not License Factorized Sampling: Auditing Conditional Style Leakag…
Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs?
Explore or Converge? Stage-Guided Per-Step Optimization for Diffusion Models
CubicQuant: Parametric Non-Uniform Codebooks for High-Throughput LLM Inference with 1-8-B…
bioMoR: Biology-Guided Mixture-of-Recursions for Effective Genomic Learning