HoloCount: A Holistic Visual Counting Benchmark for MLLMs
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering
Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors
TS-RAG: Retrieval Augmented Generation for Time Series Forecasting
EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement …
MicroEvo: Knowledge-Guided LLM Sampling for Efficient Microarchitecture Design Space Expl…
Schema-Guided Hierarchical Information Extraction and Semantic Evaluation Using Generativ…
iARCS: Iterative Agentic RL for Controllable 3D Scene Generation
CogVis: Must Open-Vocabulary Change Detection Perceive the Scene Anew for Every Query?
FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial…
Contextual Information Policy Optimization for Search Agents
Poli-Bias: Understanding and Measuring Large Language Model Biases in International Polit…
Mind the Gaps: Mixture-of-Minds for Human Simulation
ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrai…
TriQua: Reconciling Granularity and Context in Factuality Evaluation
Project2Task: Graph-Guided Project-Level Planning for Autonomous Research
Small Foundation Models of Human Cognition and Behaviour
When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distill…
PD-GS: Phoneme-Driven 3DGS for Audio-Driven Talking Heads
SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents
Abstract Event Causal Rules: Induction and Application
From Continuous Predictors to Clinical Thresholds: Early Evidence on Performance Trade-of…
Woodpecker Distillation: Weak Models Diagnose Reasoning Bugs in Strong Models
Evaluating Investment Logic in Large Language Models: A Real-World Benchmark Towards Pers…
Signal or Spurious Cue? A Randomized Audit of Survey-Country Metadata in LLM Social Infer…
Integrating Implicit and Explicit Relational Biases through Graph-Based Multiple Instance…
SCP-NL2TL: Selective Conformal Prediction with Semantic Verification for Natural Language…
Adaptive Arena-based Contestable Argumentative Network-of-Experts for Open-Ended Care Pla…
DoctorAgents: an agentic framework to iteratively refine AutoML pipeline for small clinic…
Counterfactual Analysis via Large Language Models