Benchmarking LLM Judges for Mobile Agent Evaluation
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under Non-Ideal Ente…
JieZi: A Large-Scale Expert-Audited Dataset and Benchmark for Ancient Chinese Character E…
Uncertainty-Aware Probabilistic Constrained Clustering from Entangled Pairwise Supervision
CLEAR: Class-wise Expert Aggregation with Structured Sampling for Long-Tailed Classificat…
Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Uni…
Synchronizing Beliefs with Second-Order Theory-of-Mind in Human-Autonomy Teams (Extended …
Governing Agentic AI in FinTech
Domain-Aware Lightweight Spectral-Grouped Convolutions for Hyperspectral Fish Freshness C…
Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge
Dion3: Full-Stack Orthogonal Updates
Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes
Is Per-Agent Policy Composition Safe? Rethinking Successor-Feature Transfer in Cooperativ…
RoadWeaver: Large-Scale Lane-Level HD Map Generation from Scratch for Autonomous Driving …
Confidence Calibration of Deep Learning Systems
Better Slots, Better Worlds: Representation Quality & Robustness in Object-Centric World …
Towards the Harness of Embodied Agents
SCOPE: Subspace Clustering with Online Per-Head Top-K Estimation for Sparse Video Attenti…
CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives
ViTOED: A Dataset for Target-Oriented Emotion Detection on Vietnamese Social Media Texts
Dual-Stream Cross-Anchor Correction Grounding Long-Form Captions and the Domain Limits of…
A Cloud-Edge System for Multimodal Clinical Screening in Resource-Constrained Rural Setti…
Error-Aware Reverse Auction Mechanism for Large Language Model Routing
What We Learned by Reproducing 2,200 papers from ICML
Drive-to-Music: Context-Aware Generative Audio for In-Vehicle Experiences
Represent, Then Generate: Multimodal-Conditioned Time-Series Generation under Irregular M…
Attribute-Conditioned Multimodal Slot Factorization for Controllable Fashion Retrieval
Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images?
DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Languag…
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses