SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscri…
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning
Borrowed Strength: Best-of-N Search over a Code EncodingBreaks Self-Check Jailbreak Defen…
The Human Utility Factor: A Computable Welfare Metric That Reframes AI Governance as a Co…
Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-…
Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio …
Property-driven Causal Abstractions for Markov Decision Processes
Anatomy Contextualized Adaption of CT Foundation Models
Improving Item Discoverability in e-Commerce Search via Related Intent Generation
The Social Cost of an AI Teammate: How an Artificial Teammate Reshapes Human-Human Commun…
APEX-Accounting
Balancing Centralized Learning and Distributed Self-Organization: A Hybrid Model for Embo…
Bridging the Gap in Ophthalmic AI: MM-Retinal-Reason Dataset and OphthaReason Model towar…
HealthSLM-Bench: Benchmarking Small Language Models for Mobile and Wearable Healthcare Mo…
Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests
BioPro: Towards Difference-Aware Gender Fairness for Vision-Language Models
How does downsampling affect needle electromyography signals? A generalisable workflow fo…
MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval
The Capability Paradox: How Smarter Auditors Make Multi-Agent Systems Less Secure
Shared Symbolic Backbones for Physically Consistent Multi-Output Symbolic Regression
Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition
A Picture Says Thousands of Words - Harnessing Dermal Exposure Data from Images through H…
GPT-Red: Automated Red Teaming via Self-Play at Scale
A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models
Beyond Block Boundaries: Multi-Block Editing for Diffusion Large Language Models
The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Stat…
TRACE-CTI: Auditable Post-Extraction Governance of TTP Claims with Knowledge Graphs
One-Frame Calibration with Siamese Network in Facial Action Unit Recognition
Defending Against Backdoor Attacks via Alignment Checking in Model-Contrastive Federated …
BayesAME: Bayesian Active Model Evaluation