Litmus: Zero-Label, Code-Driven Metric Specification for Evaluating AI Systems
Explorar
Noticias de IA
22318 elementos — filtrados, clasificados y sin duplicados
From Speech to Text Corpora: Evaluating ASR-Based Data Acquisition for Low-Resource Fongb…
EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning
A Matter of Time: Towards a General Theory of Agency
Context-Aware Generative AI for Automated Telecom Test Script Generation
Contrastive and Adaptive Multi-modal Masked Autoencoder for Spatial Transcriptomics
From Handcrafted Features to Functional Edge Learning: Evolution of EEG Seizure Detection…
Can Reasoning Models Detect Changes to their Chains of Thought?
When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation an…
CFAgentBench: A Reproducible Environment and Benchmark for Autonomous Construction-Financ…
Human vs Machine Mathematical Difficulty on Project Euler: An Experimental Analysis
Holmes: Multimodal Agentic Diagnosis for Mixed-Language Mobile Crashes at Industrial Scale
Measuring What Persists: Conditioning Mechanisms and a Geometric Framework for AI Agent I…
Benchmarking Robot Memory Under Interference
Reliability-Guided Adaptive Ensembling for Robust Test-Time Adaptation
On the Sparsity-Storage-Accuracy Tradeoff in Parsimoniously Activated Dictionary Learning
Structured Hyperedge Adaptation for Parameter-Efficient Fine-Tuning of Vision Transformers
Reinforcement learning to improve large language model-based automated code compliance sy…
Gold Points Sniper: Self-guided Visual Reasoning in VLM for Fine-grained Action Understan…
DreamUV: Unwrap Artist-like UV by End-to-End Flow Matching
CASPER in the Machine: Insights into Character Variety in LLM-Generated Stories
Not All Claims Are Equally Risky: FACTOR for Adaptive Verification in Factual Long-Form G…
Human and AI collaboration for pulmonary nodule segmentation
An LLM-Orchestrated Agent for Directional-Coupler Design with Self-Consistent Eigenmode a…
Enabling Cloud-Level Accuracy in Edge AI through IoT Data Preprocessing
Fed-CausalDiff: Decoupled Synchronization for Federated Do-Simulation and Policy Evaluati…
Training the Orchestrator: A Supervised Approach to End-to-End PDDL Planning with LLM Age…
Training-Free Semantic Correction for Autoregressive Visual Models
Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do
Concept-Constrained Prompt Learning for Few-Shot CLIP Adaptation