Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams
Explorar
Noticias de IA
21010 elementos — filtrados, clasificados y sin duplicados
HexEval: An Evidence-Driven Hexagonal Framework for Multidimensional Scholar Assessment
A Forced-Structure Reduction and Verifiable Bounds for Conway's 99-Graph
Credo: Declarative Control of LLM Pipelines via Beliefs and Policies
SteeringSafety: Benchmarking Representation Steering in LLMs Across Safety Perspectives
VLM2Rec: Resolving Modality Collapse in Vision-Language Model Embedders for Multimodal Se…
Causal Agent based on Large Language Model
OpenAg: Democratizing Agricultural Intelligence
RecSys Factory: Bounding LLM Agent Autonomy to Decision Points in the Industrial Recommen…
The Off-Support Barrier: Why Semantic Safety Constraints Are Not Learning-Problem Invaria…
MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning
ReXrank: A Public Leaderboard for AI-Powered Radiology Report Generation
User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware Autoscaling
HYDRA: Hyperbolic Dynamic Representation Architecture for Kolmogorov-Arnold Networks
Two-Stage Deformable-Convolutional Inverse Design of Nanophotonic Absorbers from Optical …
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG Serving
DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation
RealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI Frameworks
Semantic Lenia: Emergence of Homeostatic Solitons within the Semantic Space of Large Lang…
Co-constructing sociotechnical AI governance: participatory system mapping using algorith…
Forecasting Side Effects of Activation Steering
Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Uni…
Domain-Aware Lightweight Spectral-Grouped Convolutions for Hyperspectral Fish Freshness C…
Accuracy and Order Sensitivity Diverge Under Label-Free Strategies
Dion3: Full-Stack Orthogonal Updates
Confidence Calibration of Deep Learning Systems
Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge
DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Languag…
Advancing MLLM-based UAV Image Understanding and Reasoning: A Benchmark and a Training-Fr…