What We are Missing in Multimodal LLM Evaluation?
Explorar
Noticias de IA
22318 elementos — filtrados, clasificados y sin duplicados
How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks?
Hierarchical Fault Detection and Diagnosis for Transformer Architectures
Mapping License Plate Recoverability Under Extreme Viewing Angles for Opportunistic Urban…
S2P-Net: A Spectral-Spatial Polar Network for Rotation-Invariant Object Recognition in Lo…
Peer-Preservation in Frontier Models
TransXion: A High-Fidelity Graph Benchmark for Realistic Anti-Money Laundering
Statistical Properties of the King Wen Sequence: An Anti-Habituation Structure That Does …
Power Couple? AI Growth and Renewable Energy Investment
MedPruner: Training-Free Hierarchical Token Pruning for Efficient 3D Medical Image Unders…
A Guideline-Aware AI Agent for Zero-Shot Target Volume Auto-Delineation
Residual RL-MPC for Robust Microrobotic Cell Pushing Under Time-Varying Flow
Delegation and Verification Under AI
VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning
ReportLogic: Evaluating Logical Quality in Deep Research Reports
VecSet-Edit: Unleashing Pre-trained LRM for Mesh Editing from Single Image
Dual-Prototype Disentanglement: A Context-Aware Enhancement Framework for Time Series For…
Metaphors are a Source of Cross-Domain Misalignment of Large Reasoning Models
Digital Twin-Driven Communication-Efficient Federated Anomaly Detection for Industrial IoT
Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Sup…
Limited Reference, Reliable Generation: A Two-Component Framework for Tabular Data Genera…
Reconstruction Alignment Improves Unified Multimodal Models
Learning to Select Maximum Clique Algorithms: From Traditional Machine Learning to a Dual…
DMSC: Dynamic Multi-Scale Coordination Framework for Time Series Forecasting
Tuning Language Models by Mixture-of-Depths Ensemble
R2D-RL: A RoboCup 2D Soccer Environment for Multi-Agent Reinforcement Learning
Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM Agents
SciFig: Towards Automating Editable Figure Generation for Scientific Papers
Human-AI Complementarity: A Goal for Amplified Oversight
Autoregressive Boltzmann Generators