Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering
Explorar
Noticias de IA
30691 elementos — filtrados, clasificados y sin duplicados
DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
CrackedPDFs: A Controlled Benchmark for Hidden Prompt Injection in PDFs
Silent Failures in Multimodal Agentic Search:A Diagnostic Taxonomy and Cross-Judge Evalua…
A Systematic Benchmark of Intensity Normalisation Methods for 3D Knee MRI Segmentation an…
Cross-Subject Semantic Decoding with Shared-Space Alignment for Generalized Neural Repres…
Simulating Eutopia: Revisiting Long-term Fairness with Outcomes, Performativity, and Dyna…
Knowledge-Centric Self-Improvement
ITPEval: Benchmarking Formal Translation Across Interactive Theorem Provers
Taming the Security-Energy Paradox: A Green AI Approach to Optimized Android Malware Dete…
Are Attributions of Consciousness to AI Chatbots Epistemically Innocent?
RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Gener…
Convergence-Latency-Aware Adaptive Modulation and Resource Allocation in RIS-Assisted Wir…
MoA-Structured Decode Attention DNF Derivation, KV-Cache Accumulation, GQA/MQA, and OpenA…
Geometry-Guided Constraint Learning for LLM Safety Classification
Scale-Aware Learning of Chaotic Dynamics on Unstructured Meshes via Binned Spectral Losses
SubQuad: Near-Quadratic-Free Structure Inference with Distribution-Balanced Objectives in…
An Intelligent-Cloud Edge Multimodal Interaction System for Robots
NeuroSymActive: Differentiable Neural-Symbolic Reasoning with Active Exploration for Know…
The Chronos Vulnerability: A Taxonomy of Temporal Persistence and Memory-Based Deception …
Leveraging Offline Supervision for Efficient and Generalizable Reinforcement Learning in …
Economic Evaluations of Language Models
Recovering Clinical Utility Under Differential Privacy: Empirical Validation of Adaptive …
Structured Latent Space Modeling over Multi-Scale Temporal Patches for Multivariate Time …
BRIM: Workload-Balanced Dual-Sided Bit-Serial Sparse Inference Accelerator
NEXUS: Structured Runtime Safety for Tool-Using LLM Agents
Measuring LLM Trust Allocation Across Conflicting Software Artifacts
FedLSG: LLM-Enhanced Semantic Calibration for Federated Graph Backdoor Defense
Overview of FinMMEval 2026 Task 1: Multilingual Financial Multiple-Choice Question Answer…
Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Expla…