← Все новости

The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models

arXiv:2604.14363v2 Announce Type: replace-cross Abstract: Multimodal language models systematically underperform on visual perception tasks, yet the structure underlying this failure remains poorly understood. We propose centroid replacement, mapping tokens to their nearest K-means centroid and removing within-cluster residual structure, as a controlled probe for modal dependence. Across seven models spanning four architecture families, post-image text replacement costs 4x more accuracy than visual replacement on six BLINK tasks; VPBench and MedBLINK preserve this aggregate ordering. Since text replacement also disrupts the task and answer interface, the gap measures relative dependence. Text centroid contrastive decoding (TCCD) exploits this asymmetry without retraining, recovering up to +16.9% accuracy on an individual task at oracle best-per-task $\alpha_\text{interp}$; every model gains on at least one task, although gains vary across tasks and models. Together, centroid replacement provides a retraining-free audit of modal dependence on labeled data, while TCCD offers a reusable inference-time intervention. Code and artifacts are available at: https://github.com/yahskapar/centroid-erasure
Читать оригинал на arXiv cs.AI →