← Todas las noticias

Hallucination in Multimodal Foundation Models: A Survey on Causes, Corrections, and Evaluations

arXiv:2410.15359v2 Announce Type: replace Abstract: Multimodal Foundation Models represent a significant leap in artificial intelligence. Among them, Large Vision-Language Models (LVLMs) serve as the typical representative of these foundation models, which integrate visual modality directly into Large Language Models (LLMs). They have demonstrated strong capabilities in information processing and generation. However, the existence of hallucinations has limited the potential and practical effectiveness of LVLM in various fields. Although lots of work has been devoted to hallucination mitigation and correction, there are few reviews to summarize them. To address this gap, this survey provides a systematic review of the hallucination landscape in LVLMs. We categorize the causes related to model architecture and data quality, and construct a comprehensive taxonomy of existing mitigation strategies. Furthermore, we critically assess current hallucination evaluation benchmarks from both discriminative and generative perspectives, highlighting the limitations of existing metrics. This survey concludes by discussing open challenges and future research directions to advance the reliability and trustworthiness of LVLMs.
Leer el original en arXiv cs.AI →