Voila: Evaluation of MLLMs for perceptual understanding and analogical reasoning
Abstract Multimodal Large Language Models (MLLMs) have become a powerful tool for
integrating visual and textual information. Despite their exceptional performance on visual
understanding benchmarks, measuring their ability to reason abstractly across multiple
images remains a significant challenge. To address this, we introduce VOILA, a large-scale,
open-ended, dynamic benchmark designed to evaluate MLLMs' perceptual understanding
and abstract relational reasoning. VOILA employs an analogical mapping approach in the …
integrating visual and textual information. Despite their exceptional performance on visual
understanding benchmarks, measuring their ability to reason abstractly across multiple
images remains a significant challenge. To address this, we introduce VOILA, a large-scale,
open-ended, dynamic benchmark designed to evaluate MLLMs' perceptual understanding
and abstract relational reasoning. VOILA employs an analogical mapping approach in the …