Voila: Evaluation of MLLMs for perceptual understanding and analogical reasoning

N Yilmaz, M Patel, L Luo, T Gokhale… - International …, 2025 - proceedings.iclr.cc
Abstract Multimodal Large Language Models (MLLMs) have become a powerful tool for
integrating visual and textual information. Despite their exceptional performance on visual
understanding benchmarks, measuring their ability to reason abstractly across multiple
images remains a significant challenge. To address this, we introduce VOILA, a large-scale,
open-ended, dynamic benchmark designed to evaluate MLLMs' perceptual understanding
and abstract relational reasoning. VOILA employs an analogical mapping approach in the …