User profiles for Geewook Kim

Geewook Kim

NAVER Cloud AI
Verified email at navercorp.com
Cited by 2499

Ocr-free document understanding transformer

G Kim, T Hong, M Yim, JY Nam, J Park, J Yim… - … on Computer Vision, 2022 - Springer
Understanding document images (eg, invoices) is a core but challenging task since it requires
complex functions such as reading text and a holistic understanding of the document. …

What is wrong with scene text recognition model comparisons? dataset and model analysis

J Baek, G Kim, J Lee, S Park, D Han… - 2019 IEEE/CVF …, 2019 - ieeexplore.ieee.org
Many new proposals for scene text recognition (STR) models have been introduced in recent
years. While each claim to have pushed the boundary of the technology, a holistic and fair …

Prometheus-vision: Vision-language model as a judge for fine-grained evaluation

S Lee, S Kim, S Park, G Kim, M Seo - Findings of the Association …, 2024 - aclanthology.org
Assessing long-form responses generated by Vision-Language Models (VLMs) is challenging.
It not only requires checking whether the VLM follows the given instruction but also …

[PDF][PDF] Donut: Document understanding transformer without ocr

G Kim, T Hong, M Yim, J Park, J Yim… - arXiv preprint arXiv …, 2021 - sangdooyun.github.io
Understanding document images (eg, invoices) has been an important research topic and
has many applications in document processing automation. Through the latest advances in …

How does vision-language adaptation impact the safety of vision language models?

S Lee, G Kim, J Kim, H Lee, H Chang… - International …, 2025 - proceedings.iclr.cc
Vision-Language adaptation (VL adaptation) transforms Large Language Models (LLMs)
into Large Vision-Language Models (LVLMs) for multimodal tasks, but this process often …

Cost-effective end-to-end information extraction for semi-structured document images

W Hwang, H Lee, J Yim, G Kim… - Proceedings of the 2021 …, 2021 - aclanthology.org
A real-world information extraction (IE) system for semi-structured document images often
involves a long pipeline of multiple modules, whose complexity dramatically increases its …

On text localization in end-to-end ocr-free document understanding transformer without text localization supervision

G Kim, S Yokoo, S Seo, A Osanai, Y Okamoto… - … on Document Analysis …, 2023 - Springer
This paper presents a simple yet effective approach for weakly supervised text localization
in end-to-end visual document understanding (VDU) models. The traditional approach in …

Do modern video-llms need to listen? a benchmark audit and scalable remedy

G Kim, M Seo - arXiv preprint arXiv:2509.17901, 2025 - arxiv.org
Speech and audio encoders developed over years of community effort are routinely excluded
from video understanding pipelines, not because they fail, but because benchmarks never …

Visually-situated natural language understanding with contrastive reading model and frozen large language models

G Kim, H Lee, D Kim, H Jung, S Park… - Proceedings of the …, 2023 - aclanthology.org
Recent advances in Large Language Models (LLMs) have stimulated a surge of research
aimed at extending their applications to the visual domain. While these models exhibit promise …

K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts

…, C Lee, K Jang, J Kim, E Kim, W Cho, S Kim - arXiv preprint arXiv …, 2026 - arxiv.org
Frontier model evaluations are shifting from foundational capabilities (eg, instruction following
and reasoning) toward compositional, agentic ones, but Korean agentic benchmarks …