User profiles for Geewook Kim
Geewook KimNAVER Cloud AI Verified email at navercorp.com Cited by 2499 |
Ocr-free document understanding transformer
Understanding document images (eg, invoices) is a core but challenging task since it requires
complex functions such as reading text and a holistic understanding of the document. …
complex functions such as reading text and a holistic understanding of the document. …
What is wrong with scene text recognition model comparisons? dataset and model analysis
Many new proposals for scene text recognition (STR) models have been introduced in recent
years. While each claim to have pushed the boundary of the technology, a holistic and fair …
years. While each claim to have pushed the boundary of the technology, a holistic and fair …
Prometheus-vision: Vision-language model as a judge for fine-grained evaluation
Assessing long-form responses generated by Vision-Language Models (VLMs) is challenging.
It not only requires checking whether the VLM follows the given instruction but also …
It not only requires checking whether the VLM follows the given instruction but also …
[PDF][PDF] Donut: Document understanding transformer without ocr
Understanding document images (eg, invoices) has been an important research topic and
has many applications in document processing automation. Through the latest advances in …
has many applications in document processing automation. Through the latest advances in …
How does vision-language adaptation impact the safety of vision language models?
Vision-Language adaptation (VL adaptation) transforms Large Language Models (LLMs)
into Large Vision-Language Models (LVLMs) for multimodal tasks, but this process often …
into Large Vision-Language Models (LVLMs) for multimodal tasks, but this process often …
Cost-effective end-to-end information extraction for semi-structured document images
A real-world information extraction (IE) system for semi-structured document images often
involves a long pipeline of multiple modules, whose complexity dramatically increases its …
involves a long pipeline of multiple modules, whose complexity dramatically increases its …
On text localization in end-to-end ocr-free document understanding transformer without text localization supervision
This paper presents a simple yet effective approach for weakly supervised text localization
in end-to-end visual document understanding (VDU) models. The traditional approach in …
in end-to-end visual document understanding (VDU) models. The traditional approach in …
Do modern video-llms need to listen? a benchmark audit and scalable remedy
Speech and audio encoders developed over years of community effort are routinely excluded
from video understanding pipelines, not because they fail, but because benchmarks never …
from video understanding pipelines, not because they fail, but because benchmarks never …
Visually-situated natural language understanding with contrastive reading model and frozen large language models
Recent advances in Large Language Models (LLMs) have stimulated a surge of research
aimed at extending their applications to the visual domain. While these models exhibit promise …
aimed at extending their applications to the visual domain. While these models exhibit promise …
K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts
Frontier model evaluations are shifting from foundational capabilities (eg, instruction following
and reasoning) toward compositional, agentic ones, but Korean agentic benchmarks …
and reasoning) toward compositional, agentic ones, but Korean agentic benchmarks …