Multimodal RAG over PDFs. Three parallel pipelines (LLM summaries, raw atomic content, CLIP visual) with HyDE expansion and cross encoder reranking, answers cited inline.
-
Updated
Jul 22, 2026 - Jupyter Notebook
8000
Multimodal RAG over PDFs. Three parallel pipelines (LLM summaries, raw atomic content, CLIP visual) with HyDE expansion and cross encoder reranking, answers cited inline.
Transcription project consisting of Python scripting and usage of AI/ML text extraction models.
RAG-based PDF intelligence system using LangChain, Hugging Face embeddings, and Pinecone for semantic document querying.
OCR + LLM-enhanced parsing FastAPI service with structured JSON outputs, confidence scoring, Docker deployment, and API docs.
Local OCR and LLM-based pipeline for extracting structured JSON from complex PDF documents.
AI decision workflows that turn enterprise documents into traceable decisions, executive artifacts and operational handoffs.
Field-level evaluation for document extraction with vision language models. Tells you whether a change made your documents read better or worse — and whether the difference is real.
Reliability-first VLM evidence platform with calibrated risk routing, constrained human review, and auditable promotion gates.
Reproducible experiments for TableFormer and TFLOP transfer to scientific table recognition
Fine-tuned open-weight LLMs (Mistral-7B, LLaMA-2) with LoRA/PEFT for document understanding — 92% accuracy, 40% GPU memory cut via mixed precision.
Source-grounded document chatbot starter with citations, configurable AI providers, and local-first setup.
医结智控 MedLedger Agent|非结构化医疗销售文档到可审计结算报表
A private, citation-aware local RAG research assistant powered by Ollama, ChromaDB, and Streamlit.
AudioPage: native iPhone and iPad read-aloud app for PDFs, books, articles, scans, and notes. On-device listening with document AI when you ask.
End-to-end receipt OCR pipeline using OpenCV and Tesseract OCR to extract structured data from images and PDFs.
Governed agentic AI workflow for motor-insurance claims with extraction, validation, policy RAG, guarded tools, human review, memory, evaluations, and end-to-end observability.
Document-extraction RAG: 95.5% weighted field-level accuracy on 28-case offline CI replay; 202-case corpus (151+51). FastAPI + pgvector + Claude.
Bangla handwritten document OCR: EAST detection, CRNN recognition, FastAPI + Gradio
Add a description, image, and links to the document-ai topic page so that developers can more easily learn about it.
To associate your repository with the document-ai topic, visit your repo's landing page and select "manage topics."