Train Models Contrastively in Pytorch
-
Updated
Mar 26, 2025 - Python
FFFF
Train Models Contrastively in Pytorch
Agentic RAG for local and self-hosted document search: hybrid retrieval, reranking and multimodal RAG on embedded LanceDB, with Docling parsing and an MCP server
Radient turns many data types (not just text) into vectors for similarity search, RAG, regression analysis, and more.
Think-on-Graph 3.0: Efficient and Adaptive LLM Reasoning on Heterogeneous Graphs via Multi-Agent Dual-Evolving Context Retrieval
Production inference for encoder models - ColBERT, GLiNER, ColPali, embeddings etc. - as vLLM plugins for online and in-process deployment
A sample app for the Multimodal Retrieval-Augmented Generation pattern running in Azure, using Azure AI Search for retrieval and Azure OpenAI large language models to power Q&A experiences.
Awesome Memory Papers in Vision-Language Models
High-performance late-interaction retrieval engine for on-prem AI. ColBERT/ColPali multi-vector search with Rust fused MaxSim, Triton GPU kernels, ROQ quantization, LEMUR routing, WAL-backed CRUD, and a FastAPI server — single machine, CPU or GPU.
Build sovereign RAG systems with MAS‑RAG, Dual‑RAG, GraphRAG, Spatial‑RAG, multimodal pipelines, and vector search directly inside Oracle AI Database 26ai and Exadata.
🧠 Multimodal Retrieval-Augmented Generation that "weaves" together text and images seamlessly. 🪡
Local multimodal RAG for PDFs: MinerU, Jina CLIP, FAISS, BM25, BGE reranking and Ollama. Runs on your hardware via CLI, web and desktop.
[NAACL 2024] Official Implementation of paper "Self-Adaptive Sampling for Efficient Video Question Answering on Image--Text Models"
🔰 A Comprehensive RAG repository covering basic vanilla RAG techniques, advanced retrieval methods, hybrid search fusion approaches, hands-on reranking techniques with code + explanation 📚✨
🚀 HAG: Next-Gen AI | Neo4j + Weaviate Fusion | Dual-Similarity Retrieval | 100% Local & Private | Graph Intelligence Meets Vector Search
OpenAI-compatible multimodal embedding server for Qwen3-VL-Embedding-2B — embed text, images, or both via a simple REST API.
[EMNLP 2025] M-LongDoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework
Benchmarking search agents for closed-corpus, multi-hop retrieval over visually rich documents.
Multimodal RAG Production
A doctor-assistive AI system that interprets medical knowledge and patient images simultaneously. It utilizes a Dual-Encoder architecture to cross-reference textbook theory with visual pathology, generating clinically grounded diagnoses.
Add a description, image, and links to the multimodal-rag topic page so that developers can more easily learn about it.
To associate your repository with the multimodal-rag topic, visit your repo's landing page and select "manage topics."