N-gram back-off POS tagger for Hindi (NLTK) trained on the UD Hindi-HDTB treebank, ~88% accuracy.
-
Updated
Aug 2, 2026 - Python
8000
N-gram back-off POS tagger for Hindi (NLTK) trained on the UD Hindi-HDTB treebank, ~88% accuracy.
Python toolkit to decode legacy Hindi font-encoded PDFs (KrutiDev, Chanakya, DevLys) into Unicode Devanagari. Built for Hindi PDF & govt document ingestion pipelines.
"Offline AI system for small manufacturers — CV defect detection, Hindi alerts, predictive maintenance & demand forecasting on one dashboard, no cloud"
Fully offline, multilingual voice AI agent that helps Indian citizens discover government welfare schemes in Hindi & English. Uses Whisper ASR, fine-tuned Qwen2.5 (QLoRA), FAISS RAG, and an MCP tool server — all designed for air-gapped deployment. Covers 2,872 schemes from myscheme.gov.in.
Multilingual RAG evaluation pipeline for Hindi and English -- $0 cost, custom eval metrics, FAISS + free LLMs
A hands-free AI voice avatar — say "Hey Vansh AI" and talk naturally in Hindi or English. Powered by Gemini, with lip-synced avatar, emotion detection, and RAG.
Estimate spoken duration of Hindi, English, and Hinglish text from syllable counts - a fast, offline proxy for TTS/voice-bot script timing
Offline AI audio/video transcriber — Whisper-powered, multi-language, parallel long-file processing, TXT/SRT/VTT/JSON output
Zero-dependency NLP toolkit for Hindi, Marathi, and Bengali. Tokenization, sentiment analysis, language detection & more.
Speak a field note, get a structured record. You define the fields; every value cites the audio it was heard in. Hindi/Hinglish, runs on a laptop for free.
Voice-enabled Hindi RAG pipeline (Sarvam STT -> BGE-M3 dense retrieval over FAISS-HNSW -> grounded answers) with measured ablations, guardrails, and sub-200ms retrieval. Deployed on HF ZeroGPU.
Benchmarking and improving Retrieval-Augmented Generation (RAG) for low-resource and code-mixed Indian languages (Hindi, Hinglish) using MuRIL/IndicBERT embeddings and open LLMs. B.Tech Minor/Major Project — Amity University.
Benchmarking NER on Naamapadam across 11 Indic languages. EDA + model training using mBERT, XLM-R, T5, FlanT5, mT5 + LLM fine-tuning (TinyLlama, Llama-3.2, Gemma, Qwen, Mistral) + 0–5 shot inference on 9 generative models.
Real-time speech-to-text for code-mixed Hindi–English (Hinglish). Self-hosted faster-whisper streaming over WebSockets, with per-script language tagging and a PostgreSQL-backed transcript dashboard with full CRUD.
Kavach: India's Hindi-first anti-digital-arrest and voice-scam shield. Real-time AASIST-hindi spoof detection (0.9919 acc, 0/100 FP), Hindi coercion analysis (8 vectors), tamper-evident + ed25519 evidence packets for 1930. Audio 31/31, stress 226/226, 29/29 real incidents replayed, 37 unit tests, full battery re-runnable. IIC 3.0 Cybersecurity
Add a description, image, and links to the hindi-nlp topic page so that developers can more easily learn about it.
To associate your repository with the hindi-nlp topic, visit your repo's landing page and select "manage topics."