BleuMacaw: GPT-2 and SentenceTransformers for Paraphrases Generation
-
Updated
Dec 14, 2023 - Python
E539
BleuMacaw: GPT-2 and SentenceTransformers for Paraphrases Generation
Machine translation comparison and BLEU quality evaluation system built with Streamlit, Transformers, and NLTK.
A local evaluation suite for Luxembourgish machine translation.
Few-shot table-to-text generation on the ToTTo dataset using Qwen models, evaluated with BLEU and BLEURT.
End-to-end multilingual text generation with mT5 across 5 languages & 3 scripts — with tokenization equity analysis for non-Latin languages (Hindi, Bengali)
Open-source translation API benchmark (FLORES + COMET) across 20 languages. Compare DeepL, Google, Azure & more.
Model evaluation harness for standardized benchmarking—comprehensive metrics (F1, BLEU, ROUGE, METEOR, BERTScore, pass@k), statistical analysis (confidence intervals, effect size, bootstrap CI, ANOVA), multi-model comparison, and report generation. Research-grade evaluation for LLM and ML experiments.
A data driven query expansion approach for image caption, implemented in cpp
Generator of data for training LLMs for the specific use case of creating SQL queries from natural language. Developed as a practical project at TUM.
NLP quality evaluation: BLEU/ROUGE metrics, translation review, language quality scoring, annotation consistency
Model-agnostic toolkit to evaluate text-summarization models — ROUGE/BLEU/BERTScore + perplexity + LLM-as-judge, MLflow tracking, CLI + library with unit-tested metrics.
Translation evaluation tool - LLM-based multilingual translation platform with 20 languages, batch processing, BLEU scoring, and MiniMax TTS
A 3-level AI testing curriculum documenting LLM evaluation, RAG testing, and agentic AI validation — built from real programme delivery at EPAM across Google, regulated banking, and healthcare.
基于 Mengzi-T5 在 DuReaderQG 上微调做中文问题生成,含三档消融与检索式 baseline 对比评测 | T5 fine-tuning for Chinese question generation
Eval-driven model router: pick the best model per task, score with multi-metric eval, track win-rates on a Postgres scoreboard.
[Working on it] Implemented the Sequence-to-Sequence LSTM architecture from Sutskever et al. (2014) for English-to-French translation, achieving BLEU score comparable to the original paper; implemented custom data preprocessing, deep LSTM encoder-decoder.
Production-grade RAG evaluation library: faithfulness, hallucination, retrieval precision, answer relevance, context coverage, and UCM, an unsupervised confidence metric requiring no ground-truth labels.
Add a description, image, and links to the bleu topic page so that developers can more easily learn about it.
To associate your repository with the bleu topic, visit your repo's landing page and select "manage topics."