A flexible utility for converting tensor precision in PyTorch models and safetensors files, enabling efficient deployment across various platforms.
-
Updated
Aug 24, 2023 - Python
8000
A flexible utility for converting tensor precision in PyTorch models and safetensors files, enabling efficient deployment across various platforms.
Systematic 24-hour benchmark study of Qwen3.6-27B inference on dual NVIDIA RTX PRO 6000 Blackwell SM120 (TP=2). 8 experiments comparing repne/vllm fork vs upstream vLLM across FP8/BF16/NVFP4/Q8_0 quants and MTP/DFlash speculative decoding. Peak: 2,083 tok/s at c=32. Quality: KLD vs BF16 = 0.0018 (noise floor).
Auto GGUF Converter for HuggingFace Hub Models with Multiple Quantizations (GGUF Format)
DeepSeek-OCR-experimental is an advanced, multi-purpose visual document intelligence and object localization sandbox. Powered by the unredacted prithivMLmods/DeepSeek-OCR-Latest-BF16.I64-v2.0 architecture, this suite is designed to deliver highly accurate, structure-aware image text extractions.
Lossless AI model compression - ~34% smaller with bit-identical weights; the autopilot profiles your machine, picks the highest fidelity that runs, and streams models bigger than your RAM.
Reproducible benchmark suite and tuned Triton fused-MoE configs for NVIDIA H20 LLM inference. 24 configs, 36 perf data points, geomean 1.09× / peak 1.74× speedup.
Distributed GPT-2 fine-tuning with PyTorch FSDP and BF16 mixed precision, INT8 post-training quantisation, a custom Triton quantisation kernel achieving 1.4x throughput over unfused PyTorch, and a full FP32 vs BF16 vs INT8 benchmark suite. 13 tests passing.
Matrix-free 3D SIMP topology optimization with fused gather-GEMM-scatter CUDA kernels on NVIDIA RTX 4090. Companion code for arXiv:2604.18020.
FP12 AI Models.
Recover the native MTP predictor missing from the 8-bit MLX Qwen3.8-27B-Uncensored package, build a BF16 sidecar, and reproduce a 15.59 → 48.75 tok/s controlled M4 Max result with MTPLX.
Add a description, image, and links to the bf16 topic page so that developers can more easily learn about it.
To associate your repository with the bf16 topic, visit your repo's landing page and select "manage topics."