E538
Skip to content
#

nvfp4

Here are 143 public repositories matching this topic...

Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.io/aeon-7/aeon-vllm-ultimate:latest container, tuned for long-context draft acceptance on DGX Spark. 6 HF variants (BF16/NVFP4/MTP/MTP-XS), docker-compose, and QuickStart.

  • Updated Jul 3, 2026
  • Python

Rust + CUDA inference engine for NVIDIA RTX PRO 6000 Blackwell and RTX 5090. Serves safetensors and GGUF over an OpenAI-compatible API, with per-device tuned defaults and speculative decode gated byte-identical to plain decode. Hosted instance: inference.tiyuvta.ai

  • Updated Aug 24, 2026
  • Rust

Serving 4-bit Qwen3.8-27B on a single DGX Spark (GB10): 75 tok/s single-stream, 246 tok/s aggregate at 8-way concurrency. NVFP4 vs MixedInt4-AutoRound vs the FP8 baseline, measured on one harness — including why the quantization advantage collapses to +0.2% by c16.

  • Updated Aug 19, 2026
  • Python

An LLM server for a single RTX 5090, built for agent workloads: tool calls, long conversations, reasoning, and many requests at once. Consistently faster than llama.cpp on the same models, with the numbers in the repo. Written end to end by Claude Code.

  • Updated Aug 24, 2026
  • Cuda

Improve this page

Add a description, image, and links to the nvfp4 topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the nvfp4 topic, visit your repo's landing page and select "manage topics."

Learn more

0