8000
Skip to content

Popular repositories Loading

  1. blackwell-geforce-nvfp4-gemm blackwell-geforce-nvfp4-gemm Public

    NVFP4 inference on Blackwell GeForce (RTX 5090/5080/5070 Ti/RTX PRO 6000) — SM120 patches for vLLM + FlashInfer + CUTLASS. 175 tok/s on Qwen3.6-35B MoE.

    Python 23 2

  2. lna-es lna-es Public

    あらゆるジャンルのテキストをLLMを使いNeo4Jグラフ化して、グラフのみのデータから意味的復元をするシステムのスターター(MCP対応予定)

    16 1

  3. gemma4-12b-vllm-sm120 gemma4-12b-vllm-sm120 Public

    Reproducible recipe: serve abliterated Gemma-4-12B (gemma4_unified) at 50-118 tok/s on no-NVLink Blackwell (SM120) via vLLM nightly + ModelOpt FP8/NVFP4 + MTP spec-decode.

    Python 14

  4. GGUF-to-NVFP4-SM120 GGUF-to-NVFP4-SM120 Public

    Lna-Lab production pipeline: GGUF -> modelopt-format NVFP4 + working MTP head for vLLM on RTX PRO 6000 Blackwell (SM120). Stages 2 (NVFP4) and 3 (MTP graft) are Lna-Lab originals; stage 1 (GGUF->bf…

    Python 10 3

  5. distill-kura distill-kura Public

    蒸留蔵 — distilled long-term memory for agents: recall by meaning, writing gated by evidence, one kura per agent mode. Ships as a DeepSeek Harness plugin and an MCP server.

    Python 9

  6. LnaLang4U LnaLang4U Public

    1M-context DeepSeek-V4-Flash inference on NVIDIA Blackwell using sglang and SSD KV cache offload.

    Python 7 1

Repositories

Showing 10 of 18 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…

0