8000
Skip to content
View yzhangcs's full-sized avatar
:octocat:
Focusing
:octocat:
Focusing

Organizations

@fla-org @MoonshotAI

Block or report yzhangcs

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

You don't write AGENTS.md. You train it with gradient descent.

JavaScript 329 15 Updated Aug 25, 2026

Pretraining Recurrent Networks without Recurrence

Jupyter Notebook 70 7 Updated Jun 7, 2026

Mixture of A Million Experts

Python 64 3 Updated Jul 30, 2024

Claude's Tokenizer, Offline, Kinda

Python 66 4 Updated Aug 25, 2026

A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM

Python 763 200 Updated Aug 25, 2026

A high performanceΒ development and dispatch library for high-performance machine learning kernels.

Python 34 4 Updated Aug 15, 2026

εˆ†δΊ«AI InfraηŸ₯θ―†&δ»£η η»ƒδΉ οΌšPyTorch、vLLM/SGLang、slime/vimeζ‘†ζžΆε…₯ι—¨βš‘οΈγ€ζ€§θƒ½εŠ ι€ŸπŸš€γ€ε€§ζ¨‘εž‹εŸΊη‘€πŸ§ γ€AIθ½―η‘¬δ»ΆπŸ”§η­‰

Jupyter Notebook 3,661 349 Updated Aug 7, 2026

Babel (Kimi Code edition): open-source AI-native chiplet design flow

Verilog 4 Updated Aug 7, 2026

[ICLR 2025] MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts

Python 278 13 Updated Oct 16, 2024

Open Frontier Intelligence

8,625 699 Updated Aug 6, 2026

MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts

Python 1,102 122 Updated Aug 13, 2026
Python 79 10 Updated Jul 27, 2026

[Tech Report] Expanded Hyper-Connections

61 1 Updated Jul 21, 2026

Create epic math and physics explainer animations with Kimi K3.

Python 141 20 Updated Jul 29, 2026

Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing (Technical Report)

Python 35 Updated Jul 10, 2026

This repositories contains the reference implementation for the Sparse Delta Memory paper.More precisely, it contains the model definition as well as triton and cuda kernels for the Sparse Delta Me…

Python 36 3 Updated Jul 9, 2026

JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Causal Parallel Tree Drafting

Python 178 7 Updated Aug 9, 2026

An early research stage expert-parallel load balancer for MoE models based on linear programming.

Python 530 43 Updated Nov 19, 2025

A lightweight inference engine supporting speculative speculative decoding (SSD).

Python 993 79 Updated May 10, 2026

DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms

Python 7,034 661 Updated Jul 9, 2026

RAT+: Train Dense, Infer Sparse - Recurrence Augmented Attention for Dilated Inference (ICML2026)

Python 10 1 Updated May 21, 2026

DeepSeek-V3.2-Exp DSA Warmup Lightning Indexer training operator based on tilelang

Python 52 3 Updated Nov 19, 2025
Python 248 35 Updated Aug 21, 2026

Agentic Kernel Optimization β€” advanced & eXtensible: a closed-loop, campaign-based multi-agent system for optimizing GPU kernels (benchmark-swappable; default flashinfer-bench).

Python 65 12 Updated Aug 17, 2026

kernelbench.com β€” GPU kernel engineering benchmarks for autonomous LLM coding agents. v3 archive + v-hard latest.

HTML 72 10 Updated Aug 23, 2026

IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse

134 11 Updated Mar 14, 2026
Next
0