-
Moonshot AI
- Shanghai
-
08:09
(UTC +08:00) - https://yzhang.site
- @yzhang_cs
Stars
You don't write AGENTS.md. You train it with gradient descent.
Pretraining Recurrent Networks without Recurrence
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
A high performanceΒ development and dispatch library for high-performance machine learning kernels.
εδΊ«AI Infraη₯θ―&代η η»δΉ οΌPyTorchγvLLM/SGLangγslime/vimeζ‘ζΆε ₯ι¨β‘οΈγζ§θ½ε ιπγ倧樑εεΊη‘π§ γAI软瑬仢π§η
Babel (Kimi Code edition): open-source AI-native chiplet design flow
[ICLR 2025] MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts
Create epic math and physics explainer animations with Kimi K3.
Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing (Technical Report)
This repositories contains the reference implementation for the Sparse Delta Memory paper.More precisely, it contains the model definition as well as triton and cuda kernels for the Sparse Delta Meβ¦
JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Causal Parallel Tree Drafting
An early research stage expert-parallel load balancer for MoE models based on linear programming.
A lightweight inference engine supporting speculative speculative decoding (SSD).
DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms
RAT+: Train Dense, Infer Sparse - Recurrence Augmented Attention for Dilated Inference (ICML2026)
DeepSeek-V3.2-Exp DSA Warmup Lightning Indexer training operator based on tilelang
Agentic Kernel Optimization β advanced & eXtensible: a closed-loop, campaign-based multi-agent system for optimizing GPU kernels (benchmark-swappable; default flashinfer-bench).
kernelbench.com β GPU kernel engineering benchmarks for autonomous LLM coding agents. v3 archive + v-hard latest.
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse