PyTorch native INT8 quantization API

INT8 tensor subclass for PyTorch enabling up to 4x memory reduction with optimized CUDA/Triton kernels. Merged into TorchAO.

Read report →

nano-vLLM

Educational LLM inference engine built from scratch. Covers PagedAttention, continuous batching, chunked prefill, and scheduling with detailed C++ implementations.

Read Part 1 →

Marin (with Stanford)

An open lab with Stanford for building foundation models together. Marin-8B beats Llama 3.1 8B on 14/19 benchmarks; Marin-32B beats OLMo 2 32B Base on 14/19 benchmarks.

Visit Marin →

60+ merged PRs across 30+ open-source projects the field runs on.

verl

11 merged PRs

View on GitHub →

DeepSeek

5 merged PRs

View on GitHub →

Google DeepMind

3 merged PRs

View on GitHub →

NVIDIA NeMo

3 merged PRs

View on GitHub →

LiteLLM

3 merged PRs

View on GitHub →

SkyPilot

3 merged PRs

View on GitHub →

THUDM Slime

3 merged PRs

View on GitHub →

UCCL

3 merged PRs

View on GitHub →

Also contributing to: PyTorch/TorchAO, vLLM, LangChain, LlamaFactory, Unsloth, SGLang, Qwen FlashQLA, Ray, Instructor, Axolotl, Mistral, OpenRLHF, NVIDIA Dynamo, Flash Linear Attention, TileLang, llm-d, Sakana AI, PQXDH, Tokenspeed, Mooncake, and more.