Research Projects
PyTorch native INT8 quantization API
INT8 tensor subclass for PyTorch enabling up to 4x memory reduction with optimized CUDA/Triton kernels. Merged into TorchAO.
nano-vLLM
Educational LLM inference engine built from scratch. Covers PagedAttention, continuous batching, chunked prefill, and scheduling with detailed C++ implementations.
Marin (with Stanford)
An open lab with Stanford for building foundation models together. Marin-8B beats Llama 3.1 8B on 14/19 benchmarks; Marin-32B beats OLMo 2 32B Base on 14/19 benchmarks.
Open Source Contributions
60+ merged PRs across 30+ open-source projects the field runs on.
Also contributing to: PyTorch/TorchAO, vLLM, LangChain, LlamaFactory, Unsloth, SGLang, Qwen FlashQLA, Ray, Instructor, Axolotl, Mistral, OpenRLHF, NVIDIA Dynamo, Flash Linear Attention, TileLang, llm-d, Sakana AI, PQXDH, Tokenspeed, Mooncake, and more.