Hi, I’m Yiyan Zhai :)
I am a first-year PhD student at Carnegie Mellon University in the Catalyst Group, advised by Prof. Tianqi Chen. I received my B.S. in Computer Science from Carnegie Mellon University. My research interests lie in building efficient and scalable ML systems.
I have been working with Prof. Tianqi Chen at CMU Catalyst Group on:
- TIRx Harness, an open compiler harness for agentic GPU programming that combines a minimal compiler foundation, a kernel knowledge base, analysis tools, and a benchmark server to help agents develop correct, fast GPU kernels. [Blog]
- FlashInfer-Bench, a kernel benchmarking loop that goes from kernel generation → evaluation → drop-in replacement in serving stacks (FlashInfer/SGLang/vLLM).
- WebLLM Assistant, which integrates Overleaf and Google Workspace with in-browser agents using WebLLM.
I am also fortunate to collaborate with Prof. Juncheng Yang at Harvard SEAS on:
- Cache replacement algorithms for real-world enterprise storage systems (VMware vSAN)
- Resilient routing for LLM inference
News 📰
- May 2026: Clock2Q+ is accepted by VLDB industrial track 2026. See you in Boston! 🎉
- May 2026: We presented FlashInfer-Bench at MLSys 2026.
- Apr 2026: I will be joining CMU as a PhD student this Fall! 🎓
- Jan 2026: FlashInfer-Bench is accepted by MLSys 2026! 🎉
