Continuous GPU and LLM Profiling Talk @ Scalable Tools Workshop

Keren Zhou presented Very Low-Overhead Continuous Profiling of GPU-Accelerated LLM Workloads at the Scalable Tools Workshop on July 28, 2026. The talk was held from 9:30–10:00 a.m. in Talk Session 5.

The presentation describes the latest progress on Proton toward continuous, production-grade profiling for GPU-accelerated LLM serving and reinforcement-learning workloads with less than 1% overhead. It covers selective and phase-based profiling, memory-efficient metadata handling, lock-free metric collection, and fine-grained visibility into graph-executed GPU workloads.

View the slides

Keren Zhou
Keren Zhou
Assistant Professor