<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>LLM Systems | CAT LAB</title><link>https://deep-learning-profiling-tools.github.io/CAT-Lab/tag/llm-systems/</link><atom:link href="https://deep-learning-profiling-tools.github.io/CAT-Lab/tag/llm-systems/index.xml" rel="self" type="application/rss+xml"/><description>LLM Systems</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Tue, 04 Aug 2026 00:00:00 +0000</lastBuildDate><image><url>https://deep-learning-profiling-tools.github.io/CAT-Lab/media/icon_hu6808975029018430273.png</url><title>LLM Systems</title><link>https://deep-learning-profiling-tools.github.io/CAT-Lab/tag/llm-systems/</link></image><item><title>Paper Accepted at SC 2026</title><link>https://deep-learning-profiling-tools.github.io/CAT-Lab/news/2026-08-04-llmprof-sc2026/</link><pubDate>Tue, 04 Aug 2026 00:00:00 +0000</pubDate><guid>https://deep-learning-profiling-tools.github.io/CAT-Lab/news/2026-08-04-llmprof-sc2026/</guid><description>&lt;p>Our paper &lt;a href="https://deep-learning-profiling-tools.github.io/CAT-Lab/CAT-Lab/publication/llmprof_sc2026/">&lt;em>LLMPROF: Identifying Performance Bottlenecks in LLM Serving Systems with Top-Down Profiling&lt;/em>&lt;/a>, by &lt;strong>Tianle Zhong, Mao Lin, Hao Wu, Keren Zhou, and Geoffrey Fox&lt;/strong>, has been accepted at &lt;strong>SC 2026&lt;/strong>, The International Conference for High Performance Computing, Networking, Storage, and Analysis (Supercomputing). Congratulations to all the authors!&lt;/p></description></item><item><title>Continuous GPU and LLM Profiling Talk @ Scalable Tools Workshop</title><link>https://deep-learning-profiling-tools.github.io/CAT-Lab/news/2026-07-28-continuous-gpu-llm-profiling-talk/</link><pubDate>Tue, 28 Jul 2026 00:00:00 +0000</pubDate><guid>https://deep-learning-profiling-tools.github.io/CAT-Lab/news/2026-07-28-continuous-gpu-llm-profiling-talk/</guid><description>&lt;p>Keren Zhou presented &lt;a href="https://dyninst.github.io/scalable_tools_workshop/petascale2026/tuesday.html" target="_blank" rel="noopener">&lt;em>Very Low-Overhead Continuous Profiling of GPU-Accelerated LLM Workloads&lt;/em>&lt;/a> at the &lt;strong>Scalable Tools Workshop&lt;/strong> on July 28, 2026. The talk was held from &lt;strong>9:30–10:00 a.m.&lt;/strong> in Talk Session 5.&lt;/p>
&lt;p>The presentation describes the latest progress on Proton toward continuous, production-grade profiling for GPU-accelerated LLM serving and reinforcement-learning workloads with less than 1% overhead. It covers selective and phase-based profiling, memory-efficient metadata handling, lock-free metric collection, and fine-grained visibility into graph-executed GPU workloads.&lt;/p>
&lt;p>&lt;a href="https://dyninst.github.io/scalable_tools_workshop/petascale2026/assets/slides/Zhou%20-%20Scalable%20Tools%20Workshop.pdf" target="_blank" rel="noopener">View the slides&lt;/a>&lt;/p></description></item></channel></rss>