Calculon
An analytical performance model and open-source tool for high-level co-design of LLM training and inference systems.
I am a computer architecture researcher and software/hardware engineer specializing in AI systems, accelerator design, and high-performance interconnects.
As a Senior Research Scientist at NVIDIA, I develop analytical models and simulation tools to evaluate LLM training and inference across compute, memory, and communication. My work spans hardware/software co-design, energy-efficient inference architectures, and electrical and optical networks. Previously, at Google and Hewlett Packard Labs, I developed network architectures, routing algorithms, transport protocols, and congestion-control systems.
I build tools that connect workload behavior to architectural decisions, including Calculon for LLM systems and SuperSim for interconnection networks.
Ph.D. in Electrical Engineering, 2016
Stanford University
M.S. in Electrical Engineering, 2012
University of Utah
B.S. in Computer Engineering, 2010
University of Utah
An analytical performance model and open-source tool for high-level co-design of LLM training and inference systems.
A parallel discrete event simulation framework.
A parallel SSH-based remote machine management system.
An easy-to-use python package for parallel task execution.
A flit-level interconnection network simulator.
A project development find and goto system.
We present Calculon, a parameterized analytical performance model and open-source tool for high-level algorithm-architecture codesign of LLM training and inference systems.
We use an analytical performance model to design systems that train LLMs of up to 128 trillion parameters at 75%+ MFU.
We present ParaGraph, an open-source toolkit that bridges applications and high-fidelity network simulators to enable hardware-software co-design of supercomputer-scale systems.
We present two practical and efficient incremental adaptive routing algorithms for HyperX.
We present the first at-scale implementation of the HyperX topology and compare it side-by-side to a traditional Fat-Tree.
We present SuperSim, an open-source flit-level interconnection network simulator for large-scale high-performance networks.
This dissertation presents Sikker, a highly-scalable high-performance distributed system architecture for secure service-oriented computing.
NVIDIA
Salt Lake City, Utah
I work on AI system architecture at NVIDIA Research, using analytical models to explore software and hardware tradeoffs for large language models (LLMs). I developed Calculon, a framework for rapidly estimating performance and resource usage across a broad range of execution configurations, and built tools for large-scale LLM training analysis, including studies of models with up to 128 trillion parameters. This work was published at SC23 and helped identify more efficient LLM optimization and system design strategies. My research also examines scale-up and scale-out interconnects and network topologies, including approaches that use optical circuit switches for efficient large-scale training. I have contributed to distributed GEMM research and software infrastructure by defining reusable multi-GPU APIs and helping integrate distributed algorithms into CUTLASS. I also develop AI accelerator design concepts that minimize energy use and data movement while maximizing inference speed across emerging model workloads.
Google
Sunnyvale, California
At Google, I developed next-generation data center network topologies, routing algorithms, transport protocols, congestion control systems, and switching architectures. This work included guiding the architecture of future accelerated networking hardware and software systems. I also developed ParaGraph, an application-simulator interface and toolkit for hardware/software co-design that connected compiler-derived workload models with simulation infrastructure to support large-scale system exploration. ParaGraph was published at ICPP 2022.
Hewlett Packard Labs
Fort Collins, Colorado
At Hewlett Packard Labs, I was a lead architect for a new high-performance network supporting large-scale high-performance computing (HPC) and massively parallel memory-driven computing (MDC) systems. I was a key designer of the Gen-Z routing architecture specification. Using data-driven simulation, I also guided the design of a novel multi-chip module (MCM) switch architecture that used co-packaged integrated photonics.
L3 Communications
Salt Lake City, Utah
At L3 Communications, I developed digital processing architectures for encryption, networking, and waveform data processing. I designed systems to meet strict requirements for high performance, low power consumption, and a small hardware footprint. Each design underwent rigorous review for reliability and signal integrity.