Calculon

Bar charts of batch time and HBM memory consumption for running GPT3 175B across 4,096 GPUs with TP=8, PP=64, and DP=8, broken down by execution phase and data type

An analytical performance model and open-source tool for high-level co-design of LLM training and inference systems.

Publications

Cite

@inproceedings{isaev2023calculon,
  title={{Calculon: a Methodology and Tool for High-Level Codesign of Systems and Large Language Models}},
  author={Isaev, Mikhail and McDonald, Nic and Dennison, Larry and Vuduc, Richard},
  booktitle={Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC)},
  year={2023},
  organization={ACM}
}

Cite

@inproceedings{isaev2023scaling,
  title={{Scaling Infrastructure to Support Multi-Trillion Parameter LLM Training}},
  author={Isaev, Mikhail and McDonald, Nic and Vuduc, Richard},
  booktitle={Proceedings of the Workshop on Architecture and System Support for Transformer Models (ASSYST)},
  year={2023}
}

Search