Skip to content

About

Experiment tracking with MLFlow on Chameleon

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

67 Commits

Folders and files

Repository files navigation

In this tutorial, we explore some of the infrastructure and platform requirements for large model training, and to support the training of many models by many teams. We focus specifically on experiment tracking (using MLFlow).

Follow along at ML Experiment Tracking with MLFlow.

Note: this tutorial requires advance reservation of specific hardware! You will need a node with 2 GPUs suitable for model training. You should reserve a 3-hour block.

You can use either:

  • a gpu_mi100 at CHI@TACC, or
  • a compute_liqid at CHI@TACC

This material is based upon work supported by the National Science Foundation under Grant No. 2230079.

About

Experiment tracking with MLFlow on Chameleon

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages