{"id":26061,"date":"2026-08-10T10:23:37","date_gmt":"2026-08-10T14:23:37","guid":{"rendered":"https:\/\/www.rc.fas.harvard.edu\/?post_type=mec-events&#038;p=26061"},"modified":"2026-08-10T10:23:37","modified_gmt":"2026-08-10T14:23:37","slug":"checkpointing-workshop","status":"publish","type":"mec-events","link":"https:\/\/www.rc.fas.harvard.edu\/events\/checkpointing-workshop\/","title":{"rendered":"Checkpointing and Restarting Computational Workflows Workshop"},"content":{"rendered":"<p data-pm-slice=\"1 1 []\"><strong>NOTE: This is an in-person session on campus.<\/strong><\/p>\n<h2>Checkpointing and Restarting Computational Workflows (in-person)<\/h2>\n<p><strong>Description:<\/strong> In-person training on checkpointing and restarting computational workflows on FASRC clusters. This workshop will introduce general checkpointing strategies and demonstrate practical techniques for Python, C\/C++, and Fortran applications, as well as AI and machine-learning workflows using PyTorch, TensorFlow\/Keras, and JAX.<\/p>\n<p>Participants will learn how to make long-running jobs resilient to failures, Slurm time limits, preemption, and planned requeue events. Hands-on exercises will cover saving and restoring application state, writing requeue-aware Slurm jobs, handling termination signals, and creating robust checkpoints for machine-learning training.<\/p>\n<p><strong>Note:<\/strong> This is an in-person session. It will not be available via Zoom and will not be recorded.<\/p>\n<p><strong>Level:<\/strong> Intermediate<\/p>\n<p><strong>Time:<\/strong> 11am - 3pm<\/p>\n<p><strong>Location:<\/strong> TBD (Northwest Labs vicinity)<\/p>\n<p><strong>Presenters:<\/strong> Plamen Krastev<\/p>\n<p><strong>Who can attend this workshop:<\/strong> Anyone with a FASRC cluster account. An active cluster account is required to participate in the hands-on exercises.<\/p>\n<h2>What you will learn:<\/h2>\n<ol class=\"ProsemirrorEditor-list\">\n<li class=\"ProsemirrorEditor-listItem\" data-list-indent=\"1\" data-list-type=\"bulleted\">\n<p>Why checkpointing is important for long-running and failure-prone computational workloads<\/p>\n<\/li>\n<li class=\"ProsemirrorEditor-listItem\" data-list-indent=\"1\" data-list-type=\"bulleted\">\n<p>General checkpointing approaches, design patterns, and restart granularity<\/p>\n<\/li>\n<li class=\"ProsemirrorEditor-listItem\" data-list-indent=\"1\" data-list-type=\"bulleted\">\n<p>How to create checkpoint-and-restart workflows in Python, C\/C++, and Fortran<\/p>\n<\/li>\n<li class=\"ProsemirrorEditor-listItem\" data-list-indent=\"1\" data-list-type=\"bulleted\">\n<p>How to write requeue-aware Slurm jobs on the Cannon cluster<\/p>\n<\/li>\n<li class=\"ProsemirrorEditor-listItem\" data-list-indent=\"1\" data-list-type=\"bulleted\">\n<p>How to handle Slurm signals and exit gracefully before a job reaches its time limit<\/p>\n<\/li>\n<li class=\"ProsemirrorEditor-listItem\" data-list-indent=\"1\" data-list-type=\"bulleted\">\n<p>How to checkpoint AI\/ML workflows using PyTorch, TensorFlow\/Keras, Lightning, and JAX<\/p>\n<\/li>\n<li class=\"ProsemirrorEditor-listItem\" data-list-indent=\"1\" data-list-type=\"bulleted\">\n<p>How to save and restore model parameters, optimizer state, training progress, and random-number-generator state<\/p>\n<\/li>\n<li class=\"ProsemirrorEditor-listItem\" data-list-indent=\"1\" data-list-type=\"bulleted\">\n<p>Strategies for distributed and large-scale ML checkpointing, including rank-0 and sharded checkpoints<\/p>\n<\/li>\n<li class=\"ProsemirrorEditor-listItem\" data-list-indent=\"1\" data-list-type=\"bulleted\">\n<p>Best practices for reliable, portable, and storage-efficient checkpoints<\/p>\n<\/li>\n<\/ol>\n<h2>Prerequisites:<\/h2>\n<ol class=\"ProsemirrorEditor-list\">\n<li class=\"ProsemirrorEditor-listItem\" data-list-indent=\"1\" data-list-type=\"numbered\">\n<p>A FASRC cluster account. If you do not have an account, see<br \/>\n<a href=\"https:\/\/docs.rc.fas.harvard.edu\/kb\/how-do-i-get-a-research-computing-account\/\"><strong>Request a FAS Research Computing Account well in advance<\/strong><\/a>. See prerequisite 2 and 3.<\/p>\n<\/li>\n<li class=\"ProsemirrorEditor-listItem\" data-list-indent=\"1\" data-list-type=\"numbered\">\n<p>Previous experience submitting batch jobs on a FASRC cluster.\u00a0<\/p>\n<\/li>\n<li class=\"ProsemirrorEditor-listItem\" data-list-indent=\"1\" data-list-type=\"numbered\">\n<p>Basic familiarity with Slurm job scripts.<\/p>\n<\/li>\n<li class=\"ProsemirrorEditor-listItem\" data-list-indent=\"1\" data-list-type=\"numbered\">\n<p>Basic programming experience in Python, C\/C++, or Fortran. Participants interested in the AI\/ML examples should have basic familiarity with at least one supported machine-learning framework.<\/p>\n<\/li>\n<\/ol>\n<p><strong>Registration:<\/strong> * Registration Link Pending *<\/p>\n<div class=\"XTranslate\" style=\"all: unset;\">\u00a0<\/div>\n","protected":false},"excerpt":{"rendered":"<p class=\"lead\">NOTE: This is an in-person session on campus. Checkpointing and Restarting Computational Workflows (in-person) Description: In-person training on checkpointing and restarting computational workflows on FASRC clusters. This workshop will introduce general checkpointing strategies and demonstrate practical techniques for Python, C\/C++, and Fortran applications, as well as AI and machine-learning workflows using PyTorch, TensorFlow\/Keras, and JAX. Participants will learn how to&hellip;<\/p>\n<p class=\"more-link-p\"><a class=\"btn btn-primary\" href=\"https:\/\/www.rc.fas.harvard.edu\/events\/checkpointing-workshop\/\">Read more<\/a><\/p>\n","protected":false},"author":126,"featured_media":0,"comment_status":"closed","ping_status":"closed","template":"","tags":[],"mec_category":[191],"class_list":["post-26061","mec-events","type-mec-events","status-publish","hentry","mec_category-training"],"jetpack_sharing_enabled":true,"_links":{"self":[{"href":"https:\/\/www.rc.fas.harvard.edu\/wp-json\/wp\/v2\/mec-events\/26061","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.rc.fas.harvard.edu\/wp-json\/wp\/v2\/mec-events"}],"about":[{"href":"https:\/\/www.rc.fas.harvard.edu\/wp-json\/wp\/v2\/types\/mec-events"}],"author":[{"embeddable":true,"href":"https:\/\/www.rc.fas.harvard.edu\/wp-json\/wp\/v2\/users\/126"}],"replies":[{"embeddable":true,"href":"https:\/\/www.rc.fas.harvard.edu\/wp-json\/wp\/v2\/comments?post=26061"}],"wp:attachment":[{"href":"https:\/\/www.rc.fas.harvard.edu\/wp-json\/wp\/v2\/media?parent=26061"}],"wp:term":[{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.rc.fas.harvard.edu\/wp-json\/wp\/v2\/tags?post=26061"},{"taxonomy":"mec_category","embeddable":true,"href":"https:\/\/www.rc.fas.harvard.edu\/wp-json\/wp\/v2\/mec_category?post=26061"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}