{"id":25101,"date":"2024-01-18T14:57:56","date_gmt":"2024-01-18T19:57:56","guid":{"rendered":"https:\/\/www.rc.fas.harvard.edu\/?p=25101"},"modified":"2024-09-16T13:22:11","modified_gmt":"2024-09-16T17:22:11","slug":"cannon-2-0","status":"publish","type":"post","link":"https:\/\/www.rc.fas.harvard.edu\/blog\/cannon-2-0\/","title":{"rendered":"Cannon 2.0"},"content":{"rendered":"<h1><img loading=\"lazy\" decoding=\"async\" data-attachment-id=\"25093\" data-permalink=\"https:\/\/www.rc.fas.harvard.edu\/blog\/cannon-2-0\/attachment\/cannon_banner-20\/\" data-orig-file=\"https:\/\/www.rc.fas.harvard.edu\/wp-content\/uploads\/2021\/02\/cannon_banner-20.jpg\" data-orig-size=\"1200,450\" data-comments-opened=\"0\" data-image-meta=\"{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;}\" data-image-title=\"cannon_banner-20\" data-image-description=\"\" data-image-caption=\"\" data-large-file=\"https:\/\/www.rc.fas.harvard.edu\/wp-content\/uploads\/2021\/02\/cannon_banner-20-1024x384.jpg\" class=\"aligncenter wp-image-25093 size-large\" src=\"https:\/\/www.rc.fas.harvard.edu\/wp-content\/uploads\/2021\/02\/cannon_banner-20-1024x384.jpg\" alt=\"cannon 2.0 January 2024\" width=\"970\" height=\"364\" srcset=\"https:\/\/www.rc.fas.harvard.edu\/wp-content\/uploads\/2021\/02\/cannon_banner-20-1024x384.jpg 1024w, https:\/\/www.rc.fas.harvard.edu\/wp-content\/uploads\/2021\/02\/cannon_banner-20-300x113.jpg 300w, https:\/\/www.rc.fas.harvard.edu\/wp-content\/uploads\/2021\/02\/cannon_banner-20-768x288.jpg 768w, https:\/\/www.rc.fas.harvard.edu\/wp-content\/uploads\/2021\/02\/cannon_banner-20.jpg 1200w\" sizes=\"auto, (max-width: 970px) 100vw, 970px\" \/><\/h1>\n<h1><span style=\"font-weight: 400;\">Summary<\/span><\/h1>\n<p><span style=\"font-weight: 400;\">FASRC is adding 216 Intel Sapphire Rapids nodes with 1TB of RAM each, 4 Intel Sapphire Rapids nodes with 2TB of RAM each and 144 A100 80GB GPUs to the Cannon cluster. The Sapphire Rapids cores will be made available in the new \u2018sapphire\u2019 partition. The new A100 GPUs will be added to the \u2018gpu\u2019 partition.\u00a0 Partitions will be reorganized to account for the larger memory of the new nodes. FASSE will gain additional \u2018fasse_bigmem\u2019 and \u2018fasse_gpu\u2019 capacity. The base <\/span><a href=\"https:\/\/docs.rc.fas.harvard.edu\/kb\/fairshare\/\"><span style=\"font-weight: 400;\">gratis fairshare<\/span><\/a><span style=\"font-weight: 400;\"> will be increased to 200 on the Cannon cluster.<\/span><\/p>\n<p>These updates will go live on January 22nd, 2024.<\/p>\n<h1><span style=\"font-weight: 400;\">Overview<\/span><\/h1>\n<p><span style=\"font-weight: 400;\">Cannon 2.0 represents the expansion of the liquid-cooled Cannon cluster, providing access to Intel's latest Sapphire Rapids processors and Nvidia's A100 GPUs to the Harvard research community. This update aims to reduce wait times and provide additional resources to the community. On the Cannon cluster, we observed that limited memory on the nodes was contributing to extended wait times for workflows requiring more memory. The additional memory in these new compute nodes will help bridge this gap.<\/span><\/p>\n<p>&nbsp;<\/p>\n<h4>Cannon 2.0 consists of:\u00a0<\/h4>\n<ul>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">CPUs: 216 nodes with 1TB of RAM and 112 cores each of Lenovo SD650 V3 direct water cooling servers. These nodes offer a total of 24,192 cores of Intel 8480+ \u201cSapphire Rapids\u201d processors. The interconnect is NDR 400 Gbps Infiniband (IB) connected to 400 Gbps IB core.<\/span><\/li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">CPUs: 4 nodes with 2TB of RAM and 112 cores each of Lenovo SD650 V3 direct water cooling servers. These nodes have 448 cores of Intel 8480+ \u201cSapphire Rapids\u201d processors. The interconnect is NDR 400 Gbps Infiniband (IB) connected to 400 Gbps IB core.<\/span><\/li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">GPUs: 36 nodes each node with four Nvidia A100 80G GPUs for a total of 144 new GPUs of Lenovo SR670 V2 servers with direct water cooling. Each GPU node has 64 CPU cores, and 1TB of RAM.<\/span><\/li>\n<\/ul>\n<p>&nbsp;<\/p>\n<h4><span style=\"font-weight: 400;\">As part of our standard process for new installs we follow a phased (formerly called tiered) testing plan:<\/span><\/h4>\n<ul>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Phase 1: Lenovo burn in, HPL benchmarking, Top\/Green 500 Runs (Done)<\/span><\/li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Phase 2: Internal Testing (Done)<\/span><\/li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Phase 3: Harvard Community Grand Challenge Runs (in progress)<\/span><\/li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Production - Jan 22nd<\/span><\/li>\n<\/ul>\n<p>&nbsp;<\/p>\n<h1><span style=\"font-weight: 400;\">Partitions<\/span><\/h1>\n<p><span style=\"font-weight: 400;\">With the advent of Cannon 2.0, we have reconsidered the organization of the partitions. All the Sapphire Rapids nodes have 1TB of memory, which exceeds the capacity of our current \u2018bigmem\u2019 partition, rendering the name of that partition somewhat misleading. Additionally, the new A100s are the 80GB variety on the NDR (faster) fabric, indicating that we cannot simply merge them into the existing GPU partition. Finally, we need to consider the future needs of FASSE.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">Thus the updated <\/span><a href=\"https:\/\/docs.rc.fas.harvard.edu\/kb\/running-jobs\/?seq_no=2#Slurm_partitions\"><span style=\"font-weight: 400;\">partitions <\/span><\/a><span style=\"font-weight: 400;\">are as follows:<\/span><\/p>\n<p>&nbsp;<\/p>\n<table style=\"width: 100%; height: 832px;\">\n<tbody>\n<tr style=\"height: 80px;\">\n<td style=\"height: 80px; width: 20.3664%;\">\n<strong>Partition<\/strong>\n<\/td>\n<td style=\"height: 80px; width: 6.03448%;\">\n<strong>Nodes<\/strong>\n<\/td>\n<td style=\"height: 80px; width: 10.9914%;\">\n<strong>Cores per Node<\/strong>\n<\/td>\n<td style=\"height: 80px; width: 22.8448%;\">\n<strong>CPU Core Types<\/strong>\n<\/td>\n<td style=\"height: 80px; width: 11.2069%;\">\n<strong>Total # of Cores<\/strong>\n<\/td>\n<td style=\"height: 80px; width: 18.1034%;\">\n<strong>Usable Mem per Node (GB)<\/strong>\n<\/td>\n<td style=\"height: 80px; width: 8.40517%;\">\n<strong>Time Limit<\/strong>\n<\/td>\n<\/tr>\n<tr style=\"height: 56px;\">\n<td style=\"height: 56px; width: 20.3664%;\">\n<b>sapphire<\/b>\n<\/td>\n<td style=\"height: 56px; width: 6.03448%;\">\n<span style=\"font-weight: 400;\">196<\/span>\n<\/td>\n<td style=\"height: 56px; width: 10.9914%;\">\n<span style=\"font-weight: 400;\">112<\/span>\n<\/td>\n<td style=\"height: 56px; width: 22.8448%;\">\n<span style=\"font-weight: 400;\">Intel Sapphire Rapids<\/span>\n<\/td>\n<td style=\"height: 56px; width: 11.2069%;\">\n<span style=\"font-weight: 400;\">21,952<\/span>\n<\/td>\n<td style=\"height: 56px; width: 18.1034%;\">\n<span style=\"font-weight: 400;\">990<\/span>\n<\/td>\n<td style=\"height: 56px; width: 8.40517%;\">\n<span style=\"font-weight: 400;\">3 days<\/span>\n<\/td>\n<\/tr>\n<tr style=\"height: 56px;\">\n<td style=\"height: 56px; width: 20.3664%;\">\n<b>shared<\/b>\n<\/td>\n<td style=\"height: 56px; width: 6.03448%;\">\n<span style=\"font-weight: 400;\">277<\/span>\n<\/td>\n<td style=\"height: 56px; width: 10.9914%;\">\n<span style=\"font-weight: 400;\">48<\/span>\n<\/td>\n<td style=\"height: 56px; width: 22.8448%;\">\n<span style=\"font-weight: 400;\">Intel Cascade Lake<\/span>\n<\/td>\n<td style=\"height: 56px; width: 11.2069%;\">\n<span style=\"font-weight: 400;\">13,296<\/span>\n<\/td>\n<td style=\"height: 56px; width: 18.1034%;\">\n<span style=\"font-weight: 400;\">184<\/span>\n<\/td>\n<td style=\"height: 56px; width: 8.40517%;\">\n<span style=\"font-weight: 400;\">3 days<\/span>\n<\/td>\n<\/tr>\n<tr>\n<td style=\"width: 20.3664%;\">\n<b>hsph<\/b>\n<\/td>\n<td style=\"width: 6.03448%;\">\n<span style=\"font-weight: 400;\">36<\/span>\n<\/td>\n<td style=\"width: 10.9914%;\">\n<span style=\"font-weight: 400;\">112<\/span>\n<\/td>\n<td style=\"width: 22.8448%;\">\n<span style=\"font-weight: 400;\">Intel Sapphire Rapids<\/span>\n<\/td>\n<td style=\"width: 11.2069%;\">\n<span style=\"font-weight: 400;\">4,032<\/span>\n<\/td>\n<td style=\"width: 18.1034%;\">\n<span style=\"font-weight: 400;\">990<\/span>\n<\/td>\n<td style=\"width: 8.40517%;\">\n<span style=\"font-weight: 400;\">3 days<\/span>\n<\/td>\n<\/tr>\n<tr style=\"height: 56px;\">\n<td style=\"height: 56px; width: 20.3664%;\">\n<b>test<\/b>\n<\/td>\n<td style=\"height: 56px; width: 6.03448%;\">\n<span style=\"font-weight: 400;\">12<\/span>\n<\/td>\n<td style=\"height: 56px; width: 10.9914%;\">\n<span style=\"font-weight: 400;\">112<\/span>\n<\/td>\n<td style=\"height: 56px; width: 22.8448%;\">\n<span style=\"font-weight: 400;\">Intel Sapphire Rapids<\/span>\n<\/td>\n<td style=\"height: 56px; width: 11.2069%;\">\n<span style=\"font-weight: 400;\">1,344<\/span>\n<\/td>\n<td style=\"height: 56px; width: 18.1034%;\">\n<span style=\"font-weight: 400;\">990<\/span>\n<\/td>\n<td style=\"height: 56px; width: 8.40517%;\">\n<span style=\"font-weight: 400;\">12 hours<\/span>\n<\/td>\n<\/tr>\n<tr style=\"height: 56px;\">\n<td style=\"height: 56px; width: 20.3664%;\">\n<b>intermediate<\/b>\n<\/td>\n<td style=\"height: 56px; width: 6.03448%;\">\n<span style=\"font-weight: 400;\">12<\/span>\n<\/td>\n<td style=\"height: 56px; width: 10.9914%;\">\n<span style=\"font-weight: 400;\">112<\/span>\n<\/td>\n<td style=\"height: 56px; width: 22.8448%;\">\n<span style=\"font-weight: 400;\">Intel Sapphire Rapids<\/span>\n<\/td>\n<td style=\"height: 56px; width: 11.2069%;\">\n<span style=\"font-weight: 400;\">1,344<\/span>\n<\/td>\n<td style=\"height: 56px; width: 18.1034%;\">\n<span style=\"font-weight: 400;\">990<\/span>\n<\/td>\n<td style=\"height: 56px; width: 8.40517%;\">\n<span style=\"font-weight: 400;\">14 days<\/span>\n<\/td>\n<\/tr>\n<tr style=\"height: 56px;\">\n<td style=\"height: 56px; width: 20.3664%;\">\n<b>bigmem<\/b>\n<\/td>\n<td style=\"height: 56px; width: 6.03448%;\">\n<span style=\"font-weight: 400;\">4<\/span>\n<\/td>\n<td style=\"height: 56px; width: 10.9914%;\">\n<span style=\"font-weight: 400;\">112<\/span>\n<\/td>\n<td style=\"height: 56px; width: 22.8448%;\">\n<span style=\"font-weight: 400;\">Intel Sapphire Rapids<\/span>\n<\/td>\n<td style=\"height: 56px; width: 11.2069%;\">\n<span style=\"font-weight: 400;\">448<\/span>\n<\/td>\n<td style=\"height: 56px; width: 18.1034%;\">\n<span style=\"font-weight: 400;\">1988<\/span>\n<\/td>\n<td style=\"height: 56px; width: 8.40517%;\">\n<span style=\"font-weight: 400;\">3 days<\/span>\n<\/td>\n<\/tr>\n<tr style=\"height: 56px;\">\n<td style=\"height: 56px; width: 20.3664%;\">\n<b>bigmem_intermediate<\/b>\n<\/td>\n<td style=\"height: 56px; width: 6.03448%;\">\n<span style=\"font-weight: 400;\">3<\/span>\n<\/td>\n<td style=\"height: 56px; width: 10.9914%;\">\n<span style=\"font-weight: 400;\">64<\/span>\n<\/td>\n<td style=\"height: 56px; width: 22.8448%;\">\n<span style=\"font-weight: 400;\">Intel Ice Lake<\/span>\n<\/td>\n<td style=\"height: 56px; width: 11.2069%;\">\n<span style=\"font-weight: 400;\">192<\/span>\n<\/td>\n<td style=\"height: 56px; width: 18.1034%;\">\n<span style=\"font-weight: 400;\">2000<\/span>\n<\/td>\n<td style=\"height: 56px; width: 8.40517%;\">\n<span style=\"font-weight: 400;\">14 days<\/span>\n<\/td>\n<\/tr>\n<tr style=\"height: 56px;\">\n<td style=\"height: 56px; width: 20.3664%;\">\n<b>gpu<\/b>\n<\/td>\n<td style=\"height: 56px; width: 6.03448%;\">\n<span style=\"font-weight: 400;\">36<\/span>\n<\/td>\n<td style=\"height: 56px; width: 10.9914%;\">\n<span style=\"font-weight: 400;\">64<\/span>\n<\/td>\n<td style=\"height: 56px; width: 22.8448%;\">\n<span style=\"font-weight: 400;\">Intel Ice Lake, A100 80GB<\/span>\n<\/td>\n<td style=\"height: 56px; width: 11.2069%;\">\n<span style=\"font-weight: 400;\">144 GPUs\u00a0<\/span>\n<\/td>\n<td style=\"height: 56px; width: 18.1034%;\">\n<span style=\"font-weight: 400;\">990<\/span>\n<\/td>\n<td style=\"height: 56px; width: 8.40517%;\">\n<span style=\"font-weight: 400;\">3 days<\/span>\n<\/td>\n<\/tr>\n<tr style=\"height: 80px;\">\n<td style=\"height: 80px; width: 20.3664%;\">\n<b>gpu_test<\/b>\n<\/td>\n<td style=\"height: 80px; width: 6.03448%;\">\n<span style=\"font-weight: 400;\">14<\/span>\n<\/td>\n<td style=\"height: 80px; width: 10.9914%;\">\n<span style=\"font-weight: 400;\">64<\/span>\n<\/td>\n<td style=\"height: 80px; width: 22.8448%;\">\n<span style=\"font-weight: 400;\">Intel Ice Lake, A100 40GB, <\/span><a href=\"https:\/\/docs.nvidia.com\/datacenter\/tesla\/mig-user-guide\/\"><span style=\"font-weight: 400;\">MIG<\/span><\/a><span style=\"font-weight: 400;\"> 3g.20gb<\/span>\n<\/td>\n<td style=\"height: 80px; width: 11.2069%;\">\n<span style=\"font-weight: 400;\"> 112 MIG GPUs<\/span>\n<\/td>\n<td style=\"height: 80px; width: 18.1034%;\">\n<span style=\"font-weight: 400;\">448<\/span>\n<\/td>\n<td style=\"height: 80px; width: 8.40517%;\">\n<span style=\"font-weight: 400;\">12 hours<\/span>\n<\/td>\n<\/tr>\n<tr style=\"height: 56px;\">\n<td style=\"height: 56px; width: 20.3664%;\">\n<b>remoteviz<\/b>\n<\/td>\n<td style=\"height: 56px; width: 6.03448%;\">\n<span style=\"font-weight: 400;\">1<\/span>\n<\/td>\n<td style=\"height: 56px; width: 10.9914%;\">\n<span style=\"font-weight: 400;\">32<\/span>\n<\/td>\n<td style=\"height: 56px; width: 22.8448%;\">\n<span style=\"font-weight: 400;\">Intel Cascade Lake<\/span>\n<\/td>\n<td style=\"height: 56px; width: 11.2069%;\">\n<span style=\"font-weight: 400;\">32<\/span>\n<\/td>\n<td style=\"height: 56px; width: 18.1034%;\">\n<span style=\"font-weight: 400;\">380<\/span>\n<\/td>\n<td style=\"height: 56px; width: 8.40517%;\">\n<span style=\"font-weight: 400;\">3 days<\/span>\n<\/td>\n<\/tr>\n<tr style=\"height: 56px;\">\n<td style=\"height: 56px; width: 20.3664%;\">\n<b>unrestricted<\/b>\n<\/td>\n<td style=\"height: 56px; width: 6.03448%;\">\n<span style=\"font-weight: 400;\">8<\/span>\n<\/td>\n<td style=\"height: 56px; width: 10.9914%;\">\n<span style=\"font-weight: 400;\">48<\/span>\n<\/td>\n<td style=\"height: 56px; width: 22.8448%;\">\n<span style=\"font-weight: 400;\">Intel Cascade Lake<\/span>\n<\/td>\n<td style=\"height: 56px; width: 11.2069%;\">\n<span style=\"font-weight: 400;\">384<\/span>\n<\/td>\n<td style=\"height: 56px; width: 18.1034%;\">\n<span style=\"font-weight: 400;\">184<\/span>\n<\/td>\n<td style=\"height: 56px; width: 8.40517%;\">\n<span style=\"font-weight: 400;\">none<\/span>\n<\/td>\n<\/tr>\n<tr style=\"height: 56px;\">\n<td style=\"height: 56px; width: 20.3664%;\">\n<b>fasse<\/b>\n<\/td>\n<td style=\"height: 56px; width: 6.03448%;\">\n<span style=\"font-weight: 400;\">42<\/span>\n<\/td>\n<td style=\"height: 56px; width: 10.9914%;\">\n<span style=\"font-weight: 400;\">48<\/span>\n<\/td>\n<td style=\"height: 56px; width: 22.8448%;\">\n<span style=\"font-weight: 400;\">Intel Cascade lake<\/span>\n<\/td>\n<td style=\"height: 56px; width: 11.2069%;\">\n<span style=\"font-weight: 400;\">2,016<\/span>\n<\/td>\n<td style=\"height: 56px; width: 18.1034%;\">\n<span style=\"font-weight: 400;\">184<\/span>\n<\/td>\n<td style=\"height: 56px; width: 8.40517%;\">\n<span style=\"font-weight: 400;\">7 days<\/span>\n<\/td>\n<\/tr>\n<tr style=\"height: 56px;\">\n<td style=\"height: 56px; width: 20.3664%;\">\n<b>fasse_bigmem<\/b>\n<\/td>\n<td style=\"height: 56px; width: 6.03448%;\">\n<span style=\"font-weight: 400;\">16<\/span>\n<\/td>\n<td style=\"height: 56px; width: 10.9914%;\">\n<span style=\"font-weight: 400;\">64<\/span>\n<\/td>\n<td style=\"height: 56px; width: 22.8448%;\">\n<span style=\"font-weight: 400;\">Intel Ice Lake<\/span>\n<\/td>\n<td style=\"height: 56px; width: 11.2069%;\">\n<span style=\"font-weight: 400;\">1,024<\/span>\n<\/td>\n<td style=\"height: 56px; width: 18.1034%;\">\n<span style=\"font-weight: 400;\">500<\/span>\n<\/td>\n<td style=\"height: 56px; width: 8.40517%;\">\n<span style=\"font-weight: 400;\">7 days<\/span>\n<\/td>\n<\/tr>\n<tr style=\"height: 56px;\">\n<td style=\"height: 56px; width: 20.3664%;\">\n<b>fasse_gpu<\/b>\n<\/td>\n<td style=\"height: 56px; width: 6.03448%;\">\n<span style=\"font-weight: 400;\">4<\/span>\n<\/td>\n<td style=\"height: 56px; width: 10.9914%;\">\n<span style=\"font-weight: 400;\">64<\/span>\n<\/td>\n<td style=\"height: 56px; width: 22.8448%;\">\n<span style=\"font-weight: 400;\">Intel Ice Lake, A100 40GB<\/span>\n<\/td>\n<td style=\"height: 56px; width: 11.2069%;\">\n<span style=\"font-weight: 400;\">16 GPUS<\/span>\n<\/td>\n<td style=\"height: 56px; width: 18.1034%;\">\n<span style=\"font-weight: 400;\">488<\/span>\n<\/td>\n<td style=\"height: 56px; width: 8.40517%;\">\n<span style=\"font-weight: 400;\">7 days<\/span>\n<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&nbsp;<\/p>\n<p><span style=\"font-weight: 400;\">The reshuffle of the Cannon 2.0 system aims to better serve the community by absorbing existing \u2018bigmem\u2019 jobs into the new higher capacity \u2018sapphire\u2019 partition and \u2018bigmem_intermediate\u2019 into the \u2018intermediate\u2019 partition. The use of <\/span><a href=\"https:\/\/docs.nvidia.com\/datacenter\/tesla\/mig-user-guide\/\"><span style=\"font-weight: 400;\">MIG mode<\/span><\/a><span style=\"font-weight: 400;\"> for \u2018gpu_test\u2019 will double the effective number of GPUs, allowing for more users.\u00a0<\/span><\/p>\n<p><span style=\"font-weight: 400;\">All the Cannon changes will go live on Jan 22nd 2024. None of these changes require a downtime for the cluster nor interruption of user workflow or jobs. Existing jobs will finish on the older nodes. FASSE changes will occur throughout the week as nodes are moved over to the secure environment.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">To take advantage of the new partitions, you will need to update your job scripts and adjust job parameters according to your needs. See the FAQ below for additional recommendations. Please free to join our office hours <\/span><a href=\"https:\/\/www.rc.fas.harvard.edu\/training\/office-hours\/\"><span style=\"font-weight: 400;\">https:\/\/www.rc.fas.harvard.edu\/training\/office-hours\/<\/span><\/a><span style=\"font-weight: 400;\"> or contact us <\/span><a href=\"https:\/\/www.rc.fas.harvard.edu\/about\/contact\/\"><span style=\"font-weight: 400;\">https:\/\/www.rc.fas.harvard.edu\/about\/contact\/<\/span><\/a><span style=\"font-weight: 400;\">\u00a0 if you have any questions. To cite use of this resource please see: <\/span><a href=\"https:\/\/www.rc.fas.harvard.edu\/cluster\/publications\"><span style=\"font-weight: 400;\">https:\/\/www.rc.fas.harvard.edu\/cluster\/publications<\/span><\/a><span style=\"font-weight: 400;\">\/\u00a0\u00a0\u00a0<\/span><\/p>\n<h1><span style=\"font-weight: 400;\">Fairshare<\/span><\/h1>\n<p><span style=\"font-weight: 400;\">This new hardware will increase our computational power on our public Cannon partitions by roughly 70% over Cannon 1.0. As a result we have recalculated our base gratis <\/span><a href=\"https:\/\/docs.rc.fas.harvard.edu\/kb\/fairshare\/\"><span style=\"font-weight: 400;\">fairshare<\/span><\/a><span style=\"font-weight: 400;\"> for all groups. The base gratis fairshare on Cannon will be changed from 120 to 200. This new base gratis fairshare score will be applied to all groups on Cannon when the new hardware is made live.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">For FASSE the new hardware does not significantly impact the computational power of that cluster. As such the base gratis fairshare for FASSE will remain the same at 100.<\/span><\/p>\n<p>&nbsp;<\/p>\n<h1><span style=\"font-weight: 400;\">Cannon FAQ<\/span><\/h1>\n<h4><span style=\"font-weight: 400;\">Q. I use <\/span><i><span style=\"font-weight: 400;\">bigmem<\/span><\/i><span style=\"font-weight: 400;\">\/<\/span><i><span style=\"font-weight: 400;\">bigmem_intermediate<\/span><\/i><span style=\"font-weight: 400;\">, do I need to update my scripts?<\/span><\/h4>\n<p><span style=\"font-weight: 400;\">Yes, you should move to the updated \u2018sapphire\u2019 and \u2018intermediate\u2019 partitions. The new partitions now have 1T of RAM.<\/span><\/p>\n<h4><span style=\"font-weight: 400;\">Q. I use ultramem, do I need to update my scripts?<\/span><\/h4>\n<p><span style=\"font-weight: 400;\">Yes, you should move to \u2018bigmem\u2019 or \u2018bigmem_intermediate\u2019. The updated partitions now have 2T of memory which is the same as what \u2018ultramem\u2019 offered.<\/span><\/p>\n<h4><span style=\"font-weight: 400;\">Q. I use <\/span><i><span style=\"font-weight: 400;\">gpu_mig<\/span><\/i><span style=\"font-weight: 400;\">, do I need to update my scripts?<\/span><\/h4>\n<p><span style=\"font-weight: 400;\">Yes, you should move to \u2018gpu_test\u2019 which is set up in <\/span><a href=\"https:\/\/docs.nvidia.com\/datacenter\/tesla\/mig-user-guide\/\"><span style=\"font-weight: 400;\">MIG mode<\/span><\/a><span style=\"font-weight: 400;\"> and allows for experimentation with that feature.<\/span><\/p>\n<h4><span style=\"font-weight: 400;\">Q. I use <\/span><i><span style=\"font-weight: 400;\">gpu_test<\/span><\/i><span style=\"font-weight: 400;\">, do I need to update my scripts?<\/span><\/h4>\n<p><span style=\"font-weight: 400;\">Maybe. \u2018gpu_test\u2019 will be moving from V100\u2019s to A100\u2019s in <\/span><a href=\"https:\/\/docs.nvidia.com\/datacenter\/tesla\/mig-user-guide\/\"><span style=\"font-weight: 400;\">MIG mode<\/span><\/a><span style=\"font-weight: 400;\">. Depending on your script you may need to adjust for the new GPU type.<\/span><\/p>\n<h4><span style=\"font-weight: 400;\">Q. I use <\/span><i><span style=\"font-weight: 400;\">test<\/span><\/i><span style=\"font-weight: 400;\">\/<\/span><i><span style=\"font-weight: 400;\">intermediate<\/span><\/i><span style=\"font-weight: 400;\">\/<\/span><i><span style=\"font-weight: 400;\">gpu<\/span><\/i><span style=\"font-weight: 400;\">, do I need to update my scripts?<\/span><\/h4>\n<p><span style=\"font-weight: 400;\">No. Changes to these partitions do not necessitate updating your script.<\/span><\/p>\n<h4><span style=\"font-weight: 400;\">Q. I use <\/span><i><span style=\"font-weight: 400;\">shared<\/span><\/i><span style=\"font-weight: 400;\">, do I need to update my scripts?<\/span><\/h4>\n<p><span style=\"font-weight: 400;\">Maybe. \u2018shared\u2019 will remain as is, but if you want to leverage the new Sapphire Rapids nodes you should consider updating your script to point to the \u2018sapphire\u2019 partition, or adding \u2018sapphire\u2019 as an additional partition (i.e. -p shared,sapphire). It is worth testing to see which partition will give you better performance.<\/span><\/p>\n<h4><span style=\"font-weight: 400;\">Q. I am part of Kempner, does the gratis base fairshare impact me?<\/span><\/h4>\n<p><span style=\"font-weight: 400;\">No. The base gratis fairshare only impacts your Cannon Slurm account, not the Kempner Slurm accounts.<\/span><\/p>\n<p>&nbsp;<\/p>\n<h1><span style=\"font-weight: 400;\">FASSE FAQ<\/span><\/h1>\n<h4><span style=\"font-weight: 400;\">Q. I use <\/span><i><span style=\"font-weight: 400;\">fasse_bigmem<\/span><\/i><span style=\"font-weight: 400;\">, do I need to update my scripts?<\/span><\/h4>\n<p><span style=\"font-weight: 400;\">No. Changes to these partitions do not necessitate updating your script.<\/span><\/p>\n<h4><span style=\"font-weight: 400;\">Q. I use <\/span><i><span style=\"font-weight: 400;\">fasse_gpu<\/span><\/i><span style=\"font-weight: 400;\">, do I need to update my scripts?<\/span><\/h4>\n<p><span style=\"font-weight: 400;\">Maybe. \u2018fasse_gpu\u2019 will be moving from V100\u2019s to A100\u2019s. Depending on your script you may need to adjust for the new GPU type.<\/span><\/p>\n<p>&nbsp;<\/p>\n<h1><span style=\"font-weight: 400;\">Resources<\/span><\/h1>\n<ul>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">FASRC self-help documentation and instructions: <\/span><a href=\"https:\/\/docs.rc.fas.harvard.edu\/\"><span style=\"font-weight: 400;\">https:\/\/docs.rc.fas.harvard.edu\/<\/span><\/a><\/li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Office hours: <\/span><a href=\"https:\/\/www.rc.fas.harvard.edu\/training\/office-hours\/\"><span style=\"font-weight: 400;\">https:\/\/www.rc.fas.harvard.edu\/training\/office-hours\/<\/span><\/a><\/li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Training: <\/span><a href=\"https:\/\/www.rc.fas.harvard.edu\/upcoming-training\/\"><span style=\"font-weight: 400;\">https:\/\/www.rc.fas.harvard.edu\/upcoming-training\/<\/span><\/a><\/li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Support: <\/span><a href=\"https:\/\/www.rc.fas.harvard.edu\/about\/contact\/\"><span style=\"font-weight: 400;\">https:\/\/www.rc.fas.harvard.edu\/about\/contact\/<\/span><\/a><\/li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Cannon cluster dashboard: <\/span><a href=\"https:\/\/dash.rc.fas.harvard.edu\/d\/000000017\/cannon-cluster-node-usage?orgId=1&amp;refresh=5m\"><span style=\"font-weight: 400;\">https:\/\/dash.rc.fas.harvard.edu\/d\/000000017\/cannon-cluster-node-usage?orgId=1&amp;refresh=5m<\/span><\/a><\/li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">FASSE cluster dashboard: <\/span><a href=\"https:\/\/dash.rc.fas.harvard.edu\/d\/tUVpWaHMk\/fasse-cluster-node-usage?orgId=1&amp;refresh=5m\"><span style=\"font-weight: 400;\">https:\/\/dash.rc.fas.harvard.edu\/d\/tUVpWaHMk\/fasse-cluster-node-usage?orgId=1&amp;refresh=5m<\/span><\/a><\/li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Other dashboards: <\/span><a href=\"https:\/\/xdmod4.rc.fas.harvard.edu\/\"><span style=\"font-weight: 400;\">https:\/\/xdmod4.rc.fas.harvard.edu\/<\/span><\/a><\/li>\n<\/ul>\n<p>&nbsp;<\/p>\n","protected":false},"excerpt":{"rendered":"<p class=\"lead\">Summary FASRC is adding 216 Intel Sapphire Rapids nodes with 1TB of RAM each, 4 Intel Sapphire Rapids nodes with 2TB of RAM each and 144 A100 80GB GPUs to the Cannon cluster. The Sapphire Rapids cores will be made available in the new \u2018sapphire\u2019 partition. The new A100 GPUs will be added to the \u2018gpu\u2019 partition.\u00a0 Partitions will be&hellip;<\/p>\n<p class=\"more-link-p\"><a class=\"btn btn-primary\" href=\"https:\/\/www.rc.fas.harvard.edu\/blog\/cannon-2-0\/\">Read more<\/a><\/p>\n","protected":false},"author":103,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_feature_clip_id":0,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_post_was_ever_published":false},"categories":[36,170,198],"tags":[],"class_list":["post-25101","post","type-post","status-publish","format-standard","hentry","category-blog","category-cannon","category-fasse"],"jetpack_sharing_enabled":true,"jetpack_shortlink":"https:\/\/wp.me\/p42YvN-6wR","jetpack_featured_media_url":"","_links":{"self":[{"href":"https:\/\/www.rc.fas.harvard.edu\/wp-json\/wp\/v2\/posts\/25101","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.rc.fas.harvard.edu\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.rc.fas.harvard.edu\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.rc.fas.harvard.edu\/wp-json\/wp\/v2\/users\/103"}],"replies":[{"embeddable":true,"href":"https:\/\/www.rc.fas.harvard.edu\/wp-json\/wp\/v2\/comments?post=25101"}],"version-history":[{"count":18,"href":"https:\/\/www.rc.fas.harvard.edu\/wp-json\/wp\/v2\/posts\/25101\/revisions"}],"predecessor-version":[{"id":25447,"href":"https:\/\/www.rc.fas.harvard.edu\/wp-json\/wp\/v2\/posts\/25101\/revisions\/25447"}],"wp:attachment":[{"href":"https:\/\/www.rc.fas.harvard.edu\/wp-json\/wp\/v2\/media?parent=25101"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.rc.fas.harvard.edu\/wp-json\/wp\/v2\/categories?post=25101"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.rc.fas.harvard.edu\/wp-json\/wp\/v2\/tags?post=25101"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}