Status and Maintenance

Computing Systems Status

Electrical work required for Bouchet expansion has been completed

Friday, July 24, 2.30 pm: The electrical work required for Bouchet expansion has been completed. 172 out of 180 Bouchet nodes affected by the expansion work are now back online and performing as normal. Eight nodes in the week partition are still operating at reduced performance. These nodes have been set to not accept new jobs, and each of these nodes will be rebooted once currently running jobs on that node finish.  As the nodes are rebooted, they will return to normal operation at full performance.  The last of these reboots should complete by the end of the day on July 29.

Description

We are working to increase overall compute capacity on Bouchet, which involves electrical work at the data center during the week of July 20-24, 2026. This work may impact the performance of Bouchet compute nodes, as it will temporarily reduce power feeds to the compute racks. Users should expect jobs to run more slowly than usual during the planned work. 

Standard job partitions will still be available, though certain compute nodes may be more affected than others.  Priority-Tier partitions will be disabled during the week.

Scope

180 Bouchet nodes will be impacted by planned data center work July 20-24:

The following will be impacted during the entire week:

  • 10 H200 nodes (gpu_h200, gpu_devel partitions) with names beginning with a1122, a1124, and a1126 
  • 12 RTX 5000 Ada nodes (gpu, gpu_devel, education_gpu partitions) with names beginning with a1128
  • 10 L40S (pi_co54 partition) nodes with names beginning with a1118
  • 100 non-GPU Intel nodes (day, bigmem, devel, education, mpi* partitions) with names beginning with a1130 and a1132

The following will be impacted only on Monday morning (July 20) and Wednesday afternoon (July 22):

  • 7 B200 nodes (gpu_b200 partition) with names beginning with a1116
  • 8 RTX Pro 6000 Blackwell nodes (gpu_rtx6000, gpu_devel partitions) with names beginning with a1112 
  • 33 non-GPU AMD (day, week, devel partitions) nodes with names beginning with a1114

Bouchet nodes outside the scope of this electric work: 60 nodes in the mpi partition and one B200 node in the gpu_devel partition. 

* Since MPI workflows can be sensitive to slow-performing nodes, we will temporarily move some affected nodes (20 non-GPU Intel nodes) from the mpi partition to the day partition to mitigate the impact of slower nodes.  The remaining 60 nodes in the mpi partition will be unaffected by the electrical work.

Performance Impact

We expect minor impact to the RTX Pro 6000 Blackwell, RTX 5000 Ada, and non-GPU AMD nodes during planned electrical work. We estimate a 2x slowdown for the L40S and non-GPU Intel nodes and H200 nodes.

We expect significant performance impact on the B200 nodes. However, the extent of performance degradation will depend on the specific workflow, ranging from minimal to significant. 

Once full power is restored to the B200 nodes, they will need to be restarted to restore normal performance. In preparation, we will create a scheduler reservation to drain all running jobs on the B200 nodes by the end of the work impacting them. During this time, the scheduler will not start jobs with requested run times that overlap with the reservation time. Once the nodes are restarted, they will resume normal operation.

Please adjust your job wall-times when submitting jobs as needed and use checkpointing when possible. Let us know if you would like assistance doing so.

The anticipated completion of scheduled work is end of day, Friday, July 24, 2026. YCRC will send an email communication when the work is completed. We are committed to keeping you as informed as possible at every stage of the work. Updates will be posted on this page. We recognize the reduced performance will impact your work, and we apologize for the inconvenience.  If you have questions, comments, or need assistance with your jobs, please contact us at ycrc@yale.edu.

Helpful tips and resources:
  • Adjust the wall-time with –time SLURM option. For more information, visit Job Scheduling with Slurm documentation page.

7/15/26; 4 pm: The faulty CDU has been repaired. 52 out of the 56 nodes affected by it are now back online.

The following four nodes remain offline and require further repair or replacement: 

  • r813u23n01 (a100 node in the commons partition)
  • r813u23n02 (pi_gerstein_gpu GPU node)
  • r813u29n03 (a regular node in the week partition)
  • r813u29n09 (a regular node in the week partition)

The issues with these nodes have been escalated with the vendor.

Please check the status of your jobs and resubmit any jobs that failed due to the outage.  Please contact us if you encounter any issues.

7/15 10 am: The CDU has gone through the overnight testing and test water samples are being collected for analysis to confirm it is functioning as expected. Number of nodes have been powered on and are undergoing testing with the goal of bringing the majority of affected equipment back online in the next several hours.

7/14/26 6 pm: Additional issues discovered during repairs to the faulty CDU required installation of replacement parts. CDU currently in testing. The timeline for completion and bringing nodes back online has been extended to Wednesday, 7/15.

7/14/26 9am: Vendors are on site and continuing with the repair work, as planned. 

7/13/26 6.45 pm: Repairs to the faulty Cooling Distribution Unit (CDU) on McCleary are in progress. 

7/13/26 9am: The vendor has arrived at the West Campus and begun the work to repair the faulty Cooling Distribution Unit (CDU) on McCleary.

7/9/2026:  Reduced capacity on McCleary.  Due to a cooling issue in one of the McCleary racks at West Campus, McCleary nodes that have names beginning with r813 have again shut down. This outage is caused by a faulty Cooling Distribution Unit (CDU), causing equipment in the rack to overheat and shut down. We are working with the CDU vendor to resolve the issue as quickly as possible.

This issue impacts the following McCleary scheduler partitions:

  • 26 of the 33 nodes in the day partition
  • 14 of the 16 nodes in the week partition
  • 3 of the 20 nodes in the gpu partition
  • 1 of the 3 nodes in the pi_gerstein_gpu partition
  • 2 of the 10 nodes in the pi_jetz partition
  • 4 of the 4 nodes in the pi_ohern partition
  • 2 of the 2 nodes in the pi_sestan partition
  • 4 of the 4 nodes in the pi_tsang partition

Jobs running on the affected nodes have been terminated due to the failure. Please check the status of your jobs.

Slurm jobs can be submitted as usual; however, any jobs that would start on the affected nodes won’t start until after the cooling issue is resolved.  

We realize this is an impact to your work and apologize for the inconvenience.  If you have any questions, comments, or concerns, please contact us at research.computing@yale.edu.

6/15/2026 - 6/18/2026:  Scheduled all-cluster maintenance complete.

6/4/2026 13:00: Hopper experienced a system issue that impacted availability. We engaged the storage vendor to resolve the issue. Hopper availability was restored by 15:30.

5/11/2026:  There was a power drop at the West Campus Data Center on the morning of 5/11/2026.  This impacted running jobs on McCleary, Grace, Misha and Milgram.  Please check your jobs to see how they were impacted.

2/13/2026: Bouchet scheduler issues have been resolved. All systems are operational. 

Cluster Maintenance

To perform critical updates and minimize downtime, regular maintenance will be performed on each cluster on a rotating schedule. During maintenance, logins will be disabled, jobs will not run, and cluster storage may be unavailable. Communication will be sent to users four weeks and one week before the maintenance period and in case of any changes.

Starting in the Spring of 2025, YCRC updated our approach to system maintenance.  Until then, each cluster had two full-downtime maintenance periods per year, each lasting three days.  With the new approach, each cluster is updated twice a year.  However, only one of these two annual maintenance periods is a full downtime.  The other involves rolling updates to a live cluster.  We are working toward a system in which annual full-downtime maintenance is performed on all YCRC-managed clusters simultaneously.  Approximately six months after this full downtime, the clusters will be patched with minor updates on a rolling basis with minimal disruption. 

The new approach has several advantages.  By consolidating the major cluster updates, YCRC is able to focus on the preparation for and execution of those major updates once a year, instead of the nearly once a month, freeing more time for supporting researchers.  Simultaneously performing maintenance on all clusters within a data center enables maintenance of subsystems that affect multiple clusters.  All clusters are kept on the same major version of the image throughout the year, resulting in a more consistent, easier-to-support environment.  For any given cluster, the number of days per year of planned total downtime is reduced.  Also, for each cluster, the number of planned total-downtime periods is reduced from two to one.  There is a second period each year of rolling updates, but these will entail limited disruption.

If you have any questions, comments, or concerns, please contact us at hpc@yale.edu

Maintenance Schedule

Upcoming:

Future dates for upcoming planned maintenance will be announced soon.