Description
We are working to increase overall compute capacity on Bouchet, which involves electrical work at the data center during the week of July 20-24, 2026. This work may impact the performance of Bouchet compute nodes, as it will temporarily reduce power feeds to the compute racks. Users should expect jobs to run more slowly than usual during the planned work.
Standard job partitions will still be available, though certain compute nodes may be more affected than others. Priority-Tier partitions will be disabled during the week.
Scope
180 Bouchet nodes will be impacted by planned data center work July 20-24:
The following will be impacted during the entire week:
- 10 H200 nodes (gpu_h200, gpu_devel partitions) with names beginning with a1122, a1124, and a1126
- 12 RTX 5000 Ada nodes (gpu, gpu_devel, education_gpu partitions) with names beginning with a1128
- 10 L40S (pi_co54 partition) nodes with names beginning with a1118
- 100 non-GPU Intel nodes (day, bigmem, devel, education, mpi* partitions) with names beginning with a1130 and a1132
The following will be impacted only on Monday morning (July 20) and Wednesday afternoon (July 22):
- 7 B200 nodes (gpu_b200 partition) with names beginning with a1116
- 8 RTX Pro 6000 Blackwell nodes (gpu_rtx6000, gpu_devel partitions) with names beginning with a1112
- 33 non-GPU AMD (day, week, devel partitions) nodes with names beginning with a1114
Bouchet nodes outside the scope of this electric work: 60 nodes in the mpi partition and one B200 node in the gpu_devel partition.
* Since MPI workflows can be sensitive to slow-performing nodes, we will temporarily move some affected nodes (20 non-GPU Intel nodes) from the mpi partition to the day partition to mitigate the impact of slower nodes. The remaining 60 nodes in the mpi partition will be unaffected by the electrical work.
Performance Impact
We expect minor impact to the RTX Pro 6000 Blackwell, RTX 5000 Ada, and non-GPU AMD nodes during planned electrical work. We estimate a 2x slowdown for the L40S and non-GPU Intel nodes and H200 nodes.
We expect significant performance impact on the B200 nodes. However, the extent of performance degradation will depend on the specific workflow, ranging from minimal to significant.
Once full power is restored to the B200 nodes, they will need to be restarted to restore normal performance. In preparation, we will create a scheduler reservation to drain all running jobs on the B200 nodes by the end of the work impacting them. During this time, the scheduler will not start jobs with requested run times that overlap with the reservation time. Once the nodes are restarted, they will resume normal operation.
Please adjust your job wall-times when submitting jobs as needed and use checkpointing when possible. Let us know if you would like assistance doing so.
The anticipated completion of scheduled work is end of day, Friday, July 24, 2026. YCRC will send an email communication when the work is completed. We are committed to keeping you as informed as possible at every stage of the work. Updates will be posted on this page. We recognize the reduced performance will impact your work, and we apologize for the inconvenience. If you have questions, comments, or need assistance with your jobs, please contact us at ycrc@yale.edu.
Helpful tips and resources: