Over the past year we have noticed that water circulating through two of the Cooling Distribution Units (CDU) in Row 8a and their compute nodes has become significantly discoloured. This impurity causes cooling problems which as a result increases chances of node failure. To remedy FASRC has scheduled a full flush of these CDUs and their attached nodes. Unfortunately to do the full flush we have to fully power down all the nodes and drain the water, which is a process that takes several days to complete.
To minimize disruption we have staggered this work over two weeks. The work on CDU5 will take place from 9/22 - 9/25, while CDU6 will take place from 9/29 - 10/2. Lists of impacted partitions are below. For those impacted we recommend using other resources on the cluster during that time such as shared and gpu_h200 or any of the requeue partitions. Note this work impacts both Cannon and FASSE. No jobs will be cancel, rather blocking reservations are in place to naturally drain the nodes.
Thank you for your patience as we work to improve cluster stability and hardware longevity.
CDU 5 (9/22 - 9/25)
Cannon:
arguelles_delgado_gpu_a100
arguelles_delgado_gpu_mixed
bigmem_intermediate
blackhole_gpu
eddy
gershman
hejazi
hernquist_ice
hoekstra
huce_ice
iaifi_gpu
itc_gpu
jshapiro
kovac
kozinsky
kozinsky_gpu
murphy_ic
ortegahernandez_ice
rivas
seas_compute
seas_gpu
siag
siag_gpu
siag_combo
sur
zhuang
FASSE:
fasse_ultramem
CDU 6 (9/29 - 10/2)
Cannon:
arguelles_delgado_h100
bigmem
dvorkin
eddy
enos
gpu
hsph
hsph_gpu
intermediate
itc_cluster
janson_sapphire
joonholee
jshapiro
olveczky_sapphire
sapphire
test
yao
yao_alphatns
yao_gpu
FASSE:
cnl