SLURM Scheduler - Cannon - 정상
SLURM Scheduler - Cannon
Cannon Compute Cluster (Holyoke) - 정상
Cannon Compute Cluster (Holyoke)
Boston Compute Nodes - 정상
Boston Compute Nodes
GPU nodes (Holyoke) - 정상
GPU nodes (Holyoke)
seas_compute - 정상
seas_compute
SLURM Scheduler - FASSE - 정상
SLURM Scheduler - FASSE
FASSE Compute Cluster (Holyoke) - 정상
FASSE Compute Cluster (Holyoke)
Kempner Cluster CPU - 정상
Kempner Cluster CPU
Kempner Cluster GPU - 정상
Kempner Cluster GPU
FASSE login nodes - 정상
FASSE login nodes
Cannon Open OnDemand - 정상
Cannon Open OnDemand
FASSE Open OnDemand - 정상
FASSE Open OnDemand
Netscratch (Global Scratch) - 정상
Netscratch (Global Scratch)
Home Directory Storage - Boston - 정상
Home Directory Storage - Boston
Tape - (Tier 3) - 정상
Tape - (Tier 3)
Holylabs - 정상
Holylabs
Isilon Storage Holyoke (Tier 1) - 정상
Isilon Storage Holyoke (Tier 1)
Holystore01 (Tier 0) - 정상
Holystore01 (Tier 0)
HolyLFS04 (Tier 0) - 정상
HolyLFS04 (Tier 0)
HolyLFS05 (Tier 0) - 정상
HolyLFS05 (Tier 0)
HolyLFS06 (Tier 0) - 성능 저하
HolyLFS06 (Tier 0)
Holyoke Tier 2 NFS - 정상
Holyoke Tier 2 NFS
Holyoke Specialty Storage - 정상
Holyoke Specialty Storage
holECS - 정상
holECS
Isilon Storage Boston (Tier 1) - 정상
Isilon Storage Boston (Tier 1)
BosLFS02 (Tier 0) - 정상
BosLFS02 (Tier 0)
Boston Tier 2 NFS - 정상
Boston Tier 2 NFS
CEPH Storage Boston (Tier 2) - 정상
CEPH Storage Boston (Tier 2)
Boston Specialty Storage - 정상
Boston Specialty Storage
bosECS - 정상
bosECS
Samba Cluster - 정상
Samba Cluster
Globus Data Transfer - 정상
Globus Data Transfer
알림 내역
9월 2026
- 완료됨9월 14, 2026 ~에서 17:00UTC완료됨9월 14, 2026 ~에서 17:00UTC유지 보수가 성공적으로 완료되었습니다
- 업데이트9월 14, 2026 ~에서 13:08UTC업데이트9월 14, 2026 ~에서 13:08UTC
The Holyoke/MGHPCC row 7c power work has been CANCELLED by the facility and will be re-scheduled.
The following work WILL NOT take pleace today:
Cancelled: MGHPCC will be upgrading power on Pod 7c Even Side on September 14th 7am-5pm. This necessitates idling half the nodes on that side of the pod. A blocking reservation has been put in place to accomplish this. No jobs will be canceled but users will notice degraded scheduling throughput due to half the nodes being closed in the following partitions:
arguelles_delgado, blackhole, conroy, davies, desai, doshi-velez, dsouza, eddy, edwards, geophysics, giribet, gpu_test, hernquist, huce_cascade, huttenhower, imasc, jacobsen2, janson_cascade, janson, ke, lukin, murphy, nguyen, ni_lab, olveczky, ortegahernandez, pehlevan, seas_compute, shared, shakhnovich, tambe, unrestricted, vishwanath, whipple, xlin, yin, zon
- 진행 중9월 14, 2026 ~에서 13:00UTC진행 중9월 14, 2026 ~에서 13:00UTC현재 유지 보수가 진행 중입니다
- 예정됨8월 31, 2026 ~에서 18:56UTC예정됨8월 31, 2026 ~에서 18:56UTC
Our maintenance tasks should be completed between 9am-1pm.
Some power work on row 7c will run 9-5 but will not affect running jobs or new jobs (see below).NOTICES:
Monday, October 12 is a university holiday (Indigenous Peoples/Columbus)
Training: Upcoming training from FASRC and other sources can be found on our Training Calendar. at https://www.rc.fas.harvard.edu/upcoming-training/
Status Page: You can subscribe to our status to receive notifications of maintenance, incidents, and their resolution at https://status.rc.fas.harvard.edu/ (click Get Updates for options).
MAINTENANCE TASKS
Cannon cluster will be paused during this maintenance?: YES
FASSE cluster will be paused during this maintenance?: YESSlurm Upgrade to 26.05.4
Audience: Cluster
Impact: The cluster will be paused during this maintenance
New: Enable Termination of User Processes on Logout
Audience: Cluster
Impact: Going forward all user processes will be terminated upon login session exit on the login nodes, excluding things running in screen and tmux. Users should leverage the cluster for non-interactive processes. This is to clean up after AI agents which tend to create orphaned processes which drag down login node performance.
New: watch command cadence limit
Audience: Cluster
Impact:The watch command will have a minimum cadence of 60s when used on commands talking to the slurm scheduler (i.e. squeue, showq, sinfo, scontrol, sdiag, lsload). Users desiring faster polling should leverage the sacct command that talks to the slurm database. In general users should not poll the scheduler more than once every minute, ideally once every 5-10 minutes. Polling more often slows the scheduler. FASRC reserves the right to ban users who who tax the scheduler with queries. For more on cluster customs and responsibilities see: https://docs.rc.fas.harvard.edu/kb/responsibilities/
Holyoke/MGHPCC row 7c power work - 9am-5pm (CANCELLED)
Audience: ClusterImpact: MGHPCC will be upgrading power on Pod 7c Even Side on September 14th 7am-5pm. This necessitates idling half the nodes on that side of the pod. A blocking reservation has been put in place to accomplish this. No jobs will be canceled but users will notice degraded scheduling throughput due to half the nodes being closed in the following partitions:arguelles_delgado, blackhole, conroy, davies, desai, doshi-velez, dsouza, eddy, edwards, geophysics, giribet, gpu_test, hernquist, huce_cascade, huttenhower, imasc, jacobsen2, janson_cascade, janson, ke, lukin, murphy, nguyen, ni_lab, olveczky, ortegahernandez, pehlevan, seas_compute, shared, shakhnovich, tambe, unrestricted, vishwanath, whipple, xlin, yin, zon
Login node reboots
Audience: All login nodes
Impact: Login nodes will be unavailable until after maintenance
OOD/Open OnDemand down/reboots
Audience: All OOD users
Impact: OOD will be unavailable until after maintenance
Gurobi license key update
Audience: Anyone who uses Gurobi software on the clusters.
Impact: Jobs running Gurobi may fail. FASRC recommends waiting until maintenance is over to run Gurobi jobs.
Netscratch 90-day retention cleanup
Audience; All netscratch users
Impact: Files older than 90 days will be removed per our scratch policy. Please note that this cleanup can happen at any time, not just during maintenance.
Thank you,
FAS Research Computing
https://docs.rc.fas.harvard.edu/
https://www.rc.fas.harvard.edu/
8월 2026
- 완료됨8월 27, 2026 ~에서 20:06UTC완료됨8월 27, 2026 ~에서 20:06UTC
OS upgrade work is mostly completed.
There are a few remaining nodes in the cluster that did not successfully receive the upgrade.
If you see nodes that are 'DOWN' for the following reasons, they are all nodes that need manual intervention for OS upgrade. FASRC staff are aware of them and will upgrade these nodes over the next week.
"wrong kernel"
"Not responding"
"unexpected reboot"
"gres/gpu count lower than configured"
"/n/sw not mounted"
We appreciate your continued patience.
- 업데이트8월 26, 2026 ~에서 20:01UTC업데이트8월 26, 2026 ~에서 20:01UTC
Today's work is complete. The final round, Cannon Part 3, takes place tomorrow.
August 26th: Cannon Part 2
arguelles_delgado
conroy
davies
doshi-velez
dsouza
edwards
geophysics
giribet
gpu_test
hernquist
huce_cascadeimasc
janson
kempner_h200
kempner_rtx
murphy
ni_lab
olveczky
ortegahernandez
pehlevan
seas_compute
shakhnovich
shared
unrestricted
xlin
yin
zon - 업데이트8월 25, 2026 ~에서 19:37UTC업데이트8월 25, 2026 ~에서 19:37UTC
Today's work is complete for Cannon Part 1 and will resume tomorrow for Cannon Part 2
August 25th: Cannon Part 1
holylogin
arguelles_delgado
blackhole
davies_gpu
davies
desai
eddy
holy-cow
holy-smokes
huce_bigmem
huce_cascade
huttenhower
jacobsen2
janson_bigmem
janson_cascade
janson
ke
lukin
nguyen
olveczky_gpu
remoteviz
seas_compute
shared
sompolinsky_gpu
tambe
vishwanath
whipple
xlin
xlin_ice
zhuang_gpu
zhuang - 업데이트8월 24, 2026 ~에서 19:07UTC업데이트8월 24, 2026 ~에서 19:07UTC
Today's portion of the upgrades are completed. Tomorrow's list (Cannon part 1) can be found in this maintenance event.
August 24th:
FASSE
boslogin
private login nodes - 진행 중8월 24, 2026 ~에서 13:00UTC진행 중8월 24, 2026 ~에서 13:00UTC현재 유지 보수가 진행 중입니다
- 예정됨8월 03, 2026 ~에서 19:35UTC예정됨8월 03, 2026 ~에서 19:35UTC
FASRC will be doing OS upgrades to the latest version of Rocky 8.10 from August 24-27th. These rolling upgrades will improve cluster security and install the latest Cuda 13.2 driver.
Upgrades will happen in stages with each stage occurring between 8a-5p each day. During that period the nodes being upgraded will be unavailable. No jobs will be canceled but jobs will stay pending if all nodes in the partition are down.
Note that during this week the cluster will be in a partially upgraded state, thus jobs that span multiple nodes may become unstable due to mismatching libraries. Users concerned about this should wait until August 28th to restart runs.
Users should plan their work accordingly.
The upgrade schedule with impacted partitions is as follows:
August 24th:
FASSE
boslogin
private login nodesAugust 25th: Cannon Part 1
holylogin
arguelles_delgado
blackhole
davies_gpu
davies
desai
eddy
holy-cow
holy-smokes
huce_bigmem
huce_cascade
huttenhower
jacobsen2
janson_bigmem
janson_cascade
janson
ke
lukin
nguyen
olveczky_gpu
remoteviz
seas_compute
shared
sompolinsky_gpu
tambe
vishwanath
whipple
xlin
xlin_ice
zhuang_gpu
zhuangAugust 26th: Cannon Part 2
arguelles_delgado
conroy
davies
doshi-velez
dsouza
edwards
geophysics
giribet
gpu_test
hernquist
huce_cascadeimasc
janson
kempner_h200
kempner_rtx
murphy
ni_lab
olveczky
ortegahernandez
pehlevan
seas_compute
shakhnovich
shared
unrestricted
xlin
yin
zonAugust 27th: Cannon Part 3
arguelles_delgado_gpu_a100
arguelles_delgado_gpu_mixed
arguelles_delgado_h100
bigmem_intermediate
bigmem
blackhole_gpu
dvorkin
eddy
enos
gershman
gpu
gpu_h200
hejazi
hernquist_ice
hoekstra
hsph_gpu
hsph
huce_ice
iaifi_gpu
intermediate
itc_cluster
itc_gpu
janson_sapphire
joonholee
jshapiro
kempner_h100
kempner_h200
kempner
kempner_interactive
kovac
kozinsky_gpu
kozinsky
murphy_ice
mweber_compute
mweber_gpu
olveczky_sapphire
ortegahernandez_ice
rivas
sapphire
seas_compute
seas_gpu
seas_gpu_perf
siag_gpu
siag_combo
siag
sur
test
yao_alphatns
yao_gpu
yao
zhuang
- 업데이트UTC업데이트UTC
A number of jobs appear to have terminated during the upgrade for some reason. Those jobs will re-queue or, if not, will need to be re-submitted,
We apologize for the inconvenience. - 해결됨UTC해결됨UTC
The update has been applied and the scheduler is returned to service. Jobs are unpaused.
- 조사 중UTC조사 중UTC
We have been manually dealing with a Slurm memory leak bug behind the scenes and now have an official fix in the form of a Slurm upgrade (26.05.3). It is necessary to get this update in place as soon as possible.
At 2pm today we will pause the cluster. Jobs will be paused, not terminated, and the scheduler will be unavailable for querying or submitting new jobs until the update is completed. Once complete, jobs will resume.
Technical information about the bug:
7월 2026
- 완료됨7월 06, 2026 ~에서 17:00UTC완료됨7월 06, 2026 ~에서 17:00UTC유지 보수가 성공적으로 완료되었습니다
- 진행 중7월 06, 2026 ~에서 13:00UTC진행 중7월 06, 2026 ~에서 13:00UTC현재 유지 보수가 진행 중입니다
- 예정됨6월 26, 2026 ~에서 13:45UTC예정됨6월 26, 2026 ~에서 13:45UTC
FASRC monthly maintenance will take place on July 6th 2026. Our maintenance tasks should be completed between 9am-1pm.
Cannon cluster will be paused during this maintenance?: NO
FASSE cluster will be paused during this maintenance?: NONOTICES:
Friday July 3rd is a university holiday (independence Day observed)
Training: Upcoming training from FASRC and other sources can be found on our Training Calendar. at https://www.rc.fas.harvard.edu/upcoming-training/
Status Page: You can subscribe to our status to receive notifications of maintenance, incidents, and their resolution at https://status.rc.fas.harvard.edu/ (click Get Updates for options).
We'd love to hear success stories about your or your lab's use of FASRC. Submit your story here.
MAINTENANCE TASKS
Domain controller replacement
Audience: Internal
Impact: None. End users should not see any impact.
Reboot drained nodes in error state
Audience: Cluster nodes with errors.
Impact: These nodes will have been drained already in preparation. No impact on jobs on the day and the affected nodes will return to service in their respective partitions after the maintenance period.
OOD/Open OnDemand reboots
Audience: All OOD users, reboot of the head nodes.
Impact: Running sessions will not be affected.
Login node reboots
Audience; All login node users.
Impact: Login nodes will reboot during the maintenance window.
Netscratch 90-day retention cleanup
Audience; All netscratch users
Impact: Files older than 90 days will be removed per our scratch policy. Please note that this cleanup can happen at any time, not just during maintenance.
Thank you,
FAS Research Computing
https://docs.rc.fas.harvard.edu/
https://www.rc.fas.harvard.edu/

