SLURM Scheduler - Cannon - 运行正常
SLURM Scheduler - Cannon
Cannon Compute Cluster (Holyoke) - 运行正常
Cannon Compute Cluster (Holyoke)
Boston Compute Nodes - 运行正常
Boston Compute Nodes
GPU nodes (Holyoke) - 运行正常
GPU nodes (Holyoke)
seas_compute - 运行正常
seas_compute
SLURM Scheduler - FASSE - 运行正常
SLURM Scheduler - FASSE
FASSE Compute Cluster (Holyoke) - 运行正常
FASSE Compute Cluster (Holyoke)
Kempner Cluster CPU - 运行正常
Kempner Cluster CPU
Kempner Cluster GPU - 运行正常
Kempner Cluster GPU
FASSE login nodes - 运行正常
FASSE login nodes
Cannon Open OnDemand - 运行正常
Cannon Open OnDemand
FASSE Open OnDemand - 运行正常
FASSE Open OnDemand
Netscratch (Global Scratch) - 运行正常
Netscratch (Global Scratch)
Home Directory Storage - Boston - 运行正常
Home Directory Storage - Boston
Tape - (Tier 3) - 运行正常
Tape - (Tier 3)
Holylabs - 运行正常
Holylabs
Isilon Storage Holyoke (Tier 1) - 运行正常
Isilon Storage Holyoke (Tier 1)
Holystore01 (Tier 0) - 运行正常
Holystore01 (Tier 0)
HolyLFS04 (Tier 0) - 运行正常
HolyLFS04 (Tier 0)
HolyLFS05 (Tier 0) - 运行正常
HolyLFS05 (Tier 0)
HolyLFS06 (Tier 0) - 运行正常
HolyLFS06 (Tier 0)
Holyoke Tier 2 NFS - 运行正常
Holyoke Tier 2 NFS
Holyoke Specialty Storage - 运行正常
Holyoke Specialty Storage
holECS - 运行正常
holECS
Isilon Storage Boston (Tier 1) - 运行正常
Isilon Storage Boston (Tier 1)
BosLFS02 (Tier 0) - 运行正常
BosLFS02 (Tier 0)
Boston Tier 2 NFS - 运行正常
Boston Tier 2 NFS
CEPH Storage Boston (Tier 2) - 运行正常
CEPH Storage Boston (Tier 2)
Boston Specialty Storage - 运行正常
Boston Specialty Storage
bosECS - 运行正常
bosECS
Samba Cluster - 运行正常
Samba Cluster
Globus Data Transfer - 运行正常
Globus Data Transfer
历史记录
查看当前状态4月 2025
- 已解决UTC已解决UTC
Cannon boslogin and FASSE login nodes are back up and operational.
All holylogin nodes are still down for repair, please see our posted incident for more updates: https://status.rc.fas.harvard.edu/cm97gyay90013dturk7fxg5pb
We apologize for the unexpected disruption.
- 调查中UTC调查中UTC
Due to a configuration error, all cluster login nodes are rebooting and are temporarily unavailable. Please save any work immediately.
- 已解决UTC已解决UTCHardware has been repaired and holyoke login nodes are back online. Thanks for your patience.
- 持续监控中UTC持续监控中UTC
Holylogin chassis repair during maintenance was unsuccessful and replacement parts have been ordered.
holyoke login nodes (holylogin05-08) are down for hardware repair
Only Boston login nodes available (ie, boslogin[05-08])
If you have holylogin hard-coded in your scripts, please update to login.rc.fas.harvard.edu or boslogin.rc.fas.harvard.edu for the time being, which will redirect you to an available login node.
As always, the best method for obtaining a login node is using
login.rc.fas.harvard.eduwhich will pick a node for you.If you require a login node in a specific data center, use
boslogin.rc.fas.harvard.edu(Boston) or (once they are back in service)holylogin.rc.fas.harvard.edu(Holyoke).See also: Command line access with Terminal (login nodes) – FASRC DOCS
- 已解决UTC已解决UTC
This incident was posted by mistake.
holylogin01-04 were replaced by holylogin05-08 some time back.
As always, the best method for obtaining a login node is usinglogin.rc.fas.harvard.eduwhich will pick a node for you.Or if you require a login node in a specific data center, use
boslogin.rc.fas.harvard.edu(Boston) orholylogin.rc.fas.harvard.edu(Holyoke).See also: Command line access with Terminal (login nodes) – FASRC DOCS
- 调查中UTC调查中UTC
Holylogin chassis repair during maintenance was unsuccessful and replacement parts have been ordered.
Audience:
All cluster users
Impact:
All holylogin** servers will be down till further notice
Only Boston login nodes available (ie, boslogin[05-08])
If you have holylogin hard-coded in your scripts, please update to login.rc.fas.harvard.edu or boslogin.rc.fas.harvard.edu for the time being, which will redirect you to an available login node.
Updates to follow as we have them.
3月 2025
- 更新三月 03, 2025 在 下午 6:00UTC更新三月 03, 2025 在 下午 6:00UTCMaintenance has completed successfully
- 已完成三月 03, 2025 在 下午 6:00UTC已完成三月 03, 2025 在 下午 6:00UTCMaintenance has completed successfully
- 进行中三月 03, 2025 在 下午 2:00UTC进行中三月 03, 2025 在 下午 2:00UTCMaintenance is now in progress
- 已计划三月 03, 2025 在 下午 2:00UTC已计划三月 03, 2025 在 下午 2:00UTC
PLEASE NOTE - New time window going forward - 9am-1pm
FASRC monthly maintenance will take place Monday March 3rd, 2025 from 9am-1pm
NOTICES
Training: Upcoming training from FASRC and other sources can be found on our Training Calendar. at https://www.rc.fas.harvard.edu/upcoming-training/
Status Page: You can subscribe to our status to receive notifications of maintenance, incidents, and their resolution at https://status.rc.fas.harvard.edu/ (click Get Updates for options).
Upcoming holidays: Memorial Day - Monday, May 26
You can subscribe to our status page using the Get Updates button in the upper right
MAINTENANCE TASKS
Cannon cluster will be paused during this maintenance?: YES
FASSE cluster will be paused during this maintenance?: YESSlurm Upgrade to 24.11.2 - Crucial Update
Audience: All cluster users
Impact: Jobs and the scheduler will be paused during this upgrade
Open Ondemand (OOD) reboots
Audience: All OOD users
Impact: All Open OnDemand (aka OOD/VDI/RCOOD) nodes will be rebooted
Login node reboots
Audience: Anyone logged into a FASRC Cannon or FASSE login node
Impact: Login nodes will rebooted during this maintenance window
bos-Isilon firmware updates
Audience: bos-isilon users
Impact: No noticeable impact for storage users
Netscratch retention/cleanup ( https://docs.rc.fas.harvard.edu/kb/policy-scratch/ )
Audience: Cluster users
Impact: Files older than 90 days will be removed. Please note that retention cleanup can and does run at any time, not just during the maintenance window.
Thank you,
FAS Research Computing
https://www.rc.fas.harvard.edu/
https://docs.rc.fas.harvard.edu/
2月 2025
- 已解决UTC已解决UTC
An emergency patch of the scheduler has resolved the Multiple Partition issue
- 调查中UTC调查中UTC
Since mid-January we've been seeing some strange issues with the scheduler which caused periodic stalls or unresponsiveness in the scheduler. We had hoped that the Slurm upgrade to 24.11.1 would resolve those issues due to various architecture changes in the communications backend. Unfortunately they did not; we have since opened an issue with SchedMD (our service vendor for the scheduler). This has since spiraled into finding several other issues with the scheduler which we are working to remediate. Below is a status report regarding these issues:
1. High Agent Load Stall (RESOLVED): This was reported in https://support.schedmd.com/show_bug.cgi?id=21975 The scheduler would stall due to being oversaturated with blocking requests. This turned out to be due to a new Slurm feature called stepmgr which we had enabled to handle jobs with many steps. Unfortunately this feature also increased the load on the scheduler for array jobs exiting at the same time which caused the stall. Since we tend not to have many users that use many steps we opted to disable the stepmgr function. This resolved the High Agent Load issue. Users that have many steps in their job may still turn on the stepmgr for their specific job by adding #SBATCH --stepmgr (https://slurm.schedmd.com/sbatch.html#OPT_stepmgr)
2. Scheduler Thrashing (MONITORING): We discovered this while working on the previous bug and continued to work on it in the same bug report: https://support.schedmd.com/show_bug.cgi?id=21975 Under high load, the scheduler would get into a thrashing state where the scheduler would effectively go heads down and ignore incoming requests in order to focus on scheduling jobs. To users this would look like the scheduler was unresponsive as the scheduler was ignoring their requests to deal with higher priority traffic. To remediate this we increased the thread count for the scheduler and implemented a throttle to slow things down so that the scheduler could respond to all the requests with out impacting scheduler throughput. This is in place now and appears to have resolved the issue. We are continuing to monitor the scheduler to tune this throttle.
3. --test-only requeue crash (RESOLVED): During this investigation we also ran into another bug reported by another group related to jobs that were submitted using --test-only that would in theory preempt other jobs (see: https://support.schedmd.com/show_bug.cgi?id=21997). This caused the scheduler to crash. Given the severity of the bug we emergency patched the scheduler on Feb 12th to resolve this issue.
4. Multiple Partition Jobs Labelled with Wrong Partition (IN PROGRESS): This is a new issue identified on 2/13 related to jobs that submit to multiple partitions at once (https://support.schedmd.com/show_bug.cgi?id=22076). When the job schedules it may run in one partition but be labelled as being in another. This can lead to job preemption issues as the jobs are labelled as being in partitions that cannot be preempted even though they were originally scheduled in partitions that could be. This was identified earlier by another group and SchedMD is working on a patch. Depending on the timing FASRC will either emergency patch the scheduler for this issue or wait for the formal release of 24.11.2. Note that this issue really only impacts preemption and the scheduler is working fine otherwise. If you see jobs that you think should be preempted but are not and are blocking your work please let us know and we will investigate.
Thank you for your patience as we work through these issues.

