Notice history

Cannon Cluster

Operational

SLURM Scheduler - Cannon

Operational

Cannon Compute Cluster (Holyoke)

Operational

Boston Compute Nodes

Operational

GPU nodes (Holyoke)

Operational

SEAS compute partition

Operational

FASSE Cluster

Operational

SLURM Scheduler - FASSE

Operational

FASSE Compute Cluster (Holyoke)

Operational

Kempner Cluster

Operational

Kempner Cluster CPU

Operational

Kempner Cluster GPU

Operational

FASSE login nodes

Operational

VDI/OpenOnDemand

Operational

Cannon VDI (Open OnDemand)

Operational

FASSE VDI (Open OnDemand)

Operational

Storage

Operational

Holyscratch01 (Global Scratch)

Operational

Home Directory Storage - Boston

Operational

HolyLFS03 (Tier 0)

Operational

HolyLFS04 (Tier 0)

Operational

HolyLFS05 (Tier 0)

Operational

Holystore01 (Tier 0)

Operational

Holylabs

Operational

BosLFS02 (Tier 0)

Operational

Isilon Storage Boston (Tier 1)

Operational

Isilon Storage Holyoke (Tier 1)

Operational

CEPH Storage Boston (Tier 2)

Operational

Tape - (Tier 3)

Operational

Boston Specialty Storage

Operational

Holyoke Specialty Storage

Operational

Samba Cluster

Operational

Globus Data Transfer

Operational

bosECS

Operational

holECS

Operational

Authentication

Operational

Virtual Machines

Operational

Networking

Operational

Software & LIcensing

Operational

Websites & Tools

Operational

External Reources - Data Centers

Operational

External Resources - Service Providers

Operational

Apr 2024

Resolved
April 15, 2024 at 1:43 PM
Resolved
April 15, 2024 at 1:43 PM
This incident has been resolved. holyscratch01 is back to normal
Investigating
April 12, 2024 at 6:55 PM
Investigating
April 12, 2024 at 6:55 PM
Holyscratch01 (/n/holyscratch01) is currently experiencing instability due to large number of jobs and high load. Please be patient and allow these jobs to cycle through which will eventually bring down the load

Resolved
April 06, 2024 at 9:35 PM
Resolved
April 06, 2024 at 9:35 PM
This incident has been resolved.
Investigating
April 06, 2024 at 7:06 PM
Investigating
April 06, 2024 at 7:06 PM
We are currently investigating this incident.

Resolved
April 01, 2024 at 3:29 PM
Resolved
April 01, 2024 at 3:29 PM
This incident has been resolved. Authentication is working normally again.
Identified
April 01, 2024 at 3:20 PM
Identified
April 01, 2024 at 3:20 PM
Login nodes are working normally, VDI nodes will need to be rebooted.
Investigating
April 01, 2024 at 2:43 PM
Investigating
April 01, 2024 at 2:43 PM
Authentication fails when logging into login.rc.fas.harvard.edu or directly to a login node.
We are investigating.

Monthly Maintenance Monday April 1st, 2024 from 7am-11am

Completed
April 01, 2024 at 3:00 PM
Completed
April 01, 2024 at 3:00 PM
Maintenance has completed successfully
In progress
April 01, 2024 at 11:00 AM
In progress
April 01, 2024 at 11:00 AM
Maintenance is now in progress
Planned
April 01, 2024 at 11:00 AM
Planned
April 01, 2024 at 11:00 AM
NOTICES
- TRAINING: 2024 training sessions are available. Topics include new user training as well as advanced topics. To see current and future training sessions, view our training alendar at: https://www.rc.fas.harvard.edu/upcoming-training/
  For these and other events such as Office Hours, view our entire events calendar at: https://www.rc.fas.harvard.edu/upcoming-events/
- SURVEY: If you have not yet, we invite you to fill out our 2024 user survey *approx. 15 minutes) and give us your feedback. The survey is anonymous and asks questions about all of our services including cluster, storage, and support. The survey will be available until April 12th.
  
  https://harvard.az1.qualtrics.com/jfe/form/SV_e3AmuOrrmBOHTCu
- STATUS PAGE: You can subscribe to our status to receive notifications of any issues and their resolution at https://status.rc.fas.harvard.edu/ (click Get Updates for options).
MAINTENANCE TASKS
Cannon cluster will be paused during this maintenance?: NO
FASSE cluster will be paused during this maintenance?: NO
Ticket System host move
-- Audience: All users
-- Impact: The FASRC ticket system will be unavailable during maintenance while we move it to a new host. Emails sent during this time should pend but still reach the ticket system once it is back online.
Login node and Open OnDemand (OOD/VDI) reboots
-- Audience: Anyone logged into a login node or VDI/OOD node
-- Impact: Login and VDI/OOD nodes will rebooted during this maintenance window
Scratch cleanup ( https://docs.rc.fas.harvard.edu/kb/policy-scratch/ )
-- Audience: Cluster users
-- Impact: Files older than 90 days will be removed. Please note that retention cleanup can run at any time, not just during the maintenance window.
Thanks,
FAS Research Computing
Dept. Website: https://www.rc.fas.harvard.edu/
Documentation: https://docs.rc.fas.harvard.edu/
Status Page: https://status.rc.fas.harvard.edu/

Mar 2024

Resolved
March 21, 2024 at 12:00 PM
Resolved
March 21, 2024 at 12:00 PM
There was a brief disruption in the connectivity from MGHPCC to the outside world/Internet. During this period outside attempts to connect to services hosted in Holyoke would time out. Internal data center network remained normal and no jobs were affected.

Resolved
March 19, 2024 at 1:35 PM
Resolved
March 19, 2024 at 1:35 PM
After making some hardware and network changes, we monitored holylabs yesterday and overnight and have determined that it is stable again.
We will continue to monitor but currently believe this issue has been resolved.
Investigating
March 18, 2024 at 1:51 PM
Investigating
March 18, 2024 at 1:51 PM
Holylabs was again unstable over the weekend. We are actively working this issue and will update as soon as possible.
Monitoring
March 15, 2024 at 1:31 PM
Update
March 15, 2024 at 1:31 PM
Holylabs locked up again overnight. We are investigating.
Monitoring
March 14, 2024 at 1:38 PM
Monitoring
March 14, 2024 at 1:38 PM
Holylabs became stuck overnight and was restarted. It is back up and we are monitoring.

Resolved
March 18, 2024 at 3:21 PM
Resolved
March 18, 2024 at 3:21 PM
holyscratch01 is currently stable. However, it continues to be under heavy utilization as a matter of course due to job loads.
There is a plan underway to replace holyscratch01, but no ETA at this time as we are still qualifying and planning purchasing. We will notify the community when an ETA is known. Until then please be aware that load on scratch may continue along this trend. Thanks for your understanding.
Monitoring
March 13, 2024 at 2:15 PM
Monitoring
March 13, 2024 at 2:15 PM
Performance has improved, but we are still monitoring some high loads across object storage units.
Investigating
March 12, 2024 at 7:54 PM
Investigating
March 12, 2024 at 7:54 PM
holyscratch01 performance degraded
We are currently investigating this incident.

Resolved
March 04, 2024 at 5:55 PM
Resolved
March 04, 2024 at 5:55 PM
VMs are returning to operation. If you receive an error (on Portal, Coldfront, etc) please wait a few minutes and try again.
Ticket system is online and starting to receive delayed emails.
Investigating
March 04, 2024 at 4:41 PM
Investigating
March 04, 2024 at 4:41 PM
The RT ticket system is currently offline due to a VM issue.
Any emails sent to the ticket system will eventually be delivered once it recovers, but until then expect delayed response.

Resolved
March 04, 2024 at 9:44 PM
Resolved
March 04, 2024 at 9:44 PM
This incident has been resolved.
Monitoring
March 04, 2024 at 4:21 PM
Monitoring
March 04, 2024 at 4:21 PM
Informational Notice
The Slurm upgrade to 23.11.4 was completed successfully during maintenance. However a complication with the automation of Slurm's cryptographic keys occurred during the upgrade which caused nodes to lose the ability to talk to the Slurm master. The Slurm master therefore viewed those nodes as down and requeued their jobs.
All jobs on Cannon and FASSE were requeued.
This is deeply regrettable but the chain of events which caused this could not be foreseen.
To check the status of your jobs, see the common Slurm commands at:
https://docs.rc.fas.harvard.edu/kb/convenient-slurm-commands/#Information_on_jobs
FAS Research Computing
https://docs.rc.fas.harvard.edu/
rchelp@rc.fas.harvard.edu

Feb 2024

Resolved
February 28, 2024 at 7:59 PM
Resolved
February 28, 2024 at 7:59 PM
The Ceph instability has been resolved. Ceph Tier2 shares, VDI, and VMs should be back to their normal state.

If your VM, /net/fs-[labname] share, or VDI session is still impacted, please contact rchelp@rc.fas.harvard.edu
Identified
February 28, 2024 at 6:24 PM
Identified
February 28, 2024 at 6:24 PM
The infrastructure behind Tier2 Ceph shares and VMs is unstable.
This also affects VDI/OOD which relies on virtual machines.

/net/fs-[labname] shares, new OOD/VDI sessions, and VMs are affected and may will be inaccessible until this is resolved.

Thanks for your patience.

Resolved
February 27, 2024 at 6:31 PM
Resolved
February 27, 2024 at 6:31 PM
This incident has been resolved.
Investigating
February 26, 2024 at 2:00 PM
Investigating
February 26, 2024 at 2:00 PM
holyscratch01 is experiencing high load. Performance may be delayed.
OOD is also affected as a result of this.
We are currently investigating this incident.

Resolved
February 08, 2024 at 8:56 PM
Resolved
February 08, 2024 at 8:56 PM
Clusters are back in production status. We will continue to monitor for any aberrant behavior, but this incident has been resolved.
Identified
February 08, 2024 at 8:07 PM
Identified
February 08, 2024 at 8:07 PM
I spoke too soon, partial recovery, I'll update here when we are sure everything is back in production, apologies.
Monitoring
February 08, 2024 at 7:48 PM
Monitoring
February 08, 2024 at 7:48 PM
All cannon nodes are back in service and slurm is resuming jobs, fasse is coming back up as well, we will monitor the situation, but anticipate a return to full normal operations shortly.
Identified
February 08, 2024 at 7:35 PM
Update
February 08, 2024 at 7:35 PM
The nodes are coming back into normal service, we anticipate this to be fairly quick
Identified
February 08, 2024 at 7:20 PM
Identified
February 08, 2024 at 7:20 PM
slurm is operational, jobs are idled and should resume as normal
Investigating
February 08, 2024 at 7:18 PM
Investigating
February 08, 2024 at 7:18 PM
We are currently investigating this incident. We will update here as we can

Resolved
February 23, 2024 at 3:10 PM
Resolved
February 23, 2024 at 3:10 PM
This incident has been resolved.
Investigating
February 08, 2024 at 6:55 PM
Investigating
February 08, 2024 at 6:55 PM
Boslfs02 (/n/boslfs02) is currently experiencing instability.
We are working on the issue and will update this incident as the situation progresses.

Resolved
February 05, 2024 at 4:37 PM
Resolved
February 05, 2024 at 4:37 PM
This incident has been resolved.
Identified
February 05, 2024 at 4:33 PM
Identified
February 05, 2024 at 4:33 PM
We are working on a fix for this incident.

Feb 2024 to Apr 2024