سجل التاريخ

جاهز للعمل

ديسمبر 2023

تم الحل
ديسمبر 29, 2023 في 19:40
تم الحل
ديسمبر 29, 2023 في 19:40
The Ceph instability has been resolved. Ceph Tier2 shares, VDI, and VMs should be back to their normal state.

If your VM, /net/fs-[labname] share, or VDI session is still impacted, please contact rchelp@rc.fas.harvard.edu
محدد
ديسمبر 27, 2023 في 20:38
محدد
ديسمبر 27, 2023 في 20:38
The infrastructure behind Tier2 Ceph shares and VMs is unstable.
This also affects VDI/OOD which relies on virtual machines.

/net/fs-[labname] shares, new OOD/VDI sessions, and VMs are affected and may will be inaccessible until this is resolved.

Thanks for your patience.

تم الحل
ديسمبر 22, 2023 في 22:26
تم الحل
ديسمبر 22, 2023 في 22:26
The Ceph instability has been resolved. Ceph Tier2 shares, VDI, and VMs should be back to their normal state.

If your VM, /net/fs-[labname] share, or VDI session is still impacted, please contact rchelp@rc.fas.harvard.edu
محدد
ديسمبر 22, 2023 في 16:05
محدد
ديسمبر 22, 2023 في 16:05
The infrastructure behind Tier2 Ceph shares and VMs is unstable.
This also affects VDI/OOD which relies on virtual machines.

/net/fs-[labname] shares, new OOD/VDI sessions, and VMs are affected and may will be inaccessible until this is resolved.

Thanks for your patience.

تم الحل
ديسمبر 18, 2023 في 15:54
تم الحل
ديسمبر 18, 2023 في 15:54
Downed nodes are back online. The cluster is now available again.

Thank you for your patience.
محدد
ديسمبر 18, 2023 في 14:50
محدد
ديسمبر 18, 2023 في 14:50
Power was restored to the affected sections and we are bringing the down nodes back up.
تحقيق
ديسمبر 18, 2023 في 12:00
تحقيق
ديسمبر 18, 2023 في 12:00
There has been a loss of street power at MGHPCC, our Holyoke data center due to the windstorm. We are awaiting further details.

This likely affects most of the Cannon/Kempner compute cluster, and all of the FASSE compute cluster.

More details as we learn them.

تم الحل
ديسمبر 15, 2023 في 14:19
تم الحل
ديسمبر 15, 2023 في 14:19
The Ceph instability has been resolved. Ceph Tier2 shares, VDI, and VMs should be back to their normal state.

If your VM, /net/fs-[labname] share, or VDI session is still impacted, please contact rchelp@rc.fas.harvard.edu
محدد
ديسمبر 15, 2023 في 13:27
محدد
ديسمبر 15, 2023 في 13:27
The infrastructure behind Tier2 Ceph shares and VMs is unstable.
This also affects VDI/OOD which relies on virtual machines.

/net/fs-[labname] shares, new OOD/VDI sessions, and VMs are affected and may will be inaccessible until this is resolved.

Thanks for your patience.

Emergency scheduler update - Security Patch

مكتمل
ديسمبر 14, 2023 في 17:00
مكتمل
ديسمبر 14, 2023 في 17:00
Maintenance has completed successfully
قيد التقدم
ديسمبر 14, 2023 في 16:00
قيد التقدم
ديسمبر 14, 2023 في 16:00
Maintenance is now in progress
مخطط
ديسمبر 14, 2023 في 16:00
مخطط
ديسمبر 14, 2023 في 16:00
SchedMD has issued a patch for a critical security vulnerability. It is imperative that we apply this patch ASAP. This will require a pause of the cluster for roughly 30-60 minutes at around 11am, after which jobs and the scheduler will continue normal operation.

A side effect of this is that we will move up to version 23.02.7 with the patch which will fix the IntelMPI and PMI2 issues some users have experienced.

Thanks for your understanding and patience.
FAS Research Computing

نوفمبر 2023

تم الحل
نوفمبر 26, 2023 في 23:20
تم الحل
نوفمبر 26, 2023 في 23:20
The Ceph instability has been resolved. Ceph Tier2 shares, VDI, and VMs should be back to their normal state.

If your VM, /net/fs-[labname] share, or VDI session is still impacted, please contact rchelp@rc.fas.harvard.edu
محدد
نوفمبر 26, 2023 في 16:04
محدد
نوفمبر 26, 2023 في 16:04
The infrastructure behind Tier2 Ceph shares and VMs is unstable.
This also affects VDI/OOD which relies on virtual machines.

/net/fs-[labname] shares, new OOD/VDI sessions, and VMs are affected and may will be inaccessible until this is resolved.

Thanks for your patience.

تم الحل
نوفمبر 22, 2023 في 20:23
تم الحل
نوفمبر 22, 2023 في 20:23
The Ceph instability has been resolved. Caeph Tier2 shares, VDI, and VMs should be back to their normal state.

If your VM, /net/fs-[labname] share, or VDI session is still impacted, please contact rchelp@rc.fas.harvard.edu
محدد
نوفمبر 22, 2023 في 17:49
محدد
نوفمبر 22, 2023 في 17:49
The infrastructure behind Tier2 Ceph shares and VMs is unstable.
This also affects VDI/OOD which relies on virtual machines.

/net/fs-[labname] shares, new OOD/VDI sessions, and VMs are affected and may will be inaccessible until this is resolved.

Thanks for your patience.

تم الحل
نوفمبر 22, 2023 في 21:20
تم الحل
نوفمبر 22, 2023 في 21:20
Cooling and power have been restored to the affected racks. The compute nodes have been resumed in Slurm and are now accepting jobs again.

This incident has been resolved.
محدد
نوفمبر 22, 2023 في 20:01
محدد
نوفمبر 22, 2023 في 20:01
We have identified the partitions that are impacted due to loss of cooling in holy7c[02-12] compute nodes. Some of these partitions are fully down, others are partially down.

blackhole
blackholepriority davies desai hucecascade
hucecascadepriority
huttenhower
janson
jansoncascade joonholee lukin seascompute
shared
tambe
test
vishwanath
whipple

Please submit to other partitions in order to run jobs.

The spart command will show you all partitions you have access to, and our Running Jobs page provides a list of publicly available partitions for all cluster users. Please see our docs page for other helpful Slurm commands.

The Holyoke MGHPCC data center is working to restore cooling, and FASRC staff are onsite to assist. No ETA.
تحقيق
نوفمبر 22, 2023 في 17:00
تحقيق
نوفمبر 22, 2023 في 17:00
Compute nodes in holy7c[02-12] have experienced a power loss and are currently down. GPUs are not impacted at this time.

The 'shared' partition is significantly impacted. Other public partitions and lab-owned partitions may be down or running at reduced capacity. Jobs are still being accepted/running, but may need to wait longer in the queue due to fewer resources being available.

We are in contact with the Holyoke MGHPCC data center to investigate further. Updates to come. No ETA at this time.

تم الحل
نوفمبر 22, 2023 في 14:33
تم الحل
نوفمبر 22, 2023 في 14:33
The Ceph instability has been resolved. Caeph Tier2 shares, VDI, and VMs should be back to their normal state.

If your VM, /net/fs-[labname] share, or VDI session is still impacted, please contact rchelp@rc.fas.harvard.edu
محدد
نوفمبر 22, 2023 في 07:37
محدد
نوفمبر 22, 2023 في 07:37
The infrastructure behind Tier2 Ceph shares and VMs is unstable.
This also affects VDI/OOD which relies on virtual machines.

/net/fs-[labname] shares, new OOD/VDI sessions, and VMs are affected and may will be inaccessible until this is resolved.

Thanks for your patience.

تم الحل
نوفمبر 21, 2023 في 03:31
تم الحل
نوفمبر 21, 2023 في 03:31
The Ceph instability has been resolved. Caeph Tier2 shares, VDI, and VMs should be back to their normal state.

If your VM, /net/fs-[labname] share, or VDI session is still impacted, please contact rchelp@rc.fas.harvard.edu
محدد
نوفمبر 20, 2023 في 21:39
محدد
نوفمبر 20, 2023 في 21:39
The infrastructure behind Tier2 Ceph shares and VMs is unstable.
This also affects VDI/OOD which relies on virtual machines.

/net/fs-[labname] shares, new OOD/VDI sessions, and VMs are affected and may will be inaccessible until this is resolved.

Thanks for your patience.

أكتوبر 2023

تم الحل
أكتوبر 27, 2023 في 00:00
تم الحل
أكتوبر 27, 2023 في 00:00
A VPN certificate for vpn.rc.fas.harvard.edu expired, which led to "Untrusted Server Blocked" messages when attempting to connect. A new certificate has already been added, and you should no longer be getting any errors when connecting to the VPN.

This incident has been resolved.

تم الحل
أكتوبر 16, 2023 في 14:16
تم الحل
أكتوبر 16, 2023 في 14:16
FASSE login and OOD have been returned to service.
المراقبة
أكتوبر 15, 2023 في 23:11
المراقبة
أكتوبر 15, 2023 في 23:11
Most resources are once again available. The Cannon (including Kempner), FASSE, and Academic cluster are open for jobs. Please note that FASSE login and OpenOnDemand (OOD) nodes are not yet available. ETA Monday morning.

Thanks for your patience through this unexpected event.
تحديث
أكتوبر 15, 2023 في 19:17
تحديث
أكتوبر 15, 2023 في 19:17
Power-up is progressing with, so far, only minor issues which we are addressing according to their impact on returning the cluster to service.

Expect some remaining effects on less-essential services into tomorrow.

Please note that login nodes will remain down until we return the cluster and scheduler to service.
تحديث
أكتوبر 15, 2023 في 15:02
تحديث
أكتوبر 15, 2023 في 15:02
MGHPCC has isolated the cause of the generator failure and will continue to look into the grid failure.

At this time they will begin re-energizing the facility. Once that is complete and we have confirmed the networking is stable we can begin powering up our resources.

Please bear with us as this is a long process given the number of systems we maintain and it must be done in stages. Watch this page for updates.
محدد
أكتوبر 14, 2023 في 22:47
محدد
أكتوبر 14, 2023 في 22:47
With an abundance of caution FASRC and other MGHPCC occupants will not attempt to rush to restoration but will wait until the facility has restored primary power and confirmed stable operation before attempting to resume normal operations.

As such, we expect to begin restoring FASRC services tomorrow (Sunday). Since all Holyoke services and resources are down, this is a lengthy process similar to the startup process after the annual power-down.

Updates will be posted here. Please consider subscribing to our status page (see 'Get Updates' up top).
تحقيق
أكتوبر 14, 2023 في 21:17
تحقيق
أكتوبر 14, 2023 في 21:17
There has been a major power event at MGHPCC, our Holyoke data center.
We are awaiting further details

This likely affects all holyoke resources including the cluster and storage housed in holyoke.

More details as we learn them.

تم الحل
أكتوبر 12, 2023 في 16:26
تم الحل
أكتوبر 12, 2023 في 16:26
The security patch has been applied, and all clusters are accepting jobs at this time.
تحقيق
أكتوبر 12, 2023 في 15:02
تحقيق
أكتوبر 12, 2023 في 15:02
SchedMD (the maintainers of Slurm) have discovered a critical security flaw in Slurm. Due to the nature and severity of the issue, we will be immediately applying this patch.

Cannon and FASSE schedulers will remain down for the duration of the patching. All running jobs will be paused, and new jobs will not be accepted until the scheduler is back up.

ETA is expected to be approximately one hour.

تم الحل
أكتوبر 03, 2023 في 18:43
تم الحل
أكتوبر 03, 2023 في 18:43
This incident has been resolved.
تحقيق
أكتوبر 03, 2023 في 18:26
تحقيق
أكتوبر 03, 2023 في 18:26
We are currently investigating this incident.

FASRC monthly maintenance Monday October 2nd, 2023 7am-11am

مكتمل
أكتوبر 02, 2023 في 15:00
مكتمل
أكتوبر 02, 2023 في 15:00
Maintenance has completed successfully
قيد التقدم
أكتوبر 02, 2023 في 11:00
قيد التقدم
أكتوبر 02, 2023 في 11:00
Maintenance is now in progress
مخطط
أكتوبر 02, 2023 في 11:00
مخطط
أكتوبر 02, 2023 في 11:00
FASRC monthly maintenance will take place Monday October 2nd, 2023 from 7am-11am

NOTICES

New training sessions are available. Topics include New User Training, Getting Started on FASRC with CLI, Getting Started on FASRC with OpenOnDemand, GPU Computing, Parallel Job Workflows, and Singularity. To see current and uture training sessions, see our calendar at: https://www.rc.fas.harvard.edu/upcoming-training/

MAINTENANCE TASKS

Cannon cluster will be paused during this maintenance?: Yes
FASSE cluster will be paused during this maintenance?: No

Cannon UFM updates
-- Audience: Cluster users
-- Impact: The cluster will be paused while this update takes place.

Login node and OOD/VDI reboots
-- Audience: Anyone logged into a login node or VDI/OOD node
-- Impact: Login and VDI/OOD nodes will rebooted during this maintenance window

Scratch cleanup ( https://docs.rc.fas.harvard.edu/kb/policy-scratch/ )
-- Audience: Cluster users
-- Impact: Files older than 90 days will be removed. Please note that retention cleanup can run at any time, not just during the maintenance window.

Thanks,
FAS Research Computing
Department and Service Catalog: https://www.rc.fas.harvard.edu/
Documentation: https://docs.rc.fas.harvard.edu/
Status Page: https://status.rc.fas.harvard.edu/

أكتوبر 2023 ألى ديسمبر 2023

FAS Research Computing - سجل التاريخ

سجل التاريخ

ديسمبر 2023

نوفمبر 2023

أكتوبر 2023