Histórico de avisos

Operacional

jan 2024

Resolvido
janeiro 31, 2024 em 20:53UTC
Resolvido
janeiro 31, 2024 em 20:53UTC
fasselogin machines are back in service
Identificado
janeiro 31, 2024 em 19:26UTC
Identificado
janeiro 31, 2024 em 19:26UTC
We have identified the underlying issue and are working on fixing the FASSE login nodes. In the meantime, please use FASSE VDI. Updates to come.
Investigando
janeiro 31, 2024 em 16:00UTC
Investigando
janeiro 31, 2024 em 16:00UTC
fasselogin01/02 are currently down. We are investigating this incident.

Resolvido
janeiro 18, 2024 em 02:36UTC
Resolvido
janeiro 18, 2024 em 02:36UTC
The Ceph instability has been resolved. Ceph Tier2 shares, VDI, and VMs should be back to their normal state.

If your VM, /net/fs-[labname] share, or VDI session is still impacted, please contact rchelp@rc.fas.harvard.edu
Identificado
janeiro 18, 2024 em 00:58UTC
Identificado
janeiro 18, 2024 em 00:58UTC
The infrastructure behind Tier2 Ceph shares and VMs is unstable.
This also affects VDI/OOD which relies on virtual machines.

/net/fs-[labname] shares, new OOD/VDI sessions, and VMs are affected and may will be inaccessible until this is resolved.

Thanks for your patience.

Resolvido
janeiro 17, 2024 em 23:31UTC
Resolvido
janeiro 17, 2024 em 23:31UTC
The Ceph instability has been resolved. Ceph Tier2 shares, VDI, and VMs should be back to their normal state.

If your VM, /net/fs-[labname] share, or VDI session is still impacted, please contact rchelp@rc.fas.harvard.edu
Identificado
janeiro 17, 2024 em 22:44UTC
Identificado
janeiro 17, 2024 em 22:44UTC
The infrastructure behind Tier2 Ceph shares and VMs is unstable.
This also affects VDI/OOD which relies on virtual machines.

/net/fs-[labname] shares, new OOD/VDI sessions, and VMs are affected and may will be inaccessible until this is resolved.

Thanks for your patience.

Resolvido
janeiro 12, 2024 em 17:34UTC
Resolvido
janeiro 12, 2024 em 17:34UTC
Coldfront is accessible again. Thanks for your patience so that would could investigate further. We've added new logging to have a better view into this should it happen again.
Investigando
janeiro 12, 2024 em 15:22UTC
Investigando
janeiro 12, 2024 em 15:22UTC
Coldfront is currently up but inaccessible. Login attempts will result in a timeout/proxy error.
We would like time to gather forensics to sort out why this is happening before restarting the service.

If you need to access Coldfront, please try again later today.

Resolvido
janeiro 11, 2024 em 13:22UTC
Resolvido
janeiro 11, 2024 em 13:22UTC
The Ceph instability has been resolved. Ceph Tier2 shares, VDI, and VMs should be back to their normal state.

If your VM, /net/fs-[labname] share, or VDI session is still impacted, please contact rchelp@rc.fas.harvard.edu
Identificado
janeiro 11, 2024 em 07:57UTC
Identificado
janeiro 11, 2024 em 07:57UTC
The infrastructure behind Tier2 Ceph shares and VMs is unstable.
This also affects VDI/OOD which relies on virtual machines.

/net/fs-[labname] shares, new OOD/VDI sessions, and VMs are affected and may will be inaccessible until this is resolved.

Thanks for your patience.

dez 2023

Resolvido
dezembro 29, 2023 em 19:40UTC
Resolvido
dezembro 29, 2023 em 19:40UTC
The Ceph instability has been resolved. Ceph Tier2 shares, VDI, and VMs should be back to their normal state.

If your VM, /net/fs-[labname] share, or VDI session is still impacted, please contact rchelp@rc.fas.harvard.edu
Identificado
dezembro 27, 2023 em 20:38UTC
Identificado
dezembro 27, 2023 em 20:38UTC
The infrastructure behind Tier2 Ceph shares and VMs is unstable.
This also affects VDI/OOD which relies on virtual machines.

/net/fs-[labname] shares, new OOD/VDI sessions, and VMs are affected and may will be inaccessible until this is resolved.

Thanks for your patience.

Resolvido
dezembro 22, 2023 em 22:26UTC
Resolvido
dezembro 22, 2023 em 22:26UTC
The Ceph instability has been resolved. Ceph Tier2 shares, VDI, and VMs should be back to their normal state.

If your VM, /net/fs-[labname] share, or VDI session is still impacted, please contact rchelp@rc.fas.harvard.edu
Identificado
dezembro 22, 2023 em 16:05UTC
Identificado
dezembro 22, 2023 em 16:05UTC
The infrastructure behind Tier2 Ceph shares and VMs is unstable.
This also affects VDI/OOD which relies on virtual machines.

/net/fs-[labname] shares, new OOD/VDI sessions, and VMs are affected and may will be inaccessible until this is resolved.

Thanks for your patience.

Resolvido
dezembro 18, 2023 em 15:54UTC
Resolvido
dezembro 18, 2023 em 15:54UTC
Downed nodes are back online. The cluster is now available again.

Thank you for your patience.
Identificado
dezembro 18, 2023 em 14:50UTC
Identificado
dezembro 18, 2023 em 14:50UTC
Power was restored to the affected sections and we are bringing the down nodes back up.
Investigando
dezembro 18, 2023 em 12:00UTC
Investigando
dezembro 18, 2023 em 12:00UTC
There has been a loss of street power at MGHPCC, our Holyoke data center due to the windstorm. We are awaiting further details.

This likely affects most of the Cannon/Kempner compute cluster, and all of the FASSE compute cluster.

More details as we learn them.

Resolvido
dezembro 15, 2023 em 14:19UTC
Resolvido
dezembro 15, 2023 em 14:19UTC
The Ceph instability has been resolved. Ceph Tier2 shares, VDI, and VMs should be back to their normal state.

If your VM, /net/fs-[labname] share, or VDI session is still impacted, please contact rchelp@rc.fas.harvard.edu
Identificado
dezembro 15, 2023 em 13:27UTC
Identificado
dezembro 15, 2023 em 13:27UTC
The infrastructure behind Tier2 Ceph shares and VMs is unstable.
This also affects VDI/OOD which relies on virtual machines.

/net/fs-[labname] shares, new OOD/VDI sessions, and VMs are affected and may will be inaccessible until this is resolved.

Thanks for your patience.

Emergency scheduler update - Security Patch

Concluído
dezembro 14, 2023 em 17:00UTC
Concluído
dezembro 14, 2023 em 17:00UTC
Maintenance has completed successfully
Em curso
dezembro 14, 2023 em 16:00UTC
Em curso
dezembro 14, 2023 em 16:00UTC
Maintenance is now in progress
Ainda não iniciou
dezembro 14, 2023 em 16:00UTC
Ainda não iniciou
dezembro 14, 2023 em 16:00UTC
SchedMD has issued a patch for a critical security vulnerability. It is imperative that we apply this patch ASAP. This will require a pause of the cluster for roughly 30-60 minutes at around 11am, after which jobs and the scheduler will continue normal operation.

A side effect of this is that we will move up to version 23.02.7 with the patch which will fix the IntelMPI and PMI2 issues some users have experienced.

Thanks for your understanding and patience.
FAS Research Computing

nov 2023

Resolvido
novembro 26, 2023 em 23:20UTC
Resolvido
novembro 26, 2023 em 23:20UTC
The Ceph instability has been resolved. Ceph Tier2 shares, VDI, and VMs should be back to their normal state.

If your VM, /net/fs-[labname] share, or VDI session is still impacted, please contact rchelp@rc.fas.harvard.edu
Identificado
novembro 26, 2023 em 16:04UTC
Identificado
novembro 26, 2023 em 16:04UTC
The infrastructure behind Tier2 Ceph shares and VMs is unstable.
This also affects VDI/OOD which relies on virtual machines.

/net/fs-[labname] shares, new OOD/VDI sessions, and VMs are affected and may will be inaccessible until this is resolved.

Thanks for your patience.

Resolvido
novembro 22, 2023 em 20:23UTC
Resolvido
novembro 22, 2023 em 20:23UTC
The Ceph instability has been resolved. Caeph Tier2 shares, VDI, and VMs should be back to their normal state.

If your VM, /net/fs-[labname] share, or VDI session is still impacted, please contact rchelp@rc.fas.harvard.edu
Identificado
novembro 22, 2023 em 17:49UTC
Identificado
novembro 22, 2023 em 17:49UTC
The infrastructure behind Tier2 Ceph shares and VMs is unstable.
This also affects VDI/OOD which relies on virtual machines.

/net/fs-[labname] shares, new OOD/VDI sessions, and VMs are affected and may will be inaccessible until this is resolved.

Thanks for your patience.

Resolvido
novembro 22, 2023 em 21:20UTC
Resolvido
novembro 22, 2023 em 21:20UTC
Cooling and power have been restored to the affected racks. The compute nodes have been resumed in Slurm and are now accepting jobs again.

This incident has been resolved.
Identificado
novembro 22, 2023 em 20:01UTC
Identificado
novembro 22, 2023 em 20:01UTC
We have identified the partitions that are impacted due to loss of cooling in holy7c[02-12] compute nodes. Some of these partitions are fully down, others are partially down.

blackhole
blackholepriority davies desai hucecascade
hucecascadepriority
huttenhower
janson
jansoncascade joonholee lukin seascompute
shared
tambe
test
vishwanath
whipple

Please submit to other partitions in order to run jobs.

The spart command will show you all partitions you have access to, and our Running Jobs page provides a list of publicly available partitions for all cluster users. Please see our docs page for other helpful Slurm commands.

The Holyoke MGHPCC data center is working to restore cooling, and FASRC staff are onsite to assist. No ETA.
Investigando
novembro 22, 2023 em 17:00UTC
Investigando
novembro 22, 2023 em 17:00UTC
Compute nodes in holy7c[02-12] have experienced a power loss and are currently down. GPUs are not impacted at this time.

The 'shared' partition is significantly impacted. Other public partitions and lab-owned partitions may be down or running at reduced capacity. Jobs are still being accepted/running, but may need to wait longer in the queue due to fewer resources being available.

We are in contact with the Holyoke MGHPCC data center to investigate further. Updates to come. No ETA at this time.

Resolvido
novembro 22, 2023 em 14:33UTC
Resolvido
novembro 22, 2023 em 14:33UTC
The Ceph instability has been resolved. Caeph Tier2 shares, VDI, and VMs should be back to their normal state.

If your VM, /net/fs-[labname] share, or VDI session is still impacted, please contact rchelp@rc.fas.harvard.edu
Identificado
novembro 22, 2023 em 07:37UTC
Identificado
novembro 22, 2023 em 07:37UTC
The infrastructure behind Tier2 Ceph shares and VMs is unstable.
This also affects VDI/OOD which relies on virtual machines.

/net/fs-[labname] shares, new OOD/VDI sessions, and VMs are affected and may will be inaccessible until this is resolved.

Thanks for your patience.

Resolvido
novembro 21, 2023 em 03:31UTC
Resolvido
novembro 21, 2023 em 03:31UTC
The Ceph instability has been resolved. Caeph Tier2 shares, VDI, and VMs should be back to their normal state.

If your VM, /net/fs-[labname] share, or VDI session is still impacted, please contact rchelp@rc.fas.harvard.edu
Identificado
novembro 20, 2023 em 21:39UTC
Identificado
novembro 20, 2023 em 21:39UTC
The infrastructure behind Tier2 Ceph shares and VMs is unstable.
This also affects VDI/OOD which relies on virtual machines.

/net/fs-[labname] shares, new OOD/VDI sessions, and VMs are affected and may will be inaccessible until this is resolved.

Thanks for your patience.

nov 2023 para jan 2024

FAS Research Computing - Histórico de avisos

Histórico de avisos

jan 2024

dez 2023

nov 2023