FAS Research Computing - سجل التاريخ

جميع الأنظمة جاهزة للعمل

Status page for the Harvard FAS Research Computing cluster and other resources.

Cluster Utilization (VPN and FASRC login required): Cannon | FASSE


Please scroll down to see details on any Incidents or maintenance notices.
Monthly maintenance occurs on the first Monday of the month (except holidays).

GETTING HELP
Documentation: https://docs.rc.fas.harvard.edu | Account Portal https://portal.rc.fas.harvard.edu
Email: rchelp@rc.fas.harvard.edu | Support Hours


The colors shown in the bars below were chosen to increase visibility for color-blind visitors.
For higher contrast, switch to light mode at the bottom of this page if the background is dark and colors are muted.

جاهز للعمل

SLURM Scheduler - Cannon - جاهز للعمل

Cannon Compute Cluster (Holyoke) - جاهز للعمل

Boston Compute Nodes - جاهز للعمل

GPU nodes (Holyoke) - جاهز للعمل

seas_compute - جاهز للعمل

جاهز للعمل

SLURM Scheduler - FASSE - جاهز للعمل

FASSE Compute Cluster (Holyoke) - جاهز للعمل

جاهز للعمل

Kempner Cluster CPU - جاهز للعمل

Kempner Cluster GPU - جاهز للعمل

جاهز للعمل

FASSE login nodes - جاهز للعمل

جاهز للعمل

Cannon Open OnDemand - جاهز للعمل

FASSE Open OnDemand - جاهز للعمل

جاهز للعمل

Netscratch (Global Scratch) - جاهز للعمل

Home Directory Storage - Boston - جاهز للعمل

Tape - (Tier 3) - جاهز للعمل

Holylabs - جاهز للعمل

Isilon Storage Holyoke (Tier 1) - جاهز للعمل

Holystore01 (Tier 0) - جاهز للعمل

HolyLFS04 (Tier 0) - جاهز للعمل

HolyLFS05 (Tier 0) - جاهز للعمل

HolyLFS06 (Tier 0) - جاهز للعمل

Holyoke Tier 2 NFS (new) - جاهز للعمل

Holyoke Specialty Storage - جاهز للعمل

holECS - جاهز للعمل

Isilon Storage Boston (Tier 1) - جاهز للعمل

BosLFS02 (Tier 0) - جاهز للعمل

Boston Tier 2 NFS (new) - جاهز للعمل

CEPH Storage Boston (Tier 2) - جاهز للعمل

Boston Specialty Storage - جاهز للعمل

bosECS - جاهز للعمل

Samba Cluster - جاهز للعمل

Globus Data Transfer - جاهز للعمل

سجل التاريخ

ديسمبر 2023

Ceph instability - Affects Boston VMs (Virtual Machines) and Tier2 Ceph shares
  • تم الحل
    تم الحل

    The Ceph instability has been resolved. Ceph Tier2 shares, VDI, and VMs should be back to their normal state.

    If your VM, /net/fs-[labname] share, or VDI session is still impacted, please contact rchelp@rc.fas.harvard.edu

  • محدد
    محدد

    The infrastructure behind Tier2 Ceph shares and VMs is unstable.
    This also affects VDI/OOD which relies on virtual machines.

    /net/fs-[labname] shares, new OOD/VDI sessions, and VMs are affected and may will be inaccessible until this is resolved.

    Thanks for your patience.

Ceph instability - Affects Boston VMs (Virtual Machines) and Tier2 Ceph shares
  • تم الحل
    تم الحل

    The Ceph instability has been resolved. Ceph Tier2 shares, VDI, and VMs should be back to their normal state.

    If your VM, /net/fs-[labname] share, or VDI session is still impacted, please contact rchelp@rc.fas.harvard.edu

  • محدد
    محدد

    The infrastructure behind Tier2 Ceph shares and VMs is unstable.
    This also affects VDI/OOD which relies on virtual machines.

    /net/fs-[labname] shares, new OOD/VDI sessions, and VMs are affected and may will be inaccessible until this is resolved.

    Thanks for your patience.

Ceph instability - Affects Boston VMs (Virtual Machines) and Tier2 Ceph shares
  • تم الحل
    تم الحل

    The Ceph instability has been resolved. Ceph Tier2 shares, VDI, and VMs should be back to their normal state.

    If your VM, /net/fs-[labname] share, or VDI session is still impacted, please contact rchelp@rc.fas.harvard.edu

  • محدد
    محدد

    The infrastructure behind Tier2 Ceph shares and VMs is unstable.
    This also affects VDI/OOD which relies on virtual machines.

    /net/fs-[labname] shares, new OOD/VDI sessions, and VMs are affected and may will be inaccessible until this is resolved.

    Thanks for your patience.

نوفمبر 2023

Ceph instability - Affects Boston VMs (Virtual Machines) and Tier2 Ceph shares
  • تم الحل
    تم الحل

    The Ceph instability has been resolved. Ceph Tier2 shares, VDI, and VMs should be back to their normal state.

    If your VM, /net/fs-[labname] share, or VDI session is still impacted, please contact rchelp@rc.fas.harvard.edu

  • محدد
    محدد

    The infrastructure behind Tier2 Ceph shares and VMs is unstable.
    This also affects VDI/OOD which relies on virtual machines.

    /net/fs-[labname] shares, new OOD/VDI sessions, and VMs are affected and may will be inaccessible until this is resolved.

    Thanks for your patience.

Ceph instability - Affects Boston VMs (Virtual Machines) and Tier2 Ceph shares
  • تم الحل
    تم الحل

    The Ceph instability has been resolved. Caeph Tier2 shares, VDI, and VMs should be back to their normal state.

    If your VM, /net/fs-[labname] share, or VDI session is still impacted, please contact rchelp@rc.fas.harvard.edu

  • محدد
    محدد

    The infrastructure behind Tier2 Ceph shares and VMs is unstable.
    This also affects VDI/OOD which relies on virtual machines.

    /net/fs-[labname] shares, new OOD/VDI sessions, and VMs are affected and may will be inaccessible until this is resolved.

    Thanks for your patience.

Partial Cannon outage
  • تم الحل
    تم الحل

    Cooling and power have been restored to the affected racks. The compute nodes have been resumed in Slurm and are now accepting jobs again.

    This incident has been resolved.

  • محدد
    محدد

    We have identified the partitions that are impacted due to loss of cooling in holy7c[02-12] compute nodes. Some of these partitions are fully down, others are partially down.

    blackhole
    blackholepriority davies desai hucecascade
    hucecascadepriority
    huttenhower
    janson
    jansoncascade joonholee lukin seascompute
    shared
    tambe
    test
    vishwanath
    whipple

    Please submit to other partitions in order to run jobs.

    The spart command will show you all partitions you have access to, and our Running Jobs page provides a list of publicly available partitions for all cluster users. Please see our docs page for other helpful Slurm commands.

    The Holyoke MGHPCC data center is working to restore cooling, and FASRC staff are onsite to assist. No ETA.

  • تحقيق
    تحقيق

    Compute nodes in holy7c[02-12] have experienced a power loss and are currently down. GPUs are not impacted at this time.

    The 'shared' partition is significantly impacted. Other public partitions and lab-owned partitions may be down or running at reduced capacity. Jobs are still being accepted/running, but may need to wait longer in the queue due to fewer resources being available.

    We are in contact with the Holyoke MGHPCC data center to investigate further. Updates to come. No ETA at this time.

Ceph instability - Affects Boston VMs (Virtual Machines) and Tier2 Ceph shares
  • تم الحل
    تم الحل

    The Ceph instability has been resolved. Caeph Tier2 shares, VDI, and VMs should be back to their normal state.

    If your VM, /net/fs-[labname] share, or VDI session is still impacted, please contact rchelp@rc.fas.harvard.edu

  • محدد
    محدد

    The infrastructure behind Tier2 Ceph shares and VMs is unstable.
    This also affects VDI/OOD which relies on virtual machines.

    /net/fs-[labname] shares, new OOD/VDI sessions, and VMs are affected and may will be inaccessible until this is resolved.

    Thanks for your patience.

Ceph instability - Affects Boston VMs (Virtual Machines) and Tier2 Ceph shares
  • تم الحل
    تم الحل

    The Ceph instability has been resolved. Caeph Tier2 shares, VDI, and VMs should be back to their normal state.

    If your VM, /net/fs-[labname] share, or VDI session is still impacted, please contact rchelp@rc.fas.harvard.edu

  • محدد
    محدد

    The infrastructure behind Tier2 Ceph shares and VMs is unstable.
    This also affects VDI/OOD which relies on virtual machines.

    /net/fs-[labname] shares, new OOD/VDI sessions, and VMs are affected and may will be inaccessible until this is resolved.

    Thanks for your patience.

أكتوبر 2023

Major power event at MGHPCC (Holyoke) data center
  • تم الحل
    تم الحل

    FASSE login and OOD have been returned to service.

  • المراقبة
    المراقبة

    Most resources are once again available. The Cannon (including Kempner), FASSE, and Academic cluster are open for jobs. Please note that FASSE login and OpenOnDemand (OOD) nodes are not yet available. ETA Monday morning.

    Thanks for your patience through this unexpected event.

  • تحديث
    تحديث

    Power-up is progressing with, so far, only minor issues which we are addressing according to their impact on returning the cluster to service.

    Expect some remaining effects on less-essential services into tomorrow.

    Please note that login nodes will remain down until we return the cluster and scheduler to service.

  • تحديث
    تحديث

    MGHPCC has isolated the cause of the generator failure and will continue to look into the grid failure.

    At this time they will begin re-energizing the facility. Once that is complete and we have confirmed the networking is stable we can begin powering up our resources.

    Please bear with us as this is a long process given the number of systems we maintain and it must be done in stages. Watch this page for updates.

  • محدد
    محدد

    With an abundance of caution FASRC and other MGHPCC occupants will not attempt to rush to restoration but will wait until the facility has restored primary power and confirmed stable operation before attempting to resume normal operations.

    As such, we expect to begin restoring FASRC services tomorrow (Sunday). Since all Holyoke services and resources are down, this is a lengthy process similar to the startup process after the annual power-down.

    Updates will be posted here. Please consider subscribing to our status page (see 'Get Updates' up top).

  • تحقيق
    تحقيق

    There has been a major power event at MGHPCC, our Holyoke data center.
    We are awaiting further details

    This likely affects all holyoke resources including the cluster and storage housed in holyoke.

    More details as we learn them.

FASRC monthly maintenance Monday October 2nd, 2023 7am-11am
  • مكتمل
    أكتوبر 02, 2023 في 15:00
    مكتمل
    أكتوبر 02, 2023 في 15:00

    Maintenance has completed successfully

  • قيد التقدم
    أكتوبر 02, 2023 في 11:00
    قيد التقدم
    أكتوبر 02, 2023 في 11:00

    Maintenance is now in progress

  • مخطط
    أكتوبر 02, 2023 في 11:00
    مخطط
    أكتوبر 02, 2023 في 11:00

    FASRC monthly maintenance will take place Monday October 2nd, 2023 from 7am-11am

    NOTICES

    New training sessions are available. Topics include New User Training, Getting Started on FASRC with CLI, Getting Started on FASRC with OpenOnDemand, GPU Computing,  Parallel Job Workflows, and Singularity. To see current and uture training sessions, see our calendar at: https://www.rc.fas.harvard.edu/upcoming-training/

    MAINTENANCE TASKS

    Cannon cluster will be paused during this maintenance?: Yes
    FASSE cluster will be paused during this maintenance?: No

    Cannon UFM updates
    -- Audience: Cluster users
    -- Impact: The cluster will be paused while this update takes place.

    Login node and OOD/VDI reboots
    -- Audience: Anyone logged into a login node or VDI/OOD node
    -- Impact: Login and VDI/OOD nodes will rebooted during this maintenance window  

    Scratch cleanup ( https://docs.rc.fas.harvard.edu/kb/policy-scratch/ )
    -- Audience: Cluster users
    -- Impact: Files older than 90 days will be removed. Please note that retention cleanup can run at any time, not just during the maintenance window.

    Thanks,
    FAS Research Computing
    Department and Service Catalog: https://www.rc.fas.harvard.edu/
    Documentation: https://docs.rc.fas.harvard.edu/
    Status Page: https://status.rc.fas.harvard.edu/

أكتوبر 2023 ألى ديسمبر 2023

التالي