FAS Research Computing - Historia powiadomień

Wszystkie systemy sprawne

Status page for the Harvard FAS Research Computing cluster and other resources.

Cluster Utilization (VPN and FASRC login required): Cannon | FASSE


Please scroll down to see details on any Incidents or maintenance notices.
Monthly maintenance occurs on the first Monday of the month (except holidays).

GETTING HELP
Documentation: https://docs.rc.fas.harvard.edu | Account Portal https://portal.rc.fas.harvard.edu
Email: rchelp@rc.fas.harvard.edu | Support Hours


The colors shown in the bars below were chosen to increase visibility for color-blind visitors.
For higher contrast, switch to light mode at the bottom of this page if the background is dark and colors are muted.

Poprawne działanie

SLURM Scheduler - Cannon - Poprawne działanie

Cannon Compute Cluster (Holyoke) - Poprawne działanie

Boston Compute Nodes - Poprawne działanie

GPU nodes (Holyoke) - Poprawne działanie

seas_compute - Poprawne działanie

Poprawne działanie

SLURM Scheduler - FASSE - Poprawne działanie

FASSE Compute Cluster (Holyoke) - Poprawne działanie

Poprawne działanie

Kempner Cluster CPU - Poprawne działanie

Kempner Cluster GPU - Poprawne działanie

Poprawne działanie

FASSE login nodes - Poprawne działanie

Poprawne działanie

Cannon Open OnDemand - Poprawne działanie

FASSE Open OnDemand - Poprawne działanie

Poprawne działanie

Netscratch (Global Scratch) - Poprawne działanie

Home Directory Storage - Boston - Poprawne działanie

Tape - (Tier 3) - Poprawne działanie

Holylabs - Poprawne działanie

Isilon Storage Holyoke (Tier 1) - Poprawne działanie

Holystore01 (Tier 0) - Poprawne działanie

HolyLFS04 (Tier 0) - Poprawne działanie

HolyLFS05 (Tier 0) - Poprawne działanie

HolyLFS06 (Tier 0) - Poprawne działanie

Holyoke Tier 2 NFS (new) - Poprawne działanie

Holyoke Specialty Storage - Poprawne działanie

holECS - Poprawne działanie

Isilon Storage Boston (Tier 1) - Poprawne działanie

BosLFS02 (Tier 0) - Poprawne działanie

Boston Tier 2 NFS (new) - Poprawne działanie

CEPH Storage Boston (Tier 2) - Poprawne działanie

Boston Specialty Storage - Poprawne działanie

bosECS - Poprawne działanie

Samba Cluster - Poprawne działanie

Globus Data Transfer - Poprawne działanie

Historia powiadomień

sie 2024

Starfish upgrade
  • Zakończono
    sierpnia 27, 2024 o 12:00UTC
    Zakończono
    sierpnia 27, 2024 o 12:00UTC

    Starfish is back up

  • Aktualizacja
    sierpnia 26, 2024 o 14:35UTC
    Aktualizacja
    sierpnia 26, 2024 o 14:35UTC

    Starfish maintenance is still ongoing, no ETA at this time.

  • W trakcie
    sierpnia 24, 2024 o 00:00UTC
    W trakcie
    sierpnia 24, 2024 o 00:00UTC
    Maintenance is now in progress
  • Planowane
    sierpnia 24, 2024 o 00:00UTC
    Planowane
    sierpnia 24, 2024 o 00:00UTC

    The Starfish Zones Dashboard will be undergoing a few upgrades and maintenance this weekend from Friday, August 23rd at 8AM until Monday, August 26th at 8AM. The dashboard will not be accessible during this time. Further details will be provided, if needed. Please email rchelp@rc.fas.harvard.edu if you have any questions or concerns.

lip 2024

Authentication issues - Related to global Crowdstrike incident
  • Rozwiązany
    UTC
    Rozwiązany

    All Crowdstrike-related resources are back up and operational.

  • Aktualizacja
    UTC
    Aktualizacja
    For FASRC resources affected by the Crowdstrike issue, most are back in full services. A few remaining issues involving the following may not be resolved until Monda: - waywiser2 - proteomics2 - tmsdb3 - lic3
  • Aktualizacja
    UTC
    Aktualizacja

    Please see HUIT Status (harvard.edu) for additional information on the global issue caused by Crowdstrike security which Harvard relies on. This is an ongoing issue university-wide.

    The systems that continue to be affected at FASRC are minimal, but some Windows-based systems managed by or connected to FASRC may still be affected.

  • Monitorowanie
    UTC
    Monitorowanie

    Authentication is back up and running. Windows machines are still in a bad state and will need remedial work to get them back in service.

  • Zidentyfikowany
    UTC
    Zidentyfikowany

    Authentication is back up and running. Windows machines are still in a bad state and will need remedial work to get them back in service.

  • Analiza
    UTC
    Analiza

    Authentication is back up and running. Windows machines are still in a bad state and will need remedial work to get them back in service.

cze 2024

FASRC websites unavailable
  • Rozwiązany
    UTC
    Rozwiązany

    This incident has been resolved. Both sites are working normally.

  • Analiza
    UTC
    Analiza

    https://www.rc.fas.harvard.edu/ and https://docs.rc.fas.harvard.edu/ are offline.

    We are currently investigating this issue.

MGHPCC Pod 8A Power Upgrade June 24 will idle some Cannon nodes
  • Zakończono
    czerwca 25, 2024 o 04:00UTC
    Zakończono
    czerwca 25, 2024 o 04:00UTC
    Maintenance has completed successfully
  • W trakcie
    czerwca 24, 2024 o 16:01UTC
    W trakcie
    czerwca 24, 2024 o 16:01UTC
    Maintenance is now in progress
  • Planowane
    czerwca 24, 2024 o 04:01UTC
    Planowane
    czerwca 24, 2024 o 04:01UTC

    MGHPCC will be performing power upgrades on Pod 8A in order to increase density and allow more nodes to be added in that Pod's rows.  Similar to the May 13th work, this means that we will be idling half the nodes in 8A on two dates: June 17 and June 24th.

    These are all day events, meaning that the nodes in question will not be available for the 24 hours of that day.  This is being accomplished via reservations. So no jobs will be canceled but nodes will be drained and users may notice that their jobs may pend longer than normal as the scheduler idles these nodes.

    Where possible, please use or include other partitions in your job scripts and plan accordingly for any new or long-running jobs during that period: https://docs.rc.fas.harvard.edu/kb/running-jobs/#Slurm_partitions

    This affects the Cannon cluster. FASSE is not affected.

    Impacted partitions are:

    arguelles_delgado_gpu

    bigmem_intermediate

    bigmem

    blackhole_gpu

    eddy

    enos

    gershman gpu

    hejazi hernquist_ice

    hoekstra hsph

    huce_ice

    iaifi_gpu

    iaifi_gpu_priority

    iaifi_gpu_requeue

    intermediate

    itc_gpu

    itc_gpu_requeue

    joonholee

    jshapiro

    jshapiro_priority

    jshapiro_sapphire

    kempner

    kempner_dev

    kempner_h100

    kempner_requeue

    kempner_reservation

    kovac

    kozinsky

    kozinsky_gpu

    kozinsky_priority

    kozinsky_requeue

    murphy_ice

    ortegahernandez_ice

    sapphire

    seas_compute

    seas_gpu siag

    siag_combo

    siag_gpu

    sur test

    yao

    yao_priority

    zhuang

cze 2024 do sie 2024

Następny