FAS Research Computing - Notishistorik

Alla system fungerar

Status page for the Harvard FAS Research Computing cluster and other resources.

Cluster Utilization (VPN and FASRC login required): Cannon | FASSE


Please scroll down to see details on any Incidents or maintenance notices.
Monthly maintenance occurs on the first Monday of the month (except holidays).

GETTING HELP
Documentation: https://docs.rc.fas.harvard.edu | Account Portal https://portal.rc.fas.harvard.edu
Email: rchelp@rc.fas.harvard.edu | Support Hours


The colors shown in the bars below were chosen to increase visibility for color-blind visitors.
For higher contrast, switch to light mode at the bottom of this page if the background is dark and colors are muted.

I drift

SLURM Scheduler - Cannon - I drift

Cannon Compute Cluster (Holyoke) - I drift

Boston Compute Nodes - I drift

GPU nodes (Holyoke) - I drift

seas_compute - I drift

I drift

SLURM Scheduler - FASSE - I drift

FASSE Compute Cluster (Holyoke) - I drift

I drift

Kempner Cluster CPU - I drift

Kempner Cluster GPU - I drift

I drift

FASSE login nodes - I drift

I drift

Cannon Open OnDemand - I drift

FASSE Open OnDemand - I drift

I drift

Netscratch (Global Scratch) - I drift

Home Directory Storage - Boston - I drift

Tape - (Tier 3) - I drift

Holylabs - I drift

Isilon Storage Holyoke (Tier 1) - I drift

Holystore01 (Tier 0) - I drift

HolyLFS04 (Tier 0) - I drift

HolyLFS05 (Tier 0) - I drift

HolyLFS06 (Tier 0) - I drift

Holyoke Tier 2 NFS (new) - I drift

Holyoke Specialty Storage - I drift

holECS - I drift

Isilon Storage Boston (Tier 1) - I drift

BosLFS02 (Tier 0) - I drift

Boston Tier 2 NFS (new) - I drift

CEPH Storage Boston (Tier 2) - I drift

Boston Specialty Storage - I drift

bosECS - I drift

Samba Cluster - I drift

Globus Data Transfer - I drift

Notishistorik

aug. 2025

SMB access to shares on the FASRC samba cluster)
  • Löst
    UTC
    Löst

    SMB access has been restored. Please disconnect and retry if you have a failed mapped drive. If you still cannot connect to a share, please contact rchelp@rc.fas.harvard.edu and let us know your username and exactly which share you are attempting to map.

  • Identifierat
    UTC
    Identifierat

    We are continuing to work on a fix for this incident. No ETA.

  • Undersöker
    UTC
    Undersöker

    Drive mapping to some shares may fail if those shares use the Samba Cluster. This includes but is not limited to share paths that begin with \\smbip.

    Known affected shares:

    anderson_lab

    arlotta_lab

    bellono_lab

    bertoldi_lab c

    apellini_lab

    dasch14

    dasch15

    dasch16

    denic_lab

    dobbie_lab

    engert_lab

    ferreira_lab

    fortune_lab

    friedman_lab

    girguis_lab

    grad_lab

    hausmann_lab

    hays_lab

    hbs_liran

    hbs_rcs huh

    illumina

    jessicacohen_lab

    lichtman_boslfs02

    mallet_lab

    mason_lab

    mckinley_lab

    mcz

    mitrano_lab

    moorcroftfs5

    murraylab

    nmr_large

    nmr_small

    novitsky_lab

    pooling

    qbrc_center

    ramachandran_lab

    schnapp_lab

    schrag_lab

    srivastava_lab

    whited_lab

    yau2_lab

juli 2025

FASRC Monthly maintenance July 7, 2025 9AM-1PM
  • Slutfört
    juli 07, 2025 kl 17:00UTC
    Slutfört
    juli 07, 2025 kl 17:00UTC
    Maintenance has completed successfully
  • Pågår
    juli 07, 2025 kl 13:00UTC
    Pågår
    juli 07, 2025 kl 13:00UTC
    Maintenance is now in progress
  • Planerat
    juli 07, 2025 kl 13:00UTC
    Planerat
    juli 07, 2025 kl 13:00UTC

    FASRC monthly maintenance will take place Monday July 7th, 2025 from 9am-1pm

    NOTICES

    • ​New Quota tool available (/usr/local/sbin/quota) - Works on all filesystem types (home directory, lustre, isilon, netscratch, etc.)
      Type quota -h to see the full instructions for usage o visit the usage doc.

    • Training: Upcoming training from FASRC and other sources can be found on our Training Calendar. at https://www.rc.fas.harvard.edu/upcoming-training/

    • Status Page: You can subscribe to our status to receive notifications of maintenance, incidents, and their resolution at https://status.rc.fas.harvard.edu/ (click Get Updates for options).

    • Upcoming holidays:​ Juneteenth - ​T​hur. June 19​ / Independence Day - Fri​. July 4

    MAINTENANCE TASKS
    Cannon cluster will be paused during this maintenance?: YES
    FASSE cluster will be paused during this maintenance?: YES

    • Slurm Upgrade to 24.11.5

      • Audience: All cluster users

      • Impact: Jobs and the scheduler will be paused during this upgrade

    • Login node ​OS ​upgrades

      • Audience: Anyone logged into a FASRC Cannon or FASSE login node

      • Impact: All login nodes will ​upgraded ​and unavailable during this maintenance window

    • ​Start of cluster OS upgrades - July 7 -10

      • Audience: All cluster users

      • Impact: Over 4 days, July 7 through 10, we will upgrade the OS on 25% of the cluster each day. During that time, total capacity will be reduced across the cluster by 1/4 each day. This will require draining each sub-set of nodes ahead of time. 

    • Netscratch cleanup ( https://docs.rc.fas.harvard.edu/kb/policy-scratch/ )

      • Audience: Cluster users

      • Impact: Files older than 90 days will be removed. Please note that retention cleanup can and does run at any time, not just during the maintenance window.

    Thank you,
    FAS Research Computing
    https://docs.rc.fas.harvard.edu/
    https://www.rc.fas.harvard.edu/

Rolling cluster OS upgrades July 7 - 10
  • Slutfört
    juli 11, 2025 kl 16:02UTC
    Slutfört
    juli 11, 2025 kl 16:02UTC

    All upgrades are complete. A small number of nodes need clean-up, but the cluster is back to normal operation with all nodes running Rocky 8.10. Thanks for your patience.

  • Uppdatering
    juli 07, 2025 kl 13:00UTC
    Uppdatering
    juli 07, 2025 kl 13:00UTC

    Cannon rolling upgrades are in progress. Not all nodes are available.

    https://www.rc.fas.harvard.edu/blog/2025-compute-os-upgrade/

  • Pågår
    juli 07, 2025 kl 13:00UTC
    Pågår
    juli 07, 2025 kl 13:00UTC

    UPDATE: 7/7/25 6M FASSE is operational.

    Please be aware that FASSE jobs cannot be launched at this time due to the upgrades.
    We will return all FASSE nodes to normal services as soon as possible.

    https://www.rc.fas.harvard.edu/blog/2025-compute-os-upgrade/

  • Planerat
    juli 07, 2025 kl 13:00UTC
    Planerat
    juli 07, 2025 kl 13:00UTC

    Cluster OS upgrades - July 7 -10

    • Audience: All cluster users

    • Impact: Over 4 days, July 7 through 10, we will upgrade the OS on 25% of the cluster each day.
      During that time, total capacity will be reduced across the cluster by 1/4 each day.
      This will require draining each sub-set of nodes ahead of time. 

    Work begins during the July 7th maintenance (login nogdes will be upgraded during the 7/7 maintenance window) and will continue through July 10th.

    Additional details and a breakdown of each phase: 2025 Compute OS Upgrade

juni 2025

holylabs - New data missing
  • Löst
    UTC
    Löst

    Allowed one week for the message to propagate. Closing this incident.

  • Identifierat
    UTC
    Identifierat

    While attempting to correct the over-quota/extra data issue on holylabs, an error in the sync command caused the deletion of newly created files since the re-open of the cluster (6/5/25 9AM) for 54 lab directories. We see no evidence that any other lab directories were affected.


    Due to the large nature of the original cleanup and the error being discovered after the fact, regretfully these deleted files cannot be recovered.

    A list follows of affected /n/holylabs lab directories. If your lab is not on that list, then it is not identified as being affected but this error:
    acc_lab
    alvarez_lab
    avillar_lab
    barnett_lab
    bertoldi_lab
    bol_lab
    brenner_lab
    cgolden_lab
    charbonneau_lab
    charrison_lab
    chetty_lab
    cnelya_lab
    dam_lab
    doshi-velez_lab
    eisenstein_lab
    enos_lab
    eps_preceptors
    glassman_lab
    hanson_lab
    hekstra_lab
    holbrook_lab
    iaifi_lab
    idreos_lab
    iebecker_lab
    imai_lab
    jacobsen_lab
    jialiu_lab
    junweil_lab
    kaxiras_lab
    kdbrantley_lab
    kempner_dev
    king_lab
    kiyoul_lab
    konkle_lab
    koumoutsakos_lab
    kozinsky_lab
    kramer_lab
    maustern_lab
    nliu_lab
    pallais_lab
    park_lab
    pierce_lab
    protopapas_lab
    pslade_lab
    shro_lab
    sitanc_lab
    smousavih_lab
    sneel_lab
    snyder_lab
    sompolinsky_lab
    tamano_lab
    ylei_lab
    zickler_lab

  • Undersöker
    UTC
    Undersöker

    We are currently investigating an issue on holylabs where some labs have noticed newly created files are missing.

    We will update this incident with more info as soon as possible.

2025 MGHPCC power downtime June 2-4, 2025
  • Slutfört
    juni 05, 2025 kl 13:00UTC
    Slutfört
    juni 05, 2025 kl 13:00UTC
    Maintenance has completed successfully
  • Pågår
    juni 02, 2025 kl 13:00UTC
    Pågår
    juni 02, 2025 kl 13:00UTC
    Maintenance is now in progress
  • Planerat
    juni 02, 2025 kl 13:00UTC
    Planerat
    juni 02, 2025 kl 13:00UTC

    The yearly power downtime at our Holyoke data center, MGHPCC, has been scheduled. 
    This year's power downtime will take place on Tuesday June 3, 2025. 

    This will require FASRC to begin shutdown of our systems beginning at 9AM on Monday, June 2nd.
    We have worked to reduce the total outage time this year.
    We will begin power-up on Wednesday June 4th with an expected return to full service by 9AM Thursday June 5th.

    • Monday June 2nd -  Power-down begins at 9AM

    • Tuesday June 3rd - Power out at MGHPCC

    • Wednesday June 4th - Maintenance tasks and then power-up begins

    • Thursday June 5th - Expected return to full service by 9AM

    Maintenance:
    During this downtime, Holylabs (/n/holylabs) will move to new hardware.
    Starfish, Coldfront, and the Portal will be unavailable during the downtime.

    For more details including a graphical timeline, please see: https://www.rc.fas.harvard.edu/events/2025-mghpcc-power-downtime/

    Updates will be posted here on our status page: https://status.rc.fas.harvard.edu/
    Note that you can subscribe to receive updates as they happen. On the status page, click Get Updates.

    Notices and reminders will also be sent to all users via our mailing lists.

juni 2025 till aug. 2025

Nästa