FAS Research Computing - Istorija obaveštenja

Svi sistemi funkcionišu

Status page for the Harvard FAS Research Computing cluster and other resources.

Cluster Utilization (VPN and FASRC login required): Cannon | FASSE


Please scroll down to see details on any Incidents or maintenance notices.
Monthly maintenance occurs on the first Monday of the month (except holidays).

GETTING HELP
Documentation: https://docs.rc.fas.harvard.edu | Account Portal https://portal.rc.fas.harvard.edu
Email: rchelp@rc.fas.harvard.edu | Support Hours


The colors shown in the bars below were chosen to increase visibility for color-blind visitors.
For higher contrast, switch to light mode at the bottom of this page if the background is dark and colors are muted.

Funkcioniše

SLURM Scheduler - Cannon - Funkcioniše

Cannon Compute Cluster (Holyoke) - Funkcioniše

Boston Compute Nodes - Funkcioniše

GPU nodes (Holyoke) - Funkcioniše

seas_compute - Funkcioniše

Funkcioniše

SLURM Scheduler - FASSE - Funkcioniše

FASSE Compute Cluster (Holyoke) - Funkcioniše

Funkcioniše

Kempner Cluster CPU - Funkcioniše

Kempner Cluster GPU - Funkcioniše

Funkcioniše

FASSE login nodes - Funkcioniše

Funkcioniše

Cannon Open OnDemand - Funkcioniše

FASSE Open OnDemand - Funkcioniše

Funkcioniše

Netscratch (Global Scratch) - Funkcioniše

Home Directory Storage - Boston - Funkcioniše

Tape - (Tier 3) - Funkcioniše

Holylabs - Funkcioniše

Isilon Storage Holyoke (Tier 1) - Funkcioniše

Holystore01 (Tier 0) - Funkcioniše

HolyLFS04 (Tier 0) - Funkcioniše

HolyLFS05 (Tier 0) - Funkcioniše

HolyLFS06 (Tier 0) - Funkcioniše

Holyoke Tier 2 NFS (new) - Funkcioniše

Holyoke Specialty Storage - Funkcioniše

holECS - Funkcioniše

Isilon Storage Boston (Tier 1) - Funkcioniše

BosLFS02 (Tier 0) - Funkcioniše

Boston Tier 2 NFS (new) - Funkcioniše

CEPH Storage Boston (Tier 2) - Funkcioniše

Boston Specialty Storage - Funkcioniše

bosECS - Funkcioniše

Samba Cluster - Funkcioniše

Globus Data Transfer - Funkcioniše

Istorija obaveštenja

Aug 2025

SMB access to shares on the FASRC samba cluster)
  • Rešeno
    UTC
    Rešeno

    SMB access has been restored. Please disconnect and retry if you have a failed mapped drive. If you still cannot connect to a share, please contact rchelp@rc.fas.harvard.edu and let us know your username and exactly which share you are attempting to map.

  • Identifikovano
    UTC
    Identifikovano

    We are continuing to work on a fix for this incident. No ETA.

  • Istražuje se
    UTC
    Istražuje se

    Drive mapping to some shares may fail if those shares use the Samba Cluster. This includes but is not limited to share paths that begin with \\smbip.

    Known affected shares:

    anderson_lab

    arlotta_lab

    bellono_lab

    bertoldi_lab c

    apellini_lab

    dasch14

    dasch15

    dasch16

    denic_lab

    dobbie_lab

    engert_lab

    ferreira_lab

    fortune_lab

    friedman_lab

    girguis_lab

    grad_lab

    hausmann_lab

    hays_lab

    hbs_liran

    hbs_rcs huh

    illumina

    jessicacohen_lab

    lichtman_boslfs02

    mallet_lab

    mason_lab

    mckinley_lab

    mcz

    mitrano_lab

    moorcroftfs5

    murraylab

    nmr_large

    nmr_small

    novitsky_lab

    pooling

    qbrc_center

    ramachandran_lab

    schnapp_lab

    schrag_lab

    srivastava_lab

    whited_lab

    yau2_lab

Jul 2025

FASRC Monthly maintenance July 7, 2025 9AM-1PM
  • Završeno
    July 07, 2025 u 5:00 PMUTC
    Završeno
    July 07, 2025 u 5:00 PMUTC
    Maintenance has completed successfully
  • U toku
    July 07, 2025 u 1:00 PMUTC
    U toku
    July 07, 2025 u 1:00 PMUTC
    Maintenance is now in progress
  • Planirano
    July 07, 2025 u 1:00 PMUTC
    Planirano
    July 07, 2025 u 1:00 PMUTC

    FASRC monthly maintenance will take place Monday July 7th, 2025 from 9am-1pm

    NOTICES

    • ​New Quota tool available (/usr/local/sbin/quota) - Works on all filesystem types (home directory, lustre, isilon, netscratch, etc.)
      Type quota -h to see the full instructions for usage o visit the usage doc.

    • Training: Upcoming training from FASRC and other sources can be found on our Training Calendar. at https://www.rc.fas.harvard.edu/upcoming-training/

    • Status Page: You can subscribe to our status to receive notifications of maintenance, incidents, and their resolution at https://status.rc.fas.harvard.edu/ (click Get Updates for options).

    • Upcoming holidays:​ Juneteenth - ​T​hur. June 19​ / Independence Day - Fri​. July 4

    MAINTENANCE TASKS
    Cannon cluster will be paused during this maintenance?: YES
    FASSE cluster will be paused during this maintenance?: YES

    • Slurm Upgrade to 24.11.5

      • Audience: All cluster users

      • Impact: Jobs and the scheduler will be paused during this upgrade

    • Login node ​OS ​upgrades

      • Audience: Anyone logged into a FASRC Cannon or FASSE login node

      • Impact: All login nodes will ​upgraded ​and unavailable during this maintenance window

    • ​Start of cluster OS upgrades - July 7 -10

      • Audience: All cluster users

      • Impact: Over 4 days, July 7 through 10, we will upgrade the OS on 25% of the cluster each day. During that time, total capacity will be reduced across the cluster by 1/4 each day. This will require draining each sub-set of nodes ahead of time. 

    • Netscratch cleanup ( https://docs.rc.fas.harvard.edu/kb/policy-scratch/ )

      • Audience: Cluster users

      • Impact: Files older than 90 days will be removed. Please note that retention cleanup can and does run at any time, not just during the maintenance window.

    Thank you,
    FAS Research Computing
    https://docs.rc.fas.harvard.edu/
    https://www.rc.fas.harvard.edu/

Rolling cluster OS upgrades July 7 - 10
  • Završeno
    July 11, 2025 u 4:02 PMUTC
    Završeno
    July 11, 2025 u 4:02 PMUTC

    All upgrades are complete. A small number of nodes need clean-up, but the cluster is back to normal operation with all nodes running Rocky 8.10. Thanks for your patience.

  • Obaveštenje
    July 07, 2025 u 1:00 PMUTC
    Obaveštenje
    July 07, 2025 u 1:00 PMUTC

    Cannon rolling upgrades are in progress. Not all nodes are available.

    https://www.rc.fas.harvard.edu/blog/2025-compute-os-upgrade/

  • U toku
    July 07, 2025 u 1:00 PMUTC
    U toku
    July 07, 2025 u 1:00 PMUTC

    UPDATE: 7/7/25 6M FASSE is operational.

    Please be aware that FASSE jobs cannot be launched at this time due to the upgrades.
    We will return all FASSE nodes to normal services as soon as possible.

    https://www.rc.fas.harvard.edu/blog/2025-compute-os-upgrade/

  • Planirano
    July 07, 2025 u 1:00 PMUTC
    Planirano
    July 07, 2025 u 1:00 PMUTC

    Cluster OS upgrades - July 7 -10

    • Audience: All cluster users

    • Impact: Over 4 days, July 7 through 10, we will upgrade the OS on 25% of the cluster each day.
      During that time, total capacity will be reduced across the cluster by 1/4 each day.
      This will require draining each sub-set of nodes ahead of time. 

    Work begins during the July 7th maintenance (login nogdes will be upgraded during the 7/7 maintenance window) and will continue through July 10th.

    Additional details and a breakdown of each phase: 2025 Compute OS Upgrade

Jun 2025

holylabs - New data missing
  • Rešeno
    UTC
    Rešeno

    Allowed one week for the message to propagate. Closing this incident.

  • Identifikovano
    UTC
    Identifikovano

    While attempting to correct the over-quota/extra data issue on holylabs, an error in the sync command caused the deletion of newly created files since the re-open of the cluster (6/5/25 9AM) for 54 lab directories. We see no evidence that any other lab directories were affected.


    Due to the large nature of the original cleanup and the error being discovered after the fact, regretfully these deleted files cannot be recovered.

    A list follows of affected /n/holylabs lab directories. If your lab is not on that list, then it is not identified as being affected but this error:
    acc_lab
    alvarez_lab
    avillar_lab
    barnett_lab
    bertoldi_lab
    bol_lab
    brenner_lab
    cgolden_lab
    charbonneau_lab
    charrison_lab
    chetty_lab
    cnelya_lab
    dam_lab
    doshi-velez_lab
    eisenstein_lab
    enos_lab
    eps_preceptors
    glassman_lab
    hanson_lab
    hekstra_lab
    holbrook_lab
    iaifi_lab
    idreos_lab
    iebecker_lab
    imai_lab
    jacobsen_lab
    jialiu_lab
    junweil_lab
    kaxiras_lab
    kdbrantley_lab
    kempner_dev
    king_lab
    kiyoul_lab
    konkle_lab
    koumoutsakos_lab
    kozinsky_lab
    kramer_lab
    maustern_lab
    nliu_lab
    pallais_lab
    park_lab
    pierce_lab
    protopapas_lab
    pslade_lab
    shro_lab
    sitanc_lab
    smousavih_lab
    sneel_lab
    snyder_lab
    sompolinsky_lab
    tamano_lab
    ylei_lab
    zickler_lab

  • Istražuje se
    UTC
    Istražuje se

    We are currently investigating an issue on holylabs where some labs have noticed newly created files are missing.

    We will update this incident with more info as soon as possible.

2025 MGHPCC power downtime June 2-4, 2025
  • Završeno
    June 05, 2025 u 1:00 PMUTC
    Završeno
    June 05, 2025 u 1:00 PMUTC
    Maintenance has completed successfully
  • U toku
    June 02, 2025 u 1:00 PMUTC
    U toku
    June 02, 2025 u 1:00 PMUTC
    Maintenance is now in progress
  • Planirano
    June 02, 2025 u 1:00 PMUTC
    Planirano
    June 02, 2025 u 1:00 PMUTC

    The yearly power downtime at our Holyoke data center, MGHPCC, has been scheduled. 
    This year's power downtime will take place on Tuesday June 3, 2025. 

    This will require FASRC to begin shutdown of our systems beginning at 9AM on Monday, June 2nd.
    We have worked to reduce the total outage time this year.
    We will begin power-up on Wednesday June 4th with an expected return to full service by 9AM Thursday June 5th.

    • Monday June 2nd -  Power-down begins at 9AM

    • Tuesday June 3rd - Power out at MGHPCC

    • Wednesday June 4th - Maintenance tasks and then power-up begins

    • Thursday June 5th - Expected return to full service by 9AM

    Maintenance:
    During this downtime, Holylabs (/n/holylabs) will move to new hardware.
    Starfish, Coldfront, and the Portal will be unavailable during the downtime.

    For more details including a graphical timeline, please see: https://www.rc.fas.harvard.edu/events/2025-mghpcc-power-downtime/

    Updates will be posted here on our status page: https://status.rc.fas.harvard.edu/
    Note that you can subscribe to receive updates as they happen. On the status page, click Get Updates.

    Notices and reminders will also be sent to all users via our mailing lists.

Jun 2025 do Aug 2025

Sledeći