FAS Research Computing - История на известията

Системата е в поддръжка

Status page for the Harvard FAS Research Computing cluster and other resources.

Cluster Utilization (VPN and FASRC login required): Cannon | FASSE


Please scroll down to see details on any Incidents or maintenance notices.
Monthly maintenance occurs on the first Monday of the month (except holidays).

GETTING HELP
Documentation: https://docs.rc.fas.harvard.edu | Account Portal https://portal.rc.fas.harvard.edu
Email: rchelp@rc.fas.harvard.edu | Support Hours


The colors shown in the bars below were chosen to increase visibility for color-blind visitors.
For higher contrast, switch to light mode at the bottom of this page if the background is dark and colors are muted.

В поддръжка

SLURM Scheduler - Cannon - В поддръжка

Cannon Compute Cluster (Holyoke) - В поддръжка

Boston Compute Nodes - В поддръжка

GPU nodes (Holyoke) - В поддръжка

seas_compute - В поддръжка

В поддръжка

SLURM Scheduler - FASSE - В поддръжка

FASSE Compute Cluster (Holyoke) - В поддръжка

В поддръжка

Kempner Cluster CPU - В поддръжка

Kempner Cluster GPU - В поддръжка

В поддръжка

FASSE login nodes - В поддръжка

В поддръжка

Cannon Open OnDemand - В поддръжка

FASSE Open OnDemand - В поддръжка

Работи

Netscratch (Global Scratch) - Работи

Home Directory Storage - Boston - Работи

Tape - (Tier 3) - Работи

Holylabs - Работи

Isilon Storage Holyoke (Tier 1) - Работи

Holystore01 (Tier 0) - Работи

HolyLFS04 (Tier 0) - Работи

HolyLFS05 (Tier 0) - Работи

HolyLFS06 (Tier 0) - Работи

Holyoke Tier 2 NFS (new) - Работи

Holyoke Specialty Storage - Работи

holECS - Работи

Isilon Storage Boston (Tier 1) - Работи

BosLFS02 (Tier 0) - Работи

Boston Tier 2 NFS (new) - Работи

CEPH Storage Boston (Tier 2) - Работи

Boston Specialty Storage - Работи

bosECS - Работи

Samba Cluster - Работи

Globus Data Transfer - Работи

История на известията

авг 2026

Rolling OS Upgrades August 24th - 27th 2026
Планиран за август 24, 2026 в 13:00 – 13:00UTC
  • Актуализация
    август 24, 2026 в 19:07UTC
    Актуализация
    август 24, 2026 в 19:07UTC

    Today's portion of the upgrades are completed. Tomorrow's list (Cannon part 1) can be found in this maintenance event.

    August 24th:

    FASSE
    boslogin
    private login nodes

  • В ход
    август 24, 2026 в 13:00UTC
    В ход
    август 24, 2026 в 13:00UTC
    Maintenance is now in progress
  • Планиран
    август 03, 2026 в 19:35UTC
    Планиран
    август 03, 2026 в 19:35UTC

    FASRC will be doing OS upgrades to the latest version of Rocky 8.10 from August 24-27th. These rolling upgrades will improve cluster security and install the latest Cuda 13.2 driver.

    Upgrades will happen in stages with each stage occurring between 8a-5p each day. During that period the nodes being upgraded will be unavailable. No jobs will be canceled but jobs will stay pending if all nodes in the partition are down.

    Note that during this week the cluster will be in a partially upgraded state, thus jobs that span multiple nodes may become unstable due to mismatching libraries. Users concerned about this should wait until August 28th to restart runs.

    Users should plan their work accordingly.

    The upgrade schedule with impacted partitions is as follows:

    August 24th:

    FASSE
    boslogin
    private login nodes

    August 25th: Cannon Part 1

    holylogin
    arguelles_delgado
    blackhole
    davies_gpu
    davies
    desai
    eddy
    holy-cow
    holy-smokes
    huce_bigmem
    huce_cascade
    huttenhower
    jacobsen2
    janson_bigmem
    janson_cascade
    janson
    ke
    lukin
    nguyen
    olveczky_gpu
    remoteviz
    seas_compute
    shared
    sompolinsky_gpu
    tambe
    vishwanath
    whipple
    xlin
    xlin_ice
    zhuang_gpu
    zhuang

    August 26th: Cannon Part 2
    arguelles_delgado
    conroy
    davies
    doshi-velez
    dsouza
    edwards
    geophysics
    giribet
    gpu_test
    hernquist
    huce_cascadeimasc
    janson
    kempner_h200
    kempner_rtx
    murphy
    ni_lab
    olveczky
    ortegahernandez
    pehlevan
    seas_compute
    shakhnovich
    shared
    unrestricted
    xlin
    yin
    zon

    August 27th: Cannon Part 3

    arguelles_delgado_gpu_a100
    arguelles_delgado_gpu_mixed
    arguelles_delgado_h100
    bigmem_intermediate
    bigmem
    blackhole_gpu
    dvorkin
    eddy
    enos
    gershman
    gpu
    gpu_h200
    hejazi
    hernquist_ice
    hoekstra
    hsph_gpu
    hsph
    huce_ice
    iaifi_gpu
    intermediate
    itc_cluster
    itc_gpu
    janson_sapphire
    joonholee
    jshapiro
    kempner_h100
    kempner_h200
    kempner
    kempner_interactive
    kovac
    kozinsky_gpu
    kozinsky
    murphy_ice
    mweber_compute
    mweber_gpu
    olveczky_sapphire
    ortegahernandez_ice
    rivas
    sapphire
    seas_compute
    seas_gpu
    seas_gpu_perf
    siag_gpu
    siag_combo
    siag
    sur
    test
    yao_alphatns
    yao_gpu
    yao
    zhuang

Emergency Slurm (scheduler) patch at 2pm ETA 1-2hours. Jobs will be paused
  • Актуализация
    UTC
    Актуализация

    A number of jobs appear to have terminated during the upgrade for some reason. Those jobs will re-queue or, if not, will need to be re-submitted,
    We apologize for the inconvenience.

  • Решен
    UTC
    Решен

    The update has been applied and the scheduler is returned to service. Jobs are unpaused.

  • Разследва се
    UTC
    Разследва се

    We have been manually dealing with a Slurm memory leak bug behind the scenes and now have an official fix in the form of a Slurm upgrade (26.05.3). It is necessary to get this update in place as soon as possible.

    At 2pm today we will pause the cluster. Jobs will be paused, not terminated, and the scheduler will be unavailable for querying or submitting new jobs until the update is completed. Once complete, jobs will resume.

    Technical information about the bug:

    https://support.schedmd.com/show_bug.cgi?id=25685

h-nfs15-p inaccessible
  • Решен
    UTC
    Решен

    XFS journal recovery has completed, and kou_lab is remounted. The lab should verify the filesystem and any data that may have been written after 11:00 AM on 8/11/26 (Tuesday).

    In simple terms, too many open files caused the server to become unresponsive, which triggered an automatic reset. After the reboot, XFS required an extended recovery period before the filesystem could be mounted again.

    This does not mean the lab caused the issue. The file descriptor limit is a system-wide kernel resource shared by all filesystems and workloads on the server, so any share or process running a large job could have contributed. kou_lab was essentially collateral damage from machine-wide resource exhaustion, a “noisy neighbor” effect on a shared system.

  • Идентифициран
    UTC
    Идентифициран

    Most shares on h-nfs15-p are back up and accessible at this time.

    kou_lab storage is still down.

    We are continuing to work on a fix for this incident. Updates to come.

  • Разследва се
    UTC
    Разследва се

    Storage shares on h-nfs15-p may be inaccessible at this time. Affected labs include:

    • bellono_lab

    • debivort_lab

    • shakhnovich_lab

    • doyle_lab

    • holbrook_lab

    • kou_lab

    • koutrakis_lab

    • brennan_lab

    We are currently investigating this incident. We will continue to update as we know more

Openauth/Two-Factor issues for new users
  • Решен
    UTC
    Решен

    This issue is resolved and normal authentification for new accounts or accounts whose tokens have been reset/revoked can now log in.

    In the event that you find your token is not working still, we recommend you revoke your token and get a new one. See the last and first sections here, repectively: https://docs.rc.fas.harvard.edu/kb/openauth/

  • Актуализация
    UTC
    Актуализация

    If you have an existing 2FA token, please do not reset it at this time.

    We are continuing to work on a fix for this incident.

  • Идентифициран
    UTC
    Идентифициран

    We have identified an issue which keeps new accounts from using their two-factor/openauth token for authentication.

    This would also affect anyone resetting their token.

    If you have a new account and are unable to authenticate to the cluster, FASRC VPN, or other FASRC services, this is why.
    Existing accounts are not affected.

    We are working to resolve this as quickly as possible, but no ETA at this time. We will update this status as things change.


    Thanks for your understanding and patience.

юли 2026

FASRC monthly maintenance Monday July 6th, 2026 9am-1pm
  • Завършен
    юли 06, 2026 в 17:00UTC
    Завършен
    юли 06, 2026 в 17:00UTC
    Maintenance has completed successfully
  • В ход
    юли 06, 2026 в 13:00UTC
    В ход
    юли 06, 2026 в 13:00UTC
    Maintenance is now in progress
  • Планиран
    юни 26, 2026 в 13:45UTC
    Планиран
    юни 26, 2026 в 13:45UTC

    FASRC monthly maintenance will take place on July 6th 2026. Our maintenance tasks should be completed between 9am-1pm.

    Cannon cluster will be paused during this maintenance?: NO
    FASSE cluster will be paused during this maintenance?: NO

    NOTICES:

    • Friday July 3rd is a university holiday (independence Day observed)

    • Training: Upcoming training from FASRC and other sources can be found on our Training Calendar. at https://www.rc.fas.harvard.edu/upcoming-training/

    • Status Page: You can subscribe to our status to receive notifications of maintenance, incidents, and their resolution at https://status.rc.fas.harvard.edu/ (click Get Updates for options).

    • We'd love to hear success stories about your or your lab's use of FASRC. Submit your story here.

    MAINTENANCE TASKS

    • Domain controller replacement

      • Audience: Internal

      • Impact: None. End users should not see any impact.

    • Reboot drained nodes in error state

      • Audience: Cluster nodes with errors.

      • Impact: These nodes will have been drained already in preparation. No impact on jobs on the day and the affected nodes will return to service in their respective partitions after the maintenance period.

    • OOD/Open OnDemand reboots

      • Audience: All OOD users, reboot of the head nodes.

      • Impact: Running sessions will not be affected.

    • Login node reboots

      • Audience; All login node users.

      • Impact: Login nodes will reboot during the maintenance window.

    • Netscratch 90-day retention cleanup

      • Audience; All netscratch users

      • Impact: Files older than 90 days will be removed per our scratch policy. Please note that this cleanup can happen at any time, not just during maintenance.

    Thank you,
    FAS Research Computing
    https://docs.rc.fas.harvard.edu/
    https://www.rc.fas.harvard.edu/

юни 2026

2026 MGHPCC power downtime June 15-18, 2026
  • Завършен
    юни 18, 2026 в 21:15UTC
    Завършен
    юни 18, 2026 в 21:15UTC

    The yearly power downtime at our Holyoke data center, MGHPCC, has completed.

    The clusters and storage are back online and login nodes and OOD nodes are now available.

    If you have an issue/need help, please send a ticket to rchelp@rc.fas.harvard.edu with details.

    IMPORTANT NOTE: Tomorrow, June 19th is a university holiday. FASRC staff will return Monday to address any lingering issues and any new tickets.

  • Актуализация
    юни 18, 2026 в 20:45UTC
    Актуализация
    юни 18, 2026 в 20:45UTC

    Power-up is nearly complete, but a delay earlier in the day has us slightly behind.

    New ETA is 6PM.

  • Актуализация
    юни 18, 2026 в 12:16UTC
    Актуализация
    юни 18, 2026 в 12:16UTC

    MGHPCC has completed their maintenance and restored power to the facility.

    FASRC will now begin the power-up process. Please be aware that this takes several hours.

    We will update this status once complete.

    NOTE: A reminder that tomorrow (Friday) is a university holiday.

  • В ход
    юни 15, 2026 в 13:00UTC
    В ход
    юни 15, 2026 в 13:00UTC
    Maintenance is now in progress
  • Планиран
    юни 15, 2026 в 13:00UTC
    Планиран
    юни 15, 2026 в 13:00UTC

    The yearly power downtime at our Holyoke data center, MGHPCC, has been scheduled by the facility. This year's power downtime will take place on Tuesday June  15th - 18th, 2025.  There will be no June monthly maintenance as a result.

    Since the facility will be powered down for two days this year, we will not be performing the usual maintenance tasks. 
    That said, networking and other key infrastructure will be doing maintenance.

    IMPORTANT NOTE: FASRC storage at both Holyoke and Boston will be affected and should not be expected to be available throughout the downtime. Please plan ahead accordingly.

    • Monday June 15th -  Power-down begins at 9AM

    • Tuesday June 16th - Power out at MGHPCC

    • Wednesday June 17th - Power out at MGHPCC

    • Thursday June 18th - Expected return to full service by 5PM

    • Friday June 19th - Please note that June 19th is a university holiday

     

    Monday June 15th -  Power-down begins at 9AM
Tuesday June 16th - Power out at MGHPCC
Wednesday June 17th - Power out at MGHPCC
Thursday June 18th - Expected return to full service by 5PM

    For more detailed information and follow-up, please see:
    https://www.rc.fas.harvard.edu/mghpcc-yearly-shutdown or this Status Page

юни 2026 до авг 2026

Следващ