FAS Research Computing - Histórico de avisos

HolyLFS06 (Tier 0) Experimentando Desempenho Degradado

Status page for the Harvard FAS Research Computing cluster and other resources.

Cluster Utilization (VPN and FASRC login required): Cannon | FASSE


Please scroll down to see details on any Incidents or maintenance notices.
Monthly maintenance occurs on the first Monday of the month (except holidays).

GETTING HELP
Documentation: https://docs.rc.fas.harvard.edu | Account Portal https://portal.rc.fas.harvard.edu
Email: rchelp@rc.fas.harvard.edu | Support Hours


The colors shown in the bars below were chosen to increase visibility for color-blind visitors.
For higher contrast, switch to light mode at the bottom of this page if the background is dark and colors are muted.

Operacional

SLURM Scheduler - Cannon - Operacional

Cannon Compute Cluster (Holyoke) - Operacional

Boston Compute Nodes - Operacional

GPU nodes (Holyoke) - Operacional

seas_compute - Operacional

Operacional

SLURM Scheduler - FASSE - Operacional

FASSE Compute Cluster (Holyoke) - Operacional

Operacional

Kempner Cluster CPU - Operacional

Kempner Cluster GPU - Operacional

Operacional

FASSE login nodes - Operacional

Operacional

Cannon Open OnDemand - Operacional

FASSE Open OnDemand - Operacional

Desempenho degradado

Netscratch (Global Scratch) - Operacional

Home Directory Storage - Boston - Operacional

Tape - (Tier 3) - Operacional

Holylabs - Operacional

Isilon Storage Holyoke (Tier 1) - Operacional

Holystore01 (Tier 0) - Operacional

HolyLFS04 (Tier 0) - Operacional

HolyLFS05 (Tier 0) - Operacional

HolyLFS06 (Tier 0) - Desempenho degradado

Holyoke Tier 2 NFS - Operacional

Holyoke Specialty Storage - Operacional

holECS - Operacional

Isilon Storage Boston (Tier 1) - Operacional

BosLFS02 (Tier 0) - Operacional

Boston Tier 2 NFS - Operacional

CEPH Storage Boston (Tier 2) - Operacional

Boston Specialty Storage - Operacional

bosECS - Operacional

Samba Cluster - Operacional

Globus Data Transfer - Operacional

Histórico de avisos

set 2026

Cooling Distribution Unit (CDU) work 9/22 - 10/2 See details for affected partitions and dates
Programado para setembro 22, 2026 em 11:00 – 11:00UTC
  • Ainda não iniciou
    setembro 09, 2026 em 19:32UTC
    Ainda não iniciou
    setembro 09, 2026 em 19:32UTC

    Over the past year we have noticed that water circulating through two of the Cooling Distribution Units (CDU) in Row 8a and their compute nodes has become significantly discoloured. This impurity causes cooling problems which as a result increases chances of node failure. To remedy FASRC has scheduled a full flush of these CDUs and their attached nodes. Unfortunately to do the full flush we have to fully power down all the nodes and drain the water, which is a process that takes several days to complete.

    To minimize disruption we have staggered this work over two weeks. The work on CDU5 will take place from 9/22 - 9/25, while CDU6 will take place from 9/29 - 10/2. Lists of impacted partitions are below. For those impacted we recommend using other resources on the cluster during that time such as shared and gpu_h200 or any of the requeue partitions. Note this work impacts both Cannon and FASSE. No jobs will be cancelled, rather blocking reservations are in place to naturally drain the nodes.

    Thank you for your patience as we work to improve cluster stability and hardware longevity.

    CDU 5 (9/22 - 9/25)

    Cannon:

    arguelles_delgado_gpu_a100

    arguelles_delgado_gpu_mixed

    bigmem_intermediate

    blackhole_gpu

    eddy

    gershman

    hejazi

    hernquist_ice

    hoekstra

    huce_ice

    iaifi_gpu

    itc_gpu

    jshapiro

    kovac

    kozinsky

    kozinsky_gpu

    murphy_ic

    ortegahernandez_ice

    rivas

    seas_compute

    seas_gpu

    siag

    siag_gpu

    siag_combo

    sur

    zhuang

    FASSE:

    fasse_ultramem

    CDU 6 (9/29 - 10/2)

    Cannon:

    arguelles_delgado_h100

    bigmem

    dvorkin

    eddy

    enos

    gpu

    hsph

    hsph_gpu

    intermediate

    itc_cluster

    janson_sapphire

    joonholee

    jshapiro

    olveczky_sapphire

    sapphire

    test

    yao

    yao_alphatns

    yao_gpu

    FASSE:

    cnl

FASRC monthly maintenance will take place on September 14th, 2026.
  • Concluído
    setembro 14, 2026 em 17:00UTC
    Concluído
    setembro 14, 2026 em 17:00UTC
    Manutenção concluída com sucesso
  • Atualizar
    setembro 14, 2026 em 13:08UTC
    Atualizar
    setembro 14, 2026 em 13:08UTC

    The Holyoke/MGHPCC row 7c power work has been CANCELLED by the facility and will be re-scheduled.

    The following work WILL NOT take pleace today:

    • Cancelled: MGHPCC will be upgrading power on Pod 7c Even Side on September 14th 7am-5pm. This necessitates idling half the nodes on that side of the pod. A blocking reservation has been put in place to accomplish this. No jobs will be canceled but users will notice degraded scheduling throughput due to half the nodes being closed in the following partitions: 

      arguelles_delgado, blackhole, conroy, davies, desai, doshi-velez, dsouza, eddy, edwards, geophysics, giribet, gpu_test, hernquist, huce_cascade, huttenhower, imasc, jacobsen2, janson_cascade, janson, ke, lukin, murphy, nguyen, ni_lab, olveczky, ortegahernandez, pehlevan, seas_compute, shared, shakhnovich, tambe, unrestricted, vishwanath, whipple, xlin, yin, zon

  • Em curso
    setembro 14, 2026 em 13:00UTC
    Em curso
    setembro 14, 2026 em 13:00UTC
    A manutenção já está em andamento
  • Ainda não iniciou
    agosto 31, 2026 em 18:56UTC
    Ainda não iniciou
    agosto 31, 2026 em 18:56UTC

    Our maintenance tasks should be completed between 9am-1pm.
    Some power work on row 7c will run 9-5 but will not affect running jobs or new jobs (see below).

    NOTICES:

    MAINTENANCE TASKS

    Cannon cluster will be paused during this maintenance?: YES
    FASSE cluster will be paused during this maintenance?: YES

    • Slurm Upgrade to 26.05.4

      • Audience: Cluster

      • Impact: The cluster will be paused during this maintenance

    • New: Enable Termination of User Processes on Logout

      • Audience: Cluster

      • Impact: Going forward all user processes will be terminated upon login session exit on the login nodes, excluding things running in screen and tmux. Users should leverage the cluster for non-interactive processes. This is to clean up after AI agents which tend to create orphaned processes which drag down login node performance.

    • New: watch command cadence limit

      • Audience: Cluster

      • Impact:The watch command will have a minimum cadence of 60s when used on commands talking to the slurm scheduler (i.e. squeue, showq, sinfo, scontrol, sdiag, lsload). Users desiring faster polling should leverage the sacct command that talks to the slurm database. In general users should not poll the scheduler more than once every minute, ideally once every 5-10 minutes. Polling more often slows the scheduler. FASRC reserves the right to ban users who who tax the scheduler with queries. For more on cluster customs and responsibilities see: https://docs.rc.fas.harvard.edu/kb/responsibilities/

    • Holyoke/MGHPCC row 7c power work - 9am-5pm (CANCELLED)

      • Audience: Cluster

      • Impact: MGHPCC will be upgrading power on Pod 7c Even Side on September 14th 7am-5pm. This necessitates idling half the nodes on that side of the pod. A blocking reservation has been put in place to accomplish this. No jobs will be canceled but users will notice degraded scheduling throughput due to half the nodes being closed in the following partitions: 

        arguelles_delgado, blackhole, conroy, davies, desai, doshi-velez, dsouza, eddy, edwards, geophysics, giribet, gpu_test, hernquist, huce_cascade, huttenhower, imasc, jacobsen2, janson_cascade, janson, ke, lukin, murphy, nguyen, ni_lab, olveczky, ortegahernandez, pehlevan, seas_compute, shared, shakhnovich, tambe, unrestricted, vishwanath, whipple, xlin, yin, zon

    • Login node reboots 

      • Audience: All login nodes

      • Impact: Login nodes will be unavailable until after maintenance

    • OOD/Open OnDemand down/reboots

      • Audience: All OOD users

      • Impact: OOD will be unavailable until after maintenance

    • Gurobi license key update

      • Audience: Anyone who uses Gurobi software on the clusters.

      • Impact: Jobs running Gurobi may fail. FASRC recommends waiting until maintenance is over to run Gurobi jobs.

    • Netscratch 90-day retention cleanup

      • Audience; All netscratch users

      • Impact: Files older than 90 days will be removed per our scratch policy. Please note that this cleanup can happen at any time, not just during maintenance.

    Thank you,
    FAS Research Computing
    https://docs.rc.fas.harvard.edu/
    https://www.rc.fas.harvard.edu/

ago 2026

Rolling OS Upgrades August 24th - 27th 2026
  • Concluído
    agosto 27, 2026 em 20:06UTC
    Concluído
    agosto 27, 2026 em 20:06UTC

    OS upgrade work is mostly completed.

    There are a few remaining nodes in the cluster that did not successfully receive the upgrade.

    If you see nodes that are 'DOWN' for the following reasons, they are all nodes that need manual intervention for OS upgrade. FASRC staff are aware of them and will upgrade these nodes over the next week.

    • "wrong kernel"

    • "Not responding"

    • "unexpected reboot"

    • "gres/gpu count lower than configured"

    • "/n/sw not mounted"

    We appreciate your continued patience.

  • Atualizar
    agosto 26, 2026 em 20:01UTC
    Atualizar
    agosto 26, 2026 em 20:01UTC

    Today's work is complete. The final round, Cannon Part 3, takes place tomorrow.

    August 26th: Cannon Part 2
    arguelles_delgado
    conroy
    davies
    doshi-velez
    dsouza
    edwards
    geophysics
    giribet
    gpu_test
    hernquist
    huce_cascadeimasc
    janson
    kempner_h200
    kempner_rtx
    murphy
    ni_lab
    olveczky
    ortegahernandez
    pehlevan
    seas_compute
    shakhnovich
    shared
    unrestricted
    xlin
    yin
    zon

  • Atualizar
    agosto 25, 2026 em 19:37UTC
    Atualizar
    agosto 25, 2026 em 19:37UTC

    Today's work is complete for Cannon Part 1 and will resume tomorrow for Cannon Part 2

    August 25th: Cannon Part 1

    holylogin
    arguelles_delgado
    blackhole
    davies_gpu
    davies
    desai
    eddy
    holy-cow
    holy-smokes
    huce_bigmem
    huce_cascade
    huttenhower
    jacobsen2
    janson_bigmem
    janson_cascade
    janson
    ke
    lukin
    nguyen
    olveczky_gpu
    remoteviz
    seas_compute
    shared
    sompolinsky_gpu
    tambe
    vishwanath
    whipple
    xlin
    xlin_ice
    zhuang_gpu
    zhuang

  • Atualizar
    agosto 24, 2026 em 19:07UTC
    Atualizar
    agosto 24, 2026 em 19:07UTC

    Today's portion of the upgrades are completed. Tomorrow's list (Cannon part 1) can be found in this maintenance event.

    August 24th:

    FASSE
    boslogin
    private login nodes

  • Em curso
    agosto 24, 2026 em 13:00UTC
    Em curso
    agosto 24, 2026 em 13:00UTC
    A manutenção já está em andamento
  • Ainda não iniciou
    agosto 03, 2026 em 19:35UTC
    Ainda não iniciou
    agosto 03, 2026 em 19:35UTC

    FASRC will be doing OS upgrades to the latest version of Rocky 8.10 from August 24-27th. These rolling upgrades will improve cluster security and install the latest Cuda 13.2 driver.

    Upgrades will happen in stages with each stage occurring between 8a-5p each day. During that period the nodes being upgraded will be unavailable. No jobs will be canceled but jobs will stay pending if all nodes in the partition are down.

    Note that during this week the cluster will be in a partially upgraded state, thus jobs that span multiple nodes may become unstable due to mismatching libraries. Users concerned about this should wait until August 28th to restart runs.

    Users should plan their work accordingly.

    The upgrade schedule with impacted partitions is as follows:

    August 24th:

    FASSE
    boslogin
    private login nodes

    August 25th: Cannon Part 1

    holylogin
    arguelles_delgado
    blackhole
    davies_gpu
    davies
    desai
    eddy
    holy-cow
    holy-smokes
    huce_bigmem
    huce_cascade
    huttenhower
    jacobsen2
    janson_bigmem
    janson_cascade
    janson
    ke
    lukin
    nguyen
    olveczky_gpu
    remoteviz
    seas_compute
    shared
    sompolinsky_gpu
    tambe
    vishwanath
    whipple
    xlin
    xlin_ice
    zhuang_gpu
    zhuang

    August 26th: Cannon Part 2
    arguelles_delgado
    conroy
    davies
    doshi-velez
    dsouza
    edwards
    geophysics
    giribet
    gpu_test
    hernquist
    huce_cascadeimasc
    janson
    kempner_h200
    kempner_rtx
    murphy
    ni_lab
    olveczky
    ortegahernandez
    pehlevan
    seas_compute
    shakhnovich
    shared
    unrestricted
    xlin
    yin
    zon

    August 27th: Cannon Part 3

    arguelles_delgado_gpu_a100
    arguelles_delgado_gpu_mixed
    arguelles_delgado_h100
    bigmem_intermediate
    bigmem
    blackhole_gpu
    dvorkin
    eddy
    enos
    gershman
    gpu
    gpu_h200
    hejazi
    hernquist_ice
    hoekstra
    hsph_gpu
    hsph
    huce_ice
    iaifi_gpu
    intermediate
    itc_cluster
    itc_gpu
    janson_sapphire
    joonholee
    jshapiro
    kempner_h100
    kempner_h200
    kempner
    kempner_interactive
    kovac
    kozinsky_gpu
    kozinsky
    murphy_ice
    mweber_compute
    mweber_gpu
    olveczky_sapphire
    ortegahernandez_ice
    rivas
    sapphire
    seas_compute
    seas_gpu
    seas_gpu_perf
    siag_gpu
    siag_combo
    siag
    sur
    test
    yao_alphatns
    yao_gpu
    yao
    zhuang

Emergency Slurm (scheduler) patch at 2pm ETA 1-2hours. Jobs will be paused
  • Atualizar
    UTC
    Atualizar

    A number of jobs appear to have terminated during the upgrade for some reason. Those jobs will re-queue or, if not, will need to be re-submitted,
    We apologize for the inconvenience.

  • Resolvido
    UTC
    Resolvido

    The update has been applied and the scheduler is returned to service. Jobs are unpaused.

  • Investigando
    UTC
    Investigando

    We have been manually dealing with a Slurm memory leak bug behind the scenes and now have an official fix in the form of a Slurm upgrade (26.05.3). It is necessary to get this update in place as soon as possible.

    At 2pm today we will pause the cluster. Jobs will be paused, not terminated, and the scheduler will be unavailable for querying or submitting new jobs until the update is completed. Once complete, jobs will resume.

    Technical information about the bug:

    https://support.schedmd.com/show_bug.cgi?id=25685

jul 2026

FASRC monthly maintenance Monday July 6th, 2026 9am-1pm
  • Concluído
    julho 06, 2026 em 17:00UTC
    Concluído
    julho 06, 2026 em 17:00UTC
    Manutenção concluída com sucesso
  • Em curso
    julho 06, 2026 em 13:00UTC
    Em curso
    julho 06, 2026 em 13:00UTC
    A manutenção já está em andamento
  • Ainda não iniciou
    junho 26, 2026 em 13:45UTC
    Ainda não iniciou
    junho 26, 2026 em 13:45UTC

    FASRC monthly maintenance will take place on July 6th 2026. Our maintenance tasks should be completed between 9am-1pm.

    Cannon cluster will be paused during this maintenance?: NO
    FASSE cluster will be paused during this maintenance?: NO

    NOTICES:

    • Friday July 3rd is a university holiday (independence Day observed)

    • Training: Upcoming training from FASRC and other sources can be found on our Training Calendar. at https://www.rc.fas.harvard.edu/upcoming-training/

    • Status Page: You can subscribe to our status to receive notifications of maintenance, incidents, and their resolution at https://status.rc.fas.harvard.edu/ (click Get Updates for options).

    • We'd love to hear success stories about your or your lab's use of FASRC. Submit your story here.

    MAINTENANCE TASKS

    • Domain controller replacement

      • Audience: Internal

      • Impact: None. End users should not see any impact.

    • Reboot drained nodes in error state

      • Audience: Cluster nodes with errors.

      • Impact: These nodes will have been drained already in preparation. No impact on jobs on the day and the affected nodes will return to service in their respective partitions after the maintenance period.

    • OOD/Open OnDemand reboots

      • Audience: All OOD users, reboot of the head nodes.

      • Impact: Running sessions will not be affected.

    • Login node reboots

      • Audience; All login node users.

      • Impact: Login nodes will reboot during the maintenance window.

    • Netscratch 90-day retention cleanup

      • Audience; All netscratch users

      • Impact: Files older than 90 days will be removed per our scratch policy. Please note that this cleanup can happen at any time, not just during maintenance.

    Thank you,
    FAS Research Computing
    https://docs.rc.fas.harvard.edu/
    https://www.rc.fas.harvard.edu/

jul 2026 para set 2026

Próximo