FAS Research Computing - تاریخچه اطلاعیه‌ها

همه سیستم‌ها عملیاتی هستند

Status page for the Harvard FAS Research Computing cluster and other resources.

Cluster Utilization (VPN and FASRC login required): Cannon | FASSE


Please scroll down to see details on any Incidents or maintenance notices.
Monthly maintenance occurs on the first Monday of the month (except holidays).

GETTING HELP
Documentation: https://docs.rc.fas.harvard.edu | Account Portal https://portal.rc.fas.harvard.edu
Email: rchelp@rc.fas.harvard.edu | Support Hours


The colors shown in the bars below were chosen to increase visibility for color-blind visitors.
For higher contrast, switch to light mode at the bottom of this page if the background is dark and colors are muted.

عملیاتی

SLURM Scheduler - Cannon - عملیاتی

Cannon Compute Cluster (Holyoke) - عملیاتی

Boston Compute Nodes - عملیاتی

GPU nodes (Holyoke) - عملیاتی

seas_compute - عملیاتی

عملیاتی

SLURM Scheduler - FASSE - عملیاتی

FASSE Compute Cluster (Holyoke) - عملیاتی

عملیاتی

Kempner Cluster CPU - عملیاتی

Kempner Cluster GPU - عملیاتی

عملیاتی

FASSE login nodes - عملیاتی

عملیاتی

Cannon Open OnDemand - عملیاتی

FASSE Open OnDemand - عملیاتی

عملیاتی

Netscratch (Global Scratch) - عملیاتی

Home Directory Storage - Boston - عملیاتی

Tape - (Tier 3) - عملیاتی

Holylabs - عملیاتی

Isilon Storage Holyoke (Tier 1) - عملیاتی

Holystore01 (Tier 0) - عملیاتی

HolyLFS04 (Tier 0) - عملیاتی

HolyLFS05 (Tier 0) - عملیاتی

HolyLFS06 (Tier 0) - عملیاتی

Holyoke Tier 2 NFS - عملیاتی

Holyoke Specialty Storage - عملیاتی

holECS - عملیاتی

Isilon Storage Boston (Tier 1) - عملیاتی

BosLFS02 (Tier 0) - عملیاتی

Boston Tier 2 NFS - عملیاتی

CEPH Storage Boston (Tier 2) - عملیاتی

Boston Specialty Storage - عملیاتی

bosECS - عملیاتی

Samba Cluster - عملیاتی

Globus Data Transfer - عملیاتی

تاریخچه اطلاعیه‌ها

مشاهده وضعیت فعلی

اکتـ 2026

اکتـ5
Monthly maintenance October 5th, 2026 9am-1pm.
تکمیل شدتعمیر و نگهداری4 ساعت
  • تکمیل شد
    اکتبر 05, 2026 در 5:00 ب.ظ.UTC
    تکمیل شد
    اکتبر 05, 2026 در 5:00 ب.ظ.UTC
    Maintenance has completed successfully
  • در حال انجام
    اکتبر 05, 2026 در 1:00 ب.ظ.UTC
    در حال انجام
    اکتبر 05, 2026 در 1:00 ب.ظ.UTC
    Maintenance is now in progress
  • برنامه‌ریزی شده
    سپتامبر 24, 2026 در 1:54 ب.ظ.UTC
    برنامه‌ریزی شده
    سپتامبر 24, 2026 در 1:54 ب.ظ.UTC

    Our maintenance tasks should be completed between 9am-1pm.

    - CDU work continues until October 2nd 8am-5pm but will not affect running jobs or new jobs.

    Phase 1 is complete. See status page for timeline and affected partitions.

    ​- Power work on row 7c October 4th 5pm - October 5th 5pm. See status page for affected partitions.

    NOTICES:

    • VPN: The VPN will be cut over to new hardware on Sunday Sept. 27th. connections will drop and you will need to reconnect.

    • Training: Upcoming training from FASRC and other sources can be found on our Training Calendar. at https://www.rc.fas.harvard.edu/upcoming-training/

    • Status Page: You can subscribe to our status to receive notifications of maintenance, incidents, and their resolution at https://status.rc.fas.harvard.edu/ (click Get Updates for options).

    MAINTENANCE TASKS

    Cannon cluster will be paused during this maintenance?: YES
    FASSE cluster will be paused during this maintenance?: YES

    • Network maintenance 

      • Audience: Clusters (Cannon and FASSE)

      • Impact: The cluster will be paused during this work.
        Some network disruption may occur.

    • Login node reboots 

      • Audience: All login nodes

      • Impact: Login nodes will be unavailable until after maintenance

    • OOD/Open OnDemand reboots

      • Audience: All OOD users

      • Impact: OOD will be unavailable until after maintenance

    • Netscratch 90-day retention cleanup

      • Audience; All netscratch users

      • Update: Due to low space on netscratch, please note that retention cleanup will be run more frequently. We are reviewing our 90 day policy with a possible age reduction.

      • Impact: Files older than 90 days will be removed per our scratch policy. Please note that this cleanup can happen at any time, not just during maintenance.

    Thank you,
    FAS Research Computing
    https://docs.rc.fas.harvard.edu/
    https://www.rc.fas.harvard.edu/upcoming-training/



اکتـ4
Power work on row 7c October 4th - October 5th
تکمیل شدتعمیر و نگهداری45 ساعت 58 دقیقه
  • تکمیل شد
    اکتبر 06, 2026 در 6:58 ب.ظ.UTC
    تکمیل شد
    اکتبر 06, 2026 در 6:58 ب.ظ.UTC

    Maintenance has completed successfully.

  • در حال انجام
    اکتبر 06, 2026 در 12:20 ق.ظ.UTC
    در حال انجام
    اکتبر 06, 2026 در 12:20 ق.ظ.UTC

    Unfortunately the power work is taking longer than expected. This work will continue on Tuesday.

  • تکمیل شد
    اکتبر 05, 2026 در 11:30 ب.ظ.UTC
    تکمیل شد
    اکتبر 05, 2026 در 11:30 ب.ظ.UTC
    Maintenance has completed successfully
  • در حال انجام
    اکتبر 04, 2026 در 9:00 ب.ظ.UTC
    در حال انجام
    اکتبر 04, 2026 در 9:00 ب.ظ.UTC
    Maintenance is now in progress
  • برنامه‌ریزی شده
    سپتامبر 24, 2026 در 2:06 ب.ظ.UTC
    برنامه‌ریزی شده
    سپتامبر 24, 2026 در 2:06 ب.ظ.UTC

    MGHPCC will be upgrading power on Pod 7c Even Side on October 4th - October 5th. ETA to end 7:30 PM Oct 5th. This necessitates idling half the nodes on that side of the pod. A blocking reservation has been put in place to accomplish this. No jobs will be canceled but users will notice degraded scheduling throughput due to half the nodes being closed in the following partitions:

    arguelles_delgado

    blackhole

    conroy

    davies

    desai

    doshi-velez

    dsouza

    eddy

    edwards

    geophysics

    giribet

    gpu_test

    hernquist

    huce_cascade

    huttenhower

    imasc

    jacobsen2

    janson_cascade

    janson

    ke

    lukin

    murphy

    nguyen

    ni_lab

    olveczky

    ortegahernandez

    pehlevan

    seas_compute

    shared

    shakhnovich

    tambe

    unrestricted

    vishwanath

    whipple

    xlin

    yin

    zon



سپتـ 2026

h-nfs11-p down
حل شداختلال جزئی144 ساعت 33 دقیقه
  • حل شد
    UTC
    حل شد

    The vendor has replaced the necessary hardware components, and FASRC engineers have installed updates and confirmed working status. No data was affected.

    All h-nfs11-p shares should be accessible again once more.

    If you notice any irregularities, please reach out at rchelp@rc.fas.harvard.edu

  • به‌روزرسانی
    UTC
    به‌روزرسانی

    The vendor has dispatched parts. The vendor has indicated that this may not get fixed until Monday, but the storage team is pressing them to accelerate that timeline. We will update here when we know more.

    Thanks for your understanding.

  • به‌روزرسانی
    UTC
    به‌روزرسانی

    After lengthy troubleshooting, h-nfs11-p is still down. The vendor is scheduling a datacenter visit for on-site repairs. We will update as we know more.

  • شناسایی شد
    UTC
    شناسایی شد

    Our engineers are diagnosing the root cause of the failure and are working with the vendor to troubleshoot further

    We are continuing to work on a fix for this incident.

  • در حال بررسی
    UTC
    در حال بررسی

    h-nfs11-p is currently down. The following shares are inaccessible on FASSE

    • ncf_cnl01

    • ncf_cnl02

    • ncf_cnl03

    • ncf_cnl04

    • ncf_gspdata

    • ncf_mri_l3

    • cnl05

    • cnl06

    • eldaief

    • ncf_apps

    • jcamprodon_lab_l3

    We are investigating this incident.

آگو 2026

آگو24
Rolling OS Upgrades August 24th - 27th 2026
تکمیل شدتعمیر و نگهداری79 ساعت 6 دقیقه
  • تکمیل شد
    آگوست 27, 2026 در 8:06 ب.ظ.UTC
    تکمیل شد
    آگوست 27, 2026 در 8:06 ب.ظ.UTC

    OS upgrade work is mostly completed.

    There are a few remaining nodes in the cluster that did not successfully receive the upgrade.

    If you see nodes that are 'DOWN' for the following reasons, they are all nodes that need manual intervention for OS upgrade. FASRC staff are aware of them and will upgrade these nodes over the next week.

    • "wrong kernel"

    • "Not responding"

    • "unexpected reboot"

    • "gres/gpu count lower than configured"

    • "/n/sw not mounted"

    We appreciate your continued patience.

  • به‌روزرسانی
    آگوست 26, 2026 در 8:01 ب.ظ.UTC
    به‌روزرسانی
    آگوست 26, 2026 در 8:01 ب.ظ.UTC

    Today's work is complete. The final round, Cannon Part 3, takes place tomorrow.

    August 26th: Cannon Part 2
    arguelles_delgado
    conroy
    davies
    doshi-velez
    dsouza
    edwards
    geophysics
    giribet
    gpu_test
    hernquist
    huce_cascadeimasc
    janson
    kempner_h200
    kempner_rtx
    murphy
    ni_lab
    olveczky
    ortegahernandez
    pehlevan
    seas_compute
    shakhnovich
    shared
    unrestricted
    xlin
    yin
    zon

  • به‌روزرسانی
    آگوست 25, 2026 در 7:37 ب.ظ.UTC
    به‌روزرسانی
    آگوست 25, 2026 در 7:37 ب.ظ.UTC

    Today's work is complete for Cannon Part 1 and will resume tomorrow for Cannon Part 2

    August 25th: Cannon Part 1

    holylogin
    arguelles_delgado
    blackhole
    davies_gpu
    davies
    desai
    eddy
    holy-cow
    holy-smokes
    huce_bigmem
    huce_cascade
    huttenhower
    jacobsen2
    janson_bigmem
    janson_cascade
    janson
    ke
    lukin
    nguyen
    olveczky_gpu
    remoteviz
    seas_compute
    shared
    sompolinsky_gpu
    tambe
    vishwanath
    whipple
    xlin
    xlin_ice
    zhuang_gpu
    zhuang

  • به‌روزرسانی
    آگوست 24, 2026 در 7:07 ب.ظ.UTC
    به‌روزرسانی
    آگوست 24, 2026 در 7:07 ب.ظ.UTC

    Today's portion of the upgrades are completed. Tomorrow's list (Cannon part 1) can be found in this maintenance event.

    August 24th:

    FASSE
    boslogin
    private login nodes

  • در حال انجام
    آگوست 24, 2026 در 1:00 ب.ظ.UTC
    در حال انجام
    آگوست 24, 2026 در 1:00 ب.ظ.UTC
    Maintenance is now in progress
  • برنامه‌ریزی شده
    آگوست 03, 2026 در 7:35 ب.ظ.UTC
    برنامه‌ریزی شده
    آگوست 03, 2026 در 7:35 ب.ظ.UTC

    FASRC will be doing OS upgrades to the latest version of Rocky 8.10 from August 24-27th. These rolling upgrades will improve cluster security and install the latest Cuda 13.2 driver.

    Upgrades will happen in stages with each stage occurring between 8a-5p each day. During that period the nodes being upgraded will be unavailable. No jobs will be canceled but jobs will stay pending if all nodes in the partition are down.

    Note that during this week the cluster will be in a partially upgraded state, thus jobs that span multiple nodes may become unstable due to mismatching libraries. Users concerned about this should wait until August 28th to restart runs.

    Users should plan their work accordingly.

    The upgrade schedule with impacted partitions is as follows:

    August 24th:

    FASSE
    boslogin
    private login nodes

    August 25th: Cannon Part 1

    holylogin
    arguelles_delgado
    blackhole
    davies_gpu
    davies
    desai
    eddy
    holy-cow
    holy-smokes
    huce_bigmem
    huce_cascade
    huttenhower
    jacobsen2
    janson_bigmem
    janson_cascade
    janson
    ke
    lukin
    nguyen
    olveczky_gpu
    remoteviz
    seas_compute
    shared
    sompolinsky_gpu
    tambe
    vishwanath
    whipple
    xlin
    xlin_ice
    zhuang_gpu
    zhuang

    August 26th: Cannon Part 2
    arguelles_delgado
    conroy
    davies
    doshi-velez
    dsouza
    edwards
    geophysics
    giribet
    gpu_test
    hernquist
    huce_cascadeimasc
    janson
    kempner_h200
    kempner_rtx
    murphy
    ni_lab
    olveczky
    ortegahernandez
    pehlevan
    seas_compute
    shakhnovich
    shared
    unrestricted
    xlin
    yin
    zon

    August 27th: Cannon Part 3

    arguelles_delgado_gpu_a100
    arguelles_delgado_gpu_mixed
    arguelles_delgado_h100
    bigmem_intermediate
    bigmem
    blackhole_gpu
    dvorkin
    eddy
    enos
    gershman
    gpu
    gpu_h200
    hejazi
    hernquist_ice
    hoekstra
    hsph_gpu
    hsph
    huce_ice
    iaifi_gpu
    intermediate
    itc_cluster
    itc_gpu
    janson_sapphire
    joonholee
    jshapiro
    kempner_h100
    kempner_h200
    kempner
    kempner_interactive
    kovac
    kozinsky_gpu
    kozinsky
    murphy_ice
    mweber_compute
    mweber_gpu
    olveczky_sapphire
    ortegahernandez_ice
    rivas
    sapphire
    seas_compute
    seas_gpu
    seas_gpu_perf
    siag_gpu
    siag_combo
    siag
    sur
    test
    yao_alphatns
    yao_gpu
    yao
    zhuang

Emergency Slurm (scheduler) patch at 2pm ETA 1-2hours. Jobs will be paused
حل شداختلال جدی3 ساعت 42 دقیقه
  • به‌روزرسانی
    UTC
    به‌روزرسانی

    A number of jobs appear to have terminated during the upgrade for some reason. Those jobs will re-queue or, if not, will need to be re-submitted,
    We apologize for the inconvenience.

  • حل شد
    UTC
    حل شد

    The update has been applied and the scheduler is returned to service. Jobs are unpaused.

  • در حال بررسی
    UTC
    در حال بررسی

    We have been manually dealing with a Slurm memory leak bug behind the scenes and now have an official fix in the form of a Slurm upgrade (26.05.3). It is necessary to get this update in place as soon as possible.

    At 2pm today we will pause the cluster. Jobs will be paused, not terminated, and the scheduler will be unavailable for querying or submitting new jobs until the update is completed. Once complete, jobs will resume.

    Technical information about the bug:

    https://support.schedmd.com/show_bug.cgi?id=25685

قبلی

آگو 2026 تا اکتـ 2026

بعدی