FAS Research Computing - Notice history

Status page for the Harvard FAS Research Computing cluster and other resources.

Cluster Utilization (VPN and FASRC login required): Cannon | FASSE


Please scroll down to see details on any Incidents or maintenance notices.
Monthly maintenance occurs on the first Monday of the month (except holidays).

GETTING HELP
https://docs.rc.fas.harvard.edu | https://portal.rc.fas.harvard.edu | Email: rchelp@rc.fas.harvard.edu


The colors shown in the bars below were chosen to increase visibility for color-blind visitors.
For higher contrast, switch to light mode at the bottom of this page if the background is dark and colors are muted.

Under maintenance

SLURM Scheduler - Cannon - Under maintenance

Cannon Compute Cluster (Holyoke) - Under maintenance

Boston Compute Nodes - Under maintenance

GPU nodes (Holyoke) - Under maintenance

seas_compute - Under maintenance

Under maintenance

SLURM Scheduler - FASSE - Under maintenance

FASSE Compute Cluster (Holyoke) - Under maintenance

Under maintenance

Kempner Cluster CPU - Under maintenance

Kempner Cluster GPU - Under maintenance

Under maintenance

Login Nodes - Boston - Under maintenance

Login Nodes - Holyoke - Under maintenance

FASSE login nodes - Under maintenance

Under maintenance

Cannon Open OnDemand/VDI - Under maintenance

FASSE Open OnDemand/VDI - Under maintenance

Under maintenance

Netscratch (Global Scratch) - Under maintenance

Home Directory Storage - Boston - Operational

Tape - (Tier 3) - Under maintenance

Holylabs - Under maintenance

Isilon Storage Holyoke (Tier 1) - Under maintenance

Holystore01 (Tier 0) - Under maintenance

HolyLFS04 (Tier 0) - Under maintenance

HolyLFS05 (Tier 0) - Under maintenance

HolyLFS06 (Tier 0) - Under maintenance

Holyoke Tier 2 NFS (new) - Under maintenance

Holyoke Specialty Storage - Under maintenance

holECS - Under maintenance

Isilon Storage Boston (Tier 1) - Operational

BosLFS02 (Tier 0) - Operational

Boston Tier 2 NFS (new) - Operational

CEPH Storage Boston (Tier 2) - Operational

Boston Specialty Storage - Operational

bosECS - Operational

Samba Cluster - Under maintenance

Globus Data Transfer - Under maintenance

Notice history

Jun 2025

2025 MGHPCC power downtime June 2-4, 2025
Scheduled for June 02, 2025 at 1:00 PM – June 05, 2025 at 1:00 PM 3 days
  • In progress
    June 02, 2025 at 1:00 PM
    In progress
    June 02, 2025 at 1:00 PM
    Maintenance is now in progress
  • Planned
    June 02, 2025 at 1:00 PM
    Planned
    June 02, 2025 at 1:00 PM

    The yearly power downtime at our Holyoke data center, MGHPCC, has been scheduled. 
    This year's power downtime will take place on Tuesday June 3, 2025. 

    This will require FASRC to begin shutdown of our systems beginning at 9AM on Monday, June 2nd.
    We have worked to reduce the total outage time this year.
    We will begin power-up on Wednesday June 4th with an expected return to full service by 9AM Thursday June 5th.

    • Monday June 2nd -  Power-down begins at 9AM

    • Tuesday June 3rd - Power out at MGHPCC

    • Wednesday June 4th - Maintenance tasks and then power-up begins

    • Thursday June 5th - Expected return to full service by 9AM

    Maintenance:
    During this downtime, Holylabs (/n/holylabs) will move to new hardware.
    Starfish, Coldfront, and the Portal will be unavailable during the downtime.

    For more details including a graphical timeline, please see: https://www.rc.fas.harvard.edu/events/2025-mghpcc-power-downtime/

    Updates will be posted here on our status page: https://status.rc.fas.harvard.edu/
    Note that you can subscribe to receive updates as they happen. On the status page, click Get Updates.

    Notices and reminders will also be sent to all users via our mailing lists.

May 2025

MGHPCC power work 5/21 - 5/23 - Some partitions will be at half capacity
  • Completed
    May 23, 2025 at 7:00 PM
    Completed
    May 23, 2025 at 7:00 PM
    Maintenance has completed successfully
  • In progress
    May 21, 2025 at 11:00 AM
    In progress
    May 21, 2025 at 11:00 AM
    Maintenance is now in progress
  • Planned
    May 21, 2025 at 11:00 AM
    Planned
    May 21, 2025 at 11:00 AM

    The MGHPCC Holyoke data center will be performing power work on May 21st -23rd. This work will take out one half (or one 'side') of the power capacity for certain rows/racks including our compute rows. Because of our power draw, one side is not enough power to keep each full rack running.

    As such, we will be adding a reservation to idle half the nodes in the partitions listed below. A reservation will cause nodes to drain as jobs complete and stop scheduling new jobs on those nodes if they cannot be completed before the outage. This will allow us to idle and power down those nodes prior to the work and avoid potential blackout/brownout on those racks.

    This will mean that these partitions will be up and available, but that half the nodes from each will be down (assuming an even number of nodes).

    This work is part of an on-going power capacity upgrade at MGHPCC. We expect this will be the last power work needed and the facility will then provide enough additional power for future expansion as well adding overhead for the current load.

    The affected partitions are:

    • arguelles_delgado

    • bigmem_intermediate

    • blackhole_gpu

    • eddy gershman

    • hejazi

    • hernquist

    • hoekstra

    • huce_ice

    • iaifi_gpu

    • iaifi_gpu_requeue

    • iaifi_priority

    • jshapiro

    • jshapiro_priority

    • kempner

    • kempner_requeue

    • kempner_h100

    • kempner_h100_priority

    • kempner_h100_priority2

    • kovac kozinsky

    • kozinsky_gpu

    • kozinsky_requeue

    • ortegahernandez_ice

    • rivas

    • seas_compute

    • seas_gpu

    • siag_combo

    • siag_gpu

    • sur

    • zhuang

Apr 2025

Login nodes temporarily down
  • Resolved
    Resolved

    Cannon boslogin and FASSE login nodes are back up and operational.

    All holylogin nodes are still down for repair, please see our posted incident for more updates: https://status.rc.fas.harvard.edu/cm97gyay90013dturk7fxg5pb

    We apologize for the unexpected disruption.

  • Investigating
    Investigating

    Due to a configuration error, all cluster login nodes are rebooting and are temporarily unavailable. Please save any work immediately.

holylogin[05-08] down
  • Resolved
    Resolved
    Hardware has been repaired and holyoke login nodes are back online. Thanks for your patience.
  • Monitoring
    Monitoring

    Holylogin chassis repair during maintenance was unsuccessful and replacement parts have been ordered.

    • holyoke login nodes (holylogin05-08) are down for hardware repair

    • Only Boston login nodes available (ie, boslogin[05-08])

    If you have holylogin hard-coded in your scripts, please update to login.rc.fas.harvard.edu or boslogin.rc.fas.harvard.edu for the time being, which will redirect you to an available login node.

    As always, the best method for obtaining a login node is usinglogin.rc.fas.harvard.edu which will pick a node for you.

    If you require a login node in a specific data center, use boslogin.rc.fas.harvard.edu (Boston) or (once they are back in service) holylogin.rc.fas.harvard.edu (Holyoke).

    See also: Command line access with Terminal (login nodes) – FASRC DOCS

  • Resolved
    Resolved

    This incident was posted by mistake.

    holylogin01-04 were replaced by holylogin05-08 some time back.

    As always, the best method for obtaining a login node is usinglogin.rc.fas.harvard.edu which will pick a node for you.

    Or if you require a login node in a specific data center, use boslogin.rc.fas.harvard.edu (Boston) or holylogin.rc.fas.harvard.edu (Holyoke).

    See also: Command line access with Terminal (login nodes) – FASRC DOCS

  • Investigating
    Investigating

    Holylogin chassis repair during maintenance was unsuccessful and replacement parts have been ordered.

    Audience:

    • All cluster users

    Impact:

    • All holylogin** servers will be down till further notice

    • Only Boston login nodes available (ie, boslogin[05-08])

    If you have holylogin hard-coded in your scripts, please update to login.rc.fas.harvard.edu or boslogin.rc.fas.harvard.edu for the time being, which will redirect you to an available login node.

    Updates to follow as we have them.

Apr 2025 to Jun 2025

Next