FAS Research Computing - سجل التاريخ

جميع الأنظمة جاهزة للعمل

Status page for the Harvard FAS Research Computing cluster and other resources.

Cluster Utilization (VPN and FASRC login required): Cannon | FASSE


Please scroll down to see details on any Incidents or maintenance notices.
Monthly maintenance occurs on the first Monday of the month (except holidays).

GETTING HELP
Documentation: https://docs.rc.fas.harvard.edu | Account Portal https://portal.rc.fas.harvard.edu
Email: rchelp@rc.fas.harvard.edu | Support Hours


The colors shown in the bars below were chosen to increase visibility for color-blind visitors.
For higher contrast, switch to light mode at the bottom of this page if the background is dark and colors are muted.

جاهز للعمل

SLURM Scheduler - Cannon - جاهز للعمل

Cannon Compute Cluster (Holyoke) - جاهز للعمل

Boston Compute Nodes - جاهز للعمل

GPU nodes (Holyoke) - جاهز للعمل

seas_compute - جاهز للعمل

جاهز للعمل

SLURM Scheduler - FASSE - جاهز للعمل

FASSE Compute Cluster (Holyoke) - جاهز للعمل

جاهز للعمل

Kempner Cluster CPU - جاهز للعمل

Kempner Cluster GPU - جاهز للعمل

جاهز للعمل

FASSE login nodes - جاهز للعمل

جاهز للعمل

Cannon Open OnDemand - جاهز للعمل

FASSE Open OnDemand - جاهز للعمل

جاهز للعمل

Netscratch (Global Scratch) - جاهز للعمل

Home Directory Storage - Boston - جاهز للعمل

Tape - (Tier 3) - جاهز للعمل

Holylabs - جاهز للعمل

Isilon Storage Holyoke (Tier 1) - جاهز للعمل

Holystore01 (Tier 0) - جاهز للعمل

HolyLFS04 (Tier 0) - جاهز للعمل

HolyLFS05 (Tier 0) - جاهز للعمل

HolyLFS06 (Tier 0) - جاهز للعمل

Holyoke Tier 2 NFS (new) - جاهز للعمل

Holyoke Specialty Storage - جاهز للعمل

holECS - جاهز للعمل

Isilon Storage Boston (Tier 1) - جاهز للعمل

BosLFS02 (Tier 0) - جاهز للعمل

Boston Tier 2 NFS (new) - جاهز للعمل

CEPH Storage Boston (Tier 2) - جاهز للعمل

Boston Specialty Storage - جاهز للعمل

bosECS - جاهز للعمل

Samba Cluster - جاهز للعمل

Globus Data Transfer - جاهز للعمل

سجل التاريخ

مايو 2025

MGHPCC power work 5/21 - 5/23 - Some partitions will be at half capacity
  • مكتمل
    مايو 23, 2025 في 19:00UTC
    مكتمل
    مايو 23, 2025 في 19:00UTC
    Maintenance has completed successfully
  • قيد التقدم
    مايو 21, 2025 في 11:00UTC
    قيد التقدم
    مايو 21, 2025 في 11:00UTC
    Maintenance is now in progress
  • مخطط
    مايو 21, 2025 في 11:00UTC
    مخطط
    مايو 21, 2025 في 11:00UTC

    The MGHPCC Holyoke data center will be performing power work on May 21st -23rd. This work will take out one half (or one 'side') of the power capacity for certain rows/racks including our compute rows. Because of our power draw, one side is not enough power to keep each full rack running.

    As such, we will be adding a reservation to idle half the nodes in the partitions listed below. A reservation will cause nodes to drain as jobs complete and stop scheduling new jobs on those nodes if they cannot be completed before the outage. This will allow us to idle and power down those nodes prior to the work and avoid potential blackout/brownout on those racks.

    This will mean that these partitions will be up and available, but that half the nodes from each will be down (assuming an even number of nodes).

    This work is part of an on-going power capacity upgrade at MGHPCC. We expect this will be the last power work needed and the facility will then provide enough additional power for future expansion as well adding overhead for the current load.

    The affected partitions are:

    • arguelles_delgado

    • bigmem_intermediate

    • blackhole_gpu

    • eddy gershman

    • hejazi

    • hernquist

    • hoekstra

    • huce_ice

    • iaifi_gpu

    • iaifi_gpu_requeue

    • iaifi_priority

    • jshapiro

    • jshapiro_priority

    • kempner

    • kempner_requeue

    • kempner_h100

    • kempner_h100_priority

    • kempner_h100_priority2

    • kovac kozinsky

    • kozinsky_gpu

    • kozinsky_requeue

    • ortegahernandez_ice

    • rivas

    • seas_compute

    • seas_gpu

    • siag_combo

    • siag_gpu

    • sur

    • zhuang

أبريل 2025

Login nodes temporarily down
  • تم الحل
    UTC
    تم الحل

    Cannon boslogin and FASSE login nodes are back up and operational.

    All holylogin nodes are still down for repair, please see our posted incident for more updates: https://status.rc.fas.harvard.edu/cm97gyay90013dturk7fxg5pb

    We apologize for the unexpected disruption.

  • تحقيق
    UTC
    تحقيق

    Due to a configuration error, all cluster login nodes are rebooting and are temporarily unavailable. Please save any work immediately.

holylogin[05-08] down
  • تم الحل
    UTC
    تم الحل
    Hardware has been repaired and holyoke login nodes are back online. Thanks for your patience.
  • المراقبة
    UTC
    المراقبة

    Holylogin chassis repair during maintenance was unsuccessful and replacement parts have been ordered.

    • holyoke login nodes (holylogin05-08) are down for hardware repair

    • Only Boston login nodes available (ie, boslogin[05-08])

    If you have holylogin hard-coded in your scripts, please update to login.rc.fas.harvard.edu or boslogin.rc.fas.harvard.edu for the time being, which will redirect you to an available login node.

    As always, the best method for obtaining a login node is usinglogin.rc.fas.harvard.edu which will pick a node for you.

    If you require a login node in a specific data center, use boslogin.rc.fas.harvard.edu (Boston) or (once they are back in service) holylogin.rc.fas.harvard.edu (Holyoke).

    See also: Command line access with Terminal (login nodes) – FASRC DOCS

  • تم الحل
    UTC
    تم الحل

    This incident was posted by mistake.

    holylogin01-04 were replaced by holylogin05-08 some time back.

    As always, the best method for obtaining a login node is usinglogin.rc.fas.harvard.edu which will pick a node for you.

    Or if you require a login node in a specific data center, use boslogin.rc.fas.harvard.edu (Boston) or holylogin.rc.fas.harvard.edu (Holyoke).

    See also: Command line access with Terminal (login nodes) – FASRC DOCS

  • تحقيق
    UTC
    تحقيق

    Holylogin chassis repair during maintenance was unsuccessful and replacement parts have been ordered.

    Audience:

    • All cluster users

    Impact:

    • All holylogin** servers will be down till further notice

    • Only Boston login nodes available (ie, boslogin[05-08])

    If you have holylogin hard-coded in your scripts, please update to login.rc.fas.harvard.edu or boslogin.rc.fas.harvard.edu for the time being, which will redirect you to an available login node.

    Updates to follow as we have them.

مارس 2025

FASRC monthly maintenance - Monday March 3rd, 2025 from 9am-1pm
  • تحديث
    مارس 03, 2025 في 18:00UTC
    تحديث
    مارس 03, 2025 في 18:00UTC
    Maintenance has completed successfully
  • مكتمل
    مارس 03, 2025 في 18:00UTC
    مكتمل
    مارس 03, 2025 في 18:00UTC
    Maintenance has completed successfully
  • قيد التقدم
    مارس 03, 2025 في 14:00UTC
    قيد التقدم
    مارس 03, 2025 في 14:00UTC
    Maintenance is now in progress
  • مخطط
    مارس 03, 2025 في 14:00UTC
    مخطط
    مارس 03, 2025 في 14:00UTC

    PLEASE NOTE - New time window going forward - 9am-1pm

    FASRC monthly maintenance will take place Monday March 3rd, 2025 from 9am-1pm

    NOTICES

    • Training: Upcoming training from FASRC and other sources can be found on our Training Calendar. at https://www.rc.fas.harvard.edu/upcoming-training/

    • Status Page: You can subscribe to our status to receive notifications of maintenance, incidents, and their resolution at https://status.rc.fas.harvard.edu/ (click Get Updates for options).

    • Upcoming holidays: Memorial Day - Monday, May 26

    • You can subscribe to our status page using the Get Updates button in the upper right

    MAINTENANCE TASKS
    Cannon cluster will be paused during this maintenance?: YES
    FASSE cluster will be paused during this maintenance?: YES

    • Slurm Upgrade to 24.11.2 - Crucial Update

      • Audience: All cluster users

      • Impact: Jobs and the scheduler will be paused during this upgrade

    • Open Ondemand (OOD) reboots

      • Audience: All OOD users

      • Impact: All Open OnDemand (aka OOD/VDI/RCOOD) nodes will be rebooted

    • Login node reboots

      • Audience: Anyone logged into a FASRC Cannon or FASSE login node

      • Impact: Login nodes will rebooted during this maintenance window

    • bos-Isilon firmware updates

      • Audience: bos-isilon users

      • Impact: No noticeable impact for storage users

    • Netscratch retention/cleanup ( https://docs.rc.fas.harvard.edu/kb/policy-scratch/ )

      • Audience: Cluster users

      • Impact: Files older than 90 days will be removed. Please note that retention cleanup can and does run at any time, not just during the maintenance window.

    Thank you,
    FAS Research Computing
    https://www.rc.fas.harvard.edu/
    https://docs.rc.fas.harvard.edu/

مارس 2025 ألى مايو 2025

التالي