FAS Research Computing - سجل الإشعارات

Status page for the Harvard FAS Research Computing cluster and other resources.

Cluster Utilization (VPN and FASRC login required): Cannon | FASSE


Please scroll down to see details on any Incidents or maintenance notices.
Monthly maintenance occurs on the first Monday of the month (except holidays).

GETTING HELP
Documentation: https://docs.rc.fas.harvard.edu | Account Portal https://portal.rc.fas.harvard.edu
Email: rchelp@rc.fas.harvard.edu | Support Hours


The colors shown in the bars below were chosen to increase visibility for color-blind visitors.
For higher contrast, switch to light mode at the bottom of this page if the background is dark and colors are muted.

يعمل بشكل طبيعي

SLURM Scheduler - Cannon - يعمل بشكل طبيعي

Cannon Compute Cluster (Holyoke) - يعمل بشكل طبيعي

Boston Compute Nodes - يعمل بشكل طبيعي

GPU nodes (Holyoke) - يعمل بشكل طبيعي

seas_compute - يعمل بشكل طبيعي

يعمل بشكل طبيعي

SLURM Scheduler - FASSE - يعمل بشكل طبيعي

FASSE Compute Cluster (Holyoke) - يعمل بشكل طبيعي

يعمل بشكل طبيعي

Kempner Cluster CPU - يعمل بشكل طبيعي

Kempner Cluster GPU - يعمل بشكل طبيعي

يعمل بشكل طبيعي

FASSE login nodes - يعمل بشكل طبيعي

يعمل بشكل طبيعي

Cannon Open OnDemand - يعمل بشكل طبيعي

FASSE Open OnDemand - يعمل بشكل طبيعي

يعمل بشكل طبيعي

Netscratch (Global Scratch) - يعمل بشكل طبيعي

Home Directory Storage - Boston - يعمل بشكل طبيعي

Tape - (Tier 3) - يعمل بشكل طبيعي

Holylabs - يعمل بشكل طبيعي

Isilon Storage Holyoke (Tier 1) - يعمل بشكل طبيعي

Holystore01 (Tier 0) - يعمل بشكل طبيعي

HolyLFS04 (Tier 0) - يعمل بشكل طبيعي

HolyLFS05 (Tier 0) - يعمل بشكل طبيعي

HolyLFS06 (Tier 0) - يعمل بشكل طبيعي

Holyoke Tier 2 NFS - يعمل بشكل طبيعي

Holyoke Specialty Storage - يعمل بشكل طبيعي

holECS - يعمل بشكل طبيعي

Isilon Storage Boston (Tier 1) - يعمل بشكل طبيعي

BosLFS02 (Tier 0) - يعمل بشكل طبيعي

Boston Tier 2 NFS - يعمل بشكل طبيعي

CEPH Storage Boston (Tier 2) - يعمل بشكل طبيعي

Boston Specialty Storage - يعمل بشكل طبيعي

bosECS - يعمل بشكل طبيعي

Samba Cluster - يعمل بشكل طبيعي

Globus Data Transfer - يعمل بشكل طبيعي

سجل الإشعارات

عرض الحالة الحالية

يوليو 2026

يوليو6
FASRC monthly maintenance Monday July 6th, 2026 9am-1pm
مكتملصيانة4 ساعات
  • مكتمل
    يوليو 06, 2026 في 17:00UTC
    مكتمل
    يوليو 06, 2026 في 17:00UTC
    اكتملت الصيانة بنجاح
  • قيد التقدم
    يوليو 06, 2026 في 13:00UTC
    قيد التقدم
    يوليو 06, 2026 في 13:00UTC
    الصيانة جارية الآن
  • مخطط
    يونيو 26, 2026 في 13:45UTC
    مخطط
    يونيو 26, 2026 في 13:45UTC

    FASRC monthly maintenance will take place on July 6th 2026. Our maintenance tasks should be completed between 9am-1pm.

    Cannon cluster will be paused during this maintenance?: NO
    FASSE cluster will be paused during this maintenance?: NO

    NOTICES:

    • Friday July 3rd is a university holiday (independence Day observed)

    • Training: Upcoming training from FASRC and other sources can be found on our Training Calendar. at https://www.rc.fas.harvard.edu/upcoming-training/

    • Status Page: You can subscribe to our status to receive notifications of maintenance, incidents, and their resolution at https://status.rc.fas.harvard.edu/ (click Get Updates for options).

    • We'd love to hear success stories about your or your lab's use of FASRC. Submit your story here.

    MAINTENANCE TASKS

    • Domain controller replacement

      • Audience: Internal

      • Impact: None. End users should not see any impact.

    • Reboot drained nodes in error state

      • Audience: Cluster nodes with errors.

      • Impact: These nodes will have been drained already in preparation. No impact on jobs on the day and the affected nodes will return to service in their respective partitions after the maintenance period.

    • OOD/Open OnDemand reboots

      • Audience: All OOD users, reboot of the head nodes.

      • Impact: Running sessions will not be affected.

    • Login node reboots

      • Audience; All login node users.

      • Impact: Login nodes will reboot during the maintenance window.

    • Netscratch 90-day retention cleanup

      • Audience; All netscratch users

      • Impact: Files older than 90 days will be removed per our scratch policy. Please note that this cleanup can happen at any time, not just during maintenance.

    Thank you,
    FAS Research Computing
    https://docs.rc.fas.harvard.edu/
    https://www.rc.fas.harvard.edu/

يونيو 2026

يونيو15
2026 MGHPCC power downtime June 15-18, 2026
مكتملصيانة80 ساعة 15 دقيقة
  • مكتمل
    يونيو 18, 2026 في 21:15UTC
    مكتمل
    يونيو 18, 2026 في 21:15UTC

    The yearly power downtime at our Holyoke data center, MGHPCC, has completed.

    The clusters and storage are back online and login nodes and OOD nodes are now available.

    If you have an issue/need help, please send a ticket to rchelp@rc.fas.harvard.edu with details.

    IMPORTANT NOTE: Tomorrow, June 19th is a university holiday. FASRC staff will return Monday to address any lingering issues and any new tickets.

  • تحديث
    يونيو 18, 2026 في 20:45UTC
    تحديث
    يونيو 18, 2026 في 20:45UTC

    Power-up is nearly complete, but a delay earlier in the day has us slightly behind.

    New ETA is 6PM.

  • تحديث
    يونيو 18, 2026 في 12:16UTC
    تحديث
    يونيو 18, 2026 في 12:16UTC

    MGHPCC has completed their maintenance and restored power to the facility.

    FASRC will now begin the power-up process. Please be aware that this takes several hours.

    We will update this status once complete.

    NOTE: A reminder that tomorrow (Friday) is a university holiday.

  • قيد التقدم
    يونيو 15, 2026 في 13:00UTC
    قيد التقدم
    يونيو 15, 2026 في 13:00UTC
    Maintenance is now in progress
  • مخطط
    يونيو 15, 2026 في 13:00UTC
    مخطط
    يونيو 15, 2026 في 13:00UTC

    The yearly power downtime at our Holyoke data center, MGHPCC, has been scheduled by the facility. This year's power downtime will take place on Tuesday June  15th - 18th, 2025.  There will be no June monthly maintenance as a result.

    Since the facility will be powered down for two days this year, we will not be performing the usual maintenance tasks. 
    That said, networking and other key infrastructure will be doing maintenance.

    IMPORTANT NOTE: FASRC storage at both Holyoke and Boston will be affected and should not be expected to be available throughout the downtime. Please plan ahead accordingly.

    • Monday June 15th -  Power-down begins at 9AM

    • Tuesday June 16th - Power out at MGHPCC

    • Wednesday June 17th - Power out at MGHPCC

    • Thursday June 18th - Expected return to full service by 5PM

    • Friday June 19th - Please note that June 19th is a university holiday

     

    Monday June 15th -  Power-down begins at 9AM
Tuesday June 16th - Power out at MGHPCC
Wednesday June 17th - Power out at MGHPCC
Thursday June 18th - Expected return to full service by 5PM

    For more detailed information and follow-up, please see:
    https://www.rc.fas.harvard.edu/mghpcc-yearly-shutdown or this Status Page

مايو 2026

Cannon cluster down
تم الحلانقطاع كبير66 ساعة 24 دقيقة
  • تم الحل
    UTC
    تم الحل

    Slurm crashed on 4:30p on Friday due to a user running a large sacct query against the Slurm database. This caused the database host to run out of memory and crash the scheduler. To prevent this from reoccurring we are reducing the time range that users are permitted to query at one time to 7 days. Thus if you need to cover a month you would need to query in four 7 day increments.

    We do ask users to be judicious in their querying of the Slurm. Only ask for those fields that you require. Please also ensure any AI agents you have running limit their queries appropriately.

  • تم التحديد
    UTC
    تم التحديد

    To temporarily stabilize the situation, we have reduced the maximum query time for sacct and other Slurm commands to be 1 day. We have filed a ticket with SchedMD to further analyze the issue.

    The cluster is back up and the scheduler is accepting new jobs.

    We will continue to monitor for emergencies over the weekend, and resume in-depth troubleshooting on Monday.

  • قيد التحقيق
    UTC
    قيد التحقيق

    The Slurm scheduler is experiencing an error which is impacting jobs. The Cannon cluster will be inaccessible while we troubleshoot.

    We are currently investigating this incident.

السابق

مايو 2026 إلى يوليو 2026

التالي