FAS Research Computing - 状态页面

HolyLFS06 (Tier 0) 目前性能下降

Status page for the Harvard FAS Research Computing cluster and other resources.

Cluster Utilization (VPN and FASRC login required): Cannon | FASSE


Please scroll down to see details on any Incidents or maintenance notices.
Monthly maintenance occurs on the first Monday of the month (except holidays).

GETTING HELP
Documentation: https://docs.rc.fas.harvard.edu | Account Portal https://portal.rc.fas.harvard.edu
Email: rchelp@rc.fas.harvard.edu | Support Hours


The colors shown in the bars below were chosen to increase visibility for color-blind visitors.
For higher contrast, switch to light mode at the bottom of this page if the background is dark and colors are muted.

holylfs06 degraded
  • 持续监控中
    UTC
    持续监控中

    holylfs06 has been experiencing ongoing performance and latency issues. This has impacted workflows and made it difficult for jobs to effectively complete. Our engineers have been troubleshooting the underlying causes, and identified that the general load on holylfs06 is greater than what the filesystem can process and deliver. It is not the misuse of any particular job or group, but rather the combination of all jobs and groups. In a shared filesystem, the high load ends up impacting all users and leading to a unresponsive server.

    A more technical explanation: Jobs are reading very large volumes of data, keeping the storage arrays continuously busy with big requests. The ldiskfs journal has to make a small synchronous write every time an object is created or destroyed, and it sits on the same disks as the data. A journal bound operation that should take ~1ms is taking ~1s, with some taking over 12s. OSTs can't open new transactions until the journal commits, so service threads pile up meaning hundreds are stuck in uninterruptible wait, some for 15+ minutes. The node comes up as unresponsive, and times out. This results in the slow/stuck filesystem that many of you have experienced.

    In order to mitigate this, we have been resetting holylfs06 servers and rebooting when possible, but this is not a sustainable solution. We are asking all groups, particularly those with large sequential reads and small-file workloads to shift their workflow to netscratch if possible.

    A visual flowchart of an optimal workflow is depicted here: https://docs.rc.fas.harvard.edu/kb/data-storage-workflow-rdm/#Data_Storage_Workflow

    We also have recommendations on job efficiency and best practices to be kinder to fileystems here: https://docs.rc.fas.harvard.edu/kb/job-efficiency-and-optimization-best-practices/

    A long term solution will be the upcoming Compute Storage which uses a different hardware (nvme) than holylfs06 (Lustre). We expect to migrate your holylfs06 data to Compute Storage in the coming months, and will send out additional communication at that time.

    Please reach out to us at rchelp@rc.fas.harvard.edu if your group needs additional help adjusting your workflow.

    Thank you again for your understanding.

  • 调查中
    UTC
    调查中

    holylfs06 is in a very degraded or stuck state.

    Expect degraded to no access until the system is back up.



运行正常

SLURM Scheduler - Cannon - 运行正常

Cannon Compute Cluster (Holyoke) - 运行正常

Boston Compute Nodes - 运行正常

GPU nodes (Holyoke) - 运行正常

seas_compute - 运行正常

运行正常

SLURM Scheduler - FASSE - 运行正常

FASSE Compute Cluster (Holyoke) - 运行正常

运行正常

Kempner Cluster CPU - 运行正常

Kempner Cluster GPU - 运行正常

运行正常

FASSE login nodes - 运行正常

运行正常

Cannon Open OnDemand - 运行正常

FASSE Open OnDemand - 运行正常

性能下降

Netscratch (Global Scratch) - 运行正常

Home Directory Storage - Boston - 运行正常

Tape - (Tier 3) - 运行正常

Holylabs - 运行正常

Isilon Storage Holyoke (Tier 1) - 运行正常

Holystore01 (Tier 0) - 运行正常

HolyLFS04 (Tier 0) - 运行正常

HolyLFS05 (Tier 0) - 运行正常

HolyLFS06 (Tier 0) - 性能下降

Holyoke Tier 2 NFS - 运行正常

Holyoke Specialty Storage - 运行正常

holECS - 运行正常

Isilon Storage Boston (Tier 1) - 运行正常

BosLFS02 (Tier 0) - 运行正常

Boston Tier 2 NFS - 运行正常

CEPH Storage Boston (Tier 2) - 运行正常

Boston Specialty Storage - 运行正常

bosECS - 运行正常

Samba Cluster - 运行正常

Globus Data Transfer - 运行正常

最近的事件

查看历史记录