Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

CRITICAL: Prolonged Cluster Partition / Split-Brain on 3-Node HPE VME 8 Cluster

Дата публикации: 24-09-2026 04:03:00

Hi,Title: CRITICAL: Prolonged Cluster Partition / Split-Brain on 3-Node HPE VME 8 Cluster (Corosync failure, GFS2 deadlock)Environment: HPE VM Essentials (VME) 8.x (3-Node Cluster)Summary of Issue: We are experiencing a severe, prolonged cluster partition/split-brain scenario on our 3-node VME 8 cluster that started on Sept 19 (4-5 days ago). I have explicitly halted all manual SSH troubleshooting due to an extreme risk of unrecoverable GFS2 data corruption.Technical Observations:Node 1 (hpevmenode1): The corosync service exited/failed completely on Sept 19.Node 2 (hpevmenode2): Corosync reports Members[1]: 2 (only seeing itself: 1 of 3 nodes) and flags the node as a "non-primary component and will NOT provide any services." Without a 2/3 node majority, Node 2 lost quorum on Sept 19 and remains isolated.Node 2 CPU Anomaly: System is detecting a massive CPU load of 15,348 continuously every 30 seconds. This strongly indicates massive I/O wait (D-state processes) waiting indefinitely on broken DLM/GFS2 locks.Cluster Impact: The cluster lost its 2-of-3 majority quorum. Consequently, DLM/GFS2 locking mechanisms have been deadlocked/malfunctioning for 4-5 days.Why Manual Troubleshooting Was Halted:The cluster hosts highly critical VMs, including the VME Manager.Executing recovery commands (e.g., restarting corosync, forcing quorum) in a partitioned 3-node state risks split-brain writes. This will likely cause permanent, unrecoverable data corruption on the GFS2 shared storage.The massive CPU load on Node 2 indicates runaway/stuck I/O processes that make standard node fencing highly unpredictable right now.Questions for the Community / HPE Support: Before we attempt any node fencing, force-quorum, or cluster merge:Has anyone encountered a prolonged 3-node cluster partition with DLM deadlocks of this magnitude on VME 8?What is the safest, HPE-validated recovery sequence to safely resync Corosync across the 3 nodes and clear GFS2 locks without corrupting shared storage metadata?

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1Re: Configure GFS in Single Node Instead of 3 minimum : HPE VME022.2323-09-2026
2Increase VM density with HPE Morpheus Software memory overcommitment022.7227-08-2026
3Mastering Cluster Quorum Management in Hyper-V Clusters: Ensuring Resilience and Availability011.1902-02-2024
4第7回|HPE Morpheus VM EssentialsにおけるVM運用管理の考え方032.5425-08-2026
5HPE Throws VM Users A Lifeline, Unifying Containers And VM Management In Cloud Stack08.6213-05-2026
6Storage Failures in Hyper-V Cluster Management013.9104-02-2024
7Re: Infoblox installed on Morpheus KVM is not starting08.0623-09-2026
8MariaDB’s parallel replication to catch up013.2309-04-2024
9MySQL + Dynimize: 3.6 Million Queries per Second on a Single VM09.3622-09-2020
10How to Speed Up Re Sync of Dropped Percona Xtradb Cluster Node04.4424-02-2021

Классификация: . Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 13.74. Источник: community.hpe.com.