When operating Kubernetes at the scale of tens of thousands of clusters on Amazon EKS, hardware failures, though individually rare, become a constant operational reality. Issues ranging from GPU connectivity disruptions to container runtime freezes and disappearing network interfaces occur multiple times daily across the fleet. Historically, the response has been a manual, time-consuming process: an operator intervenes, diagnoses the problem via dashboards, connects to the affected node via SSH, isolates it, drains its workloads, terminates the instance, and awaits a replacement. This human-paced toil can lead to prolonged periods of degraded service, particularly when failures occur during off-hours. To address this critical gap, AWS developed the EKS Node Monitoring Agent, an open-source tool designed to automate the detection of node failures and signal these issues to Kubernetes’ auto-scaling and node management components, such as Karpenter, for prompt resolution. This agent is a key component within a broader system aimed at achieving fully automated Kubernetes cluster infrastructure management.
The advent of Amazon EKS Auto Mode signifies a significant leap towards this automation, encompassing compute provisioning, scaling, networking, storage, OS patching, and security hardening. This allows teams to shift their focus from the complexities of cluster operations to application development. EKS Auto Mode dynamically selects optimal EC2 instances, including high-performance GPU families like P5, P6, and G6, scales resources based on workload demand, consolidates underutilized nodes, and ensures the operating system remains patched. Crucially, automatic node repair is a default feature of EKS Auto Mode. This integrated approach handles detection, severity classification, and Karpenter-driven node replacement out-of-the-box, requiring no additional installation, configuration, or custom repair policies from users.
This article delves into the journey of building this automatic node repair system, the strategic design decisions that shaped its architecture, and the hard-won lessons derived from operating it at a demanding GPU scale.
Six Lessons from Building Self-Healing Kubernetes Nodes at Scale
After extensive operation across thousands of clusters, the core learnings from this endeavor can be distilled into a concise set of principles. These lessons are not unique to AWS’s implementation; similar patterns emerge in other node problem detection systems like Node Problem Detector (NPD), NVSentinel, Azure Kubernetes Service (AKS) Periscope, and Google Kubernetes Engine’s (GKE) auto-repair features. They represent a collective body of knowledge that should serve as a foundational checklist for anyone building or managing resilient Kubernetes infrastructure. The open-source EKS Node Monitoring Agent repository reflects these lessons, from the stability guarantees of its reason codes API to the specific jitter implementations developed to mitigate GPU workload interference.
Node Health Detection in Kubernetes: Unforeseen Traps
At its core, every node health agent within the Kubernetes ecosystem performs a similar translation: it takes noisy, low-level signals from a physical or virtual machine and transforms them into structured Kubernetes primitives, primarily NodeConditions, Kubernetes Events, and occasionally Custom Resource Definitions (CRDs). While the output appears deceptively simple—a NodeCondition comprises a type, a status (True, False, or Unknown), a reason code, and a descriptive message—the underlying decisions made during this translation process are critical and can profoundly impact whether a repair action proves beneficial or detrimental.
Reason Codes as a Public API Contract: A pivotal lesson learned through direct experience is that reason codes function as a public API. In version 1.6.2 of the EKS Node Monitoring Agent, a change was made to reclassify NvidiaDeviceCountMismatch from a "Warning" severity to "Fatal." The technical rationale was sound: a GPU that detaches from the PCIe bus typically does not recover without a node reboot or replacement. Leaving this as a "Warning" would lead to GPU workloads being scheduled onto compromised nodes, wasting valuable accelerator resources. However, this seemingly straightforward fix caused downstream automation to break. Customers had repair configurations keyed to the old severity level, and dashboards filtering for "Warning" no longer displayed the fault. Conversely, automation that only acted on "Fatal" conditions began unexpectedly draining nodes. This experience underscored the importance of treating every reason code addition as a new feature and any rename or severity change as a breaking change, necessitating careful communication and migration strategies for users.
The "Absent" Signal and Health Status: A crucial design decision revolved around how to represent the status of a disabled monitor. When per-monitor configurability was introduced in version 1.6.0, the team had to decide on the output for a monitor that was intentionally turned off. The options were to report "True" (falsely indicating health), "Unknown" (potentially triggering repair depending on downstream logic), or to omit the condition entirely. The safest and most logical choice was to emit nothing. This principle, while seemingly obvious in retrospect, is not universally implemented. Node Problem Detector, for instance, achieves a similar outcome through compile-time disabling via build tags. NVSentinel delegates this to operator-authored CEL rules. The Kubernetes specification defines "Unknown" status, but if repair automation interprets "Unknown" as actionable, it can lead to unnecessary node replacements. The EKS Node Monitoring Agent’s hard contract is to emit no condition when a monitor is disabled, ensuring clarity and predictable behavior.
Detection Latency Bounded by the Source: An initial communication error involved stating, "We detect kernel panics within 30 seconds." This was factually incorrect. While the agent’s internal detection time was indeed under 30 seconds, the actual bottleneck was the flush cadence of journald, the system’s logging daemon. If journald took 45 seconds to write a kernel panic to disk, the agent’s rapid detection was rendered irrelevant by the source’s latency. For GPU faults, the situation is more nuanced, involving multiple detection paths with varying latencies. Critical faults, such as double-bit ECC errors, XID errors, NVLink failures, page retirements, and thermal or power violations, are detected through DCGM’s (Data Center GPU Manager) push-based policy violation channel. DCGM notifies the agent immediately upon detecting a violation, achieving near-instantaneous detection (sub-second in practice) without polling intervals. A separate path monitors NVSwitch fabric health, Fabric Manager status, and clock-throttle reasons using a 5-minute field-value window. While this window dictates the minimum detection time for these specific metrics, it does not apply to the critical GPU faults that trigger immediate automatic repair. The critical takeaway is that customer-facing Service Level Objectives (SLOs) must account for the latency of the data source, and different signal paths within the same subsystem can have vastly different minimum latency floors.
Two Severities, One Switch: Automating Node Replacement Decisions
In addition to the kubelet-reported DiskPressure, MemoryPressure, and PIDPressure conditions, the EKS Node Monitoring Agent introduces five new conditions covering critical domains beyond the kubelet‘s purview: kernel health, container runtime, networking, storage, and accelerated hardware. Each detection is assigned one of two severities: "Condition" or "Event." The severity acts as the primary switch determining whether the automatic repair cycle is initiated.
Condition Severity: The Terminal Fault Trigger: When a condition is classified with "Condition" severity, it signifies a terminal fault. This action flips the corresponding NodeCondition to False, making the node immediately eligible for automatic repair. Examples include GPU device-count mismatches, critical XID and double-bit ECC errors, NVLink and NVSwitch fabric failures, a missing Fabric Manager, and Neuron DMA and HBM uncorrectable errors. On the networking and runtime fronts, this severity is assigned to issues like the VPC CNI process being down, IPAMD being unable to reach the API server, fork failures due to PID exhaustion, and pods becoming stuck in a terminating state behind a broken container runtime. These are faults that will not self-resolve. For GPU nodes, a single degraded accelerator can corrupt training checkpoints or result in thousands of dollars in wasted compute capacity per hour, necessitating swift intervention.
Event Severity: Informational Signals of Imminent Trouble: "Event" severity, conversely, is purely informational. It generates a Kubernetes Event, leaving the NodeCondition status as True. This approach provides operators with visibility into developing issues without triggering disruptive repair actions. Examples include bandwidth ceilings, connection-tracking limits, Amazon Elastic Block Store (Amazon EBS) IOPS throttling, I/O delays, filesystem fragmentation, clock drift, liveness and readiness probe failures, kube-proxy anomalies, GPU thermal and power warnings, PCIe link degradation, and page-retirement thresholds. These signals indicate trouble brewing before it escalates to a terminal state.
Misclassifying severity in either direction carries significant consequences. An overly aggressive approach can lead to the termination of healthy nodes and unnecessary workload displacement. Conversely, an overly conservative stance can allow degraded nodes to continue serving traffic for extended periods, with a GPU exhibiting memory bank degradation potentially corrupting critical training checkpoints. The guiding principle for classification is straightforward: if a failure is deterministic and infrastructure-owned—meaning hardware has physically failed, firmware has crashed, or a physical link has gone down—it triggers a replacement. If the signal could be application-induced or transient, it remains informational. The system is designed to avoid terminating a healthy node due to a misbehaving pod saturating a resource.
The DiskPressure, MemoryPressure, and PIDPressure conditions serve as canonical examples of this principle. Every major auto-repair system, including GKE, AKS, and NPD, has independently converged on the same answer: do not automatically replace nodes based solely on these conditions. These are workload-driven issues, not fundamental node-level faults. Replacing the node merely relocates the problematic workload to a new machine, where it is likely to consume resources excessively again. The appropriate response for these scenarios is kubelet-level pod eviction, not node replacement. Establishing this boundary early and communicating it clearly is paramount for any team building a node-health system.
The Agent That Hurt What It Was Protecting: GPU Workload Interference
The most challenging lesson emerged from a customer running large-scale distributed GPU training workloads. Their job relied on NCCL (NVIDIA Collective Communications Library) collectives across hundreds of GPU nodes, where every node in a communication group must complete its step before any can proceed. Consequently, a single slow node effectively halts progress for the entire group.
The customer discovered that the EKS Node Monitoring Agent itself was intermittently causing slowdowns. The agent’s monitors operated as independent goroutines, and when their polling intervals coincidentally aligned, dozens of these goroutines would wake up simultaneously, creating a burst of activity across multiple CPU cores. While this might be imperceptible on a general-purpose web service, in a distributed training job, even microseconds of jitter on one node can cascade across the entire GPU cluster, leading to measurable throughput degradation.
The customer’s decision to disable the EKS Node Monitoring Agent entirely resulted in an immediate performance improvement. This was the worst possible outcome for AWS: a health agent that actively interferes with the workload it is intended to protect is worse than having no agent at all.
The solution, once the problem was understood, was surprisingly straightforward. A startup jitter was introduced to each monitor’s polling interval. Each goroutine now delays its first execution by a random offset (up to 20% of its base interval), effectively staggering wake times and preventing simultaneous execution on boot. System calls that accessed /proc on every poll were cached. Handlers that shared an interval were consolidated into a single sequential work queue, reducing the overall goroutine count for monitors that did not require dedicated threads. The outcome was an agent whose CPU profile became flat and predictable, rather than bursty.
This experience generalized into a crucial principle: if a health agent runs on the same host as the workload, its resource consumption pattern is as critical as its total consumption. A process utilizing 0.5% CPU spread evenly across time is generally invisible. However, a process consuming the same 0.5% CPU in concentrated bursts can disrupt latency-sensitive distributed GPU workloads in ways that manifest as lost training time rather than simple CPU alarms.
This reinforces the importance of per-monitor configurability. Not every monitor is relevant to every workload. A dedicated GPU training cluster with a single pod per node and minimal pod churn does not require monitoring of IPAMD or environment scanning. The ability to disable individual monitors allows customers to retain the necessary health coverage without incurring the overhead of features they do not need.
The Repair Cycle: Seamless Automation with Karpenter
Karpenter, the compute controller responsible for provisioning and scaling EKS Auto Mode nodes, plays a central role in the automated repair process. By already managing the lifecycle of every node it launches, integrating NodeCondition consumption for repair is a natural extension of its responsibilities. This approach eliminates the need for a separate repair backend, sidecar controllers, or complex webhook chains. The same system that creates the node is also responsible for replacing it when necessary.
Karpenter’s AWS cloud provider integration allows for the definition of repair policies. Each policy pairs a specific NodeCondition type with a status that triggers node replacement. These policies incorporate toleration windows designed to prevent premature reactions to transient glitches. For instance, a policy might specify that a node will only be considered for replacement if a critical condition persists for at least 10 minutes.
The automated repair workflow unfolds as follows:
- Detection: The EKS Node Monitoring Agent detects a critical hardware or software fault.
- Condition Update: The agent updates the corresponding NodeCondition to
Falsewith an appropriate reason code. - Karpenter Notification: Karpenter, which continuously watches NodeConditions, identifies the
Falsestatus. - Toleration Window: Karpenter waits for the configured toleration window to expire, ensuring the issue is not transient.
- Replacement Provisioning: Upon expiration of the toleration window, Karpenter initiates the provisioning of a new, healthy replacement node.
- Node Termination: The unhealthy node is cordoned and drained of its workloads before being terminated.
- Workload Resumption: Once the replacement node is provisioned and registered, workloads are rescheduled onto it.
In testing, the complete cycle from fault injection to the replacement node running workloads has been observed to take under 12 minutes. Detection of critical GPU faults, utilizing DCGM’s push-based policy channel, occurs in under a second. This is followed by the 10-minute toleration period, and then approximately 90 seconds for the replacement node to launch and register.
The Surprise: Detection and Diagnosis Are Distinct Problems
A significant realization during the development process was the distinction between node failure detection and subsequent diagnosis. The auto-repair system is designed to handle the common scenario: a broken node is replaced, and the workload continues to run with minimal interruption. However, answering the question, "Why did that node fail?" is a fundamentally different problem. Initially, AWS attempted to integrate diagnostic capabilities directly into the detection path, which proved to be a misstep.
Detection’s primary function is to answer, "Is this node healthy?" It must operate continuously with minimal overhead and react extremely rapidly; a condition flip that takes five minutes to produce translates to five minutes of degraded workload performance. Diagnosis, on the other hand, aims to answer, "What went wrong?" This requires collecting detailed artifacts such as complete journald output, container runtime state, network configuration, dmesg logs, and GPU driver logs. In testing, this artifact collection typically completes within about 7 seconds and generates a compressed log bundle. Incorporating this comprehensive data collection into the high-speed detection path would have inevitably slowed down the system’s reaction time, compromising the very aspect customers value most.
The solution was to architect detection and diagnosis as separate concerns, albeit sharing the same agent binary. The NodeDiagnostic CRD provides a mechanism to request a full log bundle from any node via kubectl, eliminating the need for SSH access. This is particularly crucial in EKS Auto Mode, where nodes are managed Amazon EC2 instances designed without direct shell access by default, making this the sole method for investigating GPU failures or other node-level faults post-event.
The user experience is streamlined: a single command, kubectl ekslogs <node-name>, initiates the process. This command creates a NodeDiagnostic resource. The agent on the target node detects this resource via a watch, collects the relevant system state into a compressed tarball, and stores it temporarily for a 10-minute window. The ekslogs plugin then downloads this bundle through the kubelet‘s Node Log Query API (KEP-2258, now generally available in Kubernetes 1.36). This entire process circumvents the need for SSH, security groups, or key pair management.
This separation of concerns offers several advantages: detection performance is not degraded by evidence collection; diagnosis does not need to be perpetually active, thus conserving node resources; and it allows for the diagnosis of a node that auto-repair has flagged but has not yet terminated. The 10-minute window for accelerated hardware faults provides sufficient time to retrieve critical logs before the node is permanently removed.
Implications for EKS Users
For users operating on Amazon EKS Auto Mode, all these automated features are enabled by default. Auto Mode comprehensively manages the entire cluster infrastructure—compute, networking, storage, patching, and security hardening—allowing teams to concentrate on application development. The Node Monitoring Agent is integrated as a systemd service within the node image, not as a user-managed DaemonSet. Karpenter seamlessly consumes its NodeConditions as part of its existing compute lifecycle management. The kubectl ekslogs command provides diagnostic access without requiring SSH. There are no components to install, configure, or operate. For GPU workloads, this translates to automatic monitoring, classification, and replacement of expensive accelerator nodes without any manual operator intervention.
For those using managed node groups or self-managed Karpenter, the same automated repair loop can be established. This involves installing the Node Monitoring Agent as an EKS add-on and opting each node group into auto-repair. The underlying architecture remains identical, but it requires a deliberate assembly of the components rather than being pre-packaged.
The EKS Node Monitoring Agent is available as an Apache 2.0 licensed open-source project on GitHub at github.com/aws/eks-node-monitoring-agent. The failure modes encountered and the solutions developed from operating this system at scale are fed back into the project, benefiting anyone who uses it. For those building their own node-health systems or utilizing the EKS Node Monitoring Agent and encountering edge cases, contributing to the project is encouraged. Engagement with the EKS public roadmap is also welcomed for discussions on future improvements and integrations.
