NVIDIA Open-Sources NodeWright to Manage GPU Clusters
NVIDIA has open-sourced NodeWright, a Kubernetes-native package manager designed to update host operating systems across GPU clusters without disrupting active AI training workloads.

NVIDIA has officially open-sourced NodeWright, a Kubernetes-native package manager designed to declaratively configure and update host operating systems across accelerated GPU clusters. Previously run internally at NVIDIA under the name Skyhook, the tool is now available under the Apache 2.0 license. NodeWright solves a major pain point for infrastructure engineers who manage GPU fleets, where traditional configuration tools like Ansible or Puppet often fail to account for active, non-interruptible AI training workloads.
To prevent workload disruption, the NodeWright operator orchestrates a strict sequence on each targeted node. It first cordons the node to prevent new workloads, waits for critical pods to finish, drains remaining pods while respecting PodDisruptionBudgets, applies the host-level packages, triggers any necessary reboots, and finally uncordons the node. This process allows administrators to execute host-level operations, such as kernel parameter tuning, CVE remediation, security agent installation, and crash dump configuration, without manual intervention or late-night maintenance windows.
For large-scale deployments, NodeWright introduces DeploymentPolicy resources. These allow operators to partition their fleet into compartments and roll out updates progressively. Practitioners can choose between three batch strategies: Fixed, which updates a constant number of nodes; Linear, which increases the batch size by a fixed delta; and Exponential, which multiplies the batch size to accelerate trusted rollouts. Built-in validation checks ensure that if a package fails to install correctly, the rollout halts immediately rather than cascading through the cluster.
NodeWright is part of NVIDIA DSX OS and integrates with other ecosystem tools. It works alongside the NVIDIA AI Cluster Runtime to apply host-level parts of version-locked recipes, as well as NVCRE for pre-workload validation and NVSentinel for runtime fault monitoring. It does not replace the NVIDIA GPU Operator or Network Operator but manages the host OS layer beneath them. Administrators can install NodeWright using Helm from its OCI registry.
This is our own summary of reporting by NVIDIA Developer Blog



