The Cluster Outlives Its Installer
An open-source toolkit that turns the pre-provisioned guide into code and keeps operating the cluster after day one.

On AHV, Nutanix Kubernetes Platform creates its own virtual machines. Everywhere else, on bare metal, on other hypervisors, or where another team owns VM creation, the supported path is pre-provisioned mode: the hosts already exist, and you prepare each one to the letter of the guide before the installer touches it. That preparation has a peculiar first hour, with the guide open in one window and a terminal in the other, trusting that the fifth worker got the treatment the first one did. In that mode the guide is a specification, and a specification is only as reliable as whoever is executing it at 6 PM.
The toolkit I released on October 9 turns that specification into code. NKP Preprovisioned Toolkit is Ansible and bash, with optional OpenTofu modules, and it installs NKP 2.18 on hosts that already exist, running Rocky Linux 9 or Ubuntu 24.04 LTS. I did not want another installer, though. An installer is finished when the dashboard comes up, and an environment is not: workers come and go, disks get recycled, hardening drifts. So the same code that builds the cluster keeps operating it, with gates that refuse to proceed, workers you can add and remove, and an audit you can run again.
It installs the layout the guide prescribes for Pro and Ultimate clusters on pre-provisioned infrastructure; what I can claim is what I have validated, and so far that is a lab on virtual machines, with a few defaults that suit a lab (a public installation password, SSH password authentication) and must change before production.
Six Phases and Two Gates
./deploy.sh install starts from a jump host, one, three or five control plane nodes and a set of workers on a single Layer 2 subnet. The jump host is the machine the NKP CLI runs from, and where the temporary bootstrap cluster lives in Docker. The run ends with a self-managed cluster: kube-vip keeps the Kubernetes API reachable on a virtual IP (the VIP) across the control planes, MetalLB hands out addresses to LoadBalancer services such as the dashboard, Rook Ceph provides storage, and Kommander is the management plane on top.
Six phases get there, and the first is a read-only preflight that checks the inventory, the operating system, sizing against the Pro and Ultimate thresholds, raw data disks on the workers, address overlaps and whether the VIP is actually free.
Two gates stop the run when a host has failed or is unreachable. The first sits after preflight, before anything is changed on any host. The second sits before cluster creation, so a cluster is never built with a node that failed preparation or the hardening audit. If I could keep only one, I would keep the second: a half-prepared node is cheap to fix before nkp create cluster runs and tedious to chase afterward.
Kommander is installed with --wait=false, which means the command returns long before the platform is ready. The final banner reports success only when the dashboard has an address and every HelmRelease, the object that tracks each platform application, is Ready. Otherwise, it prints how many applications are ready out of how many.
Disks by Identity
The guide is most specific about storage, and hand work drifts fastest there. NKP’s default Ceph configuration expects block-mode volumes, which the static local provisioner of a pre-provisioned cluster is not designed to hand out, so the guide prescribes host storage mode: Ceph takes the node disks that match a filter, and the placement of its storage daemons (OSDs, one per disk) keeps Rook off the control planes.
Each worker carries an OS disk, one raw disk of at least 40 GiB for Ceph, and four disks for local volumes, each larger than 100 GiB because Prometheus alone claims a 100 GiB volume. The toolkit generates the Ceph configuration, formats and mounts the volume disks under /mnt/disks/, and never partitions a disk.
That Ceph is reserved for the platform applications, and Kommander’s Rook runs without the CSI driver, so workload volumes come from the local volume provisioner: a whole disk per claim, bound to its worker, with nothing replicated, resized or snapshotted by the storage. Fast, and adequate for the platform and for workloads that replicate their own data; a stateful workload that expects failover from the disk needs a CSI driver added after the install, the platform’s or an agnostic one, and the toolkit installs none.
On one Proxmox worker in my lab the kernel named the Ceph disk sdc and a volume disk sdb. After a reboot, the same worker called the Ceph disk sdb. Kernel names follow detection order, not the hypervisor slot, and with letter-based defaults Ceph would have landed on the wrong disk. The toolkit accepts and recommends /dev/disk/by-id/ paths for ceph_osd_device and local_volume_devices, and the OpenTofu modules emit an inventory snippet with those paths already filled in.
BlueStore Labels Survive wipefs
Day-2 runs through Cluster API, the Kubernetes project that manages cluster machines as declarative objects, and NKP’s pre-provisioned provider for it: ./deploy.sh add-worker and ./deploy.sh remove-worker. Removal cleans the node as well, purging the Ceph OSD, wiping the disk and unmounting the volumes. The first time I ran a remove and add cycle with the real disk layout, wipefs alone turned out to be insufficient. Ceph 19 uses BlueStore as its on-disk format, and BlueStore writes copies of its label deeper into the disk, so the returning worker picked up the old OSD, which then would not start.
The cleanup now zeroes every copy of the label, and the cycle passes on all four OS profiles with a new OSD and Ceph back to HEALTH_OK. Anyone recycling Ceph disks by hand will meet this eventually, or rather will meet its symptom, an OSD that refuses to start on a disk that looks clean.
Hardening Before the Cluster Exists
With a -cis profile, hardening runs as its own phase between node preparation and cluster creation, on every host including the jump host. It covers nine groups of controls aligned to CIS Level 1: unused kernel modules, sysctl, the auditd, chrony and cron services, auditd rules, file permissions, account defaults, banners, sshd and the host firewall. The deviations Kubernetes needs are applied only to cluster nodes, so the jump host gets the stricter treatment. A failed audit stops the install at the second gate, which means a cluster is born on hardened nodes or it is not born at all.
Hardening and audit read the same definitions, so the audit verifies exactly what the hardening sets, and ./deploy.sh cis-audit can be run again months later to catch drift. The host firewall is the one group left unapplied on cluster nodes by default, because the NKP guide asks for firewalld to be disabled there; the audit reports it as “NOT APPLIED (by design)” instead of hiding it or failing on it. An opt-in restricts traffic to the node subnet, the Pod CIDR and trusted CIDRs.
All of this works at the operating system level, so it does not care whether the host is a VM or a physical server, and bare metal is where a hardened pre-provisioned node makes the most sense to me. It is a subset of the benchmark, though. A passing audit is a statement about those controls only, and compliance remains a job for CIS-CAT or OpenSCAP.
The License Follows the Cluster API Object
The toolkit never applies or modifies the NKP license, because automation and entitlement are separate matters and the second one stays between you and Nutanix. At the end of the run it prints the Cluster UUID the license is requested for, and the steps to activate it in the dashboard under Global, Settings, Licensing. Pre-provisioned clusters run on Pro or Ultimate, since Starter is supported only on Nutanix infrastructure.
The UUID that matters is the UID of the Cluster API Cluster object created by the CLI, the one Nutanix KB 18100 asks for:
kubectl get clusters.cluster.x-k8s.io -A -o jsonpath='{.items[0].metadata.uid}'
The UID of the kube-system namespace, which NKP uses for monitoring, is easy to mistake for it. The KB says a key issued with the wrong UUID still licenses the cluster, with Pulse as the part that may stop working, and a CI check now keeps the code from reading the wrong identifier again.
The name is the strict one: the key I requested with the wrong cluster name came back rejected by Kommander’s admission webhook, which could not decrypt it for this cluster: “License key is not valid for this cluster” and an HTTP 406. A cluster rebuilt from scratch gets a new UUID and needs a new key, which is worth knowing before you plan a lab around a single license request.
Four Hypervisors Behind One Contract
The Ansible side does not know which hypervisor created the machines. Four OpenTofu modules cover Proxmox VE, AHV through Prism Central, vSphere, and standalone ESXi, the last one over SSH with vmkfstools and vim-cmd so that it works with the free license too. All four implement the same contract: the same cloud image, cloud-init with static addressing, the guide’s disk layout, and the inventory snippet as output.
./deploy.sh tofu-verify tests that contract in place of a full install, which would take hours per hypervisor. It checks disks and by-id names, the interface name and cloud-init completion, then runs a Layer 2 test of the VIP: it raises the address on one control plane with ip addr add, pings it from a worker, and removes it. On VMware, vSwitch policies such as forged transmits, MAC changes and promiscuous mode decide whether kube-vip and MetalLB can do their job, and it is better to learn that before the install than in the middle of it.
Validated, and Not Yet Validated
Each OS profile was validated with a complete install in the lab (with the lab shortcuts, Ceph on a loop device; the one run with the guide’s disk layout, ubuntu-cis on Proxmox, closed in 53 minutes), ending with every Kommander HelmRelease Ready, Ceph at HEALTH_OK with 4 OSDs and no pending PVC.
| Profile | Full install |
|---|---|
ubuntu |
62 min |
ubuntu-cis |
67 min |
rocky |
155 min |
rocky-cis |
164 min |
The gap is almost entirely in the control plane. NKP provisions those three nodes one after another on both distributions, but on Ubuntu each one took 5 to 7 minutes and on Rocky 20 to 30. Same hypervisor, same VM sizing, same disk layout: the difference is inside the guest, in the package installation step, and I have not profiled it on Rocky yet. After the -cis installs, the audit closed at 65 PASS and 0 FAIL on both distributions.
On Nutanix, ESXi and vSphere the contract is validated and a full NKP install on those VMs is in the public backlog. Bare metal, hosts below the Pro and Ultimate thresholds, the opt-in host firewall on cluster nodes and SSH on a port other than 22 are not validated either. I would rather publish that list than have someone discover it during a customer proof of concept.
The project is independent, released under Apache-2.0, and not affiliated with or supported by Nutanix. It ships no Nutanix software: you download the NKP CLI bundle from the Support Portal with your own entitlement, and the toolkit verifies it, as downloaded, against the SHA-256 the portal publishes. The bundle and not the extracted binary: the 2.18.0 bundle on the portal today holds a different build from the one I downloaded in July, same size to the byte, same nkp version, different Go build ID. A release can be republished under the same version number, so the hash of the executable is not a stable reference and the hash of the bundle is.
A Guide You Can Diff
Once the prerequisites live in code they can be linted, tested and argued with in a pull request, and a preflight can say yes or no before a single host is modified. That matters most on someone else’s infrastructure, where the hosts arrive from another team and the first question is whether they match what was asked for. It keeps mattering after the install, because the same inventory and the same commands serve the cluster later: adding a worker, removing one and leaving its disks clean, running the audit again.
What Comes Next
Upgrade is the lifecycle operation still missing, for a plain reason: the toolkit targets NKP 2.18, and today there is no later release to upgrade to, so there is nothing I could validate in the lab. It is the next operation I want to cover when one ships.
The pieces it needs are already there: a self-managed cluster driven through Cluster API, an inventory that stays the source of truth after day one, a preflight that can check the hosts again before anything moves, and the CLI installed one directory per release on the jump host, because an upgrade runs with the CLI of the target release. That is the part I expect to make an administrator’s life simpler, more than the first install does. With a CSI driver of your own and that upgrade path, the same inventory serves a production cluster; what changes is only what I can claim to have verified.
The inventory is also why none of the OpenTofu modules are mandatory. It is a list of hosts that answer on SSH, and the playbooks have no opinion on how they got there.
Nothing in the design stops you from running three control plane VMs on a hypervisor and joining physical servers as workers, as long as everything shares one Layer 2 subnet, for reasons that range from hardware you already own to workloads that want the whole machine.
I have not validated that mixed layout yet, and it is the one I am most curious about, because it is the case pre-provisioned mode exists for.