Skip to main content
Your cluster is running and now it needs to change: more workers, a bigger instance type, a new Kubernetes version, a pause overnight, or a clean teardown. This page covers those tasks for every cluster Ankra builds on your own account - the Ankra Managed clusters in the create dialog, on Hetzner, OVHcloud, UpCloud, DigitalOcean, Scaleway, AWS EC2, Ankra Cloud, Proxmox VE and HPE Morpheus. Where a provider behaves differently, the table in each section says so, and the provider’s own guide carries the details.
Clusters whose control plane your cloud runs (EKS, GKE, AKS, DOKS, UKS, OVHcloud MKS, Kapsule) are Cloud Managed; their node pools, upgrades and deletion are on Managed Kubernetes.

Before you start

  • The cluster shows Online and no operation is running on it. Day-2 changes are queued one at a time; follow them under Operations or with ankra cluster operations list.
  • In the dashboard, most controls live under Nodes in the cluster sidebar (tabs Node groups, Control plane, Bastion & VMs) and under Settings → General. Nodes describes every control on those tabs.
  • On the CLI, most commands find the provider from the cluster, so they are the same everywhere. The rest take the provider name: hetzner, ovh, upcloud, digitalocean, scaleway, aws, proxmox or morpheus. Ankra Cloud clusters are operated from the dashboard and the API; the released CLI has no Ankra Cloud commands yet.
  • Every command has an API equivalent under /api/v1/clusters/{provider}/{cluster_id}/... - see the API reference. Writes answer 202 Accepted and run in the background; pass --wait on the CLI to block until one finishes.

Node groups

A node group is a set of workers with one instance type, a count, and optional labels and taints. Each group scales, changes type and carries its labels independently. A group holds 0 to 100 nodes; scaling a group to 0 keeps its definition and removes its servers. In the dashboard: open Nodes → Node groups. Each card has the group’s controls, Add node group opens the size picker with live prices, and Max out sizes a group to what your cloud account still has room for. On the CLI:
  • Instance type changes go one way. A group moves to an equal or larger type only; to get smaller nodes, add a new group with the smaller type and delete the old one. How the change is applied depends on the provider - most power each node off, resize it and power it on again, while AWS replaces each node - so check the provider guide before you resize a group that runs stateful workloads.
  • Labels and taints apply to every node in the group. --clear removes them all; a taint without an effect gets NoSchedule.
  • Removing nodes drains them first. A scale-down, a group delete and a replacement drain each node, honouring PodDisruptionBudgets.
  • Deleting a group removes all its servers. Workloads on them are evicted.
  • Zones. On OVHcloud 3-AZ regions (--availability-zone), UpCloud zone pools (--zone) and multi-zone AWS clusters, a group can be pinned to one zone - do that for groups that run zonal storage.
  • Autoscaling and first-boot scripts. Set a group’s autoscaling range with ankra cluster node-group autoscaling set - see Cluster Autoscaling. OVHcloud groups can carry a cloud-init document - see Node Group Cloud-init User Data.
  • A stopped cluster. Adding or changing a node group on a stopped cluster brings its infrastructure back first, so the cluster comes online. Use Start cluster when you want the full saved topology back.

Legacy worker scaling

Clusters also keep a single default worker pool that predates node groups. ankra cluster scale <cluster> <count> sets its size, and ankra cluster <provider> workers <cluster> shows it. Prefer node groups for anything new.

Control plane

In the dashboard: Nodes → Control plane shows the controller count and instance type, and says for each change whether it runs live or offline, or why it is refused right now.
  • Growing the count is live. Going from 1 to 3 controllers (3 to 5 on Proxmox VE) provisions the new controllers and joins them to the running cluster.
  • Reducing the count is offline. Stop the cluster (its state is kept and restored on start), apply the change, then start it.
  • The instance type changes live, one controller at a time, when the cluster has three or more controllers. A single controller is resized offline, because resizing it takes the Kubernetes API down while it reboots.
  • A cluster spread across several zones or hosts keeps at least three controllers.

Restart a node

Restart one node - a control plane, a worker or the bastion - without waiting for a reconciliation. The restart is a tracked operation, and workloads on the node are briefly unavailable while it reboots. In the dashboard: open Nodes, find the machine in the Machines table and choose Restart VM from its ⋯ menu. On the CLI: find the node ID, restart it, and follow the operation:
The node must be up with no restart already in flight. If a node never joined, ankra cluster <provider> nodes cloud-init-log <cluster> <node-id> shows its cloud-init status and the end of its log, read over the bastion (not on Proxmox VE or HPE Morpheus). You can also ask Ankra’s AI, for example “restart worker-2 on my-cluster”.

Bastion

The bastion is the machine Ankra and you reach the cluster through: every SSH hop into the private network goes via it. On most providers it carries no workload traffic, so workloads keep running while it is down, but Ankra cannot provision, scale, upgrade or reconcile the cluster until it is back. In the dashboard: Nodes → Bastion & VMs shows its state and the last health verdict, with Diagnose over SSH, Restart VM and Resize. The ⋯ menu on the bastion row in the Machines table also has Resize bastion.
A resize powers the bastion off, changes its type and powers it on again, so SSH access is interrupted briefly. Where the bastion is also the NAT for the nodes (AWS bastion_nat egress), node egress stops until it is back.

SSH access and keys

Settings → Access shows copy-paste SSH commands through the bastion for the cluster’s own addresses, and the SSH key credentials attached to the cluster. To change the keys:
The new keys are written to every node on the next reconciliation; resync forces it. The key Ankra itself uses is never removed. For everyday kubectl you need no SSH at all - see Accessing Clusters with kubectl.

Upgrade Kubernetes

In the dashboard: Settings → General shows the current Kubernetes version. Pick a target and click Upgrade. The same card has Auto upgrade patch versions, which moves the cluster to the newest patch of its current minor version on its own; minor versions are never upgraded automatically. On the CLI:
  • Nodes upgrade one at a time, control plane first, then workers. Each node is cordoned, drained honouring PodDisruptionBudgets, upgraded, and must be Ready at the target version before the next one starts.
  • An etcd snapshot is taken before the control plane upgrade. With external etcd, the dedicated etcd members are upgraded first, one at a time, each saving a snapshot.
  • A drain blocked by a PodDisruptionBudget stops the rollout; --force proceeds anyway.
  • One minor version at a time (1.33 to 1.34, not 1.33 to 1.35), and no downgrades.

Stop and start

Stopping releases the cluster’s compute and keeps its configuration, stacks and credentials in Ankra, which is how you park a cluster you do not need right now. Starting re-provisions it and reconciles it back to running. What a stop does to the servers depends on the provider: Stop cluster explains the snapshot, the pause mode and the restore in full. In the dashboard: Settings → General → Danger Zone → Stop cluster, or Stop cluster / Start cluster in the cluster’s ⋯ menu. They also run on a timetable with Power Schedules.
  • A stop never deletes the volumes your workloads provisioned through the CSI driver, forced or not; they keep billing while the cluster is stopped. --force also deletes the cluster’s load balancers, and never captures a snapshot.
  • Stop and start run in the background. A start is refused with 409 while a stop or terminate is still running.
  • While the cluster is stopped, ankra cluster <provider> nodes list still shows the saved machines that the next start re-provisions.

Terminate a cluster

Terminating deletes the cluster’s servers and the network Ankra created, then removes the cluster from Ankra. It cannot be undone. In the dashboard: Settings → General → Danger Zone → Terminate. The dialog lists the persistent volumes the teardown deletes, and the button stays disabled until you accept that.
  • Volumes are deleted only once you accept it. --yes skips the teardown confirmation but never the volume one. The API lists the volumes with GET .../deprovision-volumes and needs ?accept_volume_data_loss=true on the DELETE; without it the answer is 409 naming them.
  • --force deletes leftover load balancers and tolerates infrastructure that no longer answers. It never stands in for accepting the volume loss. On Proxmox VE it finishes the teardown when the host or jumphost is unreachable, leaving the VMs behind.
  • A stopped cluster can be terminated too: the volumes recorded at stop time are still known and named.
Persistent volumes covers every case, including volumes Ankra cannot prove belong to the cluster.

Choices fixed at create time

A few choices cannot be changed on a running cluster. They are made in the create wizard (the Kubernetes step) or in the create request, and they shape the upgrade and networking tasks above.
  • Distribution. kubeadm (the default everywhere: wizard, CLI and API) is upstream Kubernetes with containerd; its versions are plain tags such as v1.33.2. k3s is a single-binary distribution; its versions look like v1.33.2+k3s1.
  • CNI (cni). kubeadm clusters always run Cilium. k3s clusters choose flannel (the default), calico or cilium; AWS defaults to Cilium for both distributions. The CNI cannot be changed later.
  • CNI features (cni_features). Cilium offers kube_proxy_replacement, hubble and wireguard_encryption; Calico offers ebpf_dataplane; flannel takes none. On k3s, kube_proxy_replacement and ebpf_dataplane need a single control plane.
  • etcd topology (kubeadm only). stacked (the default) runs etcd on the control plane nodes. external runs it on 3 or 5 dedicated machines (etcd_node_count), sized by a provider-specific field:

Next steps