Clusters whose control plane your cloud runs (EKS, GKE, AKS, DOKS, UKS, OVHcloud MKS, Kapsule) are Cloud Managed; their node pools, upgrades and deletion are on Managed Kubernetes.
Before you start
- The cluster shows Online and no operation is running on it. Day-2 changes are queued one at a time; follow them under Operations or with
ankra cluster operations list. - In the dashboard, most controls live under Nodes in the cluster sidebar (tabs Node groups, Control plane, Bastion & VMs) and under Settings → General. Nodes describes every control on those tabs.
- On the CLI, most commands find the provider from the cluster, so they are the same everywhere. The rest take the provider name:
hetzner,ovh,upcloud,digitalocean,scaleway,aws,proxmoxormorpheus. Ankra Cloud clusters are operated from the dashboard and the API; the released CLI has no Ankra Cloud commands yet. - Every command has an API equivalent under
/api/v1/clusters/{provider}/{cluster_id}/...- see the API reference. Writes answer202 Acceptedand run in the background; pass--waiton the CLI to block until one finishes.
Node groups
A node group is a set of workers with one instance type, a count, and optional labels and taints. Each group scales, changes type and carries its labels independently. A group holds 0 to 100 nodes; scaling a group to 0 keeps its definition and removes its servers. In the dashboard: open Nodes → Node groups. Each card has the group’s controls, Add node group opens the size picker with live prices, and Max out sizes a group to what your cloud account still has room for. On the CLI:- Instance type changes go one way. A group moves to an equal or larger type only; to get smaller nodes, add a new group with the smaller type and delete the old one. How the change is applied depends on the provider - most power each node off, resize it and power it on again, while AWS replaces each node - so check the provider guide before you resize a group that runs stateful workloads.
- Labels and taints apply to every node in the group.
--clearremoves them all; a taint without an effect getsNoSchedule. - Removing nodes drains them first. A scale-down, a group delete and a replacement drain each node, honouring PodDisruptionBudgets.
- Deleting a group removes all its servers. Workloads on them are evicted.
- Zones. On OVHcloud 3-AZ regions (
--availability-zone), UpCloud zone pools (--zone) and multi-zone AWS clusters, a group can be pinned to one zone - do that for groups that run zonal storage. - Autoscaling and first-boot scripts. Set a group’s autoscaling range with
ankra cluster node-group autoscaling set- see Cluster Autoscaling. OVHcloud groups can carry a cloud-init document - see Node Group Cloud-init User Data. - A stopped cluster. Adding or changing a node group on a stopped cluster brings its infrastructure back first, so the cluster comes online. Use Start cluster when you want the full saved topology back.
Legacy worker scaling
Clusters also keep a single default worker pool that predates node groups.ankra cluster scale <cluster> <count> sets its size, and ankra cluster <provider> workers <cluster> shows it. Prefer node groups for anything new.
Control plane
In the dashboard: Nodes → Control plane shows the controller count and instance type, and says for each change whether it runs live or offline, or why it is refused right now.- Growing the count is live. Going from 1 to 3 controllers (3 to 5 on Proxmox VE) provisions the new controllers and joins them to the running cluster.
- Reducing the count is offline. Stop the cluster (its state is kept and restored on start), apply the change, then start it.
- The instance type changes live, one controller at a time, when the cluster has three or more controllers. A single controller is resized offline, because resizing it takes the Kubernetes API down while it reboots.
- A cluster spread across several zones or hosts keeps at least three controllers.
Restart a node
Restart one node - a control plane, a worker or the bastion - without waiting for a reconciliation. The restart is a tracked operation, and workloads on the node are briefly unavailable while it reboots. In the dashboard: open Nodes, find the machine in the Machines table and choose Restart VM from its ⋯ menu. On the CLI: find the node ID, restart it, and follow the operation:up with no restart already in flight. If a node never joined, ankra cluster <provider> nodes cloud-init-log <cluster> <node-id> shows its cloud-init status and the end of its log, read over the bastion (not on Proxmox VE or HPE Morpheus). You can also ask Ankra’s AI, for example “restart worker-2 on my-cluster”.
Bastion
The bastion is the machine Ankra and you reach the cluster through: every SSH hop into the private network goes via it. On most providers it carries no workload traffic, so workloads keep running while it is down, but Ankra cannot provision, scale, upgrade or reconcile the cluster until it is back. In the dashboard: Nodes → Bastion & VMs shows its state and the last health verdict, with Diagnose over SSH, Restart VM and Resize. The ⋯ menu on the bastion row in the Machines table also has Resize bastion.bastion_nat egress), node egress stops until it is back.
SSH access and keys
Settings → Access shows copy-paste SSH commands through the bastion for the cluster’s own addresses, and the SSH key credentials attached to the cluster. To change the keys:resync forces it. The key Ankra itself uses is never removed. For everyday kubectl you need no SSH at all - see Accessing Clusters with kubectl.
Upgrade Kubernetes
In the dashboard: Settings → General shows the current Kubernetes version. Pick a target and click Upgrade. The same card has Auto upgrade patch versions, which moves the cluster to the newest patch of its current minor version on its own; minor versions are never upgraded automatically. On the CLI:- Nodes upgrade one at a time, control plane first, then workers. Each node is cordoned, drained honouring PodDisruptionBudgets, upgraded, and must be
Readyat the target version before the next one starts. - An etcd snapshot is taken before the control plane upgrade. With external etcd, the dedicated etcd members are upgraded first, one at a time, each saving a snapshot.
- A drain blocked by a PodDisruptionBudget stops the rollout;
--forceproceeds anyway. - One minor version at a time (1.33 to 1.34, not 1.33 to 1.35), and no downgrades.
Stop and start
Stopping releases the cluster’s compute and keeps its configuration, stacks and credentials in Ankra, which is how you park a cluster you do not need right now. Starting re-provisions it and reconciles it back to running. What a stop does to the servers depends on the provider:
Stop cluster explains the snapshot, the pause mode and the restore in full.
In the dashboard: Settings → General → Danger Zone → Stop cluster, or Stop cluster / Start cluster in the cluster’s ⋯ menu. They also run on a timetable with Power Schedules.
- A stop never deletes the volumes your workloads provisioned through the CSI driver, forced or not; they keep billing while the cluster is stopped.
--forcealso deletes the cluster’s load balancers, and never captures a snapshot. - Stop and start run in the background. A start is refused with
409while a stop or terminate is still running. - While the cluster is stopped,
ankra cluster <provider> nodes liststill shows the saved machines that the next start re-provisions.
Terminate a cluster
Terminating deletes the cluster’s servers and the network Ankra created, then removes the cluster from Ankra. It cannot be undone. In the dashboard: Settings → General → Danger Zone → Terminate. The dialog lists the persistent volumes the teardown deletes, and the button stays disabled until you accept that.- Volumes are deleted only once you accept it.
--yesskips the teardown confirmation but never the volume one. The API lists the volumes withGET .../deprovision-volumesand needs?accept_volume_data_loss=trueon theDELETE; without it the answer is409naming them. --forcedeletes leftover load balancers and tolerates infrastructure that no longer answers. It never stands in for accepting the volume loss. On Proxmox VE it finishes the teardown when the host or jumphost is unreachable, leaving the VMs behind.- A stopped cluster can be terminated too: the volumes recorded at stop time are still known and named.
Persistent volumes covers every case, including volumes Ankra cannot prove belong to the cluster.
Choices fixed at create time
A few choices cannot be changed on a running cluster. They are made in the create wizard (the Kubernetes step) or in the create request, and they shape the upgrade and networking tasks above.- Distribution.
kubeadm(the default everywhere: wizard, CLI and API) is upstream Kubernetes with containerd; its versions are plain tags such asv1.33.2.k3sis a single-binary distribution; its versions look likev1.33.2+k3s1. - CNI (
cni). kubeadm clusters always run Cilium. k3s clusters chooseflannel(the default),calicoorcilium; AWS defaults to Cilium for both distributions. The CNI cannot be changed later. - CNI features (
cni_features). Cilium offerskube_proxy_replacement,hubbleandwireguard_encryption; Calico offersebpf_dataplane; flannel takes none. On k3s,kube_proxy_replacementandebpf_dataplaneneed a single control plane. - etcd topology (kubeadm only).
stacked(the default) runs etcd on the control plane nodes.externalruns it on 3 or 5 dedicated machines (etcd_node_count), sized by a provider-specific field:
Next steps
- Nodes - every control on the Nodes tabs, including standalone VMs.
- Cluster Autoscaling - let pod demand size your node groups.
- Power Schedules - stop and start on a timetable.