Skip to main content
A Stack deploy or add-on upgrade failed and you need to know which part broke and whether it is safe to run again. Every change Ankra makes to a cluster runs as an operation, and each operation is split into steps - one per add-on install, manifest apply, GitOps sync and so on. This guide shows how to find the failed step, read its error and retry the operation.

Prerequisites

  • Access to the cluster in the Ankra portal, or the Ankra CLI logged in to your organisation.

Find the failed operation

1

Open the cluster's operations

Open the cluster and select Operations in the cluster sidebar. The list shows every operation with its status and duration. A failed step also raises an Execution step failed card in your Activity & Inbox, which links to the same operation.
2

Open the operation

Select the failed operation. Its steps are laid out in Pending, Active and Completed columns (the portal labels the steps Jobs), with the total duration at the top. Use Search jobs… to find a step by name.
3

Read the failed step

Select the step marked as failed. The panel shows the Stack, add-on, chart and namespace it targeted, its Status events, and the Tasks it ran with their output and error. Steps after a failure show Not Executed: they were skipped because an earlier step failed, not because they are broken themselves.
4

Ask the AI if the error is not obvious

Select Ask AI about this operation to open the AI Assistant with the operation already in context.
Fix the cause first - a wrong value in the add-on configuration, a missing Secret, a namespace that does not exist - then retry.

Retry or cancel

  • Retry operation is available on a failed operation. A retry starts a new operation that re-plans only the resources that did not finish; the failed one stays in the history as it was. Retry is refused while the cluster has another operation running, while the cluster is offline, and when a newer operation of the same type already exists - in that case, look at the newer one instead.
  • Cancel operation is available while an operation is running. It stops the operation.
There is no pause or resume.

Verify

Open Operations again. The retry appears as a new operation at the top; wait for it to finish as successful. The original failed row keeps its failed status, and once the resources it failed on are healthy it shows a Resolved label.

Failed history versus open problems

The operations list is an audit log, not a to-do list. A failed operation keeps its failed status for as long as the record is kept, and nothing acknowledges or deletes it, so a rollout that needed a few attempts leaves its failed attempts in the history for good. What tells you whether a failure still needs someone is its attention state: The state is worked out each time you read the list, so a successful retry resolves the operation it retried. In the portal a resolved failure carries the Resolved label; in the CLI filter with --attention open or --attention resolved. The cards in Activity & Inbox are the attention list: they clear on their own once the underlying problem recovers, or when you dismiss them.

Next step

If a retry keeps failing on the same step, open a support ticket with the operation link, or ask the AI Assistant about it.