
TL;DR
- Treat the target as an exact AOS, AHV, Prism Central, NCC, LCM, firmware, and product-service combination. “AOS 7 and AHV 10” is not specific enough for a change record.
- Validate the supported upgrade path and compatibility matrix before downloading software.
- Upgrade Prism Central first only when the validated dependency plan requires it. Back it up before changing it.
- Refresh the supported LCM framework, perform inventory, update NCC, run NCC, and complete an independent LCM precheck before the maintenance window.
- Require green data resiliency, healthy cluster services, no active rebuild or unresolved critical alert, and enough CPU, memory, and storage headroom to tolerate one node in maintenance.
- Use the order selected by LCM and the applicable release notes. A practical planning sequence is Prism Central if required, LCM, NCC, AOS, validation, AHV, validation, firmware, dependent services, and final evidence.
- Do not describe rollback as a generic downgrade. Define safe stop points, Prism Central recovery, support-directed component recovery, and application-level failback separately.
- The change is complete only after infrastructure health, workload transactions, data protection, monitoring, and audit evidence all pass.
What This Runbook Accomplishes
The objective is to move a production Nutanix AHV cluster from an approved source state to an approved AOS 7.x and AHV 10.x target while preserving data availability, workload service, manageability, and recovery options.
The runbook assumes:
- The cluster is running Nutanix AHV.
- Prism Element is available and the cluster may be registered to Prism Central.
- LCM is the authorized lifecycle mechanism. Nutanix requires LCM for AOS upgrades to AOS 6.8 or later. [3]
- The operator has administrative access, a tested support path, an approved maintenance window, and named application validators.
- The source and target releases are supported on the exact server model and hardware configuration.
This is not a universal procedure for single-node clusters, two-node clusters, Metro Availability pairs, Nutanix Cloud Clusters, or environments with special device-backed workloads. Those designs require their version-specific Nutanix procedures in addition to this baseline. LCM supports some software operations on single-node clusters, but Nutanix warns about data-loss and service-disruption risk, so a normal rolling-upgrade assumption is inappropriate. [4]
Why AOS, AHV, Prism Central, and Firmware Need Separate Gates
The platform components have different functions and failure domains:
| Component | Operational role | Upgrade concern |
|---|---|---|
| Prism Central | Central management and services across clusters | Version compatibility with registered Prism Element clusters, service dependencies, backup, and management-plane availability |
| LCM | Inventory, dependency evaluation, prechecks, software and firmware orchestration | Framework compatibility, current metadata, network or dark-site bundle access, and task history |
| NCC | Platform health assessment | Current compatible version, unresolved findings, and repeatable before-and-after results |
| AOS | distributed storage and cluster services running through CVMs | CVM health, data resiliency, active rebuilds, service state, and rolling component behavior |
| AHV | Hypervisor on each node | Host evacuation, maintenance mode, VM mobility, device-backed workloads, capacity, and host restart behavior |
| Firmware | BIOS, BMC, NIC, HBA or controller, drive, and other platform firmware | Hardware model support, Redfish or BMC access, OEM bundles, reboot duration, and dependency order |
| Product services | Files, Objects, Flow, NDB, NKP, DR, and other services | Separate interoperability matrices, upgrade guides, and application-facing validation |
LCM can calculate dependency-aware component order for selected updates, but that does not remove the need to validate the whole environment or define gates between major components. [5]
The workflow below shows the sequence that matters. A failed gate returns to remediation. It does not flow into the next component.

Define the Change Before Opening LCM
A change ticket should name the exact desired state. At minimum, record:
| Field | Required value |
|---|---|
| Cluster | Name, UUID, site, node count, hardware model, and fault-tolerance design |
| Current stack | Prism Central, Prism Element or AOS, AHV, LCM, NCC, Foundation, firmware, and product-service versions |
| Target stack | Exact approved version for every component being changed |
| Upgrade path | Supported hop or sequence returned by Nutanix Upgrade Paths |
| Compatibility evidence | Saved result from the Compatibility and Interoperability Matrix |
| Release notes | Reviewed AOS, AHV, Prism Central, LCM, and hardware bundle notes |
| Service impact | Expected management, migration, VM, and application behavior |
| Stop conditions | Specific technical and application conditions that halt the workflow |
| Recovery owners | Nutanix platform, network, storage, backup, application, business, and support contacts |
| Validation | Infrastructure checks, application transactions, monitoring checks, and observation period |
Avoid a target such as “latest AOS 7.” A valid target looks like an exact maintenance release paired with an exact AHV build and a compatible Prism Central release. The target must also account for the server model, disk and controller firmware, NIC firmware, GPU or vGPU stack, Files or Objects versions, backup integrations, and DR relationships.
Pre-Upgrade Readiness Gates
Validate the Supported Path and Product Compatibility
Use three separate checks because they answer different questions:
- Lifecycle status: Is the proposed target still maintained and supported?
- Upgrade path: Can the current version move directly to the target, or are intermediate hops required?
- Interoperability: Are Prism Central, AOS, AHV, firmware, hardware, and dependent services supported together?
LCM showing an available update is useful evidence, but it is not a substitute for reviewing the target release notes and the formal compatibility data. This matters when the environment includes products that LCM does not fully model, such as backup software, guest drivers, third-party monitoring, GPU host drivers, security agents, or application certification requirements.
Record the result, date, and operator. Compatibility pages change over time, so an audit record should preserve what was approved for the change rather than relying on a future lookup.
Resolve Prism Central Dependencies
Prism Central and the registered Prism Element clusters must remain compatible during every transition state, not only after the final upgrade. Nutanix provides a PC-to-PE compatibility health check, and the software compatibility matrix remains the authoritative planning source. [6]
Use this decision sequence:
- If Prism Central is already compatible with both the current and target AOS releases, retain it unless release notes or a required service say otherwise.
- If the target AOS requires a newer Prism Central release, upgrade Prism Central before the managed cluster.
- If Prism Central hosts Self-Service, Flow, DR, automation, or other services, validate each service against the proposed PC release.
- If Prism Central itself runs on the cluster being upgraded, confirm that its VMs can tolerate host evacuation and that local Prism Element access remains available.
- Do not unregister a cluster merely to bypass a compatibility finding. Registration can carry service, policy, and protection dependencies.
Back up Prism Central using its supported backup mechanism and verify that the backup is current. Nutanix documents One-Click Recovery for protected Prism Central instances, which makes the backup a real recovery control rather than a checkbox. [7]
Capture:
- Prism Central version and deployment size
- Registered clusters and connection state
- Identity provider and RBAC health
- Enabled Prism Central services
- Backup type, destination, last successful timestamp, and recovery prerequisites
- Current critical alerts and failed services
Refresh LCM and Perform Inventory
LCM inventory identifies the software and firmware entities visible to the framework and the bundles available to it. In a connected site, inventory can also update the LCM framework. In a dark site, the correct framework and payload bundles must be staged using the method supported by the installed LCM and platform versions. [8]
Perform inventory before the maintenance window, then verify:
- Every expected node and component appears.
- Current versions match the change record.
- The proposed target versions are available or the correct offline bundles are staged.
- No inventory operation is stuck or reporting stale data.
- BMC or Redfish credentials work where firmware orchestration requires them.
- DNS, NTP, proxy, firewall, and repository paths required by LCM are healthy.
- LCM operations history does not contain an unresolved prior lifecycle task.
An unexpected missing component is a stop condition. Do not assume LCM will safely ignore hardware or software it cannot inventory.
Update NCC and Run the Full Health Check
NCC is a prerequisite, not a post-failure diagnostic. Nutanix explicitly instructs administrators to run NCC before an upgrade and resolve results other than INFO or PASS before proceeding. [9]
Update NCC to the latest version supported by the current stack, then run the complete check from Prism or from a CVM. The following read-only baseline is intentionally small so it can be repeated after each major stage:
# Run from one CVM in each Prism Element cluster as the nutanix user. date -u ncc --version cluster status ncc health_checks run_all
The commands record the UTC time, NCC version, cluster service state, and complete NCC result. Copy the raw output into the change evidence repository. Do not reduce the record to a screenshot of a green summary.
Successful execution means:
- Cluster services expected for that release are up.
- NCC completes without unresolved findings outside INFO or PASS.
- Any accepted informational finding has a documented disposition.
- The same commands can be rerun after AOS and AHV without relying on shell history.
NCC is point-in-time evidence. A passing result from the previous week does not authorize tonight’s change.
Verify CVM, Cluster, and Data Resiliency Health
Before any rolling activity, verify:
- All CVMs are powered on and reachable.
- Cluster services are healthy and stable.
- Data Resiliency Status is green.
- No node, disk, storage pool, or metadata rebuild is active.
- No critical hardware, storage, network, CVM, or hypervisor alert is unresolved.
- Replication and protection jobs are current.
- Time synchronization and name resolution are healthy.
- Cluster performance is within its normal baseline.
Nutanix procedures use green Data Resiliency Status and NCC as prerequisites before taking a node out of service. That same principle belongs in an upgrade gate. [10]
Do not proceed while the platform is already consuming fault-tolerance capacity. In an RF2 cluster, for example, multiple simultaneous node outages can exceed the design’s protection boundary. The exact allowable failure state depends on cluster design, failure domains, RF, and current data placement.
Prove Capacity for One Node in Maintenance
AHV upgrades place hosts into maintenance mode and process them sequentially. The cluster must be able to run the remaining workload while a node is unavailable. Nutanix documents maintenance mode as a required state for host and firmware operations. [11]
Validate all three capacity planes.
Compute capacity
- Remaining hosts can absorb the CPU and memory demand of the evacuated node.
- HA reservation and admission policies remain satisfied.
- Memory overcommit, NUMA placement, and oversized VMs will not block evacuation.
- Peak demand, not only current demand, fits within the remaining cluster.
Storage capacity
- Free and rebuild-reserved capacity are healthy.
- No container is near its critical threshold.
- The cluster can tolerate normal upgrade I/O plus workload I/O.
- Capacity-tier and data-placement behavior are stable.
Network capacity
- Management, CVM, storage, and live-migration paths are healthy.
- Uplinks, bonds, VLANs, MTU, and upstream switch ports are stable.
- The network can absorb migration traffic without starving production flows.
There is no safe universal free-capacity percentage for every Nutanix cluster. Node count, RF, block awareness, storage media, failure domains, workload shape, and reserved rebuild capacity all affect the decision. Use the live resiliency and capacity state for the exact cluster.
Identify Workloads That May Not Evacuate Cleanly
Inventory VMs with placement or device constraints before the AHV stage:
- Host affinity or anti-affinity rules
- Large memory reservations
- PCI passthrough, SR-IOV, GPU, or vGPU configurations
- Special network functions or latency-sensitive appliances
- VMs with mounted media or unusual device state
- Workloads restricted by license, application clustering, or node identity
- Prism Central or infrastructure services hosted on the cluster
AHV 10 includes configuration-specific live-migration rules. Some capabilities improved with AOS 7 and AHV 10, so it is equally unsafe to assume every device-backed VM can migrate or that none can. Validate each configuration against the AHV 10 documentation and current host-driver matrix. [12]
For each exception, define one approved treatment: live migrate, shut down and restart, move to another cluster, suspend the dependent service, or exclude the host from the change. “Handle during the window” is not a plan.
Validate Firmware and Hardware Compatibility
Firmware should be planned as its own risk domain even when LCM can orchestrate it in the same workflow. Record the current and target versions for:
- BIOS or UEFI
- BMC
- Storage controller or HBA
- NIC and adapter firmware
- NVMe, SSD, and HDD firmware where applicable
- GPU firmware and host drivers
- Foundation and platform support components
Confirm the exact server model and component revisions are supported by the target AOS and AHV releases. OEM platforms may have different bundle, credential, and sequencing requirements. If LCM’s dependency resolver or the hardware release notes prescribe a different order than the generic sequence in this article, the validated vendor order wins.
Verify Backup, Replication, and Application Recovery
An infrastructure upgrade plan is incomplete without application recovery. Confirm:
- Prism Central backup is current and recoverable.
- Application backups completed successfully.
- Protection domains, recovery plans, and replication schedules are healthy.
- Metro, NearSync, asynchronous replication, and witness dependencies are explicitly reviewed where used.
- Backup proxies, agents, and storage integrations support the target releases.
- Application owners know the validation transaction and the point at which application failback is invoked.
Do not create an improvised rollback mechanism by snapshotting or manipulating CVMs, PCVMs, or AHV hosts outside a documented Nutanix procedure. Platform recovery should use supported product mechanisms or Nutanix Support direction.
Run LCM Prechecks as a Separate Dry Run
LCM can run prechecks independently of the upgrade, which helps separate readiness failures from failures that occur after software execution begins. [13]
Run the precheck early enough to remediate findings, then run it again close to the maintenance window. A valid precheck record includes:
- Selected cluster and component
- Source and target versions
- Date and time
- LCM task or operation identifier
- Complete findings
- Remediation owner and disposition
- Rerun result
Do not waive a failed precheck because an earlier cluster passed or because the same issue was harmless in a lab. A failed precheck is a no-go until the exact finding is resolved or Nutanix Support provides written direction for that environment.
Recommended Upgrade Order
The following is a planning sequence, not a command to override LCM. LCM uses compatibility information and dependencies to select update order. Release notes, Upgrade Paths, the compatibility matrix, and the order presented by LCM take precedence.
| Order | Stage | Gate before continuing |
|---|---|---|
| 1 | Prism Central, if the target plan requires it | PC services healthy, backup verified, registered clusters connected, identity and required services validated |
| 2 | LCM framework and inventory | All expected entities visible, metadata current, target bundles available, no stuck task |
| 3 | NCC | Supported NCC installed, full health check passes, raw result saved |
| 4 | AOS | Cluster services healthy, data resiliency green, no rebuild, capacity and application baseline captured |
| 5 | AOS validation | Versions correct, services stable, NCC clean, storage and application checks pass |
| 6 | AHV | Every host can enter maintenance, constrained VMs have an approved disposition, remaining capacity is sufficient |
| 7 | AHV validation | All hosts on target build, no host left in maintenance, VMs healthy, network and storage checks pass |
| 8 | Firmware | Model-specific compatibility and credentials validated, software stack stable, reboot expectations approved |
| 9 | Dependent services | Each service’s own upgrade path, matrix, backup, and validation plan approved |
| 10 | Final evidence | All infrastructure, application, monitoring, protection, and audit gates pass |
For connected environments, selecting multiple updates may be appropriate when LCM presents and validates the complete dependency plan. For high-risk production changes, separate validation gates between AOS, AHV, and firmware make failure isolation and change control clearer.
For dark sites, follow Nutanix’s version-specific dark-site method and recommended order. Do not assume a bundle prepared for one LCM framework or catalog state is valid for another.
Execute the Upgrade in Controlled Stages
Upgrade Prism Central When Required
If the approved dependency plan includes Prism Central:
- Confirm the most recent PC backup completed.
- Capture PC health, services, alerts, connected clusters, identity, and version.
- Start the supported Prism Central upgrade workflow.
- Monitor the task until it reaches a completed state.
- Confirm the new version and service health.
- Validate login, RBAC, identity provider integration, registered clusters, and required applications.
- Rerun the PC-to-PE compatibility check.
Do not begin the cluster upgrade while Prism Central is degraded unless the approved plan explicitly removes that dependency and Nutanix Support agrees with the recovery path.
Upgrade AOS
Before selecting AOS, verify the change record one final time. Confirm the exact target and review LCM’s proposed plan.
During the AOS stage:
- Watch the LCM operation, Prism alerts, cluster services, CVM reachability, storage latency, and application telemetry.
- Do not make unrelated storage, network, CVM, or cluster configuration changes.
- Do not restart several CVMs or services in response to a slow step.
- Do not start AHV or firmware work manually while AOS is still running.
- Record the operation identifier and the start and completion time.
When LCM reports completion, pause. Rerun the baseline commands, verify Data Resiliency Status, review new alerts, confirm the target AOS version on all nodes, and run the agreed application smoke tests.
Only the approved AOS validation gate authorizes the AHV stage.
Upgrade AHV
The AHV stage is host-disruptive even when workloads remain available through migration. LCM typically handles hosts sequentially, including maintenance entry, workload evacuation, update, restart, and maintenance exit.
Monitor:
- Which host is active in the workflow
- VM migration and placement
- CVM state on that node
- Maintenance-mode entry and exit
- Host management and storage connectivity
- Network uplinks and virtual switches
- Application latency and error rate
If a host cannot evacuate, do not force the cluster into a second simultaneous maintenance event. Identify the blocking VM or constraint, apply the approved exception plan, and rerun the supported precheck or task.
After AHV completes, verify every host reports the exact target build and no host remains in maintenance mode. Confirm all CVMs and user VMs are in their intended state.
Upgrade Firmware and Dependent Services
Begin firmware only after the AOS and AHV stack is stable, unless the dependency plan explicitly requires another sequence. Firmware operations may involve longer host restarts, BMC communication, and OEM-specific behavior.
Upgrade product services only with their own compatibility and recovery checks. A green AOS and AHV result does not prove Files, Objects, NDB, NKP, Flow, backup, DR, or third-party integrations are ready for their next version.
Stop Conditions and Failure Handling
The operator needs authority to stop the change before the window begins. Use explicit stop conditions such as:
- Data Resiliency Status is not green.
- A new critical hardware, storage, CVM, hypervisor, or network alert appears.
- A cluster service is down or repeatedly restarting.
- AOS completes on only part of the cluster and LCM does not present a supported continuation.
- A host cannot enter or exit maintenance mode.
- A VM cannot migrate and no approved exception exists.
- Storage latency, packet loss, application error rate, or transaction time crosses the approved threshold.
- Prism Central loses required services or cluster connectivity.
- Replication, backup, or protection state becomes unhealthy.
- The maintenance window no longer contains enough time for validation and recovery.
Use the response matrix below rather than improvising.
| Failure point | Immediate response | Do not do |
|---|---|---|
| Inventory or precheck failure | Stop, preserve findings, correct the dependency, rerun inventory and prechecks | Start the upgrade because the cluster “looks healthy” |
| Package download or staging failure | Verify bundle integrity, repository access, proxy, DNS, certificates, and space; restage through the supported method | Repeatedly upload different bundles without confirming target metadata |
| Prism Central upgrade failure | Preserve task details, validate local clusters through Prism Element, collect PC evidence, use the supported PC recovery path | Unregister clusters or rebuild PC without dependency analysis |
| AOS step fails | Do not proceed to AHV; verify cluster services and resiliency, preserve LCM history, collect logs, engage Nutanix Support | Restart multiple CVMs, run an undocumented downgrade, or clear task state manually |
| AHV host is stuck | Stop at the safe boundary, identify VM or host constraint, validate CVM and host state, follow support guidance | Put another host into maintenance or power-cycle several nodes |
| Firmware update fails | Keep the stable software state, preserve BMC and LCM evidence, use model-specific recovery | Retry power or firmware operations blindly |
| Infrastructure is green but application test fails | Freeze further lifecycle work, invoke the application incident and recovery plan, compare pre-change telemetry | Close the change based only on Prism health |
LCM’s Stop Update function chooses a safe point rather than interrupting an unsafe operation immediately. Nutanix also documents that if AOS and AHV are selected together and a stop is requested after the AOS update begins, LCM allows AOS to finish and cancels before AHV. [14]
That behavior reinforces an important distinction: stop, rollback, retry, restore, and failback are different operations.
Rollback Is a Recovery Model, Not a Downgrade Button
A credible plan defines several recovery levels.
Abort Before Change
If a readiness gate fails before software execution, cancel the change. No platform rollback is required. Preserve the evidence, remediate the finding, and reschedule.
Stop at a Safe LCM Boundary
If the change begins but risk rises, request the supported LCM stop operation. Expect LCM to stop at its next safe point. A stop may leave one component successfully upgraded while later components remain unchanged.
For example, a completed AOS upgrade is not automatically reversed because the AHV stage was canceled. Validate whether the resulting AOS and AHV combination is supported, stabilize it, and obtain support guidance before deciding the next action.
Retry or Resume the Failed Component
When LCM or the product guide presents a supported retry or resume operation, use it only after identifying the original failure and verifying that cluster health is stable. Preserve the original task identifier and evidence. Repeated retries without causal remediation can turn a contained problem into an outage.
Restore Prism Central
If Prism Central cannot be recovered in place, use the documented Prism Central backup and One-Click Recovery process. This recovers the management platform and important service configuration. It is not an AOS or AHV rollback.
Recover the Application Service
If infrastructure is healthy but the business service fails, use the application-specific recovery plan. That may mean reverting an application change, restarting a clustered service, failing traffic back, restoring data, or activating DR. The correct action depends on the application’s consistency and recovery design.
Escalate for Component-Level Recovery
Manual AOS downgrade, CVM image replacement, AHV reinstallation, task-database edits, or forced multi-node restarts should not appear as routine operator rollback steps. Those are support-directed recovery actions when applicable.
Post-Upgrade Validation
Validation should prove the desired service state, not merely the absence of a red banner.
Rebuild the LCM Inventory
Perform a fresh inventory and verify:
- Every node and component reports the approved version.
- No unexpected component remains on the source release.
- Firmware inventory matches the change record.
- No lifecycle operation remains running or failed.
- The next recommended updates are understood but not accidentally included in the current change.
Export LCM operations history and inventory for the audit record. Nutanix provides an export workflow specifically for this evidence. [15]
Rerun Platform Health Checks
Repeat the same commands used for the baseline:
# Post-upgrade check from one CVM in each cluster. date -u ncc --version cluster status ncc health_checks run_all
Compare the results with the baseline. Resolve or formally disposition every new finding.
Then verify in Prism:
- Data Resiliency Status is green.
- All CVMs and hosts are healthy.
- No node remains in maintenance mode.
- No disk, node, or metadata rebuild is unexpectedly active.
- Storage pools and containers have expected capacity.
- Cluster latency, IOPS, throughput, CPU, and memory remain within normal range.
- Alerts are reviewed, not simply acknowledged in bulk.
Validate Workload and Network Behavior
Confirm:
- Expected VMs are powered on and placed correctly.
- HA and affinity policies remain intact.
- A representative live migration succeeds if the change plan authorizes an active test.
- Virtual networks, uplinks, bonds, VLANs, MTU, routing, DNS, and load balancer paths work.
- Device-backed and exception workloads operate as designed.
- Guest time, network, and storage behavior are stable.
Validate Management, Protection, and Integrations
Confirm:
- Prism Central sees the cluster and reports the correct versions.
- SSO, directory integration, RBAC, certificates, and API access work.
- Backup jobs, snapshots, protection domains, recovery plans, and replication schedules resume.
- Monitoring, logging, SIEM, ticketing, and alert forwarding receive current data.
- Files, Objects, NDB, NKP, Flow, and other installed services pass their own checks.
- Hardware management and BMC connectivity remain healthy.
Validate the Application, Not Only the VM
Application owners should execute the same transactions captured before the upgrade. Examples include:
- User authentication
- API request and response
- Database read and write
- File creation and retrieval
- Message publish and consume
- Batch or scheduled job completion
- North-south and east-west application flows
A powered-on VM is not evidence that the service is healthy.
Post-Upgrade Evidence Package
The following table can be copied into the change record.
| Evidence | Before | After | Owner | Result |
|---|---|---|---|---|
| Prism Central version and health | Attached | Attached | Platform | Pass or fail |
| LCM inventory export | Attached | Attached | Platform | Pass or fail |
| AOS and AHV versions by node | Attached | Attached | Platform | Pass or fail |
| Firmware versions by node | Attached | Attached | Hardware | Pass or fail |
| NCC version and full result | Attached | Attached | Platform | Pass or fail |
cluster status output | Attached | Attached | Platform | Pass or fail |
| Data Resiliency Status | Attached | Attached | Platform | Pass or fail |
| Capacity and performance baseline | Attached | Attached | Operations | Pass or fail |
| Alerts and event review | Attached | Attached | Operations | Pass or fail |
| Backup and replication state | Attached | Attached | Data protection | Pass or fail |
| Network and migration validation | Attached | Attached | Network and platform | Pass or fail |
| Application transactions | Attached | Attached | Application owner | Pass or fail |
| LCM operation IDs and timestamps | Not applicable | Attached | Change owner | Complete |
| Exceptions and support case | Documented | Documented | Change owner | Closed or tracked |
Keep the change open through the approved observation period. A technically completed LCM task is a milestone, not the definition of done.
Common Upgrade Mistakes
Treating a Major Version as the Target
“AOS 7” and “AHV 10” are search terms. Production changes need exact maintenance releases and builds.
Assuming Prism Central Is Always First or Always Optional
The correct answer depends on PC-to-PE compatibility and enabled services. Validate the transition state.
Running NCC Only After Something Fails
NCC belongs before the change, after AOS, after AHV, and after final remediation when the risk warrants it.
Combining AOS, AHV, and Firmware Without Gates
LCM may support a combined dependency-aware workflow, but operational gates make ownership and failure isolation clearer.
Equating Rolling with Risk-Free
Rolling operations reduce planned interruption. They do not eliminate workload, migration, capacity, network, hardware, or application risk.
Calling Every Recovery Action a Rollback
Safe stop, retry, Prism Central restore, application failback, and support-directed component repair solve different problems.
Closing on Green Infrastructure Alone
The final gate is the application transaction plus protection, monitoring, and evidence, not the color of the Prism dashboard.
Frequently Asked Questions
Should Prism Central always be upgraded before AOS?
No. Upgrade Prism Central first when the validated source-to-target compatibility plan, Upgrade Paths, release notes, or an enabled service requires it. If the installed PC release supports both the current and target cluster versions, a PC upgrade may not be necessary for that change.
Can AOS and AHV be selected together in LCM?
LCM can present a dependency-aware combined plan. For production control, review the exact plan and retain a validation gate after AOS. If a stop is requested after AOS begins, LCM may complete AOS and cancel before AHV at a safe boundary.
Does a rolling Nutanix upgrade guarantee zero downtime?
No. Rolling orchestration reduces planned service interruption, but constrained VMs, capacity pressure, migration failures, network faults, firmware behavior, or application dependencies can still cause impact.
Can an AOS or AHV upgrade be rolled back with one click?
Do not plan on a universal downgrade button. Define pre-change abort, safe LCM stop, supported retry or resume, Prism Central recovery, application failback, and Nutanix Support escalation as separate controls.
Should firmware be upgraded in the same maintenance window?
Only when the compatibility plan, available time, hardware procedure, and recovery plan support it. Firmware has a different failure domain and often deserves a separate gate or separate change window.
Conclusion
The safest Nutanix upgrade is not the one with the fewest clicks. It is the one with the clearest dependencies, the strongest readiness gates, and the most objective evidence.
For AOS 7.x and AHV 10.x, that means resolving an exact supported target, protecting Prism Central, refreshing LCM inventory, updating and running NCC, proving CVM and cluster health, preserving node-maintenance capacity, validating hardware and workload mobility, and separating AOS, AHV, firmware, and product services with explicit gates.
Rollback should be treated as a recovery model rather than a promise of simple downgrade. When the change defines safe stop points, support escalation, management-plane recovery, application failback, and measurable validation before execution begins, LCM becomes what it should be: the orchestration engine inside a disciplined lifecycle process.
Official References
Some Nutanix Support & Insights pages may require an authenticated Nutanix account.
- Nutanix Software End of Life Information
- Nutanix Compatibility and Interoperability Matrix
- Prism 7.0: AOS Upgrade
- LCM 3.3: LCM Limitations
- LCM: Life Cycle Manager Overview
- NCC Health Check: PC and PE Version Compatibility
- Prism Central Backup, Restore, and Migration
- LCM 3.3: LCM Inventory
- Prism 7.0: Running NCC from Prism Element
- Nutanix Node Shutdown Precheck
- AHV 10.3: Node Maintenance Mode
- AHV 10.0: Live Migration Restrictions
- LCM: Performing Pre-Upgrade Checks
- LCM 3.3: Stop Firmware and Software Upgrades
- LCM 3.3: Exporting Operations History and Inventory
TL;DR A Nutanix CVM incident should be investigated as an evidence chain, not as a race to restart services. Begin with cluster...