DAY 2 ARCHITECTURE

Operate with evidence. Change with control. Recover with confidence.

Practical runbooks and decision frameworks for keeping hybrid platforms, private cloud, and production AI supportable through failures, upgrades, incidents, and transformation.

Choose the operational decision in front of you

Start with readiness, controlled change, or recovery. Each pathway is grounded in evidence, ownership, and testable outcomes.

OPERATIONAL READINESS

Prove the platform is ready before change

Connect service ownership, observability, capacity, backup, identity, dependencies, and recovery evidence before a change window or incident begins.

Prepare the platform ↓

CONTROLLED CHANGE

Execute with explicit stop conditions

Plan upgrades, migrations, and AI releases around validation gates, fallback boundaries, and accountable go or no-go decisions.

Plan a safe change ↓

CONTAINMENT AND RECOVERY

Recover the service and preserve evidence

Classify the failure, limit blast radius, restore the minimum viable control plane, and prove the service is healthy.

Start with recovery ↓

Operations and resilience guides

Six practical starting points for readiness, lifecycle change, dependency failures, recovery, transformation, and AI incidents.

VCF UPGRADE RUNBOOK

Introductory visual for VCF 5.2.x to 9.1 Upgrade Runbook: Exact Sequence, Dependencies, Downtime, and Validation.

VCF 5.2.x to 9.1 Upgrade Runbook

Follow a dependency-controlled sequence with readiness gates, maintenance impact, validation evidence, ownership, and fallback limits.

Read the VCF runbook →

NUTANIX UPGRADE RUNBOOK

Introductory visual for Nutanix AOS 7 and AHV 10 Upgrade Runbook: LCM Prechecks, Order, Rollback, and Validation.

Nutanix AOS 7 and AHV 10 Upgrade Runbook

Use lifecycle prechecks, safe execution order, stop and retry boundaries, application validation, and explicit recovery decisions.

Read the Nutanix runbook →

AZURE LOCAL DEPENDENCIES

Introductory visual for What Fails When Azure Local Loses Azure? Arc Resource Bridge, Connectivity, Updates, and.

What Fails When Azure Local Loses Azure?

Understand the boundaries between Azure connectivity, Arc Resource Bridge, lifecycle services, and the local workload plane.

Read the dependency model →

RECOVERY CONTROL PLANE

Architecture showing protected workloads and recovery control-plane services feeding dependency-aware VCF 9.1 recovery across management and application layers.

Protecting the Recovery Control Plane

Recover the management services, credentials, certificates, networks, backups, and external dependencies applications rely on.

Read the recovery runbook →

PLATFORM TRANSFORMATION

Introductory visual for Recovery During Platform Transformation.

Recovery During Platform Transformation

Protect workloads while versions, hypervisors, migration tools, and recovery mechanisms are changing around them.

Read the transformation guide →

AI INCIDENT RESPONSE

Workflow diagram illustrating The Agent Incident Response Flow for How to Roll Back AI Agents: Incident Response, Circuit Breakers, and Recovery Patterns.

How to Roll Back AI Agents

Build containment, blast-radius analysis, tool shutdown, policy rollback, evidence preservation, and a safe return to service.

Read the AI recovery guide →

Turn operational claims into tested evidence

Use these guides to improve restore assurance, change control, migration execution, identity recovery, degraded management, and platform operations.

Hands-on troubleshooting guides

Use these references when an operational signal needs to become a structured diagnosis and recovery decision.

NETWORK PATH AND PACKETS

Diagram showing Traceflow lets you inject and trace synthetic packets through the NSX-T fabric, visualizing every hop - distributed firewall, logical switches, Tier-0/1 routers, edge nodes - and highlighting where packets are delivered or...

NSX-T Traceflow and Port Mirroring Troubleshooting Guide

Separate policy and path failures from packet-level problems using synthetic tracing and live traffic capture.

Open the NSX-T guide →

HCI SERVICE HEALTH

Illustration representing Nutanix Node and CVM.

Nutanix CVM Architecture and Troubleshooting Guide

Understand CVM services, operational commands, sizing factors, resilience, and safe troubleshooting.

Open the Nutanix CVM guide →

Make the next incident less improvised.

Choose one critical platform or AI service. Map its dependencies, owners, evidence, stop conditions, and recovery path, then test the runbook before the next change window.