Architecture showing protected workloads and recovery control-plane services feeding dependency-aware VCF 9.1 recovery across management and application layers.

Protecting the Recovery Control Plane: A VCF 9.1 Management-Component Backup and Fleet DR Runbook

TL;DR Protecting workload virtual machines does not automatically protect the VMware Cloud Foundation services needed to discover, authorize, network, orchestrate, and validate their recovery. A complete VCF 9.1 recovery strategy needs several distinct mechanisms: native file-based backups for components such as SDDC Manager, vCenter Server, and NSX Manager; image-based protection for VCF Operations; backup and … Explore: Protecting the Recovery Control Plane: A VCF 9.1…

An asynchronous VKS package from the Broadcom release source is manually added through Supervisor Management so the version becomes available for upgrade.

Newer VKS Versions Missing from vCenter: Finding and Registering Asynchronous Releases

TL;DR A newer VMware vSphere Kubernetes Service version may be fully released and still not appear in the vCenter upgrade dropdown. Broadcom KB 439327 explains that this is expected for asynchronous VKS releases. The dropdown automatically shows the VKS versions embedded in the installed vCenter build, while later asynchronous releases must be registered manually by … Explore: Newer VKS Versions Missing from vCenter: Finding and…

Broadcom VKS release package flows through the vCenter Supervisor Services UI to service activation on a target Supervisor and availability for new clusters and ClusterClasses.

KB 439327: How to Find and Register Newer VKS Releases in vCenter

TL;DR A newer VMware Kubernetes Service version may be missing from the vCenter upgrade interface even when the environment is functioning correctly. VKS versions bundled with the installed vCenter build appear automatically. Newer asynchronous VKS releases must first be downloaded from the Broadcom Support Portal and registered by uploading the release’s package.yaml file. Registration does … Explore: KB 439327: How to Find and Register Newer…

Workflow diagram illustrating The Runbook at a Glance for The vCenter Log Partition Runbook: Find Growth, Preserve Evidence, Restore Headroom.

The vCenter Log Partition Runbook: Find Growth, Preserve Evidence, Restore Headroom

A full /storage/log partition on a vCenter Server Appliance is not just a housekeeping problem. It is a management-plane risk. In a standalone vSphere environment, it can interrupt administration, log collection, patching, and service stability. In VMware Cloud Foundation, the blast radius is larger because vCenter is tied into SDDC Manager workflows, workload domain lifecycle … Explore: The vCenter Log Partition Runbook: Find Growth, Preserve…

Workflow diagram illustrating EAM Trust Flow at a Glance for EAM Certificate Trust Failures: Why vSphere Extensions Break After Certificate Changes.

EAM Certificate Trust Failures: Why vSphere Extensions Break After Certificate Changes

Certificate changes in vSphere environments rarely fail in only one place. The obvious place to look is the browser warning, the expired certificate alarm, or the service that recently had its Machine SSL certificate replaced. But in production VMware Cloud Foundation and vSphere environments, certificate changes can also break something less visible: the extension and … Explore: EAM Certificate Trust Failures: Why vSphere Extensions Break…

Conceptual diagram illustrating The Mental Model: vCLS Is Cluster Plumbing, Not Workload for vCLS Retreat Mode: When to Use It, What It Breaks, and How to Exit Cleanly.

vCLS Retreat Mode: When to Use It, What It Breaks, and How to Exit Cleanly

Disable vCLS on a Cluster via Retreat Mode KB 316514 vSphere Cluster Services usually stay in the background until they get in the way of something operational. Most teams first notice vCLS when a cluster task is blocked, a warning appears after maintenance, or a handful of small system VMs show up and someone asks … Explore: vCLS Retreat Mode: When to Use It, What…

Workflow diagram illustrating DVS Upgrade Flow at a Glance for DVS Upgrade Guardrails: What Can Break When Old Distributed Switches Move Forward.

DVS Upgrade Guardrails: What Can Break When Old Distributed Switches Move Forward

A vSphere Distributed Switch upgrade can look deceptively simple in the vCenter UI. Select the switch, choose the target version, confirm the warning, and move on. That is not how it should be treated in a brownfield environment. The risk is not that a DVS upgrade is always dangerous. The risk is that old distributed … Explore: DVS Upgrade Guardrails: What Can Break When Old…

Diagram illustrating Scenario for VM Network Troubleshooting from Guest OS to Uplink: A Layer by Layer VMware Runbook.

VM Network Troubleshooting from Guest OS to Uplink: A Layer by Layer VMware Runbook

Virtual machine network problems rarely arrive with a clean label. The ticket usually says something like “the VM is unreachable,” “the application cannot connect,” “ping fails,” “internet access is down,” or “VMs on different hosts cannot talk.” The underlying cause might be inside the guest OS, on the VM’s virtual NIC, in the port group, … Explore: VM Network Troubleshooting from Guest OS to Uplink:…

Workflow diagram illustrating Triage Flow at a Glance for “No Healthy Upstream” Is Often a Certificate Problem.

vCenter “No Healthy Upstream”: Causes, Diagnosis, and Recovery

Quick answer: In vCenter, No Healthy Upstream means the vSphere Client front end cannot route the request to a healthy backend service. Broadcom describes it as an alternate HTTP 503 symptom, not a diagnosis. Common causes include stopped services, expired certificates, disk-space or memory pressure, session exhaustion, a changed /etc/hosts file, and Lookup Service problems. … Explore: vCenter “No Healthy Upstream”: Causes, Diagnosis, and Recovery

Workflow diagram illustrating Workflow at a Glance for Why Large VM vMotion and Clone Tasks Fail.

Why Large VM vMotion and Clone Tasks Fail: Device Limits, Config Hygiene, and PowerCLI Prechecks

Large VM migrations usually fail at the worst possible time: late in the change window, after the task has already consumed hours of storage, network, and operator attention. When the error is something like “Invalid configuration for device ‘##’,” the first instinct is to look for a broken virtual NIC, a missing port group, an … Explore: Why Large VM vMotion and Clone Tasks Fail:…

Diagram illustrating VMCA as a Trust Factory, Not Just a Certificate Store for The VMCA Reset Decision: When Regenerating vSphere Certificates Is the Right Move.

The VMCA Reset Decision: When Regenerating vSphere Certificates Is the Right Move

Certificates in vSphere are easy to underestimate until they become the reason vCenter will not authenticate, services will not start cleanly, NSX loses trust in its Compute Manager, or SDDC Manager stops interacting with the management domain the way it should. That is why a VMCA reset should not be treated as a generic “renew … Explore: The VMCA Reset Decision: When Regenerating vSphere Certificates…

Conceptual diagram illustrating The Safer Operating Model for From Fixcerts to vCert: A Safer vCenter Certificate Recovery Path.

From Fixcerts to vCert: A Safer vCenter Certificate Recovery Path

Quick answer: Fixcerts or vCert? Broadcom has deprecated the Fixcerts certificate-replacement script and directs operators to vCert. Identify the affected certificates and confirm your backup and rollback plan before remediation. The current vCert guidance says the tool is intended for use under Broadcom Global Support direction. Start with prerequisites and safety checks, then review post-recovery … Explore: From Fixcerts to vCert: A Safer vCenter Certificate…