The Minimum Viable VMware Cloud Foundation 9.1: Hosts, Storage, NSX, Fleet Services, Licensing, and Operational Overhead

Introduction

The most common sizing question about VMware Cloud Foundation is still framed as a host-count question: How many ESX hosts are required before VCF can be deployed?

That question is useful for checking a deployment prerequisite, but it is not sufficient for an architecture decision. A platform can satisfy a minimum host count and still have no credible capacity for maintenance, no useful failure reserve, no protected management services, no operational telemetry, no backup path, no routable north-south services, and no staff prepared to run the environment.

VCF 9.1 makes that gap more visible. The platform now includes VCF Management Services hosted on VCF Services Runtime, VCF Operations as a central operational plane, optional but consequential VCF Automation and log-management services, NSX management and optional Edge capacity, lifecycle services, identity integration, licensing services, and a larger collection of IP, DNS, certificate, storage, and recovery dependencies than an ESX-only bill of materials suggests [1], [2].

The right question is therefore not, “Can four hosts deploy VCF?” It is, “What must be funded, installed, integrated, protected, and staffed before VCF 9.1 can responsibly host its first production business workload?”

TL;DR

Four hosts remain an important VCF planning reference, especially for a greenfield vSAN-backed management domain, but four hosts describe a deployment floor rather than a comfortable production baseline. A practical small production design normally needs five or six hosts when management and business workloads share infrastructure, or a dedicated four-to-six-host management cluster plus a separate workload cluster when operational isolation matters.

The hardware is only the visible layer. Before the first production workload, architects must also account for VCF Management Services and VCF Services Runtime, VCF Operations, vCenter, SDDC Manager, NSX Manager, optional NSX Edge nodes, VCF Automation, log retention, external DNS and NTP, identity, PKI, backup repositories, software-depot access, IP address pools, lifecycle headroom, and staffing.

A small VCF deployment can be technically valid without being operationally complete. The production threshold is reached only when the platform can tolerate a host outage, evacuate a host for maintenance, recover its management plane, prove its licensing state, monitor itself, and still leave enough capacity for the business workload it was purchased to run.

The Four Host Question Is the Wrong Sizing Question

The four-host discussion persists because it is concrete. Hosts are easy to count, and older VCF mental models were strongly associated with a four-node vSAN management domain. VCF 9 widened the principal-storage choices, which naturally raised the question of whether external storage changes the minimum [3].

It changes some deployment pathways, but it does not remove the underlying availability problem.

A management domain has to run the services that build, operate, license, monitor, and update the rest of the private cloud. If one host is unavailable, those services must remain powered on. If another host needs maintenance, the remaining hosts must absorb its virtual machines without creating unacceptable CPU, memory, or storage pressure. If the same cluster also runs business workloads, the management plane must win every contention event.

This creates three different meanings of minimum:

Minimum typeWhat it provesWhat it does not prove
Installer minimumThe selected deployment workflow accepts the topologyProduction capacity, resilience, or maintainability
Technical minimumThe platform can start and basic functions workA credible service level under failure or maintenance
Production minimumThe platform can operate through expected failures and lifecycle eventsLong-term growth, site recovery, or enterprise scale

The first two are useful in a lab. The third is what an enterprise should fund.

Four hosts can be a valid floor and still be a poor target

A four-host cluster has only three hosts available after one failure. During maintenance, it may temporarily operate with three hosts even when nothing has failed. If a second problem occurs during that window, the design has little room to maneuver.

The issue is not simply whether vSphere HA can restart a virtual machine. The issue is whether all management services, NSX components, storage-policy requirements, workload reservations, and operational tools can run with acceptable contention while the platform is degraded.

For a small production environment, the DTD planning recommendation is:

  • Use five or six hosts when management services and business workloads share the first cluster.
  • Use four to six dedicated management hosts plus a separate workload cluster when lifecycle isolation, predictable maintenance, or stronger fault containment matters.
  • Treat a four-host design as a constrained footprint that requires explicit admission control, tested evacuation, conservative workload placement, and a documented growth trigger.

This is not a replacement for the Broadcom Planning and Preparation Workbook. It is the operational interpretation that should be applied before the workbook values are turned into a purchase order [4].

Translating Consolidated and Standard Architecture into VCF 9.1

Architects coming from VCF 4.x and 5.x often use two familiar terms: consolidated architecture and standard architecture. Those labels remain useful for translation, but Broadcom’s current 9.1 design language emphasizes fleet and site blueprints, including VCF Fleet in a Single Site with Minimal Footprint and broader single-site designs [5], [6].

The current minimal-footprint model

The minimal-footprint model is operationally similar to what many architects previously called consolidated architecture. Management services and selected business workloads share a smaller infrastructure footprint. The value is reduced hardware entry cost. The tradeoff is that management capacity, tenant capacity, maintenance reserve, and failure reserve compete within the same physical boundary.

This model fits when:

  • The workload count is modest.
  • The organization accepts a smaller failure envelope.
  • The same operations team owns platform and workload infrastructure.
  • Growth is predictable and cluster expansion can happen quickly.
  • Separate workload-domain isolation is not yet justified.

It becomes risky when the cluster is sized by adding only the expected business workload to the installer minimum. The management plane must be treated as a protected tenant with non-negotiable capacity.

The separated single-site model

The separated model is the modern equivalent of the old standard architecture. The management domain runs on dedicated hosts, while one or more workload domains run on separate clusters. This increases initial cost but improves lifecycle isolation, resource predictability, failure containment, and organizational clarity.

This model fits when:

  • Business workloads have production service-level objectives.
  • Maintenance windows must be predictable.
  • Different teams own platform and application operations.
  • Separate storage, network, security, or lifecycle policies are required.
  • Multiple workload classes or tenants need independent capacity boundaries.
  • The environment is expected to grow beyond one cluster.

The important point is not the label. It is the placement decision: Will business demand ever be allowed to reduce the recoverability or maintainability of the VCF management plane? If the answer is no, separate the infrastructure or reserve the capacity with equivalent discipline.

Current VCF 9.1 terminology guardrail

VCF terminology has moved faster than many architecture diagrams. Use the current name in the design, then retain the older term only as a translation aid.

Familiar or transitional termCurrent VCF 9.1 term or interpretationArchitecture implication
Fleet Services or a separate fleet-management appliance layerVCF Management Services hosted on VCF Services RuntimeSize a service runtime and its management services, not only legacy appliances
Consolidated architectureVCF fleet in a single site with minimal footprintShared infrastructure lowers entry cost but increases resource and failure-domain coupling
Standard architectureSeparated single-site design with dedicated management and workload capacityHigher initial footprint buys lifecycle and capacity isolation
Aria OperationsVCF OperationsTreat Operations as part of the VCF operational plane and lifecycle workflow
Aria AutomationVCF AutomationTreat Automation as a governed consumption platform with organizations, projects, policies, and catalogs
vRealize Log Insight or Aria Operations for LogsVCF log management or VCF Operations log-management servicesSize by ingest, retention, replication, and evidence requirements

This translation matters because an old component list can materially understate a new VCF 9.1 deployment.

What Must Exist Before the First Workload

The platform becomes production-ready through a dependency chain. The following diagram shows why the first business VM sits at the end, not the beginning, of the VCF deployment process.

What matters in this diagram is sequence. A business workload can technically be powered on before every operational dependency is mature. That does not make the platform production-ready. Production readiness exists when the dependencies below the workload have named owners, tested procedures, monitored health, and sufficient reserve.

Management Domain Design

The management domain is not administrative overhead that happens to consume a few virtual machines. It is the control environment that determines whether the rest of the private cloud can be provisioned, monitored, patched, licensed, recovered, and governed.

VCF Management Services and VCF Services Runtime

VCF 9.1 introduces a more explicit VCF Management Services layer. Broadcom documents multiple availability models, and the simple model deploys one control-plane node with multiple worker nodes while relying on vSphere HA as its availability mechanism [7], [8].

This creates an important sizing change. The management plane is no longer accurately represented by a legacy list containing only SDDC Manager, vCenter, NSX Manager, and an operations appliance. VCF Services Runtime consumes its own CPU, memory, storage, network addresses, DNS records, certificates, and recovery attention.

The smallest current VCF Management Services model is approximately 40 vCPU and 82 GB of memory before its surrounding platform components are counted. That number is useful as a directional check, but it must not be treated as the complete management-domain requirement. It represents allocated virtual resources, not the physical capacity required after availability, overcommit policy, storage overhead, and maintenance reserve are applied [8], [22].

vCenter and SDDC Manager

vCenter remains the core vSphere management plane, while SDDC Manager retains domain lifecycle and platform orchestration responsibilities. Their individual resource profiles are not the largest items in the design, but their operational criticality is high.

They require:

  • Protected placement and anti-affinity where applicable.
  • File-based backup to an external target.
  • Verified DNS and time synchronization.
  • Certificate and password lifecycle management.
  • Recovery-order documentation.
  • Capacity to restart after a host failure.
  • Connectivity to VCF Operations, software depots, and other management services.

A platform that can run vCenter but cannot restore it is not production-ready.

NSX Manager

VCF 9.1 supports different NSX Manager and control-plane models. The simple model reduces infrastructure consumption, but it also lowers management-plane availability. A production high-availability design normally uses a three-node NSX Manager cluster; a single-node model should be treated as an explicit reduced-availability decision, not as the default meaning of “small” [9].

NSX Manager is required for VCF networking and management functions even when the first business workload uses a simple VLAN-backed port group. NSX Edge nodes, however, are use-case dependent.

VCF Operations

VCF Operations is not merely a dashboard added after deployment. In VCF 9.1 it participates in fleet visibility, health, capacity, configuration, lifecycle workflows, and inventory synchronization. Broadcom documents simple and high-availability VCF Operations models because the operational plane itself has an availability design [10].

At minimum, architects should include:

  • The VCF Operations node or cluster.
  • Required collectors or proxies.
  • Capacity and retention for metrics.
  • Integration credentials and service accounts.
  • Alert routing and ownership.
  • Backup and recovery requirements.
  • Growth caused by additional vCenters, hosts, VMs, logs, and network telemetry.

VCF Automation

VCF Automation is not required to power on a traditionally managed VM through vCenter. It is required before the organization can claim that it has delivered governed self-service, catalog consumption, policy-based provisioning, tenant or project controls, and repeatable application delivery.

Broadcom documents simple and high-availability deployment models. A simple small deployment is materially smaller than a three-node high-availability deployment, but it also creates a different service commitment [11].

This leads to a useful readiness distinction:

Target outcomeVCF Automation position
First administrator-provisioned production VMCan be deferred if governance and provisioning remain in vCenter processes
Production self-service VM catalogRequired before service launch
Multi-tenant or project-based consumptionRequired with identity, quota, policy, and catalog design
Application blueprints and governed automationRequired with content lifecycle and integration ownership

Deferring Automation is a scope decision. Omitting its future resource and staffing impact from the platform design is a planning error.

Management Platform Resource Consumption

Broadcom’s Planning and Preparation Workbook and component sizing guidance remain the authoritative tools for an actual deployment [4]. The following values are DTD planning envelopes, not a vendor bill of materials. They are intended to expose the order of magnitude that should be reserved before detailed sizing begins.

The ranges include allocated virtual resources for the major VCF management services and typical optional platform layers. They do not include business-workload capacity. Physical host requirements will vary with processor generation, CPU overcommit policy, memory reservations, storage efficiency, telemetry retention, and availability model.

Reference envelopePlatform patternApproximate management-platform allocationPhysical-host patternIntended use
SmallMinimal-footprint site, simple management models, basic Operations, optional simple Automation, modest logging140-260 vCPU, 500-900 GB RAM, 10-20 TB provisioned storage5-6 shared hosts recommendedSmall production estate with controlled growth and one operations team
MediumDedicated management domain, HA NSX and Operations, HA or production Automation, log cluster, two or more Edges300-600 vCPU, 1.2-2.5 TB RAM, 25-60 TB provisioned storage4-6 management hosts plus 6-12 workload hostsDepartmental or regional private cloud with service catalog and stronger SLOs
EnterpriseExpanded HA management services, multiple collectors, longer retention, network analytics, multiple workload domains and sites600-1,200+ vCPU, 2.5-6+ TB RAM, 60-200+ TB provisioned storage6-8 management hosts plus multiple 4-16+ host workload clustersLarge fleet, multiple tenants, regulated workloads, multi-site recovery

The broad ranges are deliberate. Log retention can dominate storage. VCF Automation availability can materially change CPU and memory. NSX Edge count changes with throughput and service design. Multi-site recovery duplicates or extends parts of the platform. A point estimate without those decisions would be false precision.

A more useful way to calculate physical capacity

Do not convert allocated vCPU directly into physical cores without an overcommit and failure model. Use a planning equation such as:

Required physical compute =
  (management steady-state demand
   + management failure reserve
   + maintenance reserve
   + workload demand
   + workload failure reserve
   + growth buffer)
  / approved overcommit ratio

Memory usually becomes the harder constraint because many management components are not good candidates for aggressive memory overcommit. Storage must include both provisioned appliance disks and actual consumed data growth, especially logs, metrics, backups, snapshots used during lifecycle work, and temporary migration or recovery space.

Storage Choices and Their Operational Consequences

VCF 9.1 supports broader principal-storage choices than a vSAN-only mental model suggests. The critical distinction is between supported storage and storage directly exposed by a greenfield deployment workflow.

Broadcom documents that iSCSI, NFS v4.1, FCoE, and NVMe over Fabrics can be used as principal storage through supported converge or import workflows even when they are not available in the standard greenfield workflow [12]. The platform can lifecycle-manage those hosts and clusters, but some day-two host and cluster changes must first be performed in vCenter and then synchronized into VCF Operations.

vSAN

vSAN remains the most integrated option for many greenfield VCF deployments. It aligns storage lifecycle with the host cluster and reduces dependency on an external array team. It also creates requirements that must be funded:

  • vSAN-compatible servers, devices, controllers, firmware, and drivers.
  • Sufficient fault-domain and storage-policy capacity.
  • Network bandwidth and consistent MTU design.
  • Capacity for rebuilds, resynchronization, maintenance mode, and growth.
  • Operational familiarity with ESA or OSA design, depending on the selected architecture.
  • Additional entitlement review when usable capacity exceeds included rights.

Raw capacity is not usable capacity. Storage-policy protection, metadata, slack space, rebuild reserve, snapshots, and operational headroom all reduce what can safely be assigned to workloads.

Fibre Channel and NFS

External storage can reduce local-disk requirements and may align with an existing enterprise storage operating model. It can also improve separation between compute lifecycle and storage lifecycle. The tradeoff is a larger dependency graph:

  • Array compatibility and microcode.
  • SAN fabric or Ethernet storage design.
  • HBA, NIC, multipathing, and driver compatibility.
  • Zoning, LUN, export, and access-control processes.
  • Independent storage monitoring and support ownership.
  • Recovery coordination between VCF and storage teams.

External storage is not “less infrastructure.” It is infrastructure placed in another team’s failure and lifecycle domain.

Storage selected through converge or import

When a principal-storage type requires a converge or import path, architects must include the operational consequence in the design. Broadcom’s guidance states that lifecycle management remains available, but selected day-two inventory operations require a vCenter action followed by Sync Inventory in VCF Operations. Missing the synchronization can block lifecycle management for affected hosts and clusters [12].

That is a governance issue, not a minor procedure. The runbook, permissions, automation, change record, and validation step must all account for the two-system workflow.

NSX Manager and Edge Requirements

NSX is part of the VCF platform footprint, but NSX Edge is not required for every workload.

When NSX Manager is enough

A small environment using VLAN-backed networks and external physical routing may not need an Edge cluster for its first VM. NSX Manager still needs to be designed, protected, monitored, backed up, and included in lifecycle planning.

The simple NSX Manager model reduces resource use but accepts lower management-plane availability. A production design that depends on NSX policy, segmentation, or overlay networking should normally use the high-availability model unless the reduced service level is explicitly accepted.

When Edge nodes become part of the minimum

Two Edge nodes become a practical production floor when the platform needs centralized north-south routing or stateful network services. Typical triggers include:

RequirementEdge implication
Overlay-backed segments or VPC connectivityEdge transport and gateway design required
Tier-0 or Tier-1 gateway servicesAt least two production Edge nodes recommended
NAT, VPN, gateway firewall, or load balancingEdge capacity, state, and failure placement required
Dynamic routing to the physical fabricBGP peers, uplink VLANs, MTU, and route-policy design required
Tenant network isolationEdge cluster and gateway ownership must match tenancy model
High throughput or service scaleAdditional or larger Edge nodes may be required

A commonly used medium Edge appliance profile consumes approximately 8 vCPU, 32 GB of memory, and 200 GB of disk per node. Two nodes therefore add meaningful management-domain demand before they carry any application traffic [13].

The physical network must also provide the uplink VLANs, routing adjacencies, failure domains, and bandwidth that make Edge high availability real. Two Edge VMs on one oversubscribed host or one physical switch do not create a resilient network service.

VCF Operations, Automation, Logging, and Observability

The production footprint is determined as much by telemetry and governance as by hypervisor count.

Operations is part of the service, not an optional report

Before the first workload, VCF Operations should be collecting enough evidence to answer:

  • Are all management components healthy?
  • Is the cluster carrying enough reserve for one host failure?
  • Which datastore, network, or host is approaching a limit?
  • Are configuration and inventory synchronized?
  • Are certificates, licenses, or software binaries approaching an operational deadline?
  • Who receives the alert, and what do they do next?

A platform that can alert but has no owner or runbook is only partially observable.

Automation adds a second operating model

VCF Automation introduces organizations, projects, catalogs, policies, images, quotas, approvals, integrations, and service accounts. Those are not just product features. They become operational objects that require version control, ownership, testing, and lifecycle management.

A small team should avoid deploying Automation merely because it is included. Deploy it when the team is ready to operate a service catalog. Conversely, a team promising self-service should not omit the CPU, memory, IP addresses, content pipeline, identity design, and support procedures that make Automation reliable.

Logging is usually a storage decision disguised as a VM decision

The compute profile for log management is only the beginning. Retention period, daily ingest, indexing overhead, replication, archive requirements, and incident-investigation needs determine the real footprint [14].

Before selecting a log profile, define:

  • Expected gigabytes per day.
  • Retention in searchable storage.
  • Archive and legal-hold requirements.
  • Replication or availability model.
  • Which VCF, NSX, host, identity, and network logs are included.
  • Alert rules and SIEM forwarding.
  • Storage growth and deletion behavior.

A thirty-day retention target and a one-year evidence-retention target are fundamentally different architectures.

Identity, DNS, NTP, PKI, Backup, and Recovery

VCF depends on services that may live outside the VCF cluster. Those services are often omitted from the platform bill of materials because they already exist elsewhere. Their existing cost does not make them optional.

External dependencyBefore-first-workload requirementFailure if omitted or weak
DNSForward and reverse records, resilient resolvers, naming standard, delegated ownershipDeployment failures, certificate mismatch, broken integrations
NTPMultiple reachable sources with consistent timeAuthentication, certificate, logging, and cluster-consistency failures
Identity providerGroup model, administrative roles, break-glass accounts, lifecycle processExcess privilege, orphaned access, unavailable administration
PKICA integration, certificate profiles, renewal ownership, emergency replacementTrust failures, outages during expiration, manual certificate sprawl
Backup targetExternal repository, credentials, retention, immutability where requiredManagement-plane recovery becomes theoretical
Recovery orchestrationComponent recovery order, dependencies, validation, communicationsRestored appliances fail because supporting services are missing
Software depotConnected, proxy, or offline download process with entitlement accessLifecycle bundles unavailable during approved windows
IPAMReserved host, appliance, runtime, Edge, and future scale-out addressesDeployment collision and blocked expansion
SMTP, paging, or ITSMAlert routing and escalation ownershipHealth issues remain visible only inside the console
Privileged accessJump host, PAM, MFA, session evidence, break-glass procedureUncontrolled or unavailable administrative access

Broadcom documents both component backup and VCF instance recovery because recovery is broader than exporting one appliance configuration [15]. Management-plane recovery must account for sequence. DNS, NTP, identity, certificates, storage, network reachability, and backup repositories must be available before dependent components can become healthy.

The minimum production standard should include:

  • Scheduled file-based backups for supported components.
  • Backups stored outside the failure domain they protect.
  • Documented credentials and recovery access.
  • A tested management-domain restore plan.
  • A recovery validation that confirms lifecycle, identity, NSX, Operations, and workload visibility after restore.
  • Recovery-time and recovery-point objectives accepted by the business owner.

Workload-Domain Design and Cluster Expansion

A separate workload domain is not required simply because the environment has business VMs. It becomes valuable when the workload needs a different lifecycle, failure, security, ownership, or capacity boundary.

Reasons to create a separate workload domain

  • A different hardware generation or host image is required.
  • A different storage platform or policy is required.
  • The workload has distinct NSX, routing, or security ownership.
  • Maintenance must be independent of management-domain maintenance.
  • The platform supports multiple tenants or regulated zones.
  • Separate vCenter scale or administrative boundaries are needed.
  • GPU, database, VDI, or other specialized clusters need different admission policies.
  • The workload has a separate recovery or site strategy.

Expansion is a workflow, not a purchase order

Adding a host requires more than rack space. Before expansion, validate:

  • Hardware Compatibility Guide status.
  • Firmware and driver alignment with the target image.
  • Network switch ports, VLANs, MTU, LACP or teaming design, and cabling.
  • Storage devices, zoning, exports, or vSAN disk groups.
  • Host naming, DNS, NTP, BMC, and IP addresses.
  • License capacity and usage reporting.
  • Lifecycle bundle availability.
  • Cluster fault-domain and storage-policy effects.
  • VCF Operations inventory and health after commissioning.

External-storage designs that use converge or import pathways may require changes in vCenter followed by inventory synchronization in VCF Operations [12]. Build that sequence into the expansion runbook before the initial deployment, not after the first emergency capacity request.

Availability, Failure Domains, and Maintenance Reserve

A minimum viable production platform must survive events that are expected, not merely rare disasters. Host maintenance, firmware updates, certificate rotation, appliance patching, switch maintenance, storage-controller replacement, and operator error are normal events.

Capacity must be reserved for both failure and maintenance

Use this planning model:

Usable production capacity =
  physical capacity
  - management steady-state reserve
  - one-host failure reserve
  - maintenance and evacuation reserve
  - storage-policy and rebuild overhead
  - platform growth buffer
  - workload growth buffer

If the result is negative after one host is removed, the cluster is not production-ready regardless of what the installer allowed.

N+1 is not always enough

N+1 answers one question: Can the cluster lose one host? It does not answer:

  • Can one host be in maintenance while another experiences a fault?
  • Can vSAN rebuild while business I/O remains acceptable?
  • Can anti-affinity rules still be satisfied?
  • Can the three NSX Managers or three-node Automation cluster remain distributed?
  • Can an Edge pair avoid a shared physical failure domain?
  • Can the platform complete a lifecycle update without consuming the business workload’s performance reserve?

For small environments, N+1 may be the accepted economic boundary. That acceptance should be explicit, documented, and tested. For medium and enterprise environments, management clusters commonly need N+2 behavior at important maintenance points or enough spare capacity to approximate it.

Rack, power, switch, and storage fault domains

Host count alone can create false confidence. Six hosts in one rack with one top-of-rack switch and one power distribution path still share major failure domains.

Production design should place management components and Edge nodes across available:

  • Hosts.
  • Racks or chassis.
  • Power feeds and PDUs.
  • Top-of-rack switches or fabric members.
  • vSAN fault domains or external-array controllers.
  • Storage fabrics.
  • Sites, where the service level requires site recovery.

Hardware Compatibility and Network-Port Requirements

VCF compatibility is a complete-stack property. The server model, CPU, NIC, HBA, storage controller, boot device, firmware, BIOS settings, drivers, array microcode, and selected VCF bill of materials must form a supported combination [16], [17].

“ESX installed successfully” is not a compatibility test.

Hardware validation checklist

Before purchase or reuse, validate:

  • Server and CPU generation against the current Hardware Compatibility Guide.
  • NIC and HBA models, firmware, and driver combinations.
  • vSAN ReadyNode or component-level compatibility when using vSAN.
  • External array and transport interoperability when using FC, NFS, iSCSI, or NVMe-oF.
  • TPM, secure boot, encryption, and attestation requirements.
  • Boot-device endurance and support.
  • Vendor hardware-support-manager integration where used.
  • BIOS, NUMA, power, and PCIe settings for the intended workload.
  • Replacement-part availability over the planned platform life.

Practical host-port and bandwidth envelope

The exact design depends on the storage and network architecture. The following is a practical production starting point, not a universal product minimum.

ConnectionSmall production baselineMedium or enterprise patternDesign concern
Out-of-band management1 dedicated BMC port per hostRedundant management fabric where requiredRecovery access independent of ESX
Converged data network2 x 10 or 25 GbE, one per fabric member2 x 25 GbE or faster, often with additional separated portsManagement, vMotion, vSAN, overlay, workload, backup contention
External IP storageShared only after load validationDedicated pair or strongly governed QoSLatency, loss, MTU, congestion, failure isolation
Fibre Channel2 HBA ports across independent fabrics2 or 4 ports based on throughput and redundancyFabric, zoning, multipath, queue depth
NSX Edge uplinksRedundant uplinks and VLANsDedicated bandwidth, BGP peers, and route policyNorth-south throughput and stateful-service failover
Backup networkShared only for low-volume estatesDedicated or scheduled capacityBackup windows competing with vMotion and storage

Ten-gigabit Ethernet can satisfy many technical baselines. Twenty-five-gigabit Ethernet is often the more defensible converged production starting point because management, vMotion, storage, NSX overlay, backup, replication, and workload traffic can coincide during failure or maintenance.

Bandwidth should be sized for degraded operation, not only steady state.

Licensing and Evaluation Considerations

Licensing must be a design workstream, not the last installer screen.

Broadcom’s current VCF licensing model includes fleet licensing and license-usage reporting workflows. The documentation states that usage reports are required at least once every 180 days to maintain licenses [18]. That requirement creates an operational responsibility for the team that owns license connectivity, reporting, evidence, and escalation.

Evaluation mode is not a production strategy

A Broadcom community discussion describes a 90-day VCF 9.x evaluation period and clarifies, for a VCF 9.0 scenario, that vSAN was not capacity-capped during evaluation. The same response also makes clear that proper core and vSAN add-on entitlement was required before the period ended [19].

That is useful implementation context, but it is not a substitute for a current order document, contract, or account-team confirmation for VCF 9.1.

Use evaluation mode for:

  • Labs.
  • Proofs of concept.
  • Deployment rehearsals.
  • Short pilot periods with an approved licensing exit plan.

Do not use it to bridge an uncertain production procurement process. A production design should validate before deployment:

  • Licensed core counts and minimums.
  • Included and add-on vSAN capacity.
  • Entitlement for required VCF components and add-ons.
  • Support coverage and download rights.
  • License-server and usage-reporting workflow.
  • Renewal dates and five-year commercial assumptions.
  • Conversion or import rights for an existing VVF or vSphere environment.

A separate community discussion about moving between VVF and VCF reinforces the boundary: technical conversion requirements and commercial licensing rights are different questions. Commercial interpretation belongs with the account team and licensing documentation [20].

Day-Two Lifecycle Management

VCF earns its platform value through lifecycle consistency. That value is not automatic. It depends on disciplined prechecks, compatible binaries, working backups, capacity for evacuation, and operators who understand the dependency order.

Broadcom’s lifecycle documentation separates binary management, component lifecycle, and upgrade workflows because each introduces prerequisites and failure paths [21].

Before the first production workload, define the day-two process for:

  • Connected, proxy, or offline software-depot access.
  • Bundle download, staging, checksum, and approval.
  • Upgrade prechecks and remediation ownership.
  • Host image, firmware, and driver compatibility.
  • VCF Management Services, Operations, Automation, NSX, vCenter, and ESX sequencing.
  • Maintenance-mode capacity and VM evacuation.
  • Snapshot restrictions and supported rollback methods.
  • Certificate and password rotation.
  • Configuration drift and inventory synchronization.
  • Log collection and support-bundle retention.
  • Change windows, business communication, and escalation.
  • Post-upgrade health, lifecycle, automation, network, and recovery validation.

The lifecycle tax of optional components

Every optional platform component creates recurring work:

ComponentDay-two tax
VCF AutomationCatalog and blueprint lifecycle, integrations, policies, tenant support, content testing
Log managementRetention tuning, storage growth, parsing, alert rules, archive, access control
Operations for NetworksData-source credentials, flow collection, storage, network-model accuracy
NSX Edge servicesRouting changes, certificates, service state, throughput, failover testing
Multiple workload domainsMore vCenters, clusters, images, certificates, backups, lifecycle windows
Multi-site recoveryReplication, recovery mappings, runbooks, evidence, recurring tests

This is why the minimum viable VCF is not a static BOM. It is an operating commitment.

Staffing and Skills

VCF combines compute, storage, networking, security, identity, automation, observability, lifecycle, and recovery. One person may be capable of learning all of those domains. One person cannot provide resilient operational coverage for all of them.

Small environment

A small environment can be operated by two or three cross-skilled people when the organization has strong vendor support and adjacent network, storage, security, and identity teams. This does not necessarily mean three full-time VCF positions. It means at least two people can execute critical procedures and one person is not the only holder of administrative knowledge.

Required coverage includes:

  • vSphere and ESX operations.
  • VCF lifecycle and SDDC Manager workflows.
  • NSX management and basic routing awareness.
  • Storage administration.
  • Backup and recovery.
  • DNS, NTP, PKI, and identity integration.
  • Monitoring, alerting, and incident escalation.

Medium environment

A medium private cloud generally needs three to five role-equivalents across platform engineering and operations, with clear escalation into network, storage, security, identity, and application teams. VCF Automation introduces a platform-product function, not merely another appliance administrator.

Enterprise environment

An enterprise fleet commonly requires five to ten or more dedicated role-equivalents across architecture, platform engineering, SRE or operations, NSX, automation, security, capacity and FinOps, backup and recovery, and service management. The number varies with site count, tenant count, support hours, compliance, and automation maturity.

The correct staffing metric is not “administrators per host.” It is coverage of failure, lifecycle, security, service-catalog, and recovery responsibilities.

Before the First Workload Resource Table

The following table consolidates the complete readiness footprint. Values are planning envelopes that must be replaced by validated workbook output, hardware sizing, and service-level decisions.

Resource or capabilitySmall minimal-footprintMedium separated productionEnterprise fleet
Physical hosts5-6 shared hosts recommended4-6 management plus 6-12 workload hosts6-8 management plus multiple 4-16+ host workload clusters
Management availabilitySimple models accepted selectively; vSphere HA; strict reserveHA NSX, Operations, Automation as requiredHA management services, multiple collectors, site recovery
Management-platform allocation140-260 vCPU; 500-900 GB RAM300-600 vCPU; 1.2-2.5 TB RAM600-1,200+ vCPU; 2.5-6+ TB RAM
Management storage10-20 TB provisioned25-60 TB provisioned60-200+ TB provisioned
Storage modelvSAN or supported external storage; modest retentionSeparate management and workload tiers; stronger recoveryMultiple tiers, domains, arrays, sites, and retention classes
NSX ManagerSimple only with accepted reduced availability; otherwise three nodesThree-node clusterThree-node clusters per scale and isolation design
NSX EdgeOptional for VLAN-only first VM; two nodes when gateway services beginTwo or more production nodesMultiple Edge clusters by site, tenant, throughput, or security zone
VCF OperationsRequired operational visibility; simple model possibleHA model, collectors, capacity and config governanceScaled cluster, multiple collectors, broader integrations
VCF AutomationDefer only for admin-provisioned scope; simple model for modest serviceHA or production model with catalog governanceHA, multi-tenant policy, content pipelines, integrations
Log managementSmall ingest and short retentionHA ingest with defined retentionTiered retention, archive, SIEM, evidence controls
IP address reserveRoughly 60-100 addresses including runtime growthRoughly 100-180200-500+ across sites and services
Physical networkDual 10/25 GbE plus BMC; FC if usedDual 25 GbE or faster; service separation as neededRedundant fabrics, scalable routing, dedicated storage/Edge paths
External servicesDNS, NTP, identity, PKI, backup, depot, alertingRedundant services and tested recoveryMulti-site resilient dependencies and formal service ownership
Failure reserveOne host plus maintenance headroomOne host minimum, N+2 behavior at critical windowsCluster, rack, fabric, storage, and site failure planning
Staffing coverage2-3 cross-skilled operators with adjacent-team support3-5 role-equivalents plus specialists5-10+ dedicated role-equivalents across platform functions
Production gateTested host evacuation and management restoreTested lifecycle, failover, catalog, and recoveryAuditable SLOs, recovery exercises, capacity and cost governance

A small design is not disqualified because it uses simple availability models. It is disqualified when the reduced availability is hidden, untested, or incompatible with the business service level.

Five Year Cost Categories Without Confidential Pricing

A credible five-year view must include more than host purchase price and VCF subscription cost. The purpose of the model is not to guess confidential pricing. It is to make every cost-bearing decision visible.

Platform acquisition and entitlement

  • VCF subscription or term licensing.
  • Licensed cores and applicable minimums.
  • Included and add-on vSAN capacity.
  • NSX, automation, operations, or other entitlement validation.
  • Server, NIC, HBA, local storage, and boot devices.
  • External arrays, switches, optics, transceivers, and cables.
  • Backup software and repositories.
  • Load balancers, firewalls, IPAM, PAM, PKI, and monitoring integrations.
  • Operating-system and database licenses for supporting services.

Facilities and infrastructure

  • Rack space.
  • Power and cooling.
  • Data-center cross-connects.
  • WAN and inter-site bandwidth.
  • Hardware sparing and replacement inventory.
  • Warranty, support, and on-site service levels.

Implementation and transition

  • Architecture and validated design.
  • Planning workbook completion.
  • Network and storage changes.
  • DNS, PKI, identity, and security integration.
  • Migration tooling and coexistence.
  • Professional services.
  • Pilot, testing, and rollback capacity.
  • Training and runbook creation.

Recurring operations

  • Platform engineering and operations labor.
  • On-call coverage.
  • Capacity and performance management.
  • Lifecycle testing and change windows.
  • Certificate and credential operations.
  • Backup, restore, and disaster-recovery testing.
  • Log storage and archive.
  • Security monitoring and evidence retention.
  • Service catalog, blueprint, and integration maintenance.
  • Vendor support and escalation management.

Refresh and growth

  • Year-three through year-five capacity expansion.
  • Hardware-generation changes and mixed-cluster constraints.
  • Storage growth and telemetry retention.
  • New Edge throughput or network services.
  • Additional workload domains, sites, or tenants.
  • License growth from added cores or storage.
  • Replacement of components leaving support.

Risk and opportunity cost

  • Planned downtime during lifecycle events.
  • Capacity stranded for availability.
  • Delayed projects caused by insufficient IP, storage, or host reserve.
  • Business impact of management-plane outage.
  • Recovery testing and cyber-recovery controls.
  • Skills concentration and staff turnover.
  • Exit, migration, or platform-transition cost at the end of the horizon.

A useful five-year model compares scenarios with the same service level. A four-host minimal footprint should not be compared with a separated, recoverable platform as though they deliver identical availability, maintainability, and operational risk.

Use a transparent model such as:

Five-year platform cost =
  acquisition and entitlement
  + facilities and infrastructure
  + implementation and transition
  + recurring operations
  + refresh and growth
  + risk allowance
  - residual value

The model should identify its baseline date, currency, workload growth, capacity assumptions, included services, exclusions, and the threshold at which a different architecture becomes less expensive.

Decision Guidance

Choose the minimal-footprint pattern when the organization has a genuinely small workload estate, a limited number of tenants, a single operations team, conservative growth, and an accepted reduced failure envelope. Plan five or six hosts, reserve management capacity explicitly, and establish a trigger for separating workloads or adding hosts.

Choose a separated management and workload design when production service levels, maintenance independence, team boundaries, specialized hardware, compliance, or growth justify the additional infrastructure. The extra hosts are not waste. They purchase lifecycle isolation and protect the control plane from business-demand contention.

Choose broader enterprise fleet and multi-site patterns when the platform must support multiple business units, regulated zones, large automation catalogs, long telemetry retention, strict recovery objectives, or site-level failure. At that point, VCF is an internal cloud product and should be funded and staffed as one.

Finally, reconsider whether full VCF is the right first step when the organization does not need the integrated networking, automation, operations, lifecycle, tenancy, or fleet model and cannot staff the dependencies. A smaller vSphere-based platform may be operationally more honest than a minimally deployed VCF stack whose included services remain unowned.

Conclusion

The minimum viable VMware Cloud Foundation 9.1 is not four ESX hosts. Four hosts can satisfy an important deployment reference, but they do not describe the full platform investment required to run production workloads responsibly.

The real minimum includes a protected management domain, VCF Management Services and VCF Services Runtime, vCenter, SDDC Manager, VCF Operations, NSX Manager, the right Edge pattern, a deliberate VCF Automation decision, logging and retention, supported storage, external identity and infrastructure services, validated licensing, backup and recovery, lifecycle reserve, and enough people to operate the platform through failure and change.

For a small environment, the practical production threshold is usually five or six shared hosts with strict capacity governance, or a dedicated management cluster plus a separate workload cluster. Medium and enterprise designs should separate management and workload failure domains, use high-availability management models where the service level requires them, and fund observability, recovery, and staffing as first-class platform components.

The best architecture decision is therefore not the one with the smallest bill of materials. It is the smallest design that can lose a host, evacuate a host, recover its management plane, complete a lifecycle update, and still deliver the business workload without improvisation.

External References

Leave a Reply

Discover more from Digital Thought Disruption

Subscribe now to keep reading and get access to the full archive.

Continue reading