Site icon Digital Thought Disruption

Who Owns the Failure? Building a Support RACI for a Multivendor Private AI Platform

TL;DR

A multivendor private AI platform is not operationally complete when the hardware is installed, the GPUs are visible, and the first model endpoint responds. It is complete when the organization knows who performs the first diagnostic action when any part of the stack fails.

The customer should retain one accountable service owner and one incident commander. Individual platform teams own first-line diagnostics for their layers. Vendors own product-specific cases only within the scope of documented entitlements, validated configurations, and support policies. Fault ownership, case ownership, diagnostic ownership, and service accountability are different responsibilities.

The practical objective is not to predict every possible root cause. It is to ensure that every incident starts with an assigned diagnostic owner, a known-good configuration baseline, a synchronized evidence package, and an escalation path that prevents the customer from becoming the message bus between vendors.

Introduction

Private AI architecture diagrams usually show how everything connects.

Dell servers provide the physical compute. GPUs provide acceleration. VMware Cloud Foundation provides virtualization, networking, lifecycle management, and Supervisor services. Kubernetes provides workload orchestration. NVIDIA software provides drivers, operators, telemetry, inference services, and GPU scheduling. Storage platforms feed data into the GPUs. Physical and virtual networks connect the entire path. Azure Local may provide another infrastructure and Kubernetes operating environment.

The diagram looks integrated.

The support model often is not.

When a workload fails, the first question quickly changes from “What is broken?” to “Who should investigate it?” The hardware team sees a Kubernetes problem. The Kubernetes team sees a driver problem. The virtualization team sees a guest operating system problem. The NVIDIA case asks for platform evidence. The infrastructure vendor asks whether the GPU is visible to the hypervisor. The storage team says latency is within its own thresholds, while the AI team continues to report GPU starvation.

This is how technically valid multivendor architectures become operationally fragile.

The answer is not to nominate one vendor as the owner of every possible failure. No vendor should be assigned responsibilities beyond its documented products, support agreements, validated configurations, and entitlement boundaries.

The answer is to design an internal support operating model that survives uncertainty.

Multivendor Architecture Is a Chain of Diagnostic Boundaries

A private AI platform is not one product. It is a dependency chain.

A production inference request might depend on:

The end user sees one failed endpoint.

The platform team sees a distributed system with more than a dozen plausible failure domains.

That means ownership cannot begin with the root cause. The root cause is not yet known. Ownership must begin with the first diagnostic action.

The Support Ownership Model at a Glance

The model below separates operational accountability from vendor product support.

The most important detail is the central incident commander.

The incident commander does not need to be the deepest GPU, storage, network, or Kubernetes expert. The incident commander owns forward motion. That includes the timeline, business-impact statement, evidence package, linked support cases, technical hypotheses, handoffs, escalation timers, and communication.

Without that role, each technical team may investigate competently while the incident itself remains unmanaged.

Fault Ownership Is Not Case Ownership

Support discussions become confused when every responsibility is described as “ownership.” A useful model separates five different concepts.

Service Accountability

The customer’s AI service owner remains accountable for the service outcome.

That accountability does not transfer to Dell, Broadcom, NVIDIA, Microsoft, an integrator, or a managed-service provider simply because a support case has been opened. The service owner decides whether the service is acceptable, whether a workaround is adequate, and whether an incident can be closed.

First Diagnostic Ownership

The first diagnostic owner is the team required to collect evidence and isolate the failure at the layer where the symptom first appears.

This responsibility exists before the root cause is known.

For example, if a Kubernetes node stops advertising GPU resources, the Kubernetes and AI platform team should perform the first diagnostic action. That does not mean Kubernetes caused the failure. It means the symptom is visible at that layer, and that team is best positioned to determine whether the GPU is absent from the operating system, the NVIDIA device plugin, the runtime, or the Kubernetes resource inventory.

Fault Ownership

Fault ownership belongs to the layer or component that caused the failure.

It may eventually belong to hardware, firmware, ESXi, NSX, an NVIDIA driver, a Kubernetes operator, a CSI implementation, a switch configuration, a storage path, an application, or a customer configuration.

Fault ownership should be assigned only after evidence establishes the causal boundary.

Case Ownership

Case ownership identifies which organization currently holds an active product support request.

A Broadcom case may investigate ESXi enumeration while an NVIDIA case investigates driver compatibility and a Dell case investigates PCIe hardware events. More than one vendor case can be valid at the same time.

The customer should still maintain one master incident record that links all vendor cases.

Remediation Ownership

The team authorized to implement the fix owns remediation.

The party that identifies the defect is not necessarily the party permitted to change production. A vendor may recommend a firmware update, but the customer change owner, platform team, and hardware team still need to approve, schedule, execute, validate, and potentially roll back that change.

A Support RACI Does Not Change Vendor Contracts

An internal RACI is an operating model, not a contractual amendment.

It does not create a vendor support commitment. It does not make an unsupported configuration supported. It does not override lifecycle policies. It does not guarantee that a vendor will accept responsibility for another vendor’s component.

Vendor engagement must still be checked against:

The RACI determines what the customer does first and how evidence is routed. The relevant contract determines what a vendor is obligated to do after engagement.

Responsibility Boundaries by Platform Layer

Dell Hardware, Firmware, BIOS, GPU, NIC, Switch, and Storage

The infrastructure team should own the initial physical-platform assessment, even when a Dell or other OEM case may eventually be required.

Its responsibilities should include:

Dell documents SupportAssist collections through iDRAC as a mechanism for gathering platform information used during troubleshooting [15]. That collection should be part of the evidence package, not an action first attempted after several hours of vendor discussion.

The hardware vendor should be a case target when the failure is isolated below the hypervisor or operating-system boundary, such as:

The hardware case should include the exact server service tag, component part numbers, slot topology, firmware baseline, support collection, timestamps, and the result of any hardware diagnostics.

Physical Switch Responsibility

The network team owns physical switch configuration and first-line fabric diagnostics.

This includes:

NVIDIA Network Operator can configure host networking components and Kubernetes resources, but it does not make the physical switch fabric someone else’s responsibility.

Storage Responsibility

The storage team owns first-line diagnostics for the storage platform and its data path.

That includes:

A storage platform can report “healthy” while still failing the AI service objective. The important question is not only whether the array is online. It is whether the complete data path is delivering the throughput, latency, concurrency, and metadata behavior expected by the workload.

VMware Cloud Foundation, ESXi, vCenter, NSX, and Supervisor

The VMware platform team should own first diagnostics when the symptom appears at the virtualization, VCF lifecycle, virtual-networking, or Supervisor layer.

Its responsibilities should include:

The Broadcom Compatibility Guide should be treated as evidence, not merely a procurement-time reference [1]. The team should preserve the exact query, result, date, device model, driver, firmware, and relevant ESXi release used during validation.

For VCF 9.1 environments, SDDC Manager and VCF Management Services Platform log collection also needs to be part of the runbook. Broadcom documents the relevant SOS log-collection paths and options for VCF environments [3].

The Broadcom case target is appropriate when evidence indicates a failure involving supported VMware software behavior, such as:

A case should not simply say “NVIDIA GPU missing.” It should state precisely where the device is present and where it disappears.

Kubernetes and AI Platform Operations

Kubernetes is the boundary where infrastructure conditions become workload-consumption problems.

The Kubernetes and AI platform team should own first diagnostics for:

This team is often the first diagnostic owner even when another layer caused the failure.

For example, a pod stuck in Pending might be caused by:

The first task is to identify which constraint is blocking placement. It is not to immediately open cases with every vendor in the architecture.

NVIDIA vGPU and NVIDIA AI Enterprise

The NVIDIA software team, or the platform team assigned to NVIDIA components, should own first diagnostics for the NVIDIA software chain.

NVIDIA vGPU Software

The diagnostic boundary includes:

NVIDIA publishes a VMware vSphere ESXi support matrix for its current vGPU releases [6]. The evidence package should record the exact hypervisor build, NVIDIA vGPU Manager version, guest driver version, GPU model, vGPU profile, guest operating system, and licensing state.

A device visible to ESXi but unavailable to a VM is not the same incident as a device visible to the VM but unavailable to CUDA.

Those are different diagnostic boundaries.

NVIDIA AI Enterprise

NVIDIA AI Enterprise now documents separate infrastructure and application layers, with independently versioned drivers, Kubernetes operators, Run:ai, and application software [5].

That makes release-branch evidence important. “NVIDIA AI Enterprise is installed” is not sufficient diagnostic information.

The support package should identify:

NVIDIA GPU Operator

The GPU Operator manages multiple dependencies, including drivers, the NVIDIA Container Toolkit, the device plugin, GPU feature discovery, and DCGM-based monitoring.

The Kubernetes and GPU platform team should first inspect:

NVIDIA’s current troubleshooting documentation recommends collecting the GPU Operator must-gather archive when standard troubleshooting does not isolate the issue [7].

That archive should be collected before configuration changes remove the original evidence.

NVIDIA Network Operator

The Network Operator boundary includes Kubernetes-side provisioning of networking components used for high-speed networking, RDMA, secondary networks, device plugins, and related host software [8].

The first diagnostic package should include:

A Network Operator case should not be used to bypass physical fabric diagnostics.

NVIDIA NIM

When a NIM endpoint fails, the AI platform team should inspect the complete startup chain:

NIM cache resources may temporarily report incomplete or failed reconciliation while large images or models are being downloaded. NVIDIA documents checking the cache job, pod, custom-resource conditions, and logs to distinguish a transient download state from a persistent failure [9].

The first case should identify whether the failure occurs during image pull, model download, cache preparation, container initialization, GPU initialization, model loading, health checking, or service exposure.

NVIDIA DCGM

DCGM should be treated as a shared diagnostic service, not only as a dashboard data source.

It provides GPU health checks, diagnostics, job-level statistics, and continuous telemetry that can help separate:

NVIDIA describes DCGM as supporting both passive health monitoring and active diagnostics [10].

The GPU platform owner should preserve DCGM evidence for the affected workload window before restarting nodes or moving workloads.

NVIDIA Run:ai

Run:ai adds another control and scheduling layer.

Its diagnostic responsibilities include:

A workload waiting for a GPU may represent correct scheduler behavior rather than a GPU platform failure.

The Run:ai owner should determine whether the workload is:

Run:ai publishes its own product support policy and version lifecycle. Those documents should be checked before assuming that every deployed release has the same support status [11].

Microsoft Azure Local Responsibilities

Azure Local introduces a similar shared-responsibility boundary.

The Azure Local platform team should own first diagnostics for:

Microsoft currently documents both Discrete Device Assignment and GPU partitioning as GPU attachment approaches for Azure Local workloads [12]. The selected method must be captured in the incident evidence because the availability, migration, isolation, and driver behavior can differ.

The hardware vendor remains the likely case target for a device that is absent from firmware or host inventory. Microsoft becomes a likely case target when the supported device is visible to the Azure Local host but fails at the Azure Local platform, Arc, VM-management, or AKS integration layer. NVIDIA may become a case target when the device is assigned successfully but the supported NVIDIA guest or Kubernetes software fails.

Azure Local provides on-demand diagnostic log collection through the Azure portal and PowerShell, subject to the documented prerequisites and feature state [13]. Azure Monitor also exposes compute, storage, and network metrics that can help correlate failures across the platform [14].

The customer should collect that evidence before changing extensions, recreating virtual machines, or resetting configuration.

Customer, Integrator, MSP, and Vendor Roles

A support RACI needs to distinguish operational roles from commercial relationships.

Customer

The customer should remain accountable for:

The customer may outsource execution, but it cannot outsource the need for clear accountability.

Systems Integrator

The integrator’s responsibilities should be defined by the statement of work.

Typical project-phase responsibilities may include:

An integrator should not be assumed to provide indefinite production support unless the contract says so.

Managed-Service Provider

The MSP may own day-two monitoring, incident response, patching, vendor cases, or platform administration, depending on the service contract.

The contract should state:

“Managed platform” is not precise enough for a multivendor AI stack.

Vendors

Each vendor should be engaged for the products and services covered by its entitlement.

Vendors may need to collaborate, but the customer should not assume that collaboration will happen automatically or that one vendor will manage another vendor’s case.

The customer’s case coordinator should maintain the shared timeline and ensure that each vendor receives the same relevant evidence.

Build Compatibility Evidence Before the Incident

Compatibility evidence should be generated during design and commissioning, not reconstructed during an outage.

LayerEvidence to PreserveWhy It Matters
Server platformModel, service tag, CPU, memory, GPU, NIC, HBA, storage controllerEstablishes the physical configuration
FirmwareBIOS, BMC, GPU, NIC, HBA, controller, switch firmwareIdentifies linked firmware dependencies
BIOSSR-IOV, virtualization, memory mapping, device settingsExplains enumeration and assignment behavior
VMwareVCF, SDDC Manager, ESXi, vCenter, NSX, Supervisor buildsEstablishes the virtualization platform state
NVIDIA vGPUvGPU Manager, guest driver, GPU profile, license stateProves host-to-guest compatibility
KubernetesDistribution, version, runtime, CNI, CSIDefines the orchestration environment
NVIDIA operatorsGPU Operator, Network Operator, NIM OperatorEstablishes operator and component versions
AI servicesNIM container, model profile, Run:ai release and policiesDefines the workload-control layer
Azure LocalAzure Local build, Arc extensions, GPU mode, AKS versionEstablishes Microsoft platform state
StoragePlatform version, CSI driver, protocol, multipathing, networkDefines the data path
NetworkSwitch model, firmware, topology, MTU, RDMA settingsDefines the workload and storage fabric

Every compatibility record should include:

A screenshot without the query parameters and date is weak evidence. A spreadsheet cell saying “supported” is weaker.

Maintain a Known-Good Configuration Baseline

A known-good baseline should define more than software versions.

It should also record the platform behavior that proved the stack worked.

A useful baseline includes:

The baseline should also contain test evidence:

A baseline is valuable because it converts “this used to work” into measurable evidence.

Collect Logs Before Opening Cases

The first evidence package should be created once and reused across vendor cases.

At minimum, it should contain:

Layer-Specific Evidence

Hardware and Dell evidence

VCF evidence

Kubernetes and NVIDIA evidence

Storage evidence

Azure Local evidence

Time synchronization is essential. Logs that differ by several minutes can create false causal sequences and send the investigation toward the wrong layer.

Reproduce the Failure at the Correct Layer

The strongest troubleshooting method is controlled reduction.

Do not reproduce the entire production workload first. Reduce the problem until only one or two layers remain.

This sequence can establish where the failure begins.

Examples include:

A reproduction is useful only when it removes variables.

Prevent Circular Vendor Referrals

Vendor ping-pong usually starts with incomplete evidence and weak handoffs.

A practical anti-referral protocol should require the following.

One Master Incident

The customer maintains one master incident record containing:

One Case Coordinator

The coordinator owns communication between vendors.

Engineers can communicate directly during technical sessions, but the coordinator ensures that decisions, requests, and conclusions return to the master incident.

A Written Handoff Standard

A team or vendor should not transfer diagnostic ownership by writing only “not our issue.”

A valid handoff should state:

A Shared Hypothesis Register

HypothesisEvidence ForEvidence AgainstOwnerNext TestStatus
GPU hardware faultPCIe event at failure timeDevice passes later health checkHardware teamRun offline diagnosticsOpen
vGPU compatibilityFailure began after ESXi updateMatrix review not completeVCF teamValidate full version chainOpen
Storage starvationGPU idle and read latency elevatedSingle-node local test healthyStorage teamCompare local and shared dataLikely
Scheduler constraintWorkload remains pendingGPU resources availableRun:ai teamReview quota and node-pool policyOpen

The register makes reasoning visible and reduces repeated tests.

Parallel Cases When Boundaries Overlap

Parallel cases are appropriate when evidence genuinely crosses product boundaries.

For example, after a VCF update, a GPU initialization failure may justify:

Parallel cases should share the same version matrix, timestamp, reproduction, and change record.

Severity Definitions and Escalation Triggers

Vendor severity definitions vary and must be checked against the applicable support agreement. Broadcom, for example, documents severity levels based on total service loss, severe degradation, minor impact, and general requests [2].

The organization should maintain its own internal severity model and map it to each vendor’s model.

Internal SeverityExample ImpactRequired Internal ActionEscalation Trigger
Severity 1Production AI service unavailable, multiple critical tenants affected, no workaroundIncident commander, continuous bridge, executive communication, immediate evidence collectionNo accepted diagnostic owner, service-loss expansion, data or safety concern
Severity 2Severe degradation, missed service objectives, partial tenant impactNamed technical lead, frequent updates, parallel diagnosticsWorkaround failing, impact increasing, unresolved cross-vendor boundary
Severity 3Limited impact, workaround available, noncritical environmentNormal support workflow and tracked investigationRepeated occurrence, growing scope, approaching change deadline
Severity 4Question, planned validation, documentation, proactive compatibility reviewPlanned case or advisory request where entitlement permitsChange blocked or risk becomes production-impacting

Severity should reflect business impact, not the number of GPUs involved.

One failed GPU in a resilient development pool may be Severity 3. One failed GPU that prevents a regulated production service from running may be Severity 1.

Escalation Triggers

Escalation should occur when:

These are internal management triggers, not claims about vendor response commitments.

Change Control Across the Linked Stack

The private AI stack should not be patched as a collection of unrelated products.

The compatibility chain may look like this:

A change at the top can affect every layer below it.

Required Change Evidence

Every linked-stack change should record:

VCF Updates and NVIDIA Compatibility

A VCF or ESXi update should not be approved solely because the VCF lifecycle workflow allows it.

The change review should also validate:

A lifecycle tool can prove that an update is available. It does not independently prove that every external dependency in the AI platform has been validated.

Maintenance-Window Coordination

A maintenance window is not only a time reservation.

It is a staffed diagnostic agreement.

For changes affecting the linked AI stack, the window plan should identify:

The team should agree on:

The worst time to discover that the only GPU specialist is unavailable is after the hosts have already been upgraded.

RACI Matrix for Common Private AI Incidents

The following roles are used in the matrix:

The matrix is a recommended customer operating model. Vendor obligations remain governed by the applicable contracts and support policies.

IncidentAccountableResponsible for First Diagnostic ActionConsultedInformed
GPU missing from a VMASOVCFHW, KAI, SI/MSP, relevant VSIC, application owner
GPU missing from a Kubernetes nodeASOKAIVCF or ALO, HW, NVIDIA platform owner, relevant VSIC, tenant owner
Distributed training performance collapseASOKAI performance leadNET, STO, HW, VCF or ALO, Run:ai owner, VSIC, application owner
NIM endpoint fails to startASOKAISTO, NET, VCF or ALO, registry owner, NVIDIA owner, VSIC, consuming application team
Storage latency causes GPU starvationASOSTOKAI, NET, VCF or ALO, storage vendor, SI/MSPIC, application owner
VCF update creates driver compatibility concernVCF service ownerVCF lifecycle leadHW, KAI, NVIDIA owner, Broadcom, Dell, SI/MSPASO, change advisory board
Azure Local node reports GPU integration problemsAzure Local service ownerALOHW, KAI, NVIDIA owner, Microsoft, SI/MSPIC, application owner
Azure Local node reports network integration problemsAzure Local service ownerALO and NETHW, KAI, Microsoft, switch vendor, SI/MSPIC, application owner

The incident commander remains responsible for incident coordination even when another team is responsible for diagnostics.

Incident Example: A GPU Is Missing from a VM or Kubernetes Node

First Diagnostic Owner

Diagnostic Sequence

Begin at the lowest observable layer:

  1. Is the GPU visible in iDRAC or the system firmware?
  2. Is it visible to ESXi or the Azure Local host?
  3. Is the correct assignment mode configured?
  4. Is the GPU attached to the intended VM?
  5. Is it visible inside the guest operating system?
  6. Does the NVIDIA guest driver initialize?
  7. Does the container runtime expose the device?
  8. Does Kubernetes advertise the expected resource?
  9. Does Run:ai or the native scheduler permit placement?

Routing Logic

Do not replace the failed node or reinstall the operator until the original evidence has been captured.

Incident Example: Distributed Training Performance Collapses

A performance collapse is a system incident, not automatically a GPU incident.

First Diagnostic Owner

The AI platform performance lead should coordinate the first diagnostic action because the symptom spans workload, scheduler, GPU, storage, and network behavior.

Evidence to Correlate

Isolation Tests

Run controlled comparisons:

Diagnostic Interpretation

The evidence should show where scaling efficiency is lost, not merely that the job is slower.

Incident Example: A NIM Endpoint Fails to Start

First Diagnostic Owner

Kubernetes and AI platform operations.

Diagnostic Sequence

  1. Was the workload admitted?
  2. Was it scheduled?
  3. Was the required GPU resource allocated?
  4. Was the container image pulled?
  5. Did registry or NGC authentication succeed?
  6. Was the model cache created or mounted?
  7. Did the persistent volume attach?
  8. Did the model profile resolve?
  9. Did the container initialize the GPU?
  10. Did the model load?
  11. Did the health probe pass?
  12. Is the service reachable?

Routing Logic

“NIM failed” is not a useful case description. The startup stage must be identified.

Incident Example: Storage Latency Causes GPU Starvation

First Diagnostic Owner

Storage operations, after the AI platform team demonstrates a correlation between workload stalls and data-path behavior.

Required Evidence

Isolation Tests

Routing Logic

GPU starvation is an outcome. The diagnostic goal is to determine whether the starvation begins in the application, file system, storage network, virtualization path, or backend platform.

Incident Example: A VCF Update Creates a Driver Compatibility Concern

This is ideally a change-risk event, not a production incident.

Accountable Owner

VCF service owner and change owner.

Responsible First Action

The VCF lifecycle lead should stop the change from progressing until the entire dependency chain has been reviewed.

Required Review

Safe Execution Pattern

  1. Export the current known-good baseline.
  2. Validate the target version chain.
  3. Rehearse in a representative environment.
  4. Update a canary host or cluster.
  5. Run the complete GPU, network, storage, Kubernetes, and NIM validation suite.
  6. Observe the canary.
  7. Continue only after evidence-based approval.
  8. Preserve the rollback point until production validation is complete.

A successful ESXi boot is not sufficient acceptance evidence for a private AI host.

Incident Example: An Azure Local Node Reports GPU or Network Integration Problems

First Diagnostic Owner

Azure Local operations, with immediate consultation from hardware and network operations.

Diagnostic Sequence

  1. Is the node healthy and supported?
  2. Is the GPU or NIC visible to the host?
  3. Are firmware and drivers aligned with the approved Azure Local baseline?
  4. Is the intended GPU mode DDA or GPU-P?
  5. Does the device assignment exist on the VM?
  6. Is the device visible inside the guest?
  7. Does the NVIDIA driver initialize?
  8. Does AKS enabled by Azure Arc advertise the GPU where applicable?
  9. Are Arc extensions and VM-management components healthy?
  10. Are virtual-switch, physical-switch, MTU, and RDMA settings correct?
  11. Do Azure Local metrics show link, storage, or compute anomalies?
  12. Has the Azure Local diagnostic bundle been collected?

Routing Logic

The evidence should make the transition point between hardware, Azure Local, guest software, and Kubernetes visible.

Operational Documentation Required Before Handoff

A private AI platform should not transition into production operations until the following artifacts are complete.

Architecture and Dependency Documentation

Ownership Documentation

Configuration Documentation

Operational Runbooks

Validation Documentation

Support Documentation

If these documents do not exist, the platform has been installed but not operationally handed over.

Joint Problem-Management Reviews

Recurring or cross-vendor incidents should move from incident management into problem management.

A joint review should include:

The review should examine:

The result should be one problem record with named actions and due dates.

Closing three vendor cases does not resolve a recurring platform problem if the customer still cannot explain why the service fails.

A Practical Implementation Sequence

A support RACI can be built without waiting for the next incident.

Define the Service Boundary

Document what the private AI service includes:

Name the Accountable Service Owner

One role must be accountable for the end-to-end service.

Assign First Diagnostic Owners

Assign an initial diagnostic team to every major symptom class.

Do not wait until the root cause is known.

Build the Version and Compatibility Register

Capture the complete version chain and validation evidence.

Create the Known-Good Baseline

Run and preserve the commissioning tests.

Create Layered Evidence Runbooks

Define what each team must collect before opening a case.

Establish the Master-Incident Process

Require one timeline, one coordinator, linked vendor cases, and written handoffs.

Map Internal Severity to Vendor Severity

Do not assume the terms are identical.

Test the RACI

Run tabletop and technical exercises for:

Review After Every Major Change

Update the RACI, baselines, evidence procedures, and escalation contacts whenever the platform architecture or support contracts change.

Conclusion

A multivendor private AI platform is not operationally complete because all its components are individually supported.

It is operationally complete when the organization knows what happens between the first symptom and the first defensible technical hypothesis.

The customer needs one accountable service owner, one incident commander, named first diagnostic owners, evidence-driven escalation, and a master incident record that connects every vendor case. Hardware, VCF, Kubernetes, NVIDIA software, storage, networking, and Azure Local each have distinct diagnostic boundaries. Those boundaries need to be documented before production failure exposes them.

Fault ownership may take hours or days to determine. First diagnostic ownership should take seconds.

That is the practical purpose of the support RACI.

The platform architecture explains how the technology works when everything is healthy. The support architecture explains how the organization works when it is not.

External References

[1] Broadcom: VMware by Broadcom Compatibility Guide
Canonical URL: https://compatibilityguide.broadcom.com/

[2] Broadcom: Best Practices for Creating Support Requests on the Broadcom Support Portal
Canonical URL: https://knowledge.broadcom.com/external/article/429027/best-practices-for-creating-support-req.html

[3] Broadcom: Collecting SDDC Manager and VMSP Logs
Canonical URL: https://knowledge.broadcom.com/external/article/385749/collecting-sddc-manager-and-vmsp-logs.html

[4] Broadcom: VMware Cloud Foundation 9.1
Canonical URL: https://techdocs.broadcom.com/us/en/vmware-cis/vcf/vcf-9-0-and-later/9-1.html

[5] NVIDIA: NVIDIA AI Enterprise Documentation
Canonical URL: https://docs.nvidia.com/ai-enterprise/release-8/latest/index.html

[6] NVIDIA: VMware vSphere ESXi Support
Canonical URL: https://docs.nvidia.com/vgpu/latest/product-support-matrix/vmware-vsphere.html

[7] NVIDIA: Troubleshooting the NVIDIA GPU Operator
Canonical URL: https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/troubleshooting.html

[8] NVIDIA: NVIDIA Network Operator v26.1.0
Canonical URL: https://docs.nvidia.com/networking/display/kubernetes2610/nvidia-network-operator-v26-1-0.pdf

[9] NVIDIA: Caching NIM Models
Canonical URL: https://docs.nvidia.com/nim-operator/latest/cache.html

[10] NVIDIA: NVIDIA DCGM Documentation Overview
Canonical URL: https://docs.nvidia.com/datacenter/dcgm/latest/user-guide/index.html

[11] NVIDIA: NVIDIA Run:ai Product Support Policy
Canonical URL: https://run-ai-docs.nvidia.com/saas/support-policy/product-support-policy

[12] Microsoft: Prepare GPUs for Azure Local
Canonical URL: https://learn.microsoft.com/en-us/azure/azure-local/manage/gpu-preparation

[13] Microsoft: Collect Diagnostic Logs for Azure Local
Canonical URL: https://learn.microsoft.com/en-us/azure/azure-local/manage/collect-logs

[14] Microsoft: Monitor Azure Local with Azure Monitor Metrics
Canonical URL: https://learn.microsoft.com/en-us/azure/azure-local/manage/monitor-cluster-with-metrics

[15] Dell Technologies: PowerEdge Export a SupportAssist Collection Using iDRAC UI or RACADM Command
Canonical URL: https://www.dell.com/support/kbdoc/en-us/000126308/export-a-supportassist-collection-via-idrac9

Exit mobile version