
TL;DR
NVIDIA Run:ai can turn a shared Kubernetes GPU cluster into a governed multi-tenant platform by organizing workloads into departments and projects, assigning guaranteed GPU quotas per node pool, and allowing controlled over-quota use when capacity would otherwise remain idle.
The design depends on four controls working together:
- Quota establishes the resource entitlement for a department or project.
- Fairshare determines how unused capacity is distributed.
- Priority orders workloads within the relevant scheduling queue.
- Preemptibility determines whether borrowed capacity can be reclaimed.
Production workloads should normally run from quota-backed projects as non-preemptible workloads. Development and experimental workloads can use preemptible over-quota capacity, provided they support checkpointing and tolerate interruption.
The objective is not to keep every GPU permanently allocated. It is to make allocation predictable, measurable, recoverable, and aligned with business ownership.
Introduction
A Kubernetes cluster can expose GPUs as schedulable resources, but that does not automatically make it a multi-tenant AI platform.
Without an additional governance layer, the team that submits first, requests the most GPUs, or leaves the most workloads running can consume a disproportionate amount of the cluster. Other teams may have technically valid workloads but no predictable path to capacity.
NVIDIA Run:ai addresses this problem by placing organizational and scheduling controls above the raw Kubernetes resource model. Departments, projects, node pools, quotas, workload policies, priority, and preemption become part of one allocation system.
The difficult part is not creating a few projects and assigning arbitrary GPU numbers. The difficult part is deciding:
- Which organizational boundary owns each quota?
- Which workloads receive guaranteed access?
- Which teams may borrow unused GPUs?
- Which workloads can be interrupted?
- How is expensive hardware separated from general-purpose capacity?
- Who approves temporary exceptions?
- How do operators know when the quota model is no longer working?
This tutorial builds a practical configuration model that answers those questions.
What You Will Build
By the end of the tutorial, you will have a design for:
- Mapping departments, teams, products, and environments into Run:ai departments and projects
- Assigning users and identity groups through scoped access rules
- Separating GPU hardware into meaningful node pools
- Defining guaranteed GPU quotas per department, project, and node pool
- Allowing controlled over-quota consumption
- Establishing workload priority and preemption rules
- Separating development and production scheduling behavior
- Measuring quota utilization and queue pressure
- Preventing a single team from monopolizing shared GPU resources
- Operating a documented exception and quota-review process
The examples assume a shared Kubernetes cluster containing multiple GPU types. The same governance model can be scaled down for a smaller environment or extended across a larger NVIDIA Run:ai deployment.
Prerequisites and Assumptions
Before configuring tenant scheduling, confirm the following:
- NVIDIA Run:ai is installed and connected to the target Kubernetes cluster.
- GPU nodes are visible and healthy.
- GPU resources are correctly advertised to Kubernetes.
- Users authenticate through the expected identity provider.
- Identity groups exist for department owners, project administrators, developers, production operators, and platform administrators.
- You have permission to create departments, projects, node pools, roles, access rules, and workload policies.
- Workload owners understand whether their applications can tolerate preemption.
- Training workloads that may be preempted have a tested checkpoint and resume process.
This tutorial does not cover the initial NVIDIA Run:ai installation, Kubernetes GPU enablement, MIG configuration, or time-slicing configuration. Those capabilities should be operational before tenant quotas are introduced.
Understand the Run:ai Scheduling Hierarchy
The administrative hierarchy and the scheduling hierarchy are related, but they are not exactly the same thing.
Departments and projects are visible organizational objects. Queues are part of the scheduler’s allocation process. A queue is not usually another tenant container that administrators must create for every team.
The scheduler evaluates project and department demand for each applicable node pool.
The following diagram shows the relationship.
There are several important operational implications.
First, quota is evaluated in the context of a node pool. A project can have a guaranteed allocation on one GPU class and only best-effort access to another.
Second, a project’s workload priority affects the ordering of work in its scheduling context. Priority should not be treated as a substitute for cross-team quota design.
Third, unused capacity can be distributed as over-quota capacity. This increases utilization, but it does not convert borrowed capacity into a permanent entitlement.
Fourth, capacity reclamation depends on workloads being preemptible. A non-preemptible workload should fit within the project’s available deserved quota.
Map the Organization Before Creating Projects
The first configuration task should happen outside the platform.
Create a mapping of business ownership, technical ownership, workload purpose, and resource requirements before creating departments or projects. Otherwise, the platform structure tends to reproduce inconsistent team names and short-lived organizational charts.
Decide What a Department Represents
A department should represent a durable governance or ownership boundary.
Good department candidates include:
- Product engineering
- Applied research
- Central data science
- Enterprise analytics
- Shared AI platform services
- Business-unit AI teams
A department is useful when several related projects should share:
- A parent quota
- Common policy
- A common department owner
- Access to the same assets
- Similar node-pool preferences
- A single budget or capacity plan
Do not create a department for every temporary experiment. That creates unnecessary administrative depth and makes quota ownership difficult to maintain.
Decide What a Project Represents
A project is the practical scheduling and access boundary for workloads.
A project can represent:
- A product
- A model family
- A team
- An application
- A development environment
- A production service
- A temporary initiative
Projects are where administrators apply resource quotas, node-pool preferences, access controls, and workload policies.
Avoid using one project for an entire department if its workloads have different availability, security, or preemption requirements.
Separate Environment and Workload Intent
Development and production should not share a project merely because the same team owns both.
A useful starting structure is:
| Department | Project | Purpose | Scheduling intent |
|---|---|---|---|
| Product AI | product-ai-dev | Interactive development and test training | Small guarantee, preemptible over-quota allowed |
| Product AI | product-ai-prod | Production inference and release training | Guaranteed capacity, non-preemptible |
| Applied Research | foundation-research | Long-running model experiments | Moderate guarantee, preemptible over-quota |
| Shared Platform | ai-platform-services | Monitoring, controllers, and shared services | Small reserved guarantee |
| Shared Platform | performance-validation | NCCL, benchmark, and validation jobs | Scheduled windows or controlled best effort |
This structure makes operational intent visible before a workload is submitted.
Build a Source-Controlled Tenant Plan
NVIDIA Run:ai objects can be configured through supported interfaces, but the approved allocation model should also exist in source control.
The following YAML is a governance worksheet. It is not presented as a native Run:ai import format. Its purpose is to give the platform team a reviewable source of truth before applying configuration through the user interface or API.
organization:
departments:
- name: product-ai
owner_group: grp-product-ai-platform
node_pool_quotas:
h100-production:
guaranteed_gpus: 6
allow_over_quota: false
max_gpus: 6
a100-shared:
guaranteed_gpus: 6
allow_over_quota: true
over_quota_weight: medium
max_gpus: 10
projects:
- name: product-ai-dev
access_group: grp-product-ai-developers
node_pool_quotas:
h100-production:
guaranteed_gpus: 0
allow_over_quota: false
max_gpus: 0
a100-shared:
guaranteed_gpus: 2
allow_over_quota: true
over_quota_weight: medium
max_gpus: 6
default_preemptibility: preemptible
default_priority: medium-low
- name: product-ai-prod
access_group: grp-product-ai-production
node_pool_quotas:
h100-production:
guaranteed_gpus: 6
allow_over_quota: false
max_gpus: 6
a100-shared:
guaranteed_gpus: 4
allow_over_quota: false
max_gpus: 4
default_preemptibility: non-preemptible
default_priority: very-high
- name: applied-research
owner_group: grp-applied-research
node_pool_quotas:
h100-production:
guaranteed_gpus: 2
allow_over_quota: true
over_quota_weight: low
max_gpus: 4
a100-shared:
guaranteed_gpus: 8
allow_over_quota: true
over_quota_weight: high
max_gpus: 14
projects:
- name: foundation-research
access_group: grp-foundation-models
node_pool_quotas:
h100-production:
guaranteed_gpus: 2
allow_over_quota: true
over_quota_weight: low
max_gpus: 4
a100-shared:
guaranteed_gpus: 8
allow_over_quota: true
over_quota_weight: high
max_gpus: 14
default_preemptibility: preemptible
default_priority: medium
The values that matter most are:
- Guaranteed GPUs per node pool
- Whether over-quota use is permitted
- The relative over-quota weight
- The maximum allocation
- Default workload priority
- Default preemptibility
- Identity group ownership
- Production versus development intent
Successful implementation means the live Run:ai configuration matches this approved model and any variance has a recorded exception.
Configure Node Pools Around Infrastructure Characteristics
Node pools group Kubernetes nodes using labels. They are useful for separating heterogeneous hardware and controlling where workloads are entitled to run.
Common node-pool dimensions include:
- GPU model
- GPU memory capacity
- NVLink or network topology
- Production certification status
- Geographic or data-residency boundary
- Maintenance lifecycle
- Cost class
- Interactive versus batch usage
- Dedicated inference versus training capacity
A practical example might include:
| Node pool | Typical hardware | Primary use |
|---|---|---|
| h100-production | H100 systems with validated high-speed fabric | Production inference and critical distributed training |
| a100-shared | A100 systems | General training, research, and burst capacity |
| l40s-interactive | L40S systems | Workspaces, visualization, development, and smaller inference |
| validation | Representative GPU nodes | Upgrade testing, NCCL validation, and platform qualification |
Avoid a Node Pool for Every Team
Node pools should normally describe infrastructure characteristics, not mirror the tenant hierarchy.
Creating one node pool per team can lead to:
- Fragmented capacity
- Stranded GPUs
- Complex label management
- Difficult maintenance
- Poor flexibility when teams change
- Excessive quota administration
Use departments and projects to represent ownership. Use node pools to represent hardware or placement characteristics.
Apply Stable Node Labels
Choose labels that describe characteristics unlikely to change during ordinary operations.
For example:
runai.node-pool=h100-production gpu.platform=nvidia-h100 service.class=production fabric.class=high-bandwidth
The exact label convention should match your platform standards. Do not rely on ad hoc labels added by individual application teams.
Review Default Access to New Pools
A newly created node pool may appear in project and department configuration with zero guaranteed GPU quota while still permitting over-quota use, depending on the effective settings.
After creating a pool:
- Review every department.
- Review every project.
- Confirm whether over-quota access is enabled.
- Confirm node-pool ordering.
- Set a maximum allocation where appropriate.
- Verify that zero quota means the intended best-effort behavior, not unrestricted accidental access.
This review is particularly important for high-value GPU pools.
Create Departments and Assign Parent Quotas
Create departments after the node-pool design is stable.
For each department:
- Assign an owner.
- Select the permitted node pools.
- Order the node pools by preference.
- Define guaranteed GPU quota for each pool.
- Decide whether over-quota use is allowed.
- Configure over-quota weighting when used.
- Define maximum GPU allocation where the release and configuration expose that control.
- Apply department-level policies.
- Record the department’s business and operational purpose.
Make Parent Quotas Internally Consistent
A department quota constrains the projects beneath it.
If project guarantees under a department total 12 GPUs on a node pool but the department has only 8 GPUs of deserved quota, all project guarantees cannot be honored simultaneously.
Use this validation:
Project quotas can still be shaped differently when not all projects require simultaneous guarantees, but that should be a deliberate capacity decision rather than an accidental oversubscription.
Reserve Capacity for Shared Platform Functions
Do not allocate the complete physical pool to user departments without considering platform services.
Capacity may be required for:
- Monitoring and diagnostics
- Validation workloads
- Shared model services
- Platform testing
- Incident response
- Upgrade qualification
- Emergency production recovery
A small platform-services department or project makes this capacity visible and prevents it from being treated as permanently free.
Create Projects and Assign Guaranteed Quotas
Each project should receive quota according to the service it is expected to provide.
Production Project Pattern
A production project commonly uses:
- Guaranteed quota sufficient for normal steady-state demand
- A hard maximum aligned with approved scale
- Non-preemptible workloads
- High or very-high workload priority
- Restricted production operator access
- Production-certified node pools
- Limited or disabled over-quota use
- Strict workload policies
- Monitoring and support ownership
Production should not depend on opportunistic capacity for its minimum service objective.
Development Project Pattern
A development project commonly uses:
- A small guaranteed quota
- Preemptible workloads
- Controlled over-quota access
- Lower workload priority
- Broader developer access
- Shared or lower-cost node pools
- Request-size limits
- Idle workload controls
- Mandatory checkpointing for long training runs
This gives developers useful access without allowing interactive experiments to displace production services.
Research Project Pattern
Research often needs a hybrid model:
- A moderate guarantee for baseline progress
- High access to unused shared capacity
- Preemptible large training runs
- Checkpointing and automatic resume
- Limits on use of scarce production GPU classes
- An exception process for time-bound, large-scale experiments
The guarantee protects ongoing research. Over-quota access lets the team use idle GPUs without turning temporary availability into a permanent allocation.
Configure Users, Groups, Roles, and Access Rules
Run:ai access control should be integrated with the enterprise identity model.
An access rule combines:
- A subject, such as a user, identity-provider group, or service account
- A role
- A scope, such as the organization, cluster, department, or project
Use groups instead of assigning large numbers of users individually.
Recommended Access Model
| Subject | Scope | Suggested responsibility |
|---|---|---|
| Central AI platform administrators | Organization or cluster | Platform configuration, node pools, scheduling policy, and support |
| Department owner group | Department | Department visibility, project governance, and quota requests |
| Project administrator group | Project | Project workload administration and local assets |
| Developer group | Development project | Submit and manage development workloads |
| Production operator group | Production project | Operate approved production workloads |
| CI/CD service account | Specific project | Submit controlled automated workloads |
| Audit or FinOps group | Organization or department | Read-only usage and allocation review |
Apply Least Privilege by Scope
A developer who only needs access to product-ai-dev should not receive a broad role at the Product AI department scope.
A project administrator should not automatically receive permission to:
- Modify department quotas
- Create unrestricted projects
- Change organization-wide policy
- Manage unrelated node pools
- Create privileged access rules
Run:ai RBAC and Kubernetes RBAC are related operationally, but they should still be reviewed as part of one access model. Direct Kubernetes access must not become a bypass around platform governance.
Use Service Accounts for Automation
Use dedicated service accounts for:
- CI/CD workload submission
- Scheduled model training
- Capacity reporting
- Configuration reconciliation
- Monitoring integration
Do not use a human administrator’s credentials in pipelines.
Each service account should have:
- A defined owner
- A narrow project or department scope
- Credential rotation
- An expiration or review date
- Audit logging
- A documented purpose
Configure Controlled Over-Quota Use
Guaranteed quota and maximum allocation solve different problems.
Guaranteed quota answers:
What capacity is this tenant entitled to receive?
Maximum allocation answers:
What is the largest amount of capacity this tenant may consume?
Over-quota configuration answers:
May this tenant temporarily borrow unused capacity between those boundaries?
A well-designed project can therefore have:
Guaranteed quota: 2 GPUs Maximum allocation: 8 GPUs Over-quota: Enabled Preemptibility: Required for borrowed capacity
The team can always make progress on two GPUs. It can consume up to eight GPUs when capacity is unused. The extra six GPUs remain reclaimable.
Use Over-Quota Weights Deliberately
When over-quota weights are enabled, they influence how unused capacity is divided between eligible departments or projects.
Use weights to represent business preference, not political importance.
For example:
| Tenant | Guaranteed quota | Over-quota weight | Interpretation |
|---|---|---|---|
| Production release validation | 2 | High | Prefer this workload when temporary capacity is available |
| General model research | 4 | Medium | Receive a normal share of excess capacity |
| Personal experiments | 0 | Low | Best-effort access only |
Do not give every project a high weight. When every tenant is exceptional, the weights no longer express meaningful policy.
Treat Zero Quota Carefully
A project with zero guaranteed quota and over-quota enabled can be useful for:
- Sandboxes
- Short experiments
- Training workshops
- Low-priority validation
- Temporary user access
It has no guaranteed entitlement. Its workloads may remain pending or be preempted when quota-backed demand appears.
That limitation should be communicated to users before the project is offered as a service.
Configure Fairness, Priority, and Preemption
Quota, fairness, priority, and preemption are separate controls.
Fairness Controls Cross-Tenant Sharing
Run:ai calculates fairshare for departments and projects per node pool.
Conceptually:
Fairshare = Guaranteed deserved quota + Eligible share of unused over-quota resources
This prevents the first tenant to consume idle GPUs from treating those resources as permanently owned.
Priority Controls Scheduling Order
Workload priority controls which workload should be considered ahead of another workload in the relevant project queue.
Current Run:ai CLI priority values include:
- very-low
- low
- medium-low
- medium
- medium-high
- high
- very-high
Changing priority does not automatically change preemptibility. Configure both fields explicitly.
Priority should reflect workload urgency and service intent:
| Workload | Suggested priority |
|---|---|
| Personal experiment | very-low or low |
| Development workspace | medium-low |
| Routine training | medium |
| Release validation | medium-high or high |
| Production batch operation | high |
| Latency-sensitive production inference | very-high |
Do not mark ordinary development jobs as very-high merely to shorten queue time. That destroys the priority model.
Preemptibility Controls Reclamation
A preemptible workload may use opportunistic capacity and may be interrupted when higher-entitlement demand requires those resources.
A non-preemptible workload:
- Must remain within the project’s available deserved quota
- Cannot rely on borrowed over-quota capacity
- Will not be interrupted through normal scheduler preemption after it starts
Use non-preemptible workloads for services that cannot safely tolerate interruption. Use preemptible workloads for restartable training, experiments, and batch work.
Checkpoint Before Allowing Preemption
Preemption without checkpointing can turn high GPU utilization into low useful throughput.
A training workload should save:
- Model state
- Optimizer state
- Scheduler state
- Current epoch or step
- Random seeds when reproducibility matters
- Data-loader progress where practical
Checkpoint intervals should balance storage overhead against the amount of work that could be lost.
Enforce Scheduling Intent with Workload Policies
Do not depend on every user selecting the correct values manually.
Workload policies can apply defaults and restrictions across system, cluster, department, or project scopes. Policies can govern workloads submitted through the user interface, CLI, API, or supported Kubernetes YAML path.
Useful policy controls include:
- Default development workloads to preemptible
- Default production inference to non-preemptible
- Restrict allowed priority values
- Limit maximum GPU requests
- Restrict node pools
- Require approved container security settings
- Prevent privileged execution
- Require labels and annotations
- Enforce storage or data-source requirements
- Apply node affinity or toleration rules
- Restrict image registries
Example Policy Intent
For product-ai-dev:
Default priority: medium-low Default preemptibility: preemptible Maximum GPUs per workload: 4 Allowed node pools: a100-shared, l40s-interactive Privileged containers: prohibited Required label: environment=development
For product-ai-prod:
Default priority: very-high Default preemptibility: non-preemptible Maximum GPUs per workload: approved production scale Allowed node pools: h100-production Privileged containers: prohibited Required label: environment=production Required owner annotation: mandatory
The exact policy schema depends on the workload type and current Run:ai release. Validate the available fields before translating this intent into an enforceable policy.
Validate Over-Quota Allocation and Preemption
Do not consider the design complete until quota, over-quota, fairness, and preemption have been tested.
The following test uses two projects with a quota of two GPUs each on a four-GPU node pool.
Test Scenario
- Submit three one-GPU preemptible workloads to
team-a. - Submit one one-GPU workload to
team-b. - Confirm
team-ais using one GPU over quota. - Submit a second workload to
team-b. - Confirm one preemptible
team-aworkload is reclaimed. - Confirm both projects receive their two-GPU deserved allocation.
The current Run:ai CLI supports explicit project, node-pool, priority, GPU request, and preemptibility settings.
runai login runai training standard submit team-a-job-1 \ -p team-a \ -i runai.jfrog.io/demo/quickstart-demo \ --gpu-devices-request 1 \ --node-pools a100-shared \ --priority low \ --preemptibility preemptible runai training standard submit team-a-job-2 \ -p team-a \ -i runai.jfrog.io/demo/quickstart-demo \ --gpu-devices-request 1 \ --node-pools a100-shared \ --priority low \ --preemptibility preemptible runai training standard submit team-a-job-3 \ -p team-a \ -i runai.jfrog.io/demo/quickstart-demo \ --gpu-devices-request 1 \ --node-pools a100-shared \ --priority low \ --preemptibility preemptible runai training standard submit team-b-job-1 \ -p team-b \ -i runai.jfrog.io/demo/quickstart-demo \ --gpu-devices-request 1 \ --node-pools a100-shared \ --priority medium \ --preemptibility preemptible
At this stage, the expected allocation is:
team-a quota: 2 GPUs team-a allocation: 3 GPUs team-a over quota: 1 GPU team-b quota: 2 GPUs team-b allocation: 1 GPU cluster allocation: 4 of 4 GPUs
Submit the second Team B workload:
runai training standard submit team-b-job-2 \ -p team-b \ -i runai.jfrog.io/demo/quickstart-demo \ --gpu-devices-request 1 \ --node-pools a100-shared \ --priority medium \ --preemptibility preemptible
The expected steady state is:
team-a allocation: 2 GPUs team-b allocation: 2 GPUs team-a-job-3: Preempted, terminating, or returned to pending team-b-job-2: Scheduled and running
The specific workload selected for preemption can depend on scheduler state and workload characteristics. Validate the entitlement outcome rather than assuming a particular pod name will always be reclaimed.
Test Development and Production Separation
A second validation should prove that development cannot displace production incorrectly.
Production Workload
runai training standard submit production-service \ -p product-ai-prod \ -i runai.jfrog.io/demo/quickstart-demo \ --gpu-devices-request 2 \ --node-pools h100-production \ --priority very-high \ --preemptibility non-preemptible
Development Workload
runai training standard submit development-training \ -p product-ai-dev \ -i runai.jfrog.io/demo/quickstart-demo \ --gpu-devices-request 2 \ --node-pools a100-shared \ --priority medium-low \ --preemptibility preemptible
Confirm that:
- The production workload only uses approved production capacity.
- The development workload cannot select the production pool.
- Production allocation does not depend on over-quota capacity.
- The development workload may consume unused shared capacity.
- Development can be reclaimed without interrupting production.
- Access rules prevent ordinary developers from operating the production workload.
Measure Quota Utilization
Quota management must be based on measured behavior rather than annual estimates.
Run:ai project and department views expose allocation relative to quota. Monitoring and telemetry can also provide GPU allocation, utilization, pending-time, and workload information.
Track at least the following metrics per project, department, and node pool.
| Metric | Why it matters |
|---|---|
| Allocated GPUs divided by guaranteed quota | Indicates how much of the entitlement is being consumed |
| Over-quota GPU hours | Shows dependence on borrowed capacity |
| Pending workload time | Reveals insufficient quota, placement constraints, or fragmentation |
| Preemption count | Shows how frequently opportunistic work is displaced |
| Lost work after preemption | Tests whether checkpointing is effective |
| Idle allocated GPUs | Identifies workloads holding GPUs without useful compute |
| Node-pool utilization | Identifies stranded or overloaded hardware classes |
| Maximum concurrent GPU allocation | Supports future quota sizing |
| Production scheduling delay | Validates whether guaranteed capacity is sufficient |
| Exception frequency | Reveals whether the standard quota model fits real demand |
Starting Review Thresholds
The following are operating-policy examples, not NVIDIA defaults:
- Guaranteed quota below 40 percent utilization for 30 days: review for reduction.
- Guaranteed quota above 85 percent utilization with sustained pending work: review for increase.
- Over-quota consumption above 30 percent of usage for several weeks: determine whether demand has become permanent.
- Repeated preemption with substantial lost training time: require improved checkpointing or revise workload placement.
- Idle allocated GPUs above 10 percent for more than 30 minutes: investigate application behavior.
- Production queue delay above the service objective: treat as a capacity or placement incident.
Use longer windows for research workloads with irregular experiments. Use tighter thresholds for production services.
Prevent One Team from Monopolizing the Cluster
No single control prevents monopolization. Use multiple boundaries.
Assign Meaningful Guarantees
Do not assign the complete shared node pool as one project’s guaranteed quota unless that project truly owns the hardware.
Use Maximum Allocations
A project may have a small guarantee and permission to borrow, but it should not necessarily be able to consume the entire pool.
Require Preemptibility for Opportunistic Work
Borrowed capacity must remain reclaimable. Otherwise, over-quota access becomes an uncontrolled reservation.
Limit GPUs per Workload
A single accidental request for all available GPUs can create unnecessary queue pressure even when project-level quotas are working correctly.
Control Expensive Node Pools
For scarce pools:
- Set zero quota for projects without entitlement.
- Disable over-quota access where necessary.
- Restrict the pool through policy.
- Require an approved project or department.
- Set explicit maximum allocations.
- Review new pool defaults immediately after creation.
Detect Idle Allocation
A workload can monopolize GPUs without performing useful work. Monitor GPU utilization, GPU memory behavior, workload phase, and idle allocation together.
Expire Temporary Exceptions
A temporary quota increase without an expiration date is usually a permanent quota increase with incomplete documentation.
Troubleshooting Common Scheduling Problems
| Symptom | Likely cause | Corrective action |
|---|---|---|
| Workload remains pending despite apparently free GPUs | Free GPUs are in a different node pool or do not satisfy affinity, topology, taint, memory, or workload-size requirements | Check selected node pools, ordering, node labels, affinity, tolerations, GPU type, and contiguous capacity |
| Non-preemptible workload cannot start above quota | Non-preemptible workloads cannot rely on over-quota capacity | Increase guaranteed quota, reduce the request, or use an approved preemptible workload |
| Project cannot use a newly created node pool | Project has zero quota, over-quota is disabled, maximum is zero, or policy blocks the pool | Review department and project configuration for the new pool |
| Development workload runs on production GPUs | Node-pool ordering, over-quota defaults, or policy permits access | Set project quota and maximum to zero for the production pool and restrict it through policy |
| Production workload is preempted | It was submitted as preemptible or inherited an incorrect default | Correct the workload policy and resubmit as non-preemptible within deserved quota |
| Team reports a quota utilization ratio over 100 percent | The project is consuming over-quota capacity | Review allocation, guaranteed quota, maximum allocation, and over-quota weight |
| One project receives most unused GPUs | Other projects are not eligible, have lower over-quota weight, are blocked by placement, or have no pending workloads | Review eligibility, weights, queue state, and node-pool constraints |
| Department guarantees appear ineffective | Project guarantees exceed the parent department quota | Reconcile department and project quota totals per node pool |
| Frequent preemption wastes training progress | Checkpoints are too infrequent, incomplete, or not restored | Validate checkpoint persistence and resume behavior before allowing large preemptible jobs |
| GPUs are allocated but utilization remains low | Workload is blocked on data, CPU, storage, network, initialization, or application logic | Correlate GPU metrics with CPU, storage, network, and workload logs |
| High-priority development jobs dominate queue order | Priority policy is too permissive | Restrict allowed priority values at the development project or department scope |
| Quota changes cause unexpected scheduling behavior | Several quota, maximum, weight, and policy settings changed together | Roll back to the previous baseline and change one control at a time |
Operate Quota Changes as Controlled Platform Changes
Changing quotas can trigger workload movement and preemption. Treat quota changes as production changes.
Use the following sequence:
- Capture current project, department, and node-pool allocation.
- Record pending workloads and current preemptible workloads.
- Confirm the parent department has capacity for the proposed project guarantee.
- Check whether maximum allocations must also change.
- Check over-quota eligibility and weight.
- Identify workloads that may be preempted.
- Verify checkpoint status for affected training jobs.
- Apply the smallest required change.
- Observe allocation, pending time, and preemption.
- Confirm that production service objectives remain satisfied.
- Record the result and expiration date.
- Roll back if the expected entitlement outcome is not observed.
Do not simultaneously change quota, node-pool order, workload policy, and access rules unless the change has been validated in a non-production environment.
Define Operational Ownership
A shared GPU platform requires named owners for both technology and allocation decisions.
| Capability | Accountable owner | Responsible operator |
|---|---|---|
| Cluster and scheduler health | AI platform owner | Platform engineering |
| Node-pool labels and membership | Infrastructure owner | Kubernetes or GPU platform team |
| Department quota | AI governance or capacity owner | Central platform administrator |
| Project quota | Department owner | Project administrator with approval |
| Workload priority policy | Platform governance owner | Run:ai administrator |
| Identity groups and access lifecycle | Security or IAM owner | Identity operations |
| Checkpointing and restart behavior | Application owner | ML engineering team |
| Production service capacity | Product service owner | Production operations |
| Utilization reporting | Capacity or FinOps owner | Platform analytics |
| Exception approval | Named governance authority | Platform service manager |
Quota ownership should not sit entirely with Kubernetes administrators. The administrator can implement a quota, but the business or service owner must justify the entitlement.
Establish an Exception Process
Exceptions are inevitable. Undocumented exceptions are optional.
A quota exception should record:
Requesting department: Requesting project: Business owner: Technical owner: Requested node pool: Current guaranteed quota: Temporary guaranteed quota: Maximum requested allocation: Over-quota requirement: Preemptibility: Workload priority: Business justification: Expected start: Expiration: Checkpoint validated: Production impact assessment: Approver: Rollback action:
Use separate exception paths for:
- Temporary guaranteed quota increases
- Access to restricted node pools
- Increased maximum allocation
- Non-preemptible research workloads
- Very-high priority workloads
- Large distributed training runs
- Temporary production recovery
- Planned benchmark or validation activity
Every exception should expire automatically or enter a mandatory review state.
Recommended Production Baseline
A defensible starting baseline is:
Production
- Separate production projects
- Guaranteed quota sized for normal demand
- Non-preemptible workloads
- High or very-high priority
- Restricted node pools
- Explicit maximum allocation
- No dependency on over-quota capacity
- Restricted access groups
- Strict workload policies
- Documented service owner
Development
- Separate development projects
- Small guaranteed quota
- Preemptible by default
- Medium-low priority
- Over-quota enabled on shared pools
- Maximum allocation enforced
- Idle workload monitoring
- No access to production-only pools
- Broader developer access
- Mandatory checkpointing for long jobs
Research
- Baseline guaranteed quota
- Preemptible over-quota use
- Medium priority
- Higher over-quota weight only when justified
- Maximum allocation below full cluster capacity
- Tested checkpoint and resume
- Time-bound exceptions for large experiments
- Usage and outcome review
Shared Platform Services
- Dedicated project or department
- Small guaranteed reservation
- Non-preemptible critical services
- Clear platform ownership
- Restricted administrative access
- Capacity protected during tenant demand peaks
Conclusion
Multi-tenant GPU scheduling is not solved by dividing the number of GPUs by the number of teams.
A workable NVIDIA Run:ai design maps durable organizational boundaries into departments, isolates practical workload boundaries into projects, and assigns quota separately for each relevant node pool. Guaranteed quota provides predictability, while controlled over-quota use improves utilization when capacity would otherwise remain idle.
Fairness and preemption make borrowed capacity recoverable. Priority helps order work, but it should never replace quota design. Production workloads should receive quota-backed, non-preemptible capacity. Development and research workloads can use preemptible excess capacity when checkpointing and restart behavior are proven.
The operating model is as important as the scheduler configuration. Departments need accountable owners, projects need access boundaries, quota changes need evidence, and exceptions need expiration dates. When those controls are in place, a shared GPU cluster can support multiple teams without becoming either permanently underutilized or operationally unpredictable.
External References
- NVIDIA Run:ai Documentation: Adapting AI Initiatives to Your Organization
Canonical URL: https://run-ai-docs.nvidia.com/saas/platform-management/aiinitiatives/adapting-ai-initiatives - NVIDIA Run:ai Documentation: Departments
Canonical URL: https://run-ai-docs.nvidia.com/saas/platform-management/aiinitiatives/organization/departments - NVIDIA Run:ai Documentation: Projects
Canonical URL: https://run-ai-docs.nvidia.com/saas/platform-management/aiinitiatives/organization/projects - NVIDIA Run:ai Documentation: Node Pools
Canonical URL: https://run-ai-docs.nvidia.com/saas/platform-management/aiinitiatives/resources/node-pools - NVIDIA Run:ai Documentation: The NVIDIA Run:ai Scheduler: Concepts and Principles
Canonical URL: https://run-ai-docs.nvidia.com/saas/platform-management/runai-scheduler/scheduling/concepts-and-principles - NVIDIA Run:ai Documentation: How the Scheduler Works
Canonical URL: https://run-ai-docs.nvidia.com/saas/platform-management/runai-scheduler/scheduling/how-the-scheduler-works - NVIDIA Run:ai Documentation: Workload Priority and Preemption
Canonical URL: https://run-ai-docs.nvidia.com/saas/platform-management/runai-scheduler/scheduling/workload-priority-control - NVIDIA Run:ai Documentation: Over Quota, Fairness and Preemption
Canonical URL: https://run-ai-docs.nvidia.com/saas/platform-management/runai-scheduler/scheduling/quick-starts/over-quota - NVIDIA Run:ai Documentation: Access Rules
Canonical URL: https://run-ai-docs.nvidia.com/saas/infrastructure-setup/authentication/accessrules - NVIDIA Run:ai Documentation: Roles
Canonical URL: https://run-ai-docs.nvidia.com/saas/infrastructure-setup/authentication/roles - NVIDIA Run:ai Documentation: Policies and Rules
Canonical URL: https://run-ai-docs.nvidia.com/saas/platform-management/policies/policies-and-rules - NVIDIA Run:ai Documentation: Metrics and Telemetry
Canonical URL: https://run-ai-docs.nvidia.com/saas/platform-management/monitor-performance/metrics - NVIDIA Run:ai Documentation: Best Practices: Checkpointing Preemptible Training Workloads
Canonical URL: https://run-ai-docs.nvidia.com/saas/workloads-in-nvidia-run-ai/using-training/checkpointing-preemptible-workloads
TL;DR MCP and A2A solve different integration problems. MCP standardizes how an agent discovers and invokes tools, APIs, resources, and data services....
