top of page

GPU Infrastructure Planning: Architecture, Operations and Governance

Writer: Mike Lupescu
Mike Lupescu
7 hours ago
16 min read

Use workload evidence to size capacity, select a deployment model, design the full data path and define how the environment will be governed and operated.


In this guide



Start with the decisions the infrastructure must support


GPU infrastructure planning turns an AI workload portfolio into an investable, operable service. It should tell the organisation how much capacity it needs, which deployment model fits each workload, what the surrounding network and storage must deliver, how users receive resources and who owns the service after launch. A hardware list cannot answer those questions on its own.


The planning record needs to work for several readers. Executive sponsors need the outcome, cost boundary and decision gates. AI and application teams need capacity that fits their models and delivery timelines. Infrastructure and security teams need a design they can integrate and control. Procurement needs comparable options. Operations needs telemetry, support procedures and an accountable service boundary.


The strongest plan distinguishes evidence from forecast. Current workload traces, queue records, dataset sizes, run histories and service metrics are evidence. Growth, new use cases and future model behaviour are assumptions. Both belong in the model, but the owner, source, confidence and review date should be visible.


For implementation support after the planning decisions are defined, explore SkyLab’s GPU infrastructure solutions.


Use a repeatable planning framework


Seven GPU infrastructure planning stages: outcome, workload, capacity, placement, architecture, operations and decision.

Stage

Question and required output

Outcome

What business or service result requires GPU capacity? Record a named outcome, owner and success measure.

Workload

Which jobs run, how do they behave and when do they matter? Produce workload classes and a demand baseline.

Capacity

How much usable capacity is required over time? Model base, peak, reserve and growth.

Placement

Where may and should each workload run? Set dedicated, consumed or hybrid placement rules.

Architecture

What complete system keeps accelerators productive? Design compute, fabric, storage, facilities and software.

Operations

How is the service allocated, observed and recovered? Define policy, service-level objectives, KPIs, runbooks and responsibilities.

Decision

What evidence permits the next commitment? Produce a phased roadmap and acceptance gates.


Teams can apply this framework to an existing cluster, a new dedicated deployment or a hybrid estate.


The level of detail should match the decision. A preliminary capacity option does not require a final port map, but it does require enough evidence to expose the risks that could invalidate cost, location or timing.


Profile workloads before choosing GPUs


Start by grouping workloads according to behaviour rather than business-unit ownership. Two teams may use the same model family but require different infrastructure because one runs scheduled training and the other serves interactive inference. The profile should capture model and dataset size, precision, memory demand, number of accelerators, scale efficiency, run duration, concurrency, arrival pattern, checkpoint behaviour, data path, latency or completion objective and software dependencies.


Training and fine tuning


Training and fine-tuning workloads are often throughput-oriented and may span several accelerators or nodes. The planner needs measured scaling behaviour, communication intensity, checkpoint frequency, restart cost and the point at which adding devices stops producing economical gains. A job that scales poorly can consume more capacity without shortening the delivery timeline proportionally.


Long-running work also changes resilience design. The cost of one failed node is not only the repair; it includes lost compute, repeated data movement and delayed experimentation. Checkpoint interval, checkpoint write performance, restart procedure and spare capacity should be part of the workload definition.


Batch and online inference


Batch inference can be scheduled for throughput and may tolerate queues or preemption. Online inference is governed by latency, throughput, availability and response-time objectives. Model loading, memory footprint, batching, cache behaviour, request variability and autoscaling boundaries affect capacity. Average request volume alone can hide the peak and tail behaviour that determines user experience.


Reasoning and agentic services can add long contexts, repeated model calls and variable output length.


Their infrastructure should be planned from measured token and request behaviour, not from a fixed ratio borrowed from a different model or application.


Research analytics and high performance computing


Scientific computing, simulation and analytics workloads may depend on high precision, tightly coupled communication, large host memory, specialised libraries or significant preprocessing. They can also share a cluster with AI workloads only if the scheduler and software environment express their resource requirements correctly. The plan should identify whether one architecture can serve both groups without creating persistent contention or support complexity.


Data preparation and adjacent CPU work


Data ingestion, cleaning, tokenisation, augmentation, feature generation and evaluation can consume CPU, memory, network and storage before or after the GPU stage. If these steps are excluded from the workload map, accelerators may wait for a pipeline that was never sized. The capacity model should therefore include the complete workflow and identify which stages can execute independently.


Build a capacity model from GPU hours and service constraints


For each workload class, a useful starting point is required GPU-hours. Calculate the number of planned runs multiplied by GPUs per run and average run duration, then sum the classes for the planning period.


Use measured distributions where possible. One average can understate long-running jobs and peak concurrency.


Required GPU-hours = sum of runs multiplied by GPUs per run multiplied by run hours. This establishes demand, not installed capacity. Installed-equivalent capacity depends on the planning window, the share of time that can be scheduled productively, maintenance, resilience, queue objectives, fragmentation and the degree to which demand can shift to another environment.


A planning conversion can divide required GPU-hours by available hours per device and an agreed schedulable-utilisation assumption. Add resilience and growth according to the service design rather than hiding them inside the utilisation factor. Every assumption should be sensitivity-tested. The result is a range that the buying committee can challenge, not a false point estimate.


Input

Measurement and planning risk

Runs and arrival pattern

Use job history, pipeline schedules and the product forecast. Averages alone hide peaks and deadlines.

GPUs and duration per run

Use representative benchmarks and production traces. Otherwise nominal demand is detached from real workloads.

Scale efficiency

Test single and multiple devices on the intended stack. Do not assume that more GPUs reduce runtime linearly.

Queue and service objective

Define acceptable wait, completion, latency and priority by class. Capacity must support a user outcome.

Failure and maintenance

Use incident history, repair assumptions, checkpoints and change windows. Otherwise usable capacity is overstated.

Growth and uncertainty

Give each scenario an owner and review date. Avoid locking a purchase to one unsupported forecast.


Worked example: a capacity range, not a purchase order


Consider an illustrative training workload of 40 runs per week, each using 4 GPUs for 6 hours. Demand is 40 × 4 × 6 = 960 GPU-hours. This example is hypothetical; it is not a SkyLab customer result or a recommended utilisation target.


With a 120-hour scheduling window per device and an assumed 60% productive scheduling factor, the baseline is 960 ÷ (120 × 0.60) = 13.33 GPU equivalents. Rounding to 14 devices describes an aggregate arithmetic minimum before the design constraints are applied. Four-GPU jobs, node boundaries and overlapping deadlines may require a different purchasable configuration.


At a 50% factor, the same demand becomes 16 equivalents; at 70%, it becomes approximately 11.43. These scenarios show why a single unsupported factor can materially change the investment. They do not establish a safe operating range. Validate the factor from representative run and queue evidence, then model the arrival pattern and scheduling policy.


Record failure reserve, maintenance and forecast growth separately so they are not counted twice. If two jobs must start together, the design must make eight compatible GPUs available at that moment even if weekly totals appear comfortable. Test the proposed topology, memory fit and service objectives before treating any rounded total as an equipment requirement.


The decision record should retain the formula, units, workload version, measurement window, assumption owner and next review date. Re-run the calculation when a model, precision, runtime or deadline changes; carry unresolved inputs forward as explicit evidence gaps.


Benchmark representative workflows


A benchmark should answer a purchasing or design question. Begin with the workload version, dataset or synthetic data method, framework and library versions, precision, batch or request profile, number of devices, topology, storage path and success criteria. Record warm-up, repeated runs, variability and any tuning that would not be available in production.


For training, measure time to a defined result, scale efficiency, checkpoint behaviour and recovery as well as samples or tokens processed. For online inference, report throughput with latency percentiles, input and output characteristics, concurrency and quality settings. For data-intensive pipelines, observe ingestion, preprocessing, model loading and checkpoint I/O so a faster accelerator does not conceal a slower end-to-end workflow.


Use the benchmark to test sensitivity. Change the number of devices, data path, batch or request mix, sharing mode and failure condition when those choices could change the design. The benchmark report should separate measured results from vendor figures and projected production demand. Procurement can then see which result supports the recommended system and which assumptions remain open.


Separate utilisation allocation and productive output


Utilisation is often used as a single health score, but it describes different things at different layers.


Allocation shows whether a scheduler has assigned a device. Active compute indicates whether processing units are busy. Memory activity, interconnect traffic and power show other parts of workload behaviour. None of these alone proves that a useful job completed on time.


A service dashboard should combine resource and workload measures. High allocation with low active compute may point to over-requesting, a data bottleneck, synchronisation or an idle interactive session.


High active compute with growing queue time may indicate efficient hardware and insufficient capacity.


Low average utilisation may be acceptable for a low-latency service that must retain headroom for peaks.


Set targets by workload class and business objective. Training queues, shared research, batch inference and online services need different policies. The aim is useful throughput, predictable service and transparent cost, not the highest possible percentage on every device.


Compare dedicated bare metal cloud and service models


The placement decision should use the same workload baseline, risk assumptions and time horizon. A dedicated system can provide control and predictable access, but it creates facility, lifecycle and operating commitments. Bare-metal consumption can provide dedicated nodes with less asset ownership, while public cloud can give fast access and a broad service ecosystem. GPU as a service can package capacity, tenancy and operating controls for users who do not need to manage the entire platform.


Model

Demand, control, cost and exit considerations

Dedicated or owned

Fits a steady strategic baseline. Offers greater direct control within the owned boundary. Account for procurement and facility lead time, capital, depreciation, support and retained operations. Exit exposure includes asset and technology lifecycle.

Consumed capacity

Can fit variable, temporary or generation-specific demand. Control depends on the service and contract; access depends on provider availability and onboarding. Include usage, reservations, data movement and service charges. Evaluate data, workflow and supplier portability.

Hybrid

Combines a stable baseline with peaks or regional needs. Shared policy spans different operating boundaries. Requires a standing design and a tested burst path. Include both commitment and coordination costs; keep placement rules and interfaces portable.


A practical model often assigns workloads rather than declaring one environment universally superior.


Sensitive or steady workloads may have a preferred dedicated location. Experiments and temporary peaks may use consumed capacity. A placement rule should state the permitted data path, performance requirement, cost owner, fallback and evidence needed before a workload moves.


SkyLab's Resources and Capacity service can support the sourcing decision once workload, architecture and commercial requirements are known. Current availability, locations, hardware, terms and lead times must be confirmed for the engagement.


The GPU as a service case study shows a service-delivery context for these decisions. Use it alongside the wider case studies when defining the questions and evidence for your own engagement.


Design the GPU cluster as a balanced system


Compute nodes


Select accelerators with the complete node. GPU memory capacity and bandwidth, supported precision, inter-device connectivity and workload performance matter, but so do CPU architecture, host memory, local storage, PCIe topology, network adapters, serviceability and software support. The planned node should be benchmarked with representative data and the intended framework, drivers, libraries and container runtime.


Compute and storage fabrics


Distributed workloads create east-west traffic between nodes. Storage and data services create another path. User, management and external service traffic create others. The architecture should identify each traffic class, its performance and security requirements, expected oversubscription, congestion behaviour, monitoring and failure domain. Reference architectures are useful starting points, but the validated design must reflect the customer's scale and workloads.


Storage and data lifecycle


Storage design should model datasets, checkpoints, model artefacts, logs, caches and retained results.


Test sequential and metadata-intensive patterns, concurrency, small files, data loading, checkpoint writes and restore. Capacity tiers, data locality, backup, archive and deletion should reflect the data lifecycle and recovery requirement.


Facilities


Confirm rack power, distribution, redundancy, cooling method, water or facility interfaces where applicable, floor loading, rack dimensions, cabling, fire and safety requirements, service clearances, delivery route and expansion. Use current OEM and facility information for the selected system. High-density designs may require liquid cooling or a different location; this decision cannot be deferred until equipment arrives.


Management and software


The software layer includes firmware, drivers, container runtime, images, schedulers, cluster management, identity, secrets, data access, monitoring, service management and user interfaces. Version and support compatibility should be recorded as a controlled matrix. Automation can reduce configuration drift, but it does not remove the need for approval, testing, rollback and an accountable owner.


Select GPU generations with dated evidence


Accelerator choice changes quickly. Keep a dated evidence register for each shortlisted configuration: the OEM specification, supported software matrix, quoted delivery window, facility requirements and workload benchmark. Distinguish a roadmap announcement from an orderable system and confirmed delivery. Recheck material assumptions at design approval and procurement; a vendor announcement alone does not establish availability for a particular market or engagement.


Compare generations using the workload's precision, memory footprint, communication pattern, scale, latency or throughput objective, software support, power and cooling, delivery date, support term and lifecycle cost. Run a representative benchmark where the decision is material. Published peak figures use specific precisions and conditions and should not be substituted for application performance.


A mixed-generation estate may be rational when workloads have different requirements or capacity is added over time. The plan must then define scheduler labels, supported software stacks, placement rules, user expectations and operational spares. Heterogeneity without policy can increase queue fragmentation and support cost.


Choose scheduling and sharing methods deliberately


Kubernetes exposes GPUs to workloads through device plugins, while Slurm can define and schedule GPUs and other generic resources through GRES. These are allocation mechanisms, not complete service policies. A production design still needs queues or classes, quotas, priorities, reservations, fair-share rules, topology awareness, maintenance handling, accounting and user support.


Whole-device allocation offers a clear boundary and suits workloads that need the full accelerator.


Supported partitioning such as Multi-Instance GPU can divide selected GPUs into isolated instances with dedicated compute and memory resources. Time-slicing can increase concurrency for suitable workloads but does not provide the same memory or fault isolation. The chosen method should be tested for framework support, predictability, security and operational recovery.


Preemption can protect urgent work and improve the use of interruptible capacity, but only for workloads designed to tolerate it. The planner should define checkpoint frequency, termination notice, retry behaviour, maximum lost work and the classes that may preempt or be preempted. Interactive sessions and production services need separate controls from fault-tolerant batch jobs.


Build security and tenancy into the architecture


Begin with authoritative identity and explicit access to users, services, devices and administrative functions. Zero trust principles avoid granting implicit trust because a workload sits on an internal network or owned cluster. Access should be limited to the resource and action required, then logged for review.


The design should cover tenant and project boundaries, namespace or partition policy, network segmentation, secrets, image and package provenance, vulnerability management, driver and firmware updates, data encryption, storage permissions, model and artefact access, administrative sessions and break-glass procedures. Multi-tenant services also need quota, billing or cost attribution and a clear process for onboarding and offboarding.


For regulated or sovereign workloads, record the responsible legal, policy, data and security owners. They define the applicable requirements and approve residual risk. Infrastructure controls should produce the evidence those owners need without claiming that a technology or provider creates compliance by itself.


Plan resilience around workload recovery


Resilience should begin with the consequence of interruption. A failed training node may require job restart from a checkpoint. An inference service may need traffic shifted to healthy replicas within an agreed time. A research queue may accept a longer recovery if capacity and job state remain protected.


Each class needs a recovery objective and a tested mechanism.


Design decisions include power and network redundancy, spare or reserved nodes, scheduler behaviour, node health checks, automated quarantine, checkpoint storage, configuration backup, control-plane recovery and supplier escalation. Not every service needs a second full cluster, but every critical dependency needs an understood failure mode, owner and recovery procedure.


Acceptance should test degraded scenarios, not only peak performance. Examples include loss of a node or fabric path, failed job restart, full storage tier, expired credential, image rollback, telemetry loss and a change during a long-running workload. The result should show whether the service behaves as designed and whether the operating team can diagnose and restore it.


Use metrics that connect infrastructure to service outcomes


Metric

Question and required context

Allocatable and allocated GPUs

How much capacity is available and reserved? Include maintenance, partitions, reservations and sharing mode.

Active compute and memory

Are assigned devices doing expected work? Consider workload phase, precision, batching and the data path.

Queue wait and start rate

Do users obtain capacity within the service objective? Segment by priority, requested size and arrival pattern.

Job completion and failure

Does work finish successfully and repeatably? Distinguish application errors, infrastructure faults and retry policy.

Inference latency and throughput

Are online service objectives met? Record request mix, input/output length, batching and percentile.

Fabric and storage performance

Does data movement limit productive compute? Include traffic class, concurrency and topology.

Cost per successful workload

What does useful output cost across the full service? Include capacity commitments, software, facilities, data and operations.


NVIDIA DCGM and its exporter can expose accelerator health and performance telemetry for supported NVIDIA environments. Kubernetes, Slurm and platform tools provide allocation and workload context. The operating model should combine these sources with application, queue, service and cost data rather than presenting a hardware dashboard as the complete service view.


Define calculation rules and retention for every KPI. For example, queue time needs a start event, end event, excluded states and a percentile or distribution. Cost per successful workload needs the cost boundary and a definition of success. Clear definitions allow finance, operations and users to discuss the same result.


Review capacity after go live


The first capacity model should become an operating process. Review actual demand, queue distributions, workload size, completion, failures, data-path constraints and cost against the approved scenarios. Separate policy problems from physical shortages. A queue can grow because capacity is insufficient, because requests are oversized or because priority and reservations do not reflect the current portfolio.


Define thresholds that trigger investigation, not automatic procurement. The review should identify the affected workload class, duration of the condition, available policy or optimisation actions, lead time for new capacity and the date a decision is required. This creates a controlled route from telemetry to investment and keeps growth assumptions current.


Model total lifecycle cost


Compare options over a period that reflects the commitment. Include hardware or service charges, network, storage, data movement, facilities, power, cooling, software, platform, implementation, support, operations, spare capacity, training, refresh and exit. Separate fixed commitments from usage-driven cost and state which cost remains with the customer.


Capacity efficiency should be measured against useful output. A lower hourly rate can be more expensive when jobs run longer, fail more often or wait for data. A premium option may not be justified when the workload cannot use its capability. Benchmarking and sensitivity analysis show which assumptions drive the decision.


The model should also value time and constraint. Delaying a product release, research programme or customer service has a business effect even when it is not booked as infrastructure expense. Conversely, overcommitting to capacity can lock funds and operations into demand that never materialises. Decision makers should see both exposures.


Sequence implementation through evidence gates


Gate one validates demand


Confirm priority workloads, data, service objectives and a demand range. Exit when the buying committee agrees which workloads justify the next design commitment and which assumptions still require measurement.


Gate two selects placement and sourcing


Compare dedicated, consumed and hybrid options against the shared baseline. Confirm current availability, facility feasibility, integration, security, operating responsibility and lifecycle cost. Exit with an approved direction and documented alternatives.


Gate three approves detailed design


Complete the compute, network, storage, facility, platform, security and management design. Record versions, dependencies, test plan, acceptance criteria and responsibility. Resolve any substitution that could change performance or support.


Gate four proves the service


Install and integrate the environment, then test representative workloads, policy, telemetry, failure and recovery. Exit only when evidence meets the agreed acceptance criteria and remaining risks have an owner.


Gate five hands over operations


Confirm monitoring, service desk, escalation, maintenance, capacity governance, security operations, documentation and training. Schedule post-launch reviews against the workload and service baseline.


Define the managed services boundary alongside the technical handover, including retained customer responsibilities and escalation ownership.


Procurement and partner evaluation checklist


  • State the workload classes, demand range, required evidence and service objectives used to compare options.

  • Confirm the exact system, accelerator, memory, interconnect, network adapter, storage and software configuration.

  • Record current product availability, location, delivery assumptions, substitutions and dependencies.

  • Define facility, connectivity, data, identity, security and integration responsibilities before award.

  • Separate equipment, capacity, platform, implementation, support, operations and retained customer cost.

  • Specify acceptance tests for representative workloads, policy enforcement, telemetry, resilience and handover.

  • Identify support hours, severity definitions, escalation, replacement process, maintenance and service exclusions.

  • Define data, configuration and workload portability, contract exit, asset retirement and evidence retention.


Questions buyers should resolve


How much GPU capacity does the organisation need?


Calculate demand from workload runs, GPUs per run and duration, then account separately for concurrency, queue objectives, maintenance, resilience, fragmentation, growth and uncertainty. Use representative traces and benchmarks. A universal utilisation target is not a reliable sizing method.


When should bare metal be preferred to cloud GPU capacity?


Bare metal or dedicated infrastructure is worth evaluating for sustained demand, strong control requirements, predictable access or tightly coupled workloads. Cloud or consumed capacity can suit variable demand, rapid access and short-lived needs. Compare both on the same workload, data, service and lifecycle-cost basis.


What is the difference between GPU allocation and utilisation?


Allocation means a scheduler has assigned a device or slice to a workload. Utilisation describes activity measured at one or more hardware layers. Neither alone shows whether a useful workload completed on time. Combine them with queue, completion, latency, failure and cost measures.


Can Kubernetes and Slurm manage the same GPU estate?


They can appear in the same wider environment, but the ownership and integration design must avoid conflicting allocation. Kubernetes uses device plugins or related resource mechanisms, while Slurm schedules GPUs as generic resources. The right pattern depends on workload, tenancy, operations and the selected platform.


How should a team compare GPU generations?


Use representative workload benchmarks and compare memory, precision, communication, scaling, software support, facility demand, availability, support term and lifecycle cost. Date-stamp vendor specifications and distinguish projected figures from shipping configurations and measured application results.


Which GPU infrastructure metrics matter most?


Track capacity and allocation, active compute and memory behaviour, queue time, job completion and failure, inference latency and throughput, data-path performance, hardware health and cost per successful workload. Define each metric by workload class and service objective.


Plan the next controlled commitment


A credible GPU plan ends with a decision, not a larger collection of specifications. The organisation should know which workloads are in scope, the capacity range, preferred placement, architecture principles, unresolved evidence, commercial exposure, operating owner and the test that permits the next commitment.


SkyLab can help turn that record into a production design and delivery path across capacity, GPU infrastructure solutions, data centre enablement, platform integration and managed operations. The engagement begins by defining the decision and the evidence already available.



 
 
bottom of page