Saudi Arabia designated 2026 the Year of Artificial Intelligence. That national initiative provides context for investment, but each infrastructure project still needs its own business case and workload measurements. Saudi Press Agency announcement.
Start with a workload, not a rack
Define the model, dataset, users, response-time target and expected concurrency. Training, fine-tuning and inference can have very different resource requirements. A small inference service and a distributed training cluster should not inherit the same hardware specification.
Record the precision and context length you intend to use, model and data licensing, and the capacity needed for growth. Test a representative workload before treating a demonstration as a production sizing result. Include document retrieval and application services in the test where they form part of the user journey.
Check the entire electrical path
NVIDIA lists a maximum system power of 10.2 kW for DGX H100/H200 systems. This is a model-specific maximum, not a typical consumption figure for every GPU server or the measured draw of your application. NVIDIA DGX specifications.
Have qualified engineers check the selected equipment against the supply, distribution, rack PDUs, UPS capacity and redundancy requirements. Include switches and storage in the load schedule. A server fitting physically into a rack does not establish that the room can power it safely.
Document maximum design load and measured workload draw separately. Both inform operating costs, but they answer different engineering questions.
Design cooling for the selected equipment
There is no single rack-power threshold that makes every air-cooled installation unsuitable. Equipment requirements, air distribution, inlet conditions, room layout and the cooling plant all matter.
Ask the equipment supplier and facilities designer to agree the cooling method, environmental limits, maintenance access and response to a cooling failure. Where liquid cooling is proposed, include its facility connections, monitoring and maintenance responsibilities in the design.
Benchmark network and storage
Network speed alone does not establish suitability for AI. Distributed jobs, dataset loading, checkpoint writes and user requests stress different paths. Test throughput and latency with the intended storage layout and workload.
Confirm switch ports, optics, redundancy and any accelerator-interconnect requirements against the supported design. Avoid purchasing a fixed 25 or 100 Gb/s configuration solely because another inference deployment used it.
Use a staged acceptance plan
- Run a representative pilot and record performance, power and temperatures.
- Check capacity when a planned redundant component is unavailable.
- Validate model access, administrative permissions and dataset handling.
- Assign owners for hardware, facilities, platform updates and monitoring.
- Agree measurable acceptance criteria and a rollback approach before installation.
Pair the equipment assessment with a private AI pilot plan so model quality, data permissions and operating responsibilities are tested with the infrastructure.
BustanTech's IT infrastructure, hardware and Red Hat AI services can form part of a scoped deployment. Discuss the workload and site requirements before specifying the equipment.