AI infrastructure

AI infrastructure

Plan infrastructure for model hosting, inference, training and fine-tuning with explicit memory and workload requirements.

Choose a configuration in Console to review its resources and estimated cost.

What to plan for

An AI workload is more than a GPU: model weights, cache, datasets, checkpoints and an application endpoint all consume different resources.

Choose the exact model revision, precision, context length and concurrency before sizing. Validate throughput with your own workload.

Configuration choices

24–32 GB class

L4, A10, RTX 4090 or RTX 5090. Consider for smaller models only after measuring total memory, including runtime and cache.

48–80 GB class

L40S, A40, A100 or H100. More memory may accommodate larger working sets; it does not guarantee concurrency or token speed.

141 GB and multi-GPU

H200, B200 or B300 and explicitly sized multi-GPU systems. Confirm exact model, precision, topology and serving framework.

Weight-only memory ≈ parameters × bytes per parameter. Runtime, KV cache, activations and optimizer state are additional. This is a sizing method, not a deployment recommendation.

Smart Infra Finder

Your workload, your requirements

Matches come from the catalog. Unclear sizing needs clarification; performance and availability are never invented.

Explore the catalog

LLM hosting

LLM hosting: model selection, memory requirements, storage and deployment choices without invented performance claims.

Explore

AI inference

AI inference: model selection, memory requirements, storage and deployment choices without invented performance claims.

Explore

AI training

AI training: model selection, memory requirements, storage and deployment choices without invented performance claims.

Explore

Fine-tuning

Fine-tuning: model selection, memory requirements, storage and deployment choices without invented performance claims.

Explore

Infrastructure for DeepSeek

Infrastructure for DeepSeek: model selection, memory requirements, storage and deployment choices without invented performance claims.

Explore

Infrastructure for Llama

Infrastructure for Llama: model selection, memory requirements, storage and deployment choices without invented performance claims.

Explore

Infrastructure for Mistral

Infrastructure for Mistral: model selection, memory requirements, storage and deployment choices without invented performance claims.

Explore

Embedding infrastructure

Embedding infrastructure: model selection, memory requirements, storage and deployment choices without invented performance claims.

Explore

Image generation infrastructure

Image generation infrastructure: model selection, memory requirements, storage and deployment choices without invented performance claims.

Explore

Speech AI infrastructure

Speech AI infrastructure: model selection, memory requirements, storage and deployment choices without invented performance claims.

Explore

Check availability: AI infrastructure

Check the current price and configuration in Console. Your choice is saved through registration. The current lookup uses mock adapters; no real capacity is reserved.

Related products and guides