24–32 GB class
L4, A10, RTX 4090 or RTX 5090. Consider for smaller models only after measuring total memory, including runtime and cache.
Machine learning infrastructure
Machine learning infrastructure: practical compute, data and recovery choices for a measured deployment plan.
Choose a configuration in Console to review its resources and estimated cost.
Separate data preparation, training and serving; these stages need different compute and storage profiles.
Use CPU for preparation, GPU when the framework supports acceleration, and persistent storage for datasets and checkpoints. Measure the training batch before choosing VRAM.
L4, A10, RTX 4090 or RTX 5090. Consider for smaller models only after measuring total memory, including runtime and cache.
L40S, A40, A100 or H100. More memory may accommodate larger working sets; it does not guarantee concurrency or token speed.
H200, B200 or B300 and explicitly sized multi-GPU systems. Confirm exact model, precision, topology and serving framework.
Weight-only memory ≈ parameters × bytes per parameter. Runtime, KV cache, activations and optimizer state are additional. This is a sizing method, not a deployment recommendation.
Check the current price and configuration in Console. Your choice is saved through registration. The current lookup uses mock adapters; no real capacity is reserved.
Compare GPU memory and architectures, from RTX to data-center accelerators, then check configurations in one console.
Explore →Cloud servers: configuration choices, sizing requirements and availability checks in GlobalAtlas AI.
Explore →Compare object, block and file storage, backups and snapshots by access pattern, durability needs and restore workflow.
Explore →Cloud backup: access model, capacity planning and recovery considerations before choosing a configuration.
Explore →