This lesson is written to make serving open-source LLMs in production easy for you.
In this lesson, we will discuss how to find the right GPU for deploying these two LLMs:
Qwen3.8-27B: A Dense model with 27 billion parameters
GLM-4.5-Air: A Mixture-of-Experts (MoE) model with 106 billion total parameters and 12 billion active parameters
We will consider the following four popular GPUs across three NVIDIA GPU generations:
L40S (Ada Lovelace generation)
H100 SXM (Hopper generation)
H200 (Hopper generation)
B200 (Blackwell generation)
Step 1: Figuring out how much space LLM parameters take
The formula to calculate the storage space that a model’s parameters will take in memory is:
LLMs are commonly served in the following precisions:
BF16/FP16: 2 bytes per parameter
FP8/INT8: 1 byte per parameter
MXFP8 (Microscaling FP8): 1.03 bytes per parameter
INT4/FP4: 0.5 bytes per parameter
MXFP4 (Microscaling FP4): 0.53 bytes per parameter
The storage space required by the models is:
Qwen3.8-27B in FP8 takes
27B x 1 byte/ parameter = 27 GBof spaceGLM-4.5-Air in FP8 takes
106B x 1 byte/ parameter = 106 GBof space
(Please note that the size of the actual model checkpoints available for download may vary depending on the quantization method used and any additional metadata.)
Step 2: Figuring out the GPU’s memory capacity
GPUs with their memory capacities are as follows:
L40S: 48GB of GDDR6
H100 SXM: 80GB
H200: 141GB
B200: 180GB
Qwen3.8-27B comfortably fits on all four GPUs, but GLM-4.5-Air fits only on H200 and B200.
On H200, GLM will have roughly 35 GB remaining, and on B200, 74 GB remaining, after storing the parameters.
But this is not the end of the story, and one needs to keep space for the following to ensure smooth inference:
KV cache
Activations
CUDA graphs
Runtime buffers
Other overheads



