Into AI

Into AI

How to choose the right GPU for LLM inference

A practical guide to figuring out the GPU requirements for LLM inference using Qwen and GLM models on NVIDIA GPUs.

Dr. Ashish Bamania's avatar
Dr. Ashish Bamania
Oct 10, 2026
∙ Paid

This lesson is written to make serving open-source LLMs in production easy for you.
In this lesson, we will discuss how to find the right GPU for deploying these two LLMs:

  1. Qwen3.8-27B: A Dense model with 27 billion parameters

  2. GLM-4.5-Air: A Mixture-of-Experts (MoE) model with 106 billion total parameters and 12 billion active parameters

We will consider the following four popular GPUs across three NVIDIA GPU generations:

  1. L40S (Ada Lovelace generation)

  2. H100 SXM (Hopper generation)

  3. H200 (Hopper generation)

  4. B200 (Blackwell generation)


Step 1: Figuring out how much space LLM parameters take

The formula to calculate the storage space that a model’s parameters will take in memory is:

\(\text{Storage Size}=\text{Number of Parameters}\times \text{Bytes per Parameter}\)


LLMs are commonly served in the following precisions:

  • BF16/FP16: 2 bytes per parameter

  • FP8/INT8: 1 byte per parameter

  • MXFP8 (Microscaling FP8): 1.03 bytes per parameter

  • INT4/FP4: 0.5 bytes per parameter

  • MXFP4 (Microscaling FP4): 0.53 bytes per parameter

The storage space required by the models is:

  • Qwen3.8-27B in FP8 takes 27B x 1 byte/ parameter = 27 GB of space

  • GLM-4.5-Air in FP8 takes 106B x 1 byte/ parameter = 106 GB of space

(Please note that the size of the actual model checkpoints available for download may vary depending on the quantization method used and any additional metadata.)


Step 2: Figuring out the GPU’s memory capacity

GPUs with their memory capacities are as follows:

  • L40S: 48GB of GDDR6

  • H100 SXM: 80GB

  • H200: 141GB

  • B200: 180GB

Qwen3.8-27B comfortably fits on all four GPUs, but GLM-4.5-Air fits only on H200 and B200.

On H200, GLM will have roughly 35 GB remaining, and on B200, 74 GB remaining, after storing the parameters.

But this is not the end of the story, and one needs to keep space for the following to ensure smooth inference:

  • KV cache

  • Activations

  • CUDA graphs

  • Runtime buffers

  • Other overheads

User's avatar

Continue reading this post for free, courtesy of Dr. Ashish Bamania.

Or purchase a paid subscription.
© 2026 Dr. Ashish Bamania · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture