Subscribe
Sign in
Home
🧁 Articles
⚡️ LLM Inference
⚙️ GPU essentials
📉 LLM training
🤝 Sponsor
📚 Tech guides
Leaderboard
About
LLM Inference
10 LLM Inference Metrics Every AI Engineer Must Know
TTFT, TPOT, ITL, Goodput, MBU, MFU, and more.
Sep 21
•
Dr. Ashish Bamania
9
1
5 LLM inference batching techniques every AI engineer should know
Static, Dynamic, and Continuous batching, Chunked prefill, and Prefill–decode disaggregation, simply explained.
Aug 22
•
Dr. Ashish Bamania
14
How does Claude watermark text?
Visually understand how Anthropic watermarks and detects Claude-generated text.
Aug 20
•
Dr. Ashish Bamania
13
1
10 LLM Inference Optimization Techniques, Simply Explained
10 techniques that make LLM inference faster and cheaper: KV caching, Quantization, FlashAttention, PagedAttention, Speculative decoding, and more.
Aug 1
•
Dr. Ashish Bamania
22
7
Arithmetic Intensity, Simply Explained
A no-jargon breakdown of Arithmetic Intensity and the Roofline model, and how they are used to optimize LLM inference.
Jul 6
•
Dr. Ashish Bamania
7
2
A hardware-level tour of how LLMs generate text
Understand what actually happens on the CPU and GPU when an LLM turns your prompt into text.
Jun 22
•
Dr. Ashish Bamania
205
4
25
Speculative Decoding, Simply Explained
Learn how Speculative Decoding works from scratch and how to use it in your AI applications for faster and cheaper inference.
Apr 30
•
Dr. Ashish Bamania
8
3
Top 4 Decoding Strategies In LLMs Explained Simply
Learn all about the magical statistics that make text generation with LLMs possible.
Oct 17, 2025
•
Dr. Ashish Bamania
7
5
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts