Into AI

Into AI

Home
🧁 Articles
⚡️ LLM Inference
⚙️ GPU essentials
📉 LLM training
🤝 Sponsor
📚 Tech guides
Leaderboard
About

LLM Inference

10 LLM Inference Metrics Every AI Engineer Must Know
TTFT, TPOT, ITL, Goodput, MBU, MFU, and more.
Sep 21 • Dr. Ashish Bamania
5 LLM inference batching techniques every AI engineer should know
Static, Dynamic, and Continuous batching, Chunked prefill, and Prefill–decode disaggregation, simply explained.
Aug 22 • Dr. Ashish Bamania
How does Claude watermark text?
Visually understand how Anthropic watermarks and detects Claude-generated text.
Aug 20 • Dr. Ashish Bamania
10 LLM Inference Optimization Techniques, Simply Explained
10 techniques that make LLM inference faster and cheaper: KV caching, Quantization, FlashAttention, PagedAttention, Speculative decoding, and more.
Aug 1 • Dr. Ashish Bamania
Arithmetic Intensity, Simply Explained
A no-jargon breakdown of Arithmetic Intensity and the Roofline model, and how they are used to optimize LLM inference.
Jul 6 • Dr. Ashish Bamania
A hardware-level tour of how LLMs generate text
Understand what actually happens on the CPU and GPU when an LLM turns your prompt into text.
Jun 22 • Dr. Ashish Bamania
Speculative Decoding, Simply Explained
Learn how Speculative Decoding works from scratch and how to use it in your AI applications for faster and cheaper inference.
Apr 30 • Dr. Ashish Bamania
Top 4 Decoding Strategies In LLMs Explained Simply
Learn all about the magical statistics that make text generation with LLMs possible.
Oct 17, 2025 • Dr. Ashish Bamania
© 2026 Dr. Ashish Bamania · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture