Blog
All the articles I've written.
-
I Nodded Along While Inference Geeks Talked GPUs — So I Wrote the Field Guide I Wish I Had
A comprehensive, from-zero guide to GPU hardware buzzwords: HBM, memory bandwidth, FP8/FP4, NVLink, H100 vs H200 vs B200 vs GB200, NVL72 racks, AMD Instinct, TPUs, Groq, Cerebras — plus what to actually run in a homelab and why datacenter GPU infra is genuinely hard.
-
Production Inference Optimization Strategy — Full Deck (All Deliverables)
The complete slide deck covering the production inference optimization strategy: architecture, benchmarks, routing, and cost analysis.
-
Architectural Specification — Production Inference Optimization Strategy
Architectural specification for a production inference optimization strategy for the agentic inference cloud.
-
Inference Optimization Mini-Benchmark: FP16 vs Q8_0 vs Q4_K_M
Hands-on benchmark of quantization levels on a 1.5B model: decode throughput, TPOT, and memory footprint compared across FP16, Q8_0, and Q4_K_M.