<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Anup Sharma</title><description>Systems Engineer — Distributed Systems &amp; AI Infrastructure</description><link>https://twilighttechie.github.io/</link><item><title>I Nodded Along While Inference Geeks Talked GPUs — So I Wrote the Field Guide I Wish I Had</title><link>https://twilighttechie.github.io/posts/gpu-hardware-field-guide/</link><guid isPermaLink="true">https://twilighttechie.github.io/posts/gpu-hardware-field-guide/</guid><description>A comprehensive, from-zero guide to GPU hardware buzzwords: HBM, memory bandwidth, FP8/FP4, NVLink, H100 vs H200 vs B200 vs GB200, NVL72 racks, AMD Instinct, TPUs, Groq, Cerebras — plus what to actually run in a homelab and why datacenter GPU infra is genuinely hard.</description><pubDate>Tue, 21 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Production Inference Optimization Strategy — Full Deck (All Deliverables)</title><link>https://twilighttechie.github.io/posts/production-inference-full-deck/</link><guid isPermaLink="true">https://twilighttechie.github.io/posts/production-inference-full-deck/</guid><description>The complete slide deck covering the production inference optimization strategy: architecture, benchmarks, routing, and cost analysis.</description><pubDate>Sat, 18 Jul 2026 00:10:00 GMT</pubDate></item><item><title>Architectural Specification — Production Inference Optimization Strategy</title><link>https://twilighttechie.github.io/posts/inference-architecture-spec/</link><guid isPermaLink="true">https://twilighttechie.github.io/posts/inference-architecture-spec/</guid><description>Architectural specification for a production inference optimization strategy for the agentic inference cloud.</description><pubDate>Sat, 18 Jul 2026 00:05:00 GMT</pubDate></item><item><title>Inference Optimization Mini-Benchmark: FP16 vs Q8_0 vs Q4_K_M</title><link>https://twilighttechie.github.io/posts/inference-mini-benchmark/</link><guid isPermaLink="true">https://twilighttechie.github.io/posts/inference-mini-benchmark/</guid><description>Hands-on benchmark of quantization levels on a 1.5B model: decode throughput, TPOT, and memory footprint compared across FP16, Q8_0, and Q4_K_M.</description><pubDate>Sat, 18 Jul 2026 00:02:00 GMT</pubDate></item><item><title>Serving the Agentic Inference Era — Optimization Strategy (Slides)</title><link>https://twilighttechie.github.io/posts/agentic-inference-presentation/</link><guid isPermaLink="true">https://twilighttechie.github.io/posts/agentic-inference-presentation/</guid><description>Slide deck on optimizing LLM inference for the agentic era: quantization, KV cache, batching, and bandwidth-bound decoding.</description><pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate></item><item><title>I Ran vLLM on a Mac Mini With No GPU — Here&apos;s Everything I Learned About Inference</title><link>https://twilighttechie.github.io/posts/vllm-cpu-inference-guide/</link><guid isPermaLink="true">https://twilighttechie.github.io/posts/vllm-cpu-inference-guide/</guid><description>A complete, beginner-friendly guide to vLLM: what inference actually is, how to build vLLM from source on an Apple Silicon Mac with no GPU, every command explained, the three errors I hit and fixed, the flags that matter, and real throughput numbers from my living room.</description><pubDate>Sat, 18 Jul 2026 00:00:00 GMT</pubDate></item><item><title>OpenCode Doesn&apos;t Have an Auto Mode. So I Built One with vLLM Semantic Router</title><link>https://twilighttechie.github.io/posts/opencode-semantic-router/</link><guid isPermaLink="true">https://twilighttechie.github.io/posts/opencode-semantic-router/</guid><description>Cursor picks the right model for you automatically. OpenCode doesn&apos;t - yet. Here&apos;s a complete, production-shaped guide to adding intelligent auto model selection to OpenCode (or any OpenAI-compatible agent) using vLLM Semantic Router and AgentGateway.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate></item><item><title>I Almost Built a Grafana Stack—Then AgentGateway Shipped Everything I Needed.</title><link>https://twilighttechie.github.io/posts/agentgateway-observability/</link><guid isPermaLink="true">https://twilighttechie.github.io/posts/agentgateway-observability/</guid><description>I Almost Built a Grafana Stack—Then AgentGateway Shipped Everything I Needed.</description><pubDate>Thu, 02 Jul 2026 00:00:00 GMT</pubDate></item><item><title>How Cassandra Compression Actually Works (Chunks, Offsets, and Reads)</title><link>https://twilighttechie.github.io/posts/cassandra-compression/</link><guid isPermaLink="true">https://twilighttechie.github.io/posts/cassandra-compression/</guid><description>How Cassandra Compression Actually Works (Chunks, Offsets, and Reads)</description><pubDate>Wed, 01 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Giving AgentGateway a Semantic Brain with vLLM Semantic Router - Inside My Homelab</title><link>https://twilighttechie.github.io/posts/agentgateway-semantic/</link><guid isPermaLink="true">https://twilighttechie.github.io/posts/agentgateway-semantic/</guid><description>Giving AgentGateway a Semantic Brain with vLLM Semantic Router - Inside My Homelab</description><pubDate>Mon, 22 Jun 2026 00:00:00 GMT</pubDate></item><item><title>I Traced Personal Agent&apos;s Source Code. Inside Was Pi... And It Dreams at 3 AM.</title><link>https://twilighttechie.github.io/posts/traced-agent/</link><guid isPermaLink="true">https://twilighttechie.github.io/posts/traced-agent/</guid><description>I Traced Personal Agent&apos;s Source Code. Inside Was Pi... And It Dreams at 3 AM.</description><pubDate>Mon, 15 Jun 2026 00:00:00 GMT</pubDate></item><item><title>I Built an AI That Decides Which AI to Talk To — Running 24/7 From My Living Room</title><link>https://twilighttechie.github.io/posts/deciding-ai/</link><guid isPermaLink="true">https://twilighttechie.github.io/posts/deciding-ai/</guid><description>I Built an AI That Decides Which AI to Talk To — Running 24/7 From My Living Room</description><pubDate>Wed, 10 Jun 2026 00:00:00 GMT</pubDate></item></channel></rss>