Tag: inference
All the articles with the tag "inference".
-
I Thought ASR Needed Its Own vLLM — Here’s How Speech Models Are Actually Served
Inside the ASR inference engine: model execution, batching, streaming state, parallelism, multilingual serving, open-source runtimes, and the breakthroughs that made modern speech AI practical.
-
Production Inference Optimization Strategy — Full Deck (All Deliverables)
The complete slide deck covering the production inference optimization strategy: architecture, benchmarks, routing, and cost analysis.
-
Architectural Specification — Production Inference Optimization Strategy
Architectural specification for a production inference optimization strategy for the agentic inference cloud.
-
Inference Optimization Mini-Benchmark: FP16 vs Q8_0 vs Q4_K_M
Hands-on benchmark of quantization levels on a 1.5B model: decode throughput, TPOT, and memory footprint compared across FP16, Q8_0, and Q4_K_M.