large-language-models

Close up of a green circuit board with microchips representing GPU compute hardware

What is Speculative Decoding in vLLM

Discover how speculative decoding accelerates large language model inference on AMD GPUs, its underlying mechanisms, benefits, and how to enable it in vLLM…

September 7, 2026 9 min read
Data center server racks powering large language model inference

Self-Hosting Alibaba Open Source Qwen 3.8-Max

Explore the hardware, architecture, and licensing considerations for self-hosting Alibaba’s open-source Qwen 3.8-Max AI model, the first of its size to be…

August 7, 2026 13 min read