GGPUStackingpustack.hashnode.dev·Jul 30 · 17 min readGPUStack Day 0 Support for Kimi-K3: vLLM vs. SGLang Inference Benchmark on 8×B300 GPUsThis article documents the deployment and benchmarking of Kimi-K3 on a single server equipped with 8×NVIDIA B300 GPUs. It compares vLLM and SGLang under 64K and 200K long-context workloads. The key fi00
GGPUStackingpustack.hashnode.dev·Jul 23 · 10 min readWhere Did Your GPU Resources Go? GPUStack Usage Tracking Gives You the Answer at a GlanceYou bought a stack of GPUs and deployed a bunch of models. Then, at the end of the month, your boss asks: “Are these GPUs actually worth the cost?” Can you answer? If not, this article is for you. GP00
GGPUStackingpustack.hashnode.dev·Jul 22 · 6 min readDay 0 Deployment Experience | Deploying GLM-5.2-FP8-DSpark on GPUStackThis article is curated from a real-world deployment experience shared by a GPUStack community member. GLM-5.2-FP8-DSpark is an enhanced version of GLM-5.2-FP8 that applies Speculative Decoding by loa00
GGPUStackingpustack.hashnode.dev·Jul 1 · 6 min readDay 0 Benchmark: Deploying DeepSeek-V4-Flash-DSpark on GPUStack Doubles ThroughputThis article is based on a community benchmark contributed by a GPUStack user. DeepSeek-V4-Flash-DSpark enhances DeepSeek-V4-Flash by adding a Speculative Decoding module. Using the same model weights00
GGPUStackingpustack.hashnode.dev·Jun 30 · 8 min readGPUStack v2.2: From Model Serving to Token Operations, from Compute Pooling to GPU-as-a-ServiceDeploying a model and bringing it online is only the starting point of AI service delivery. As large language model applications move into scaled production, AI infrastructure is entering an inevitabl00