Qwen3.6-27B-FP8 on One RTX 6000 Ada: Fast TTFT, 668 tok/s Peak Throughput [Benchmark]
Overview
We benchmarked Qwen/Qwen3.6-27B-FP8 using vLLM 0.19 on a single RTX 6000 Ada 48GB GPU.
The goal was to evaluate realistic chat-serving behavior with:
Streaming enabled
OpenAI-compatible /v1
blog.hexgrid.cloud13 min read