I Tested ROCm vs Vulkan for AMD Local LLMs — One Backend Silently Tanked by 6x
The same model, the same AMD GPU, two different compute stacks. One hit 143 tokens per second. The other — on a slightly bigger model — quietly collapsed to a sixth of its neighbour's speed, and the s
rosluk.hashnode.dev8 min read