MSMIG serversinmigservers.hashnode.dev·1d ago · 5 min readAMD EPYC 9006 "Venice" CPUs: Architecting the Agentic AI StackModern AI infrastructure has rapidly transitioned from static, monolithic training and inference clusters to highly dynamic Agentic AI pipelines. Unlike traditional models that rely on predictable com00
TThrottleinthrottle.hashnode.dev·3d ago · 4 min readThe FP8 trap: my GPU bill dropped 47% because the model was printing "!!!!!!"I had one hour on an AMD MI300X and one question: what does a token actually cost on it? One GPU, vLLM's ROCm build, Qwen2.5 at 7B, 32B and 72B. Thirty-two requests in flight, 256 output tokens max, p00
VJVedant Jadhavinvedantjadhav.hashnode.dev·6d ago · 7 min readWhat 57 LLM Benchmark Runs on an AMD Instinct MI300X Taught Me About VRAM, Throughput, and QuantizationHow much VRAM does a model actually need? How does model size affect generation speed? Does quantization always reduce memory? Does lower precision always make inference faster? I wanted to answer the13A
JGJosh Greeninjoshgreen.hashnode.dev·Aug 1 · 5 min readA 30B Model Crawled on My Mini PC, and the Weights Were InnocentThere is a comfortable assumption in the local LLM world: if a model only activates a few billion parameters per token, it will run fast on modest hardware. Sparse mixture-of-experts models are sold o21R
MSManu Shuklainecorpit.hashnode.dev·Jul 28 · 14 min readROCm vs CUDA in 2026: can enterprises actually escape CUDA lock-in?ROCm vs CUDA in 2026: can enterprises actually escape CUDA lock-in? Summary. CUDA has been NVIDIA's real moat for 18 years, and 2026 is the first year that moat has visible cracks. AMD's ROCm 7, relea00
DDustinVKinundefinedbehavior.hashnode.dev·Jul 12 · 19 min readI Gave My Coding Agent a Free tierThere is a TL;DR section near the bottom if you just want the summary of setting it all up minus the story. Why host a model? The world of AI coding is moving at lightspeed in 2026. Every week I learn00
BKBrian Kinginsolodev.app·Jul 9 · 9 min readInstalling ROCm & Llama.cpp for CPUs & APUs.Abstract. I document my complete workflow for installing AMD ROCm drivers on Ubuntu Desktop 24.04 LTS and building the llama.cpp library with HIP (GPU) backend support for my AMD Strix Halo APU with i00
BKBrian Kinginsolodev.app·Jul 9 · 12 min readInstalling vLLM onto EVO-X2.Abstract. I download a vLLM Docker image that supports AMD, and the ROCm GPU driver that runs on the Strix Halo platform (AI Max+ 395 APU). I connect the running Docker container to local GGUF AI mode00
RMRaghul Minblog.raghul.in·Jul 8 · 5 min readThe Hidden Softwares Behind Every AI GPU : Why Hardware Alone Isn't EnoughMost people think NVIDIA, AMD, or Intel won the AI race by building faster GPUs. That's only half the story. The real winner isn't just the company with the most powerful hardware it's the company wit42J
Kkuroappworksinkuroappworks.hashnode.dev·May 13 · 10 min read[Solved] Windows 11 Failed to Wake from Sleep and Restarts: Fixing a 4-Year-Old Bug with a BIOS UpdateEnvironment Component Model / Version CPU AMD Ryzen 5 5600G Motherboard ASUS TUF GAMING X570-PLUS Graphics Card ASUS DUAL-RX9060XT-16G OS Windows 11 [Conclusion] Summary of the Cause an00