AAAhmed Adawyinadawy.hashnode.dev·3d ago · 5 min readDemystifying LLM Serving Infrastructure: How PagedAttention and Continuous Batching Scale InferenceMoving a Large Language Model (LLM) from a local prototype in a Jupyter notebook to a high-throughput, multi-tenant production environment is a brutal awakening. While data scientists spend months opt00
AAAhmed Adawyinadawy.hashnode.dev·5d ago · 3 min readLLM Inference Engineering: Overcoming the KV-Cache Bottleneck and Maximizing Production ThroughputLLM Inference Engineering: Overcoming the KV-Cache Bottleneck and Maximizing Production Throughput When moving a Large Language Model (LLM) from a prototype to a production environment, you immediatel00
MPMikhail Panfilovinmnwa.hashnode.dev·Sep 29 · 22 min readCan safe Rust ever beat Google's C Brotli?Brotli, SIMD, bounds checks, and what it takes to compete with Google's C implementation. At one point, mbrotli's greedy match finder had the right AVX2 token and no native vector compares. The genera00
GAGonzalo Aguilaringonzalosimon.hashnode.dev·Sep 28 · 12 min readWhen Personalisation Costs 15 Seconds: Redesigning an LLM Feature for ReliabilityHow I changed personalised vocabulary practice so learners could start a session without waiting for generated sentences. Parlato already had a usable exercise sentence, yet loading a learning session22EB
Ggayatrikakumanu25ingayatrikakumanu.hashnode.dev·Sep 24 · 17 min readFigma-to-Code at Scale: What Actually Drives Cost, Quota, and QualityOr: why the "smart" way to fetch Figma designs turned out to be the expensive way. TL;DR: The "fetch a lightweight outline first" advice is backwards — it cost almost as much as fetching everything 00
GYGulshan Yadavinmrgulshanyadav.hashnode.dev·Sep 22 · 12 min readReducing Model Size Without Losing Accuracy: QuantizationA practical deep dive into model quantization — the precision-ladder trick that takes a 14GB model down to 3.5GB, what you lose, what you keep, and the code to measure both. Six months ago I was tryin00
CHCédric Hervetincedric-hervet.hashnode.dev·Sep 15 · 9 min readWhy the hard part of route optimization is the modeling, not the algorithmEveryone benchmarks the algorithm. Almost nobody talks about the layer that actually breaks projects: modeling. Route optimization projects rarely fail on the algorithm. Between a business rule and a 20
RZRameen Zafarinoracle-impact.hashnode.dev·Sep 14 · 4 min readSQL Query Performance Tuning: 5 Golden Rules for Slow Joins & Un-indexed TablesIs your database running slow? Are users staring at endless loading spinners in your web application? In 90% of cases, the bottleneck isn't the application code or server hardware,it’s un-optimized SQ00
Mmarketingintoolscase.hashnode.dev·Sep 8 · 7 min readHow to Reduce PDF File Size Without Ruining Document QualityLarge PDF files can become surprisingly inconvenient. A report filled with screenshots may be difficult to email. A portfolio may exceed an upload limit. Scanned documents can consume tens or even hun00
Mmarketingintoolscase.hashnode.dev·Sep 8 · 5 min readHow to Optimize Images for the Web Without Sacrificing QualityImages are an essential part of modern websites. They make landing pages more engaging, explain products visually, support ecommerce experiences, and help content communicate more effectively. But ima00