AKArthur Kazainarthurkaza.com·Sep 18 · 4 min readFeeding the Beast: why your TPU is bored, and how to fix itPicture a restaurant that hired the best chef in the city. She can plate ten thousand dishes an hour. In the back, one person peels potatoes with a butter knife. The owner looks at the dining room, se00
SPSamu Physicistinsamu-physicist.hashnode.dev·Sep 16 · 10 min readImplementing Kolmogorov-Arnold Networks as Dense GEMMs in JAX: A Neural QMC ExperimentUpdate (Sep 2026): After publishing, I found problems in the design described here, and I have revised this post. The ε-manifold mapping (originally Sections 1–2) was unnecessary. Real solid harmonics00
SHSanskriti Harmukhinvultr.hashnode.dev·Sep 16 · 3 min readInstalling JAX with ROCm Acceleration on Ubuntu 24.04JAX is an open-source library for high-performance numerical computing and machine learning research, offering tools for automatic differentiation, GPU/TPU acceleration, and just-in-time compilation. 00
WKWesley Kambaleinkambale.dev·Aug 7 · 16 min readNever lose your progress: Checkpointing with OrbaxPicture this. You have spent six hours training a model. The loss curve looks beautiful, accuracy is climbing, and you are one epoch away from a result worth writing home about. Then the power goes ou10
WKWesley Kambaleinkambale.dev·Jul 29 · 13 min readFeeding the Beast: Data Pipelines with GrainOver the last seven weeks we have built models, written training loops, composed optimizers, bulletproofed our code with Chex, and learned to debug inside JIT-compiled functions. Every one of those ar22N
WKWesley Kambaleinkambale.dev·Jul 18 · 13 min readWhen JIT hides your errors: Debugging JAXLast week, Chex gave us a firewall against shape mismatches and NaN gradients. That firewall catches the bugs that announce themselves; a tensor with the wrong shape, a gradient that explodes into inf20
WKWesley Kambaleinkambale.dev·Jul 12 · 14 min readCatching bugs before they catch youHere's a scenario every ML engineer eventually lives through: you kick off a training run before leaving the office. Three hours later, you check back in. The loss curve looks like a flat line at NaN.00
WKWesley Kambaleinkambale.dev·Jul 1 · 11 min readOptax: Optimizers You Can Compose Like LEGOIn our last article on this blog in March, we built a complete training loop. We used optax.adamw() to create an optimizer, wrapped it in nnx.Optimizer, and watched our model learn. But we barely scra20
AAshiteshinashitesh.me·Jun 11 · 4 min read Scaling Quantum State-Vector Simulation to 36 Qubits on Google Cloud TPU The Problem Simulating quantum circuits on a classical computer is exponentially hard. An n-qubit state vector holds 2ⁿ complex amplitudes — at 30 qubits that's 8 GB, at 33 qubits 64 GB, at 36 qubits 11C
AAshiteshinashitesh.me·Jun 2 · 5 min readSimulating Quantum Computers at Scale: JAX, GPUs, and Cloud TPUsNote: The detailed research paper and full experimental results for this project have already been published. This post provides a high-level summary of the core system design, key benchmarks, and eng10