DKDivyajot Kaurinai-beginners-journey.hashnode.dev·6d ago · 5 min readPart 23: Types of Gradient Descent: Batch, Mini-Batch & StochasticIn the previous blog, we learned about Gradient Descent—an optimization algorithm that helps a machine learning model reduce its error by gradually adjusting its weights. But there is an important que00
WRWei Ruiinintuitiveml.hashnode.dev·Sep 24 · 12 min readA Software Engineer's Intuition for Gradient Descent, Part 2: The Step SizeIn Part 1 we built gradient descent around one update rule: $$\theta := \theta - \alpha \nabla f(\theta)$$ Almost all of that article was about the gradient ∇f(θ), the part that tells us which way is 00
WRWei Ruiinintuitiveml.hashnode.dev·Sep 24 · 9 min readA Software Engineer's Intuition for Gradient DescentFor those of us who have spent years building deterministic software systems, the machine learning realm can feel like a parallel universe. In traditional software engineering, we write the exact rule00
DKDivyajot Kaurinai-beginners-journey.hashnode.dev·Sep 21 · 5 min readPart 22: Gradient Descent: How Neural Networks Find Better WeightsImagine you are standing on top of a mountain, but it's completely dark. 🌙 Your goal is to reach the lowest point in the valley. You can't see the entire path, so what would you do? You might take a 00
TTharunintharunpandya.hashnode.dev·Aug 22 · 10 min readPolynomial Regression: Teaching a Line to BendTake our best straight-line model so far, the one that fit our five-order table perfectly. Now widen the range a bit. Some orders close by, a couple kilometers away. Some far, ten or twelve kilometers10
TTharunintharunpandya.hashnode.dev·Aug 19 · 10 min readFeature Scaling in Machine Learning: Why It MattersTake the exact same model, the exact same data, the exact same code. Change one thing: write the distance column in meters instead of kilometers. Nothing else moves. Run it, and the model doesn't just10
TTharunintharunpandya.hashnode.dev·Aug 16 · 9 min readMultiple Linear Regression: When One Input Isn't EnoughPicture two different orders. Both are exactly 3 kilometers away from the restaurant. One arrives in 18 minutes. The other takes 35 minutes. If distance is the only thing our model looks at, it has no10
EJEva J Patelinevapatel123.hashnode.dev·Aug 15 · 33 min readHow Gradient Descent Works: The Math Behind Machine LearningWhen you train a machine learning model, you might write something as simple as: model.fit(X, y) and suddenly your model can make predictions. But there's a big question hiding behind that one line: 10
TTharunintharunpandya.hashnode.dev·Aug 13 · 10 min readGradient Descent: Walking Downhill BlindfoldedIn the last post, we built a way to measure exactly how wrong a guess is, using the cost function. Given any w and b, we can now compute a single number that tells us how badly that line fits our data10
CPChai Planetinthe-tech-inside.hashnode.dev·Jul 24 · 9 min readInside the Transformer Training LoopNeo: I know kung fu. Morpheus: Show me. (Neo steps forward, his body moving in a blur of perfect, calculated strikes. He doesn't have to think about it. The knowledge is crystallized in his muscles. H00