HKHarshavardhan Katkaminharshakatkam.dev·21h ago · 24 min readDay 00a: Six ideas that explain every AI agent, before any codeSomewhere in the first hour of reading about AI agents, you meet a diagram. A box labelled Thought, an arrow to Action, an arrow to Observation, and an arrow curling back to the start. The captions sa10
TTivardintivard.hashnode.dev·9h ago · 10 min readPrompt Engineering Isn't Enough: Enforcing Deterministic Boundaries on Probabilistic WorkflowsPart I: The Silent Failure Mode There is an expensive failure mode in production LLM pipelines that traditional monitoring often misses. It isn’t a crash, a timeout, or an obvious hallucination. It is11M
HKHarshavardhan Katkaminharshakatkam.dev·21h ago · 7 min readDay 00: Learn AI agents by building one, in 30 daysSomeone hands you an empty file and says: build an agent. Perhaps you have read the explainers and can recite that an agent is "an LLM in a loop with tools". Perhaps you have read none of them, and th01I
NSNagappan Sinnagspidey.hashnode.dev·22h ago · 9 min readEvals: How Do You Actually Know Your AI Feature Works?You ship a new AI feature. It works great in your testing — you type in a few prompts, the answers look right, you demo it to your team, everyone nods. Three weeks later, a user reports that the assis04AAI
PCPeesh Choprainpeeshchopra.hashnode.dev·9h ago · 7 min readPricing AI Features When Every Request Costs MoneyA startup adds an AI assistant to its $40 per seat plan. Adoption is strong, the demo closes deals, and support tickets drop. Three months later the finance lead opens the model provider invoice and f01A
BDBhautik Dalwadiinbhautik-ai-tech.hashnode.dev·21h ago · 20 min readGoogle OKF vs RAG: What’s the Difference, and How Do They Work Together?RAG has become one of the most common ways to connect large language models with external knowledge. But as AI systems become more agentic, another question becomes important: how should that knowledg02IA
HGHarrison Guoinharrisonsec.hashnode.dev·12h ago · 7 min readLLM-as-Judge Is Not a Score. It Is a Reasoning Contract.Most teams use an LLM judge the same way. Write a prompt that ends in "rate this from 1 to 10", call the model, read the number, gate on it. Ship. That is not a judge. That is a vibe with a number sta02I
ARAleksei Romanovingfactor.hashnode.dev·7h ago · 14 min readMulti-Reward RL, Part 2: Benchmarking GRPO, DAPO, and CISPO on Unseen TasksFollow-up: Part 3 scales the CISPO + REPO-R recipe to Qwen3.8-27B and 600 steps, with a one-change-per-run holdout ladder and an advantage-floor failure we found in a harsher environment. Part 1 analy00
MKMuhammad Kamraninjudgemyai.hashnode.dev·12h ago · 1 min readReviewer agreement is not the same as evaluation accuracyTwo reviewers can agree and both be wrong. Exact label agreement answers one narrow question: how often did they choose the same label? For a small teaching exercise, give both reviewers the same evid00
DKDaniel Krydynskiindanielkrydynski.hashnode.dev·9h ago · 26 min readOuroboros: A Recursive Dev Loop Where AI Improves Code — SafelyOuroboros — the recursive dev loop, a walkthrough Built by Daniel Krydynski, with Kiko. A harness that lets AI agents continuously improve a codebase — safely. This is the inside-out explanation of th00