Jjasmineparkinjas-blogs.hashnode.dev·3d ago · 14 min readTrace sampling at 10 percent: a $238 invoice against my $82 estimate A product manager asked what one feature had cost us in April. I pulled the number from our trace backend, divided by the sample rate, and gave her a figure just over eighty dollars. She came back a w00
Jjasmineparkinjas-blogs.hashnode.dev·Aug 6 · 7 min readA 4% cache hit rate was costing us money. Here is the arithmetic I should have run first. We turned on prompt caching for our document-QA service and the invoice went up. Not dramatically. About 5%. Enough that I assumed it was traffic growth for the first two weeks, and it was not. This p10
Jjasmineparkinjas-blogs.hashnode.dev·Aug 6 · 10 min readThree days of overage after a four-minute revert The change went out at 12:00 on a Monday. Nothing paged. Our spend-rate alert fires at 3x the trailing hourly median, which is built for a tenant running away with the bill, and this was a steady 50 p00
Jjasmineparkinjas-blogs.hashnode.dev·Jul 29 · 5 min readOur p50 latency SLO was green all quarter. Nearly 1 in 10 sessions hit a wall anyway. The dashboard said 1.9s p50 against a 2.5s target. Green. It stayed green the entire quarter. Meanwhile churn in one segment crept up and the support inbox filled with "the assistant is so slow" from 00
Jjasmineparkinjas-blogs.hashnode.dev·Jul 24 · 5 min readCost per action: the number your LLM spend dashboard cannot produceTL;DR. Cost dashboards for LLM systems report spend per model, per provider, per thousand calls. Finance asks what the support copilot costs per resolved ticket, and the dashboard cannot answer becaus00