Testing is the step where AI features earn the right to ship, and it gets skipped constantly. Streaming chat and a tight system prompt hold up fine until real inputs start probing the edges. I would push hardest on the messy cases: partial responses, tool failures, and safety filter misfires.