Really enjoyed this deep dive! What stood out to me is that you didn’t just focus on whether the model could make a decision—you actually looked at the less glamorous parts that matter in production: calibration, cold starts, prompt formatting, cost, and how the model behaves on data it hasn’t seen before.
The “declared mass” detail was especially interesting. It’s a great example of how a model can appear confident while actually responding to something different than what we intended.
I also appreciate the honesty around the limitations of the 4B model. Showing where the approach worked, where it failed to generalize, and why some of the improvements didn’t hold up on the held-out data makes the experiment much more useful than a simple “we built it and it works” post.
The privacy/control trade-off with running the model inside your own AWS environment is probably the biggest takeaway for me. Great technical write-up!