We run something adjacent for this exact kind of job -- scripted browser actions (Playwright launch plus saved storage state) for platform engagement, no LLM per click, so the marginal token cost of a run is basically zero once the script exists. The interesting tradeoff your numbers surface is between that (zero marginal cost, but someone has to hand-write and maintain a script per site) and Jev's approach (goal-only calls, no per-site script, but you pay the decision model on every run). Curious how the decision-model cost compares once you factor in the manual-fallback re-runs after a captcha or blocking modal, since that's exactly the case where a hand-written script doesn't degrade at all.