Great read on the local-first architecture! One thing I'm wondering about though: doesn't shifting to locally hosted OCR VLMs and serving models locally significantly increase your infrastructure and inference costs compared to lightweight APIs, especially if you have to provision dedicated GPUs?