This is a genuinely useful reframe — "treat your publishing site like a software product" is the right mental model, and the honest-disqualification point is the part most affiliate/review sites get backwards. Readers can tell within a paragraph whether a site is optimizing for their trust or for a click, and the disqualification examples you gave (pricing tier jumps, who shouldn't buy) are exactly the signal that's missing from 95% of "best X" content.
I build RAG/LLM pipelines and scraping/automation systems for a living, so the data-aggregation section is close to my daily work. One thing I'd add to the requests/BeautifulSoup/pandas pipeline you described: once you're structuring scraped spec/pricing data into tables, that same structured store is a natural fit for a lightweight RAG layer on top — instead of just generating comparison matrices once, you can let an LLM answer specific long-tail questions directly against your own verified data ("does X support Y at this price tier") rather than writing a page for every possible query combination in advance. That compounds well with your high-intent/long-tail point, since it lets you cover the long tail of questions you didn't anticipate, grounded in data you've actually verified rather than the model guessing.
On the revenue side, since you're already treating this as a software product: the scraping+structuring pipeline itself is a sellable asset independent of the content site. Once it reliably tracks pricing/spec changes for a niche, a small paid API or data feed for that niche (other publishers, affinity marketers, or even the vendors themselves wanting competitive tracking) is a second revenue line that doesn't depend on search rankings at all — which fits neatly with your "own your distribution" point, since it's revenue that isn't exposed to a core algorithm update either.
Solid piece — the framing of content-as-engineering-problem is the right one, and more people writing in this space should think in these terms.