Splitting the question by workstream is the right move, and the backend migration versus recommendation engine example makes it concrete. A migration has a clear before and after, a known end state and a test for done. A recommendation engine is ongoing learning about your own users. Those deserve different staffing models because they are different kinds of problems, not because one matters more.
The test I find most useful is whether the piece of work can be written as one problem with a fixed outcome. If it can, for example "move these three services to containers with zero downtime and hand over the runbooks", it is a good candidate for outside delivery at a fixed scope. If the outcome can only be described as "keep improving", it usually belongs with the people who will live with it.
On the question in the responses about who owns bugs after a handoff: in my experience that has to be agreed before the work starts, with a defined warranty window and a written list of what the internal team will own afterwards. When that is left vague, the operations cost quietly lands on whoever answers the pager first.