Making image_model scene data removes the hardcoded-engine assumption. The illustrated _loaded_pipelines dictionary introduces a separate memory risk, though: every newly selected checkpoint remains resident, so an episode using several overrides can recreate the original OOM even when each model fits on its own.
I would give the cache an explicit residency budget and define when an inactive pipeline can be evicted. An active generation needs a lease so its pipeline cannot disappear halfway through a scene. A sequence alternating three checkpoints would be a useful regression case, recording resident memory after each switch rather than testing only a fresh process with one model.