The idea of measuring the sequence instead of the prettiest frame is probably the most important takeaway here. A multi-scene video is really a system, and a single great generation can still fail if continuity, pacing, or the narrative handoff breaks.
I also like the failure-labeling approach. Once you start tracking things like identity drift, camera mismatch, and prompt ambiguity, QA stops being subjective and becomes a feedback loop for improving the generation process. That’s a much more scalable way to work than simply regenerating until something “looks right.”