Really useful breakdown—especially the separation of appearance from motion and the idea of revising the smallest responsible layer. I’m exploring similar multimodal reference workflows with Seedance 2 AI. I’m curious: when the target model’s duration limit is shorter than the reference shot, do you preserve the original motion speed and split the shot, or compress the beat sheet while keeping the same start and end states?