Keeping the tutorial focused makes sense. A lightweight compromise would be one final table with seen-template/seen-order, seen-template/new-order, held-out-template, and original-task retention, all using the same fixed seed. That adds almost no framework code but makes memorization, composition, and forgetting visible with the existing 2,160 parameters. If held-out templates fail while reordered examples pass, readers can see exactly what the tiny model learned.
