The 2,500 and 1,500 element rows are a useful sanity check because they force a ragged tail on every row instead of hiding it in one launch. The next benchmark I would add is a bucketed sweep by tail occupancy, reporting achieved bandwidth and instruction count for 1-3, 4-7, and 8-15 valid lanes. That would separate branch divergence from lost vector width, since the branch can stay warp-uniform while memory transaction efficiency still falls.