That deployment article could make the boundary very concrete with one fixture that records the resize ratio, padding offsets and expected boxes in original coordinates. I would compare the outputs before suppression as well as the final detections, so a threshold or NMS difference does not hide a coordinate mismatch.
Keeping the fixture in the exported runtime’s integration checks would also make it useful after model updates, when the application mapping may remain unchanged but the output layout or precision changes.
The export section is where I would add a coordinate-level check. ONNX or TensorRT execution can produce convincing detections while the application maps boxes incorrectly after resizing and padding. The normalized and pixel-coordinate representations you explain earlier make a useful test fixture for that boundary.
Run a few non-square images through both the original and exported pipelines, then compare boxes in the original image coordinates with identical preprocessing and suppression settings. Include objects near the padded edges and predictions close to the confidence threshold. That catches integration mistakes that an aggregate speed measurement will miss, and separates export differences from postprocessing differences.