The export section is where I would add a coordinate-level check. ONNX or TensorRT execution can produce convincing detections while the application maps boxes incorrectly after resizing and padding. The normalized and pixel-coordinate representations you explain earlier make a useful test fixture for that boundary.
Run a few non-square images through both the original and exported pipelines, then compare boxes in the original image coordinates with identical preprocessing and suppression settings. Include objects near the padded edges and predictions close to the confidence threshold. That catches integration mistakes that an aggregate speed measurement will miss, and separates export differences from postprocessing differences.