Keeping floating-point precision because INT8 flattened the near field is a more useful result than assuming quantization is always the right optimization. For the Sonar mode, I would be careful about turning the average depth of each half into a distance-like cue: relative monocular depth does not establish a stable metric scale by itself.
A nearby narrow obstacle can also disappear into that average when the background occupies most pixels. Testing a lower-depth percentile within the walking corridor, alongside explicit frame-age and invalid-output checks, would expose that failure mode. The heatmap can look convincing while the audible cue misses the object that matters most.
Keeping floating-point precision because INT8 flattened the near field is a more useful result than assuming quantization is always the right optimization. For the Sonar mode, I would be careful about turning the average depth of each half into a distance-like cue: relative monocular depth does not establish a stable metric scale by itself.
A nearby narrow obstacle can also disappear into that average when the background occupies most pixels. Testing a lower-depth percentile within the walking corridor, alongside explicit frame-age and invalid-output checks, would expose that failure mode. The heatmap can look convincing while the audible cue misses the object that matters most.