The async sleep comparison makes the event-loop problem easy to see. Moving blocking work to a plain def endpoint is useful, but I would measure thread-pool saturation too: it changes where requests wait rather than making the blocking operation disappear.
For the chat pipeline, a load test should separate connection-pool wait, model latency and response serialization, while recording event-loop lag. A bounded inference queue with an explicit overload response would keep many slow model calls from occupying the application indefinitely. That gives the framework benchmark a clear boundary from the end-to-end service behaviour you are describing.