Extraction on partial transcripts plus reasoning, CRM write and synthesis running concurrently is the right shape, and the admission about chasing the wrong fixes first is the part other people will actually learn from.
Two things that tend to bite at the next stage. Partial transcripts get revised, so an extraction fired on a segment the ASR later corrects can write a wrong value confidently. Worth deciding explicitly whether fields are overwritten on revision or held until a segment is stable.
And once the steps run concurrently the CRM write is no longer last, so it needs to be idempotent. Otherwise a retry after a partial failure duplicates the record. Both are cheap now and unpleasant to retrofit.
Post-call CRM entry is a good problem to pick because the output is structured and checkable - name, intent, follow-up date are fields you can validate, not a summary someone has to judge. That also makes the failure mode manageable: a wrong date is visible to the rep, unlike a subtly wrong paragraph. Two minutes to three seconds is a big enough jump that the interesting detail is where the time was going. In pipelines like this it is usually not the model but the serial arrangement around it - waiting for a full transcript before extraction starts, one request per field, or a synchronous write at the end. Curious which of those dominated for you, since that is the transferable part of the result.