Good catch, check-then-create never gave you exclusivity, two requests can both read "nothing here yet" a few microseconds apart. Making the claim itself atomic fixes it: insert the key under a unique constraint, so the second writer hits a conflict instead of a false green light, and gets "in progress" back instead of falling through to a second create.
Crash after the external effect is the part that key claim can't fix on its own, though. You still need something that checks actual state with the downstream system before deciding what "in progress" turned into.
On the retention window, no, I wouldn't read an expired key as proof nothing happened. It just means the provider stopped promising to dedupe. I keep my own case record with its own retention, separate from whatever TTL the provider uses. If the key's gone and a retry shows up, that internal record is what I trust, not the provider's memory.
One test I would add: release two requests with the same key at the same instant. A check-then-create path can let both workers see an empty record, even though a later sequential retry works.
I would claim the key atomically before the side effect, bind it to the operation and payload, and give the second worker an explicit in-progress result. A crash after the external effect still needs your reconciliation state; the database claim alone cannot settle that outcome.
How do you handle recovery after the provider retention window expires? I would avoid reusing an expired key as proof that the original action never happened.