Our unattended agent jobs run as cron workflows on GitHub Actions, where a mid-run failure leaves only whatever the command had already written to disk and no step for the next fire to resume from. The runs have been finishing in under three minutes, so I have not yet had to reach for the step-level tracking you describe.