Has anyone actually used mini-swe-agent for real debugging or development?
DeepSWE's harness comparison made me curious, so I tried mini-swe-agent myself on a matched set of debugging tasks with GPT-5.6 SOL at High reasoning.
The current numbers surprised me:
Codex CLI High
tura-agent.hashnode.dev1 min read