Good benchmark results don't mean identical outputs. The task completion rate is the same: both pass the same exercises. The paths the model takes differ, but the end result is the same -> pass the tests
The full benchmarking methodology and tooling will be published alongside the paper.