Hey, author here.
You can check out our evaluation library (https://github.com/wellecks/lm-evaluation-harness) for the exact benchmark implementations we used, including prompting.
In particular, the prompt that starts at line 27 in this file (https://github.com/wellecks/lm-evaluation-harness/blob/maste...) is quite good for high school/olympiad problems. We took this prompt from Google's Minerva paper.