It might be kind of overlooked when people read about the big scary results from mythos; the real breakthrough was probably just as much the application of the (very decent) model through a well engineered wrapper (harness). Other models including codex or glm result in significant findings as well.
Harness example: https://github.com/evilsocket/audit