i think i actually need to dig for a broader breadth of benchmark scores. right now, taking from the model page means i get the 'best of' results. instead of tests where the model sucks - glm is bad at biology for instance.
im going to aggregate as many benchmarks from as many sources as i can