For math, even the frontier has shortcomings, and there is a steep drop from GPT 5.5 xhigh to anything else. The time wasted by less-than-SotA just isn't worth it.
Thank you for this comment. It is great to hear an inside take.
Idle curiosity, but what NLP tools evaluate translation quality better than a person? I was under the (perhaps mistaken) impression that NLP tools would be designed to approximate human intuition on this.
Their results are also highly biased. Most senior researchers aren't going to waste their time filling this out (90% of people did not fill it out). They almost certainly got very junior people and those with an axe to grind. Many of the respondents also have a conflict of interest, they run AI startups.
The survey does address the points above a bit. Per Section 5.2.2 and Appendix D, the survey had a response rate of 15% overall and of ~10% among people with over 1000 citations. Respondents who had given "when HLMI [more or less AGI] will be developed" or "impacts of smarter-than-human machines" a "great deal" of thought prior to the survey were 7.6% and 10.3%, respectively. Appendix D indicates that they saw no large differences between industry and academic respondents besides response rate, which was much lower for people in industry.