In section 4.2 they have a figure with the models that each of the studies used that they analyzed. Overwhelmingly (28/36) it was a GPT model. Top 4 use cases for the apps were self-directed information seeking, co-authoring stories with AI, teaching science through dialogue, and eliciting children's emotional expression.
HN user
wpenman
new SWE, formerly Princeton and Rhetoric PhD building philosophically interesting LLM benchmarks https://willpenman.com
I'm pretty lost, but it seems like a really interesting project.
Thanks for the feedback, I should revise the copy there. I think when I started the contest I wasn't fully sure myself. Here are the goals that stand out to me. 1. Is auto-scoring an essay really possible? (Meaning, cost-effective, reliable, valid) This is more of a personal motivation - I worked hard on the 190 items and receiving responses will help me figure out if they're full of shit. 2. It would be super cool to see an “ideal” essay on this! AI safety is so topical. The contest asks entrants, in part, to say: Is Klara and the Sun in favor of AI safety efforts? If no, that’s awkward! If yes, on what grounds? You would think since I made up the rubric that I know what the "right" answer is here. But there's a lot of ambiguity in the text. 3. (For people writing their essays via AI) What is the value of prompting if/when we get to ASI? For most prompt engineering-style entries, I figure they aren't also a scholar in the humanities. So entering the contest is a test case of trying to prompt AI to do something "better" beyond what they can already do themselves. You could spin that up as prototypical of what life in general might be like in some amount of time. What prompt techniques even survive when you can't really assess what "Now make it better!" should look like? That's interesting! 4. (For people writing their essays as humans) Maybe this contest format is more humane than typical peer review? You always know where you stand. You get a score back in a few minutes, as soon as the scorer finishes going through the rubric. And the revision process is private, with no loss of face. In contrast, scholars in the humanities sometimes have to wait up to a year to hear back about a submission.