HN user

radioactivist

611 karma
Posts0
Comments99
View on HN
No posts found.

I am somewhat skeptical of this.

First, the headline result of 0.7*sigma improvement is the output of a statistical based on lessons/reviews they engaged with and their mid-term score, with that shift being for "full engagement". Based on their tables something like ~16 students (11% of the group) actually reached that level of engagement

Second, trying to incorporate past grades into their modelling is not a substitute for a randomized trial.

Third, the headline engagement number of 90% is for "engaging with the platform, via Module Review or Lesson Quizzes, at least once". I don't know why much of that couldn't just be attributed to novelty. Or even partly a professor with all sorts of enthusiasm for the platform.

Fourth, the "full dosage" effectiveness is measured based the final exam scores. Were these exam questions produced independently from the "Phosphor" materials? (e.g. by blinding?) Were they checked for direct overlap with those materials? The 0.7 sigma shift is 3 points on a 24 point exam; if even a few of the questions on that exam were very similar to those materials it could account for almost all of it. This is not clear to me from the manuscript.

If this was the case, then it's a question less of "is AI effective" vs. "did the students look at the materials". You could still argue that the AI platform got them to read, but that is a somewhat different statement than the AI helped them learn.

While I echo some of your points, [1] is bad example (as a Canadian).

Research money in Canada is harder to come by; a basic research grant is roughly ~5x-10x lower than a comparable American grant (students are cheaper here, so its not completely proportional, but equipment, travel, etc doesn't scale).

The example for money for poaching international researchers also comes with the asterisk that while they found ~$2B for this, they also are cutting the base funding of the federal granting agencies by a few percent at the same time, atop of that funding being anemic for decades at this point. A big "fuck you" to the Canadian research community in my opinion.

Also a physicist here -- I had the same reaction. Going from (35-38) to (39) doesn't look like much of a leap for a human. They say (35-38) was obtained from the full result by the LLM, but if the authors derived the full expression in (29-32) themselves presumably they could do the special case too? (given it's much simpler). The more I read the post and preprint the less clear it is which parts the LLM did.

Prism 6 months ago

Is anyone else having trouble using even some of the basic features? For example, I can open a comment, but it doesn't seem like there is any way to close them (I try clicking the checkmark and nothing happens). You also can't seem to edit the comments once typed.

Prism 6 months ago

In my circles the killer features of Overleaf are the collaborative ones (easy sharing, multi-user editing with track changes/comments). Academic writing in my community basically went from emailed draft-new-FINAL-v4.tex files (or a shared folder full of those files) to basically people just dumping things on Overleaf fairly quickly.

Seems like the someone dug something up from the literature on this problem (see top comment on the erdosproblems.com thread)

"On following the references, it seems that the result in fact follows (after applying Rogers' theorem) from a 1936 paper of Davenport and Erdos (!), which proves the second result you mention. ... In the meantime, I am moving this problem to Section 2 on the wiki (though the new proof is still rather different from the literature proof)."

This is a comparison between a new and interactive medium (+ slides, mind-maps, etc) and a static PDF book as a control. How do we know that a non-AI based interactive book wouldn't give similar (modest) increases in performance without any of the personalization AI enables?

It happens again in the next video. It says:

The team came up with a use case the teaching team hadn’t thought of – using AI to critique the team’s own hypotheses. The AI not only gave them criticism but supported it with links from published scholars. See the demo here:

But the video just shows Claude giving some criticism but then just says go look at some journals and talk to experts (doesn't give any references or specifics).

At one point this states:

Claude was also able to create a list of leaders with the Department of Energy Title17 credit programs, Exim DFC, and other federal credit programs that the team should interview. In addition, it created a list of leaders within Congressional Budget Office and the Office of Management and Budget that would be able to provide insights. See the demo here:

and then there is a video of them "doing" this. But the video basically has Claude just responding saying "I'm sorry I can't do that, please look at their website/etc".

Am I missing something here?

I'm not the person you're replying to, but in my subfield (scientist is such a broad term) I would say in my opinion at least half of those key problems that are listed in the article are basically non issues. Things really are quite different field to field.

OpenAI o4-mini-high

   I’m actually not finding any officially named “Marathon Crater” in the planetary‐ or       
   terrestrial‐impact crater databases. Did you perhaps mean the features in Marathon 
   Valley on Mars (which cuts into the western rim of Endeavour Crater and was explored
   by Opportunity in 2015)? Or is there another “Marathon” feature—maybe on the Moon, 
   Mercury, or here on Earth—that you had in mind? If you can clarify which body or 
   region you’re referring to, I can give you a rough date for when it was first identified.

Most of their categories have straightforward interpretations in terms of students using the tool to cheat. They don't seem to want to/care to analyze that further and determine which are really cheating and which are more productive uses.

I think that's a bit telling on their motivations (esp. given their recent large institutional deals with universities).

Some hard problems have remain unsolved in basically every field of human interest for decades/centuries/millennia -- despite the number of intelligent people and/or resources that have been thrown at them.

I really don't understand the level optimism that seems to exist for LLMs. And speculating that people "secretly hate LLMs" and "feel threatened by them" isn't an answer (frankly, when I see arguments that start with attacks like that alarm bells start going off in my head).

The data set quality seems a really spotty based on looking a few random problems (I looked at about a dozen in the "Physics" subcategory). Several problems had no clear question (or answer) and seemed to be clipped from some longer resource and thus had back references to Sections and Chapters that the models clearly couldn't follow. Worse is that the verification of the answer seems to be via an LLM and not all that reliable; I saw several where the answer was marked correct when it clearly wasn't and several that were correct but not in the precise form given as "the" answer and thus were labelled as incorrect.

There are caveats there too. Generally topological qubits can be immune to all kinds of noise (i.e. built-in error correction) but Majorana zero modes aren't exact the right kind of topological for that to be true. They only enjoy protection on most operations, but not all. So there is a still a need for error correction here (and all the complication that entails) it is just hopefully less onerous since only essentially one operation requires it.

I understand the claim and what they are trying to do (and they've been trying to do it for 20 years now). It's an interesting approach and it is orthogonal enough from other efforts that it is absolutely worthwhile to pursue scientifically (I'm in an adjacent field in condensed matter physics).

But they are doing a full court press in the media (professionally produced talking head videos, NYT articles/other media, etc, etc) claiming all of those things you've just said are right around the corner. And that's going to confuse and mislead the public. So there needs to push back on what I think is clear bullshit/spin by a company trying to sell itself using this development.

The ideas that underpin their device have been around for some time and aren't called by that name in the literature -- it appears to be entirely a branding exercise. A clear signal to me they don't seriously think it is a good name is that don't use the name outside this article (it appears nowhere in their Nature paper or anywhere else for that matter).

A few things to keep in mind, given how hard of a media push this is being given (which should immediately set off alarm bells in your head that this might be bullshit)

- Topological phases of matter (similar, but not identical to the one discussed here) have been known for decades and were first observed experimentally in the 1980s.

- Creating Majorana quasiparticles has a long history of false starts and retracted claims (discovery of Majoranas in related systems was announced in 2012 and 2018 and both were since retracted).

- The quoted Nature paper is about measurements on one qubit. One. Not 100, not 1000, a single qubit.

- Unless they think they can scale this up really quickly it seems like its a very long (or perhaps non-existent) road to 10^6 qubits.

- If they could scale it up so quickly, it would have been way more convincing to wait a bit (0-2 years) and show a 100 or 1000 qubit machine that would be comparable to efforts from Google, IBM, etc (which have their own problems).

I'm a bit skeptical of this study given how it is unpublished, from a (fairly junior) single author and all of the underlying details of the subject are redacted. Is there any information anywhere about what this company in the study was actually doing? (the description in the article are very vague -- basically something to do with materials)