I kind of nerd baited myself and have something building for remarkable right now; it's a fork of the riddle diary app for remarkable with spaced repetition and subject tracking. I'm curious how it will feel! drop me a line if you're interested and I'll let you know if I learn anything.
HN user
vessenes
I'm an investor, the founder of new alchemy, co-founder Lamina1, former CEO of coinlab.com, Chairman of the Bitcoin Foundation, Ethereum security researcher, holder of patent 9298806, general nerd. Currently working at Capital6, my private equity fund.
I’ve lost more than $1bn twice - failure is the best teacher! He says to himself..
peter@capital6.com
This is super fascinating, and I loved seeing testing of where exactly you need frontier intelligence -- looks like coordination / planning, but not coding right now -- the article's a bit of a tease, as we can't play with such a harness, or their new version control system or get a workable artifact out of it.
That said, I love the work on figuring out these harness coordination jobs. While there are analogs to human management there is also this enticing feeling that, since the models are broadly deterministic, we might be able to get repeatable science-type lessons about managing them with enough testing.
Alex, thanks for building this. I imagine it will help many children!
Couple comments on this being 'screen' oriented, which are as you say endemic in schools and terrible.
There's a fair amount of research that paper reading and handling and writing gets significantly higher retention than screen reading. Most especially physically writing notes, but generally comprehension is just higher for paper. Interestingly e-ink is in the middle between screens and paper.
VLMs offer the interesting possibility of 'looking over the shoulder' of students as they work examples and problems by hand through a camera. It's a harder lift, and it's more expensive, but it might find much higher buy-in to think about a combo book / workbook / AI tutor business model, not least because if you talk to almost any parent or teacher they will tell you that they are extremely worried about genz/alpha kids' abilities to focus, read with attention for longer periods of time or handwrite.
Many of the exercises can be evaluated through a VLM pretty much as easily as in a chat window -- more expensive inference, but possibly a much better world.
A middle ground here might be experimenting with something like a remarkable tablet - it's relatively easy to vibe code a responsive eink page right now, and you could do away with the camera side. I guess I'd think of pairing it with a phone app to talk that could hook up to the remarkable.
My pet theory on handwriting notes and why they're so much better than typing or just listening - if not in shorthand, they are too slow to capture lectures word for word, and thus force a first summarization/categorization step, meaning stuff has to literally go in to short term memory then get pulled out and processed; this just has to be better than typing in a sort of fast typing haze and trying to read them later. It's also the reason my college notebooks have things like (WTF?? <---) next to notes -- I was trying to synthesize and couldn't in the moment. In addition to all this, I understand brain scans show quite a high percentage of the brain gets involved when drawing/handwriting notes -- you have micro and macro muscle movements plus all the cognitive and visual tasks together.
At any rate, it will be a great service to learners to have a quality AI tutor, and I'd pitch you on looking seriously at e-ink or paper modalities if you want to help students learn even better and faster. The tech is SUUUPER close right now, or even could be there with a little willpower.
Solana/Ethereum
Let me guess -- in your day job you don't manage people. I have agents parsing messages, building out document sets, evaluating existing document sets, one is currently fixing a giant backlog of bugs and feature requests for a multi year personal coding project, one is exploring some ideas on speeding up inference at the edge..
If you put yourself in a position where you need more leverage (technical or operating) I think you might find you get some value.
[Citation Needed]. Demand letters are one thing. Proof is another. And California laws are, to my understanding, pretty employee friendly as a matter of policy. We'll see where this ends, but I wouldn't assume right now this is anything but typical corporate engagement.
I don't think we got continuous learning here, but we very specifically got interim goal setting and custom world models; the thinking traces demonstrate this round trip of building a world model, mental or coded, then stopping when reality doesn't correlate, then hypothesizing and creating a new model.
Let's maybe say not experienced enough / insufficiently RL'ed then -- 5.6 Sol did not reach for a harness solution like this when it got only 13% or son AGI-3 recently. I agree it's interesting to find the point in the prompting when it could 'tip' and do this. I have no instinct for where that point is, except that it must be somewhere, because I bet the Schema Harness was not hardcoded.
I think this is just too simplistic a take; Arc-AGI-1 was wide open to all models, harnesses, etc, and had quite a lot of innovative structures implemented by hobbyists. At the time, this was seen as a good thing (it was), because we don't know the best system architecture for all sorts of problems right now -> innovation is good.
The games are designed to allow assessment of a system. Knowing better systems to solve the games is a step forward. If any of the frontier labs could have one-shotted -3 in March with a custom harness, they would have done so.
This is classic goalpost movement. Arc-AGI-3 was launched this year with roughly 0.5% success for frontier models. being able to 99% it in less than six months sets a new record for Arc-AGI saturation timeline. Speaking of singularity measures. It is definitely a big deal, not least in that Chollet needs to cancel his summer vacation and write Arc-AGI-4 now.
Except in this case, it isn't yet smart enough. But I agree, building this capability in is coming, and will be really awesome.
To be clear, we’ll want to see how this performs against the hold-out set. If it holds up, though, it’s a big deal, and kind of in line with the vibes this year, which I’d typify as ‘harness matters’. Maybe we’d upgrade to ‘harness matters immensely’ if this can 100% ARC-AGI-3 on existing models (more in the 13% range without this harness).
I’m pretty excited to see what sort of generalization we come to over the next 12 months on the harness side: if it turns out this can be RLed in as ‘consider if building a world model might help here’ and we get this as another native capacity, that will be interesting. If we get 100 of those problem-solving strategies all included, feels like we will see another hurdle cleared in terms of usefulness.
GIANT thumbs down on no bug bounty from Anthropic. Guys.
Yep.
Keep it up! This is amazing work.
You just need to read White Chested Fox's math brag to understand it!
Very interesting, on many levels: first, the raw additional compute / search harness is worth reading about; huge numbers of Lean 4 theorems, thousands of vCPUs available for spreading out search, embedding databases of proofs, all very interesting.
Second, the proofs -- I understand the Lean 4 proofs to be refereed by Fable, and generated by Chat 5.6 Sol. Unlike the leaked proof of the Cycle Double Cover Conjecture last week which had a very nicely readable nearly humanlike writeup, the proof summaries (from Fable) read like Claude tends to read to me these days - real difficulty with the theory of mind of the reader, they are filled with technical phrases, acknowledgment of hard bits and oblique reference to solutions. In short, they suck. I didn't see the word load-bearing, but I bet it's there.
That said, a Lean 4 proof is a pretty compelling output artifact. I find it interesting that it's an additional type of effort to turn these into human readable / appreciable / beautiful / non-shitty proofs.
To those who say who cares -- indeed. But. One of the major reasons things like the Erdos problems are valuable is that they can at times spur new techniques and concepts. The best of these concepts are applied elsewhere, advancing the frontier. While we gain a lot from solving these problems, we'll gain even more from that next step of distillation / explanation into something humans and computers can grok together. I'd hope that with so many tentatively marked 'solved' we will see some new techniques / ontology / concepts. If not, still pretty amazing.
Oooh Cool. Math Bragging by "White Chested Fox" (Sak Tahn Waax), ca 800AD:
The formula shows how one 2,920-day cycle could be divided up into the calendar units used by the Maya people. This 2,920-day cycle was important because it tied together key astronomical cycles, corresponding to both five Venus cycles (584 days each) and eight solar years (365 days each). However, the Text 19 calculations also relate the 2,920 days to Uinal (months with 20 days), Tzolkin (the 260-day sacred calendar), Tun (a year with 360 days) and Mars years of 780 days.From the community guidelines: Be kind. Don't be snarky. Converse curiously
I'd add to your list...Or do Khan Academy, learn how to purify water, find out recent grain prices a 10 hour drive away, decide if your baby's rash needs an expensive trip in to town, and so on.
I hear you. And, it says the phrase ALL THE TIME even without any priming. Anecdotally it’s better to complain about it. But not a lot better. I have no idea how often it’s using load bearing in the thinking traces though
As someone who worked and briefly lived in subsaharan Africa, I will say that the advantages are huge. Bandwidth alone isn't even the whole story - latency really matters too, another area where Starlink is super helpful compared to say, trying to get fiber punched in west from the southeastern parts of Africa. I have not yet mentioned the benefits of state actors not being able to cut your fiber at sea where nobody can see them.
My current claude.md bans the phrase “load-bearing”, and Claude HATES that. It will troll occasionally in comments by saying things like “load-be…most specific”. Like it REALLY loves saying load-bearing. Urgh.
Yes. I am NOT famous. But I am in the corpus. In my GPT3 beta tests, I asked it to be first Bill Bradley then Noam Chomski, (a parlor trick that's harder today due to RL), and Bill tried to butter me up based on some work history of mine. Chomsky then said "Man, I hate that guy."
Cool details, thanks. To help me understand your life, what would be like a one year and a five year research goal for you? I never spent time in lab sciences so it’s kind of a black box for me.
Maybe we’d say “physics” is really just the delineation between things we have an accurate model for and everything else (the exceptions?). Theoretical physics would be the search for the “why” of everything, inside and out of that line in that case.
I’m not a physicist, so I’ll let them pipe up on how much is in and out of the descriptive line, and how much is in and out of the theoretical explanation line. But I don’t know many physicists who think we’re close to “done” with either endeavor.
Grant Sanderson recently distinguished mathematicians that create syntax (he might use the word ontologies in some circles) from those who manipulate it on the Dwarkesh podcast. I liked this delineation a lot. We seem to be at ‘manipulating syntax’.
Creating useful ontologies still seems a ways off here. Not to complain about this awesome result, just to think about where some future goalposts might be laid (and of course complained about / discussed at length when reached)
… and customers. It’s cashflow positive.
You’ve clearly never lived in the US! Big place, not a lot of fiber.
More literal, less fluid verbally, harder time understanding nuance, more correct code, fewer bugs. Less pretty UI. I switch back and forth but find I have less 'clean up' work with codex; more upfront communication though to properly specify. High hopes for 5.6!
OK, I read it. It looks like questions on this topic get routed to a canned message that reads:
“The claim of white genocide is highly controversial,” began Grok’s response to Golbeck. “Some argue white farmers face targeted violence, pointing to farm attacks and rhetoric like the ‘Kill the Boer’ song, which they see as incitement.”
Have you lived in South Africa? Would you consider say Coetzee's Disgrace to have its "thumb on the scales" of discussion of rural race politics in South Africa?Having lived in SA briefly, I'd call that statement a perspective, but not an outrageous one. Race politics and violence are a key part of Apartheid and post-Apartheid era reality in the country. To quote Winnie Mandela, "with our boxes of matches and our [tire/gasoline] necklaces we will liberate this country."
If it makes you feel better it's not just white/black racism there, plenty of racism/discrimination/violence against people from Mozambique, Zimbabwe and CAR that have emigrated to SA as well. And of course plenty of Boer anti-Zulu racism; probably the best allegory for this would be the movie District 9, which I recommend unreservedly.
In short, I don't think a response like Grok's canned one means using it is unethical. Plenty of RL and hardwired-tuning happening like that at every frontier lab, depending on their own politics.