I actually think people who are great at understanding problems, coming up with requirements and designing solutions (all things I would expect someone who is good at churning out MVPs would be good at) are exactly the people most empowered by the current batch of LLMs. Its the people who are only good at working on small chunks of problems that I'm concerned about..
HN user
neonbjb
So it made me wonder. Is Brainf*ck the ultimate test for AGI?
Absolutely not. Id bet a lot of money this could be solved with a decent amount of RL compute. None of the stated problems are actually issues with LLMs after on policy training is performed.
Yes we do. If you worked at Google you know moma. Our moma is an internal version of chat. It is very good.
Also work at OpenAI. Every tender offer has made full payouts to previous employees. Sorry to ruin your witch hunt..
I work for openai.
o4-mini gets much closer (but I'm pretty sure it fumbles at the last moment): https://chatgpt.com/share/680031fb-2bd0-8013-87ac-941fa91cea...
We're pretty bad at model naming and communicating capabilities (in our defense, it's hard!), but o4-mini is actually a _considerably_ better vision model than o3, despite the benchmarks. Similar to how o3-mini-high was a much better coding model than o1. I would recommend using o4-mini-high over o3 for any task involving vision.
You're missing the fact that requests are batched. It's 70 tokens per second for you, but also for 10s-100s of other paying customers at the same time.
You wouldn't do that to this model. It finds its own mistakes and corrects them as it is thinking through things.
I'm James Betker.
Of course architecture matters in this regard lol. Comparing a CNN to a transformer is like comparing two children brought up in the same household but one has a severe disability.
What I meant in this blog post was that given two NNs which have the same basic components that are sufficiently large and trained long enough on the same dataset, the "behavior" of the resulting models is often shockingly similar. "Behavior" here means the typical (mean, heh) responses you get from the model. This is a function of your dataset distribution.
:edit: Perhaps it'd be best to give a specific example: Lets say you train two pairs of networks: (1) A Mamba SSM and a Transformer on the Pile. (2) Two transformers, one trained on the Pile, the other trained on Reddit comments. All are trained to the same MMLU performance.
I'd put big money that the average responses you get when sampling from the models in (1) are nearly identical, whereas the two models in (2) will be quite different.
As an employee of OpenAI: fuck you and your condescending conclusions about my peers and my motivations.
I don't think the plan is an occasionally rocket launch long term. I think the plan is to launch rockets as fast as humanly possible.
Its cool that this is starting to approach real time video territory (30 images per second, this claims close to 1 image/sec).
@dooraven - I also work in ML (including recently working at Google) and I agree with @whimsicalism.
You seem to be under the mistaken belief that: 1. Google has competent high-level organization that effectively sets and pursues long term goals. 2. There is some advantage to developing a highly capable LLM but not releasing it.
(2) could be the case if Google had built an extremely large model which was too expensive to deploy. Having been privy to what they had been working on up until mid-2022 and knowing how much work, compute and planning goes into extremely large models, this would very much surprise me.
Note: I did not have much visibility into what deepmind was up to. Maybe they had something.
You can make almost anything work in DL if you try hard enough, that doesn't mean it is the correct thing to do. Convolutions have inductive biases which are the cause of many of the problems associated with deep learning over the last 10 years. Researchers don't "love the ViT". They use it because it is simply better in every way, in every application.
The only reason convolutions are still used in modern (intelligently designed) ML systems is because it is not known how to build a sparse attention algorithm that achieves 2D and 3D locality and is also compatible with modern accelerators. Swin is an attempt at that, but it is something of a hack.
I theorize (but cannot prove) that the processes underpinning creativity in the human mind are exactly the same statistical processes that ML models use.
Think about it: you live your life. You experience things. You experience art, and experience emotions or have interactions with other humans grounded in that art. You form connections with certain styles or techniques.
If you then turn around to create art, you form in your mind a general idea of what you want to create. You then draw on your past experiences to actually create the physical art. What process other than statistical extraction from your mind could it come from?
For sure I believe there are things that we don't understand about the human mind. I think the impact of drug use on art creation is very interesting, for example. It indicates that random chemical processes in our brains can play a large determining role in the actions we take (and in this case, the things that we create).
But to say that humans do not use some sort of inbaked statistical world model in the creative process seems wrong to me.
AGI consists of a set of problems: 1. Finding an algorithm which is capable of learning anything. 2. Building the computers that can run said algorithm. 3. Collecting, sorting and filtering the data that the algorithm learns from.
DeepMind claims (with good cause, IMO), that MuZero can be such an algorithm. Showing that this one algorithm can tackle disparate problems is a way of proving this.
I think the questions that still stand are: is it even possible to build computers that could drive a scaled up MuZero to AGI? And is there a more efficient way to get there? I suspect the answer to both questions is yes.
Still, I think it is pretty incredible that we've managed to build computer programs that can totally adapt to arbitrary datasets and perform arbitrary tasks.
This is wrong IMO - the fed has NO tools in its hands. Its last tool was inflation rates, and it burned that during the great recession. It is now caught in a pickle: raise interest rates, causing a mass exodus of wealth from securities to savings accounts and CDs and causing the stock market to tumble (like it arguably always should have done since 2008), or let inflation take its course.
I am not an economist, and I'd love to be proven wrong. I just don't see how this ends in a good way for the economy.
There are other reasons to farm indoors (or in vertical farms) than electricity and water too.
Pests are one such reason. They are extremely difficult to control outdoors, but far simpler to do so inside. Both fertilizer and pesticide treatments (if necessary) can be far more specific - e.g. less wasteful and environmentally damaging) when done indoors. Similarly, inclement weather generally does not affect indoor farms.
There's also my favorite argument: if we're ever going to try to colonize space or other planets, we better be damned good at growing plants artificially. IMO every dollar spent improving this space gets us one step closer to unlinking our future from Earth's.