Do those adjacent organisations inspire confidence?
HN user
9q9
Agreed, Friston's Bona Fides are impressive. (Aside: His fame in neuroscience comes from him having written important FMRI software that everybody cites.)
That's also why I worked with his team and read a lot of his papers for a while. His principal idea was originally that neurons perform free energy minimisation. This idea makes a lot of sense, once you understand what free energy means. But, to the best of my knowledge, it has not at all been empirically verified for neurons (I'd be delighted to be proven wrong in this belief). So he went the route of generalising the free energy principle: "the free energy principle asserts that any “thing” that attains a nonequilibrium steady state can be construed as performing an elemental sort of Bayesian inference". Terms "can be construed" and "elemental sort of Bayesian inference" do a lot of work here. Updating and generalising one's research hypothesis is legitimate (albeit one could be more explicit about this), but it weakens the claim being made. Anyway, under a charitable interpretation of those terms, I agree that this is true, but, at the same time doesn't say much. Indeed, under the charitable interpretation it basically equates doing free energy minimisation with existence. Friston has lately said that the FEP is not falsifiable. Take it from the horse's mouth (i.e. a Verses employee): "the free energy principle just applies to stones, it applies to birds, it applies to any kinds of animals" on Machine Learning Street Talk [1].
Here is my current position: from a mathematical principle this general one cannot derive scalable ML algorithms!
he seems to have joined only in 2022, they were 4 years old at the time.
The company founders have a cryptocurrency and (later) metaverse background.
Benefit of doubt is a great concept. The other extreme is: extraordinary claims require extraordinary evidence. Why should we restrict ourselves to a binary choice? Can we not think in a more nuanced fashion, in Bayesian terms? In other words look at all available evidence and assign probabilities?
"We are the next DeepMind" is easy to say ... The DeepMind founders had a stellar predigee in computer games, AI and neuroscience, the Verses founders have a cryptocurrency background. Verses also released [1] last month. What both the Atari and the Mastermind announcements have in common is the lack of details, including code. Why do they not show their code? How do we know their figures are real? We've just had the OpenAI vs FrontierMath discussion [2, 3]. Presumably, being able to play Pong, a 1972 computer game, is unlikely to be their moat ...
Interesting also their 2024 MLST presentation [4]. Does that inspire confidence? It was that video that made my priors on Friston having had a breakthrough in ML change downwards dramatically ... But do not take my word for it, please make up your own mind.
[1] https://www.verses.ai/blog/genius-outperforms-openai-model-i...
[2] https://techcrunch.com/2025/01/19/ai-benchmarking-organizati...
So I'm not the only who wonders about the hyperbole emanating from Friston et al! Some more morsels:
- The CEO is an "International Bestselling Author" [1].
- The company blog states that Friston has "successfully [decoded] the underlying mechanisms of intelligence as it functions in the brain and biological systems" [2].
But they got $10M investment from G42, an Emirati VC [3]. Note that G42 have also invested in Cerebras and OpenAI [4]. So their PR works.
[1] https://www.linkedin.com/in/gabriel-ren%C3%A9-0201902/
[2] https://www.verses.ai/blog/blogs/letter-from-the-ceo
[3] https://21624003.fs1.hubspotusercontent-na1.net/hubfs/216240...
Geoff Hinton was denied an academic position at the University of Sussex's CS department where he had done postdoc work (That department is now 'famous' for consciousness studies and integrated information theory https://osf.io/preprints/psyarxiv/zsr78. I bet they are kicking themselves now ...)
"Academia will one day wake up, and realize that"
Charlie Munger famously said, "Show me the incentive and I'll show you the outcome" ...
expertise either mad or catatonic.
Yes.Unfortunately it still works really well with students, who are not yet knowledgable enough to be able to recognise the Fristonian ideas for what they are--hot air--and who waste their time with active learning before realising that it's currently mostly hot air. I've seen this play out several times by now with my (current and former) students.
Regarding average: which average in the sense of: average over what time window? Any specific choice here needs to be justified as happening in the brain.
Regarding "we will find brain models with observable activity that follows the FEP?": abstractly you are saying that your prediction for theory T is that we will eventually confirm T. This does not exclude anything, I can state this for any theory T whatsoever. (For fun, try to instantiate T with outlandish theories, e.g. with "We will eventually find weapons of mass destruction in Irak", or with plausible theories that have failed so far, e.g. "We will eventuallly see supersymmetry". Does your prediction rule anything out?)
Regarding variational Bayes, that was not invented by the Free Energy millieu.
Gambling etc is not the dark-room argument, I've explicitly left out the dark-room.
Coincidentally, Friston's treatment [1] of the dark room is not convincing, but it nicely illustrates Friston's tendency to make ad-hoc adjustments, for example in [1] he talks about "average" surprise, but there are many ways you can average. Which one is it? How for example do the 302 neurons of C elegans average? Saying this is a difficult task is correct given our understanding of neurons in 2020, but the fact that Friston seems to think Free Energy accomodates all possibilities means it in "not even wrong" territory. In it's current shape, Free Energy does not make interesting predictions for neuroscience, and none of the progress in AI/ML has come from the Free Energy millieu either.
If "surprise is used in a very technical statistical sense" means something concrete, precise, for example minimising KL-divergence of states, the question becomes: show me that this is what the brain does. Or build an AI that does something that is competitive with other forms of contemporary AI.
[1] K. Friston et al, Free-Energy Minimization and the Dark-Room Problem https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3347222/
I have explicitly stated that I am using a simplistic interpretation.
I am neither seeing that Friston has (A) produced anything even remotely resembling a testable framework this "kind of surprise that is being minimized in the free-energy framework" and (B) pointed to any plausible mechanisms in the brain that should that this is in fact "the kind of surprise that is being minimized". He just handwaves.
What clearcut evidence can you give me that humans minimise this "kind of surprise"? What evidence would you accept as falsifying this? Where does Friston make clear that "secondary motivations" don't count? Also making a super vague, unquantified statement like "large contingent of people who are desperately trying to enact conservatism ..." in defense of Friston / free-energy doesn't give me a lot of confidence in the social milieu that this theory comes from. All the more so, since my OPs explicitly criticised Friston for vagueness.
Thanks.
I looked at that link but could not find the code that lets me reproduce the results in the paper. Maybe I didn't look hard enough?
I cannot see the "these routines [being] available in the development version of the next SPM release". The development version is a 111 MB zip file [1]. When I uncompress the file I get a big flat directory with 100s of files. Which of those is is the software used in the paper? I have a bad feeling about this. I don't see how the authors are displaying intellectual integrity by not releasing, concurrently with the paper, software for such an important problem public health issue.
ideas are hardly vague.
Hard to understand
The core intuition is easy to understand: brain predicts its observations including observations about itself (proprioception) and acts in a way to minimise surprise. This can be seen as a form of self-supervised learning in the terminology of contemporary machine learning. Lots of people have said somewhat similar things before at a similar level of vagueness. Nobody disagrees that "somehow" the brain learns about the world by prediction and interaction. The interesting question is to go beyond this vagueness: what exactly is the brain doing? Where exactly is the brain minimising 'free energy'? Can I have a testable prediction please?If read literally, Friston's core intuition is false: people regularly and deliberately expose themselves to surprise, e.g. gambling, watching sports, speed dating. Now there are various ad-hoc fixes to save free-energy-minimisation, which should make the theory more testable, but Friston then has to state clearly which of the many conflicting ad-hoc fixes are in place, and explain how they manifest themselves in the brain! Friston has been confronted with those problems many times, but he basically ignores them.
Short answer: no!
Friston's work is (in)famous for being so vague as to be completely untestable. He has been criticised over the extreme vagueness of his ideas many times, and he has never given a good answer. Sometimes some of his followers try to make it testable as neuroscience or useful for AI. Both have failed so far. As far as I can see, leading working neuroscientists don't take this Friston / free-engery stuff seriously. He's even got a parody Twitter account now: https://twitter.com/farlkriston
He now claims to have the best COVID model based on free-energy. Can I please see the code and form my own opinion? Has anybody seen this code?
Thanks!
The C FFI worries me because if I choose Rust for our project I'll need to interface a huge number of C / C++ libraries with the Rust core. OTOH, Mozilla seems to be doing fine mixing Rust with C++.
7
3+4
2nd-smallest double Mersenne prime
Number of continents
has _a lot_ of issues
I value your perspective on programming languages. Would you mind sharing those issues? I am in a position where I can and do influence PL design. Macron and Merkel learn
Javascript and ...
Why not?
The prime minister of Singapore is an active C++ programmer, and has shared
source code on his Facebook page, asking for bug reports [1].
Code is at [2].[1] https://arstechnica.com/information-technology/2015/05/prime...
[2] https://drive.google.com/drive/u/0/folders/0B2G2LjIu7Wbdfjha...
I agree with most of your points, and the huge unified, and rather homogeneous market is a core advantage of the US in certain product categories. Since you mention SAP: clearly, SAP is successful in a space where Europe's heterogeneity should be a problem -- different legal systems, different accounting rules etc ... and yet SAP succeeded. Maybe it was because SAP was founded in 1972, half a century ago, when European decline was not as pronounced as it is in today?
France has been world-leading in verification, e.g. CompCert and Coq come from INRIA, model-checking was co-invented in France. This stuff is largely language independent. Yet the big sellers of this kind of stuff (e.g. EDA software from Synopsys, Cadence, and Mentor) is in the US.
I work with a lot of French IT people, they are all from the Grande écoles, and amazing on average. None of them work in France. What a loss!
That is true for customer-facing softwre, but would not be relevant in e.g. semiconductors, compilers, formal verification.
In terms of innovation in semiconductors, should also learn from the success of China, South Korea, Taiwan!
"something great"
Also: this "something great" must be something they (EU politicians deciding about funding) have heard about! And what have they heard about? Something US companies have already been commercialising, hence started a big PR offensive.That's why the EU started big funding of research on cloud computing after Amazon made money with it, started big funding of research on search engine when Google made money with it, started big funding of research on AI/ML after Google made money with it ...
I predict that the EU will go all in on supporting research on quantum computing when Google sells it in a big way.
Terra paper
I didn't mean to push Terra in particular, and I agree that doing
meta-programming is easier in the same language you do normal programming in. I
just mentioned the Terra paper since it provided a rationale for using
meta-programming in high-performance programming (I could have pointed
to others making the same argument).Rust's procedural macros are standard compile-time meta programming.
proc macros do AST->AST folds)
Last time I looked (I have not used Rust for a while) procedural
macros operate over tokens, not ASTs. Maybe that changed?Thanks. Amazing overview.
abstract away what they
cannot know
Here I'm a bit surprised. For theoretical reasons (I've never gotten
my hands dirty with GPU implementations, I'm afraid to admit), I would
have expected that it is highly useful to have a language interface
that allows run-time (or at least compile-time) reflection on:- Number of levels (e.g. how many layers of caches).
- Size of levels.
- Cost of memory access at each level.
- Cost of moving data from a level to the next level up/down.
The last 3 numbers don't have to be absolute, I imagine, but can be relative, e.g.: size(Level3) = 32 * size(Level2). This data would be useful to decide how to partition compute jobs as you describe, and do so in a way that is (somewhat) portable. There are all manner of subtle issues, e.g. what counts as cost of memory access and data movements (average, worst case, single byte, DMA ...), and what the compiler and/or runtime should do (if anything) if they are are violated. In abstract terms: what is the semantics of the language representation of the memory hierarchy. Another subtle but important question is: what primitives should a language provide to access a memory level, and which ones to move between levels. An obvious choice is to treat each level as an array, and have DMA-like send/receives to move blocks of data between levels. Is that a good idea?
Equally subtle, and I already alluded to this above, is when to make this information available. Since the processor doesn't change during computing, I imagine that using a multi-stage meta-programming setup (see e.g. [1] for a rationale), might be the right framework: you have a meta-program, specialising the program doing the compute you are interested in. C++ use template for program specialisation, but C++'s interface for meta-programming is not easy to use. It's possible to do much better.
As you wrote above, programming "in those languages is hard and error prone", and the purpose of language primitives is to catch errors early. What errors would a compiler / typing system for such a language catch, ideally without impeding performance?
[1] Z. DeVito, J. Hegarty, A. Aiken, P. Hanrahan, J. Vitek, Terra: A Multi-Stage Language for High-Performance Computing. https://cs.stanford.edu/~zdevito/pldi071-devito.pdf
The 6 requirements you list for doing a dot-product on the GPU can be phrased in abstract as a constraint solving problem where the number of thread blocks, the cost of communication etc are parameters.
Interestingly, a lot of programmers (who program for the CPU) worry about caches a great deal, despite:
- Programmers being unable to control caches, at least directly, and reliably.
- Languages (e.g. C/C++) having no direct way of expressing memory constraints.
This suggests to me that even in CPU programming there is something important missing, and I imagine that a suitable explict representation of the memory hierarchy might be it. A core problem is that its unclear how to abstract a program so it remains perfomant over different memory hierarchies.
Why do you think most languages avoid exposing the memory hierarchy? What is the core problem that has not yet been solved?
As far as I understand, the idea in Regent is that the memory hierarchy is used dynamically, since the programmer cannot know about much of the memory hierarchy anyway at program writing time. So dynamic scheduling is used to construct the memory hierarchy and schedule at run-time. The typical use case for Regent is physicists writing e.g. weather / particle physics simulations, they are not aware of the details of L1/2/3/Memory / network/cluster ... sizes.
This is probably quite different from your use case.
Thanks. [1] contains an interesting discussion of Sequoia, and why it was considered a failed experiment, leading to Legion ][2], its successor.
Thanks.
How to abstract "hardware primitives" in a way that can be instantiated to many GPU architectures without performance penalty and be useful for higher-level programming is not so clear. How would you, to take an example from the CPU world, fruitfully abstract the CPU's memory model? As far as I'm aware that's not a solved problem in April 2020, and write papers on this subject are still appearing in top conferences .