What is not well done in this study?
HN user
dimask
1) Then more math should get formalised in lean.
2) How is a solution by LLMs supposed to be verified without such a formalisation?
The spatial reasoning on reading code does not happen on the dimensions of the literal text, at least not only on these. It happens in how we interpret the code and build relations in our minds while doing so. So I think that the problem is not about the spatial reasoning of what we literally see per se, but if the specific representation helps in something. I like visual representations for the explanatory value they can offer, but if one tries to work rigorously on a kind of spatial algebra of these, then this explanatory power can be lost after some point of complexity. I guess there may be contexts where a visual language may be working well. But in the contexts I have encountered I have not found them helpful. If anything, the more complex a problem is, the more cluttered the visual language form ends up being, and feels overloading my visual memory. I do not think it is a geometric feature or advantage per se, but about how brains of some people work. I like visual representations and I am in general a quite visual thinker, but I do not want to see all these miniscule details in there, I want to them to represent what I want to understand. Text, on the other hand, serves better as a form of (human-related) compression of information, imo, which makes it better for working on these details there.
I do not think what they say is that it is hard to visualise it, but that it does not offer much utility to do so. A "for" loop like that is not that complicated to understand and visualising it externally does not offer much. The examples the article gives is about more abstract and general overviews of higher level aspects of a codebase or system. Or to explain some concept that may be less intuitive or complicated. In general less about trying to be formal and rigorous, and more about being explanatory and auxiliary to the code itself.
At least they do not name them after themselves.
Well I precisely talked about things I have engaged professionally. Obviously this cannot cover everything one may do, eg I do not build chatbots for customer service or stuff like that, thus I obviously cannot speak for all possible applications of LLMs and how useful they may be. I am pretty sure there will be useful applications in fields I am not and will not be engaged in as nobody engages with everything. However, some other things that I have tried (eg copilots, summarising scientific articles) imo create much more hype than real value. They can be a bit useful if you know what to actually use them for and what their limits are, but nowhere close to the hype they generate, and I just find myself just googling again tbh. They are absolutely horrible especially with more niche subjects and areas. On the other hand, data extraction and structuring has a quite universal application, has already demonstrated usefulness and potential, and seems a quite realistic, down to earth application that I am happy to see other people and startups working on. Not as fancy, and harder to build hype upon, but very useful regardless.
Thanks for putting all this work and sharing it in such a detail! Data extraction/structuring data is the only serious application of LLMs I have actually engaged in for real work and found useful. I had to extract data from experience sampling reports which I could not share online, thus chatgpt etc was out of question. There were sentences describing onsets and offsets of events and descriptions of what went on. I ran models through llama.cpp to turn these into csv format with 4 columns (onset, offset, description, plus one for whether a specific condition was met in that event or not which had to interpreted through the description). Giving some examples of how I want it all structured in the prompt, was enough for many different models to do it right. Mixtral 8x7b was my favourite because it ran the fastest in that quality level on my laptop.
I am pretty sure that a finetuned smaller model would be better and faster for this task. It would be great to start finetuning and sharing such smaller models: they do not really have to be really better than commercial LLMs that run online, as long as they are not at least worse. They are already much faster and cheaper, which is a big advantage for this purpose. There is already need for these tasks to be offline when one cannot share the data with openai and the like. Higher speed and lower cost also allow for more experimentation with more specific finetuning and prompts, with less care about token lengths of prompts and cost. This is an application where smaller, locally run, finetunable models can shine.
I do not think that such a conversation is done in productive way (especially when cutting phrases in half to make them appear making no sense) but will try get a couple of points across:
- I interpreted the comment in the context of the answer to a specific article/interview. In the linked content, for example, there is a video of a 2.5yo doing a "flashcard class". While I do not think there is anything inherently harmful or anything, it is not a way that 2.5yos learn about the world, and even if it is not harmful it is not needed for 2.5yos to sit on a chair and doing a class to learn about the world. Their curiosity and own exploration drive is enough for pulling them into learning, and this is what I mean by parents should feed, ie see what their kids are most curious and interested in and feeding them inputs to that direction. The comment you answered to was referring to this article, and I interpreted your answer in that context. If I misinterpreted anything, I can only see the context that is shared here, not in your mind.
- To reiterate and clarify more on the context, "Speaking and reading to children is a natural activity" is _not_ what OP was about. What OP was about is applying a specific strategy for kids at 2+ years, ie to learn to read using a specific exploitation-based approach. If that is all you meant by your previous comment, then you may want to reread the comment you answered to from that perspective. Nobody here is saying "leave the kids do what they want and do not care about interacting with them much/talking around them" that you seem to suggest. When I say parents building upon kids' own curiosity and exploration drive I mean seeing what sort of inputs their kids become more curious and interested in at a certain time and feeding them inputs like that. When a kid starts being interested in sounds and music, feed them with sounds and music and sound/music-related books and toys. There is no handbook that is gonna say which month and day exactly this should happen for a specific kid.
- I may miss a lot of knowledge indeed, but I still find setting goals of "maximising language exposure" and "maximising IQ" weird and unclear. No, I have never read or heard this way of approaching development and learning. Parents doing their best and being mindful of the importance of language exposure is different than "maximising" anything. Maximising with respect to which parameters? Even defining this as an optimisation problem, any complex optimisation problem like this is a tradeoff between different parameters and outcomes. What happens to the other parameters and outcomes when you optimise on just one?
- If "you do not speak the jargon" is what you prefer to focus, just say that and any more discussion will not be needed.
I think you are confusing "memory" with strategies based on memorisation. Yes memorising (ie putting things into memory) is always involved in learning in some way, but that is too general and not what is discussed here. "Compression is understanding" possibly to some extent, but understanding is not just compression; that would be a reduction of what understanding really is, as it involves a certain range of processes and contexts in which the understanding is actually enacted rather than purely "memorised" or applied, and that is fundamentally relational. It is so relational that it can even go deeply down to how motor skills are acquired or spatial relationships understood. It is no surprise that tasks like mental rotation correlates well with mathematical skills.
Current research in early mathematical education now focuses on teaching certain spatial skills to very young kids rather than (just) numbers. Mathematics is about understanding of relationships, and that is not a detached kind of understanding that we can make into an algorithm, but deeply invested and relational between the "subject" and the "object" of understanding. Taking the subject and all the relations with the world out of the context of learning processes is absurd, because that is in the exact centre of them.
I work in human developmental research and have never heard or read anybody make such claims that you consider "bog standard child development science", and some of what you say are definitely not supported by the current understanding of human development.
For example
child language development milestones that are waymarked by age down to the month
is totally false. It is quite known that developmental milestones are acquired by children in different times and even in different orders and sequences. This "down to the month" is pure non-sense for most of the milestones.
Young children are better served to be guided by their own curiosity, interest and exploration drives and which parents feed with variable inputs and building upon, rather than by anxious parents feeding them with whatever terabytes of exploitation-intended information they think is gonna "serve to maximize IQ".
Yes, reading to kids in certain ways (using numbers/spatial relationships/theory of mind stuff/interactively) has been found in some studies to correlate with certain outcomes but there is nothing to suggest a totally linear relationship such that talking to a kid 24/7 since the womb is gonna produce the next Einstein.
It is not "just more text". That is an extremely reductive approach on human cognition and experience that does favour to nothing. Describing things in text collapses too many dimensions. Human cognition is multimodal. Humans are not computational machines, we are attuned and in constant allostatic relationship with the changing world around us.
How many homework questions did your entire calc 1 class have? I'm guessing less than 100 and (hopefully) you successfully learned differential calculus.
Not just that: people learn mathematics mainly by _thinking over and solving problems_, not by memorising solutions to problems. During my mathematics education I had to practice solving a lot of problems dissimilar what I had seen before. Even in the theory part, a lot of it was actually about filling in details in proofs and arguments, and reformulating challenging steps (by words or drawings). My notes on top of a mathematical textbook are much more than the text itself.
People think that knowledge lies in the texts themselves; it does not, it lies in what these texts relate to and the processes that they are part of, a lot of which are out in the real world and in our interactions. The original article is spot on that there is no AGI pathway in the current research direction. But there are huge incentives for ignoring this.
Claims of isomorphisms are really strong claims to not be backed up with some kind of evidence.
Would an intelligent but blind human be able to solve these problems?
Blind people can have spatial reasoning just fine. Visual =/= spatial [0]. Now, one would have to adapt the colour-based tasks to something that would be more meaningful for a blind person, I guess.
When they can outperform human infants in learning, eg data required to learn and versatility, we can talk business.
Not all world is "big data".
The problem is that there is little, incremental progress last 1 year or so after the big chatGPT boom to justify the hype, technically wise. Most of the "progress" going on is basically marketing, and making the models respond in ways humans like, or being more useful in certain practical applications. The basic, fundamental issues/limitations remain unanswered and unaddressed. As products, they have improved a lot and most probably are gonna improve more. But if we are talking for going towards AGI or more complex applications, I do not see evidence on that except as toys.
That's not where they belong.
Where do they belong? When we set forever chemical loose in the environment, it is expected that a quantity of them will reach the ones in the top of the food chain, which humans are. Where are forever chemicals supposed to end up when we decide it is less costly economic-wise to use them?
You do not teach git/github, you teach students what the best practices (ie version control) are as part of working on their projects.
To be fair people in the private sector typically make more money at least and have better work-related perks. Academia advertises mostly 1. you research stuff you find interesting and 2. you will get job security some point after a couple of postdocs. Currently neither 1 nor 2 apply for the vast majority of academic work.
oops
It was a disaster waiting to happen. I had entirely disabled automatic updates for my iphone apps specifically this kind of risk with this specific app. An OTP app is sold to some totally shady guy, what can go wrong. Though tbh I must say I did not expect extortion but rather I was more afraid of malware that would just steal the TOTPs and sell them in the dark web.
one that somehow synced the clipboard when copying an OTP code on your mobile app, so it was immediately available on the Mac.
Iphones and macs already do that (syncing their clipboards) as long as they are on the same network or sth. It usually works when they are on the same wifi, or if the iphone is connected to the mac through usb.
Only that the app was never really under open source license, at least not since 2019. It used to be under CC BY-NC 4.0 but then it changed to a "source available" type of status [0], where any modification, redistribution etc was prohibited. The previous author explained "Unfortunately I had to apply these restrictions in the license, as people started to redistribute my app to the appstore" [1]. I do not judge the fact of stopping open sourcing it, but he did not stop advertising it as "open source". Even now, that the source code is not even available any more, its website calls it "open source" [2].
[0] https://github.com/raivo-otp/ios-application/commit/03791edd...
You _can_ disable automatic updates and install updates manually. This is what I actually do, ironically for this very exact risk about the Raivo app. I like raivo, but since it was bought by this shady guy, I thought that by blocking automatic updates I could use the app without risk of some shady update coming. The only issue is that this is all or none; meaning that from time to time I have to go through my list of apps and manually tap on the upgrade button. So glad for that decision now.
Anybody who is concerned and wants to learn what personal info of them ticketmaster holds and hence may have been compromised, and lives in a country where GDPR applies, they can send them a GDPR subject access request, eg something like [0]. Their emails in general look like privacy@ticketmaster.XXX where XXX the country code but probably also double check to make sure.
[0] https://www.linkedin.com/pulse/nightmare-letter-subject-acce...
Template for writing GDPR subject access requests [0] for whoever it may concern. It is a bit too harsh, but it may be useful for one to know what info may have been compromised.
https://www.linkedin.com/pulse/nightmare-letter-subject-acce...
MAO-B and A inhibitors (they inhibit the enzymes that break down dopamine, serotonin and other amines) are not too uncommon to be used as nootropics in the nootropic communities, and selegiline has probably some more nootropic effects. It is also used for ADHD. While definitely indication of risky behaviour (disrupting body's homeostatic and allostatic mechanisms is quite risky, people can die when taking this together with foods high in tyrosine) I doubt he used it specifically to increase impulsivity.
It should be enough to be cheaper if you include the environmental cost that fossil fuels cause to society and planet.
Video tl;dr:
Psychtoolbox is a popular open source tool for creating vision and neuroscience experiments. As typical of open source projects, the project mostly runs by a single, underpaid developer's time (Mario Kleiner). They tried to be funded through "support membership" donations, it did not work, and now they announce that a paid license will be required for running their compiled binaries "on some operations systems".
The relevant part where they describe the funding situation of the project starts around 6:00 as in https://youtu.be/05gpkP_EMoc?t=363 (sorry for messing up the link in the post)
In a chain of thought manner, as every proper AI, of course.