HN user

chewxy

5,014 karma

You can contact me here: chewxy [at] gmail dot com

My personal blog is http://blog.chewxy.com . Be warned. Lots of nonsense in there.

Posts172
Comments1,103
View on HN
blog.djnavarro.net 1y ago

When good pseudorandom numbers go bad

chewxy
71pts38
arxiv.org 2y ago

Chinchilla Scaling: A Replication Attempt

chewxy
2pts0
arxiv.org 3y ago

Are emergent abilities of large language models a mirage?

chewxy
154pts130
blog.chewxy.com 4y ago

Quickly Annotate Your Machine Learning Dataset with One Weird Trick (It's Lisp)

chewxy
2pts0
whoo.ps 5y ago

You merely adopted dark mode

chewxy
3pts0
dgraph.io 5y ago

Manual Memory Management in Go using jemalloc

chewxy
50pts46
dgraph.io 5y ago

Manual Memory Management in Go using jemalloc

chewxy
38pts6
www.atlasobscura.com 5y ago

The Mesmerizing Geometry of Malaysia's Most Complex Cakes

chewxy
1pts0
commandcenter.blogspot.com 6y ago

Computational Reproducibility: Some Challenges

chewxy
7pts1
www.ams.org 7y ago

The Grothendieck I Knew: Telling, Not Hiding, Not Judging [pdf]

chewxy
64pts10
www.youtube.com 7y ago

Biology as Information Dynamics (2017)

chewxy
2pts0
www.infoq.com 7y ago

The Why of Go

chewxy
1pts1
incredible.pm 7y ago

The Incredible Proof Machine (2016)

chewxy
79pts16
news.ycombinator.com 7y ago

Ask HN: Good books on computation

chewxy
5pts3
www.fhi.ox.ac.uk 7y ago

AI Governance: A Research Agenda [pdf]

chewxy
3pts0
quillette.com 8y ago

Political Moderates Are Lying

chewxy
1pts0
probablydance.com 8y ago

A new fast hash table in response to Google’s new fast hash table

chewxy
390pts115
aeon.co 8y ago

You Don't Have a Right to Believe Whatever You Want To

chewxy
2pts1
programmersatwork.wordpress.com 8y ago

Programmers at Work: Bill Gates (1986) – 10 year anniversary of republication

chewxy
2pts0
medium.com 8y ago

Psychology of Code Readability

chewxy
2pts0
www.rand.org 8y ago

An AI in Our Image: The Risks of Bias and Errors in Artificial Intelligence

chewxy
2pts0
medium.com 8y ago

Effects of CPU Caches

chewxy
99pts23
github.com 8y ago

CSPFuck: Brainfuck, with Go-style channels

chewxy
4pts0
aeon.co 8y ago

In the 1950s everyone cool was a little alienated. What changed?

chewxy
1pts0
lukeoakdenrayner.wordpress.com 8y ago

The Philosophical Argument for Using ROC Curve

chewxy
4pts0
news.ycombinator.com 8y ago

Ask HN: Best Outsourcing Practices

chewxy
2pts1
medium.com 8y ago

Amazon Neptune isn't that great - discussions regarding vertical scaling

chewxy
6pts3
www.youtube.com 8y ago

Rob Pike's Fantastic Intro to Upspin

chewxy
19pts5
www.cse.chalmers.se 8y ago

Parsing mixfix operators [pdf]

chewxy
1pts0
www.michaelburge.us 8y ago

Injecting a Chess Engine into Amazon Redshift

chewxy
91pts8

So I have been practicing writing fiction the past year or so. It identifies a fiction piece I wrote as Greg Egan[0]. Another paragraph from another piece was identified as China Mieville[1]. The accompanying blog posts explaining the making of the fiction pieces were identified as me.

Both pieces have never been published. Neither have the blog posts.

[0] in https://blog.chewxy.com/2026/04/01/how-i-write/ this is the story titled "there is no constant non-zero derivative in nature". It does not read like Egan at all.

[1] in https://blog.chewxy.com/2026/04/01/how-i-write/ this is the story titled "The Case of the Liquidated Corps". I use a lot of biological metaphors. Once again, nothing like Mieville.

If only I could write like them! These pieces were all rejected by the major scifi mags

Maybe read the paper first?

This study asked whether Large Language Models (LLMs) understand sentences in the minimal sense of representing “who did what to whom”. In Experiment 1, we found that the overall geometry of LLM distributed activity patterns failed to capture this information: similaritiesbetween sentences reflected whether they shared syntax more than whether they shared thematic role assignments. Human judgments, in contrast, were strongly driven by this aspect of meaning.

In Experiment 2, we found limited evidence that thematic role information was available even in a subset of hidden units. Whereas activity patterns in subsets of hidden units often allowed for significant classification of whether sentence pairs had shared vs. opposite thematic role assignments, the effect sizes were small; even the best-performing case appeared to lag behind humans, and its representation of thematic roles did not seem robust across syntactic structures.

However, thematic role information was reliably available in a large number of attention heads, demonstrating LLMs have the capacity to extract thematic role information. In some cases, information present in attention heads descriptively exceeded human performance.

Tree Calculus 2 years ago

Barry Jay's got an upcoming paper at PEPM regarding typed tree calculus. Good read too.

I'm working on my scifi novel. I had started writing it when LLMs started taking off - I had been doing AI for two decades and I was well-placed to be in a good position to profit with the rise of LLMs, but I ended up gaining nothing much and I was depressed about it - so I started writing instead. Been picking at it for about a year before befriending an editor who encouraged me to keep writing. He's helped me developmentally edit it to a point I am now ready to work on my second draft.

It's a hard scifi novel with mild existential horror tones that is borne mostly of maths jokes. At one point the main character tries to escape the matrix (reality). But the matrix is defective, so the best way out was to orthogonalize the subspace and reduce the matrix to its eigenbasis instead. Most of the scenes are based on similar maths jokes.

Tentative name is Diagonalization of the Meta (I had previously called it The Metaverse).

I told a variant of the original Little Mermaid story as part of a school outreach program. The kids came to the conclusion that God wasn't a fair being because he didn't give mermaids souls. I walked away satisfied that my little counterprogramming against catholic school indoctrination might have worked. I wasn't invited back (at least for school year 2024).

Most LLMs out there are very American in their writing mannerisms. The article even alludes to this:

The AI did, however, try to sound like someone. It was folksy and upbeat, talky and pretend-excited

I've not seen an LLM, even when fine tuned that doesn't actually do that (Chinese LLMs excepted). There's something just inherently American about the instruction datasets that these LLMs are instructed with.

I've been running GolangSyd for about 8 years now, and previously Sydney Python for about the same amount of time. I found finding sponsors to be one of the hardest things. One of the last Sydney Pythons in which I gave a talk on machine learning had ~250 people attending (PyConAU at that time had roughly as many people I think). So large that the pizza bill was way more than the host (who was the sponsor as well) had anticipated, so we were no longer welcome at that venue. And said host is a well known billion dollar Aussie company.

Thankfully for GolangSyd, my cohosts have been extremely talented with finding sponsors. This gives us a lot more opportunity to do weirder things like https://gogogogogo.casa .

Running meetups are hard work, full stop. On the other hand, I've gotten to know some people very well and some of my best collabs have been thru these meetups. Running a meetup was a way for me to overcome my own reluctance to socialize.

A good thing about living in a small house. I can hear when my rice cooker is done (plays a tune), when my washing machine is done (plays a tune), when my dryer is done (plays a tune). :) (note that that hasn't prevented me from hooking them up to smart plugs the way terrence did)

Go (the game). I wrote a clone of AlphaGo in Go (the programming language) 8 years ago. Along the way I learned to play Go.

I've been using a combination of my own AI, LeelaZero and KataGo to teach myself in Go. For 8 years I've languished at the same level of play. Then I met a real human teacher who taught me Go at the end of 2023. And since then my game improved. I beat my own AI (which was intentionally trained to be impoverished in skill) for the first time in January.

Learning Go is teaching me all sorts of new ideas in pedagogy and putting a dampener in any enthusiasm that involves LLMs in education.

I agree. There are other types of AIs with different applications that do not need to be trained on the internet. The examples you have given however, are examples where the deep nets are extremely data hungry.

Take computer vision for example - a "hello world" version of object recognition would use ImageNet, which is 14 million hand annotated images. Or Cifar10 which is 80 million images. That of course but sets the stage for training data differentiation. Google's image recognition algorithm is far superior to other search engines'. Why? Because of Google's data set.

Any Tom Dick and Harry can go create their own image recognition AI and train it based on all the public datasets (COCO, CIFAR, ImageNet) but that's considered pretty baseline nowadays. The differentiator is what _other_ datasets you have.

Different datasets yield different results. It doesn't matter the network. More data is better (usually).

It was good enough in 2015/2016 for me to run a startup that allowed people to program in natural language. We even had paying clients though eventually none could stomach the $2000 per month for incremental/on-line training costs.

The only real difference between then and now is that OpenAI's models are significantly better than my models from 2015, and they have that because well, they can afford to pile on more data. TBH, I never even considered using a large proportion of the whole internet as a training set as even remotely possible due to the sheer mind boggling costs.

Even now, to go through about 10% of The Pile would cost me way too much money.

GP mentioned that the current slate of transformer based AIs are not transformative in the same way the Internet was. Rather it's more of a triumph of data engineering practices.

OP disagrees with GP. OP's main thesis is that AI enables a lot new applications. OP claims that GP is simply looking at it as if it were training data.

I stated that current AI techniques ARE indeed just reflections of the data used in training. I agree with GP that the current "AI"s are simply not transformative in the same way the Internet was.

If you change the training data for the current generation of AI, you get different behaviours. The training data forms a manifold - which you can think of as a landscape with features forming valleys and hills. What the current generation of AI does is that it tries to find a shape that fits the landscape - think of it like taking a very large sheet of cloth to cover a landscape. The stiffer the cloth, the less well the cloth fits to the landscape. The "stiffness" of the cloth is the amount of parameters that a neural network has. Modern deep nets are highly overparameterized - imagine a very soft pliable cloth - of course it fits to a landscape well.

So if you have a different training data - the neural network will fit to this different landscape as well. Hence the response will be different.

It's unfortunate that the training data is the entire internet for a few reasons:

1. Only the rich can train a vaguely competent AI. You're at the whims of those well-resourced enough. 2. There's no "alternate" training dataset anymore. (Though a clever thing people at OpenAI are doing are Mixture of Experts models, where you train multiple NNs using different subsets of the full training set, so you get multiple competencies)

I talked to Matt about not owning our own data after GopherConSG where he gave this talk. It was enlightening how complicated the issue is - there's a lot of legal liabilities on the end of the data provider (the company that monitors the glucose) so I can understand why larger corps are a bit hesitant to open up.

On the other hand, it seems quite heinous that users don't have access to data that is rightfully theirs that they can action on it.