HN user

thomasahle

6,801 karma

https://thomasahle.com

Posts48
Comments2,140
View on HN
danluu.com 3mo ago

DeWitt Clauses

thomasahle
3pts1
juleflet.dk 7mo ago

Show HN: Make your own Danish Julehjerter (Braided hearts)

thomasahle
3pts0
news.ycombinator.com 8mo ago

Show HN: Trace.taxi – easy agent messages visualization

thomasahle
3pts1
arxiv.org 1y ago

Breaking the Sorting Barrier for Directed Single-Source Shortest Paths

thomasahle
2pts0
old.reddit.com 1y ago

Entire Subreddits Full of Bots

thomasahle
3pts4
tensorcookbook.com 1y ago

The Tensor Cookbook

thomasahle
2pts1
www.cognition.ai 1y ago

A review of OpenAI o1 and how we evaluate coding agents

thomasahle
34pts2
tensorcookbook.com 1y ago

The Tensor Cookbook

thomasahle
3pts0
blog.normalcomputing.ai 2y ago

Supersizing Transformers: Going Beyond Rag with Extended Minds for LLMs

thomasahle
7pts0
github.com 2y ago

Better Gzip Language Model from Beam Search

thomasahle
5pts1
github.com 2y ago

ZipLm with Beam Search, BPE and Progressive Compression

thomasahle
6pts1
github.com 3y ago

New Ann Benchmarks Run

thomasahle
2pts0
dudo.ai 3y ago

Show HN: Liar's Dice AI

thomasahle
2pts2
11011110.github.io 4y ago

Maybe powers of π don't have unexpectedly good approximations?

thomasahle
75pts36
www.youtube.com 4y ago

Unveiling Transformers with Lego

thomasahle
2pts1
medium.com 4y ago

Lair’s Dice by Self-Play

thomasahle
7pts3
dudo.ai 4y ago

Show HN: Liar's Dice AI from reinforcement learning

thomasahle
3pts5
dudo.ai 4y ago

Show HN: Simple Liar's Dice AI Using CFR

thomasahle
1pts1
twitter.com 5y ago

You Can't Sample from a Dictionary in Python

thomasahle
2pts1
d2l.ai 5y ago

Dive into Deep Learning (Updated 2021)

thomasahle
4pts1
www.eff.org 5y ago

How We Saved Dot Org

thomasahle
1187pts151
contemporary-home-computing.org 5y ago

Prof. Dr. Style

thomasahle
1pts0
thomasahle.com 6y ago

Netflix Recommendations and the Coronavirus Outbreak

thomasahle
7pts0
www.youtube.com 6y ago

AlphaStar Final vs. Serral

thomasahle
2pts0
arxiv.org 7y ago

Provably Robust Deep Learning

thomasahle
2pts0
www.newscientist.com 7y ago

Famed mathematician claims proof of 160-year-old Riemann hypothesis

thomasahle
3pts1
docs.google.com 8y ago

Leela Chess Zero Progress

thomasahle
3pts0
github.com 8y ago

Show HN: Codenames AI Using Word Embeddings

thomasahle
5pts1
web.stanford.edu 8y ago

A Simple Alpha(Go) Zero Tutorial

thomasahle
3pts0
drops.dagstuhl.de 8y ago

Theory and Applications of Hashing [pdf]

thomasahle
3pts0
Kimi Work 2 days ago

duping is a great way to show people that even the app layer can be commoditized

Everyone is already cloning the app layer. There are 30+ claude-code clone, and codex work already cloned cowork, grok and gemini is doing the same.

Mistral OCR 4 29 days ago

I used to part time for the (Danish) mail service. The only sorting that was done automatically was the post codes. That was enough to get the letter to the right post office. The rest was done by the mailmen/women early in the morning. It was a lot of fun trying to figure out what was meant by some of the addresses. The older people in particular often knew the story of why certain places were sometimes addressed in certain ways, or could guess the addresses based on the names of the people living there.

After reading the article, the main "case against geometric algebra" I could find in there was that the author does not like the people using/doing research in geometric algebra

Mathematics is a social activity. The research cultures of different branches matter.

Ban it from the dataset, add it to the analysis. You can choose your own flavor of noise.

Not sure exactly what you're proposing, but if the noise is added independently to different people, you can just buy multiple copies to reduce it.

There are a lot of ways to do this wrong, which is why so much analysis has gone into differential privacy.

I recently built a very large test bench for System Verilog.

I ran a bunch of different compilers on it, including some open source ones.

Some of them failed some tests, and it was natural to have my LLM (Claude Fable 5) root-cause the issues, and to double-check my test bench wasn't to blame.

But now I stood with all these patches that I couldn't just throw at the upstream maintainers all at once. I ended up just filing a few issues and moved on to other things.

It felt weird to just file issues when my LLM had already spent a lot of time root-causing and fixing the issues. But then, maybe they could just have their LLMs do the same.

Still not sure if it was the right call?

The rate of fundamental, broad-based breakthroughs lifting all LLM applications has clearly slowed with many of the most impactful recent discoveries being in scaling, optimization, tuning and productization toward specific domains.

To me it definitely feels like it's still accelerating, with the most impactful recent discovery being RL training reasoning models (late '24, early '25).

There's an interesting article called "sigmoids won't save you" https://www.astralcodexten.com/p/the-sigmoids-wont-save-you which argues that (unless you have privileged information) you should always assume a process will continue about as long as it’s continued already. (Lindy's Law)

With that in mind the current disruption should last another 10-15 years (assuming it started in '10 or '17.)

I'm currently choosing between the right formalization for a big hardware project.

I'm considering between SVA, TLA+ and Lean. With the former being more domain specific and the later more general.

Do you think we'll move towards "Lean for everything" or do domain specific formalisms still make sense?

The human savant will remember where they read it and give you credit. It might lead more people to read your work, and ultimately you make money.

The AI won't even know where the page of text it's seeing came from, and people will avoid your book as they can just ask the AI. So you make less money. (Talking about specialized technical books here.)

DeWitt Clauses 3 months ago

In 1983 David DeWitt (https://en.wikipedia.org/wiki/David_DeWitt) published benchmarking results showing poor performance for Oracle databases. Larry Ellison wasn't happy with the results and it's said that he tried to have DeWitt fired.

Given how difficult it is to fire professors when there's actual misconduct, the probability of Ellison sucessfully getting someone fired for doing legitimate research in their field was pretty much zero. It's also said that, after DeWitt's non-firing,

Larry banned Oracle from hiring Wisconsin grads and Oracle added a term to their EULA forbidding the publication of benchmarks. Over the years, many major commercial database vendors added a license clause that made benchmarking their database illegal.

See also: https://web.archive.org/web/20160719145221/http://sqlmag.com...

This is crazy car-centric legislation.

Now, instead of letting car owners pay for the public space they use (street parking), you are forcing anyone without a car to waste their own private space, in case somebody wants to park there.

We scaled on "virtually all RL tasks and environments we could conceive." - apparently, they didn't conceive of pelican SVG RL.

I've long thought multi-modal LLMs should be strong enough to do RL for TikZ and SVG generation. Maybe Google is doing it.