HN user

fastball

12,598 karma

I used to argue on the internet too much.

Engineering at Known (https://known.com)

Co-founder and tinkerer at Supernotes (https://supernotes.app)

Posts16
Comments6,589
View on HN
supernotes.app 2y ago

Are We Losing Our Minds to Artificial Intelligence?

fastball
10pts1
twitter.com 2y ago

DearMoon, the first private circum-lunar flight project, will be cancelled

fastball
27pts15
twitter.com 2y ago

Semrush purchases website, rewrites reviews to praise themselves

fastball
2pts0
supernotes.app 2y ago

Show HN: Supernotes 3 – Offline-Ready, Collaborative and Cross-Platform Notes

fastball
9pts5
supernotes.app 3y ago

Should notes be end-to-end encrypted?

fastball
140pts146
supernotes.app 4y ago

Show HN: Supernotes 2 – a fast, Markdown notes app for journalling and sharing

fastball
251pts164
news.ycombinator.com 5y ago

Ask HN: Google support is the worst, good alternatives to G-Suite?

fastball
7pts6
republic.co 5y ago

Gumroad raises $5M from Crowd SAFE in first 24 hours

fastball
99pts70
investors.virgingalactic.com 5y ago

Virgin Galactic Raises $460M in Public Offering

fastball
4pts0
supernotes.app 6y ago

Show HN: Supernotes – The Social Notework for collaborative knowledge management

fastball
13pts4
supernotes.app 6y ago

Show HN: Supernotes – a better way to collect your thoughts

fastball
22pts5
www.thelostogle.com 6y ago

OKC-based company wants to keep employees’ $1,200 stimulus payments

fastball
44pts20
supernotes.app 6y ago

Productivity apps for remote teams and startups

fastball
2pts0
www.bloomberg.com 8y ago

Trump Expels 60 Russian Diplomats for U.K. Attack

fastball
3pts0
www.digitaltrends.com 10y ago

OLO – the $99 box that transforms any smartphone into a 3D printer

fastball
69pts45
www.fastcodesign.com 12y ago

I Tasted BBQ Sauce Made By IBM's Watson, And Loved It

fastball
19pts1

The most interesting result for me is that they apparently prompted the models to optimize for SSIM, but many of the models trend worse over time. I suppose because viewing the canvas always comes after drawing, and they didn't give "revert to previous" capability as part of the toolkit.

Which in turn kinda jives with my experience of using these models for code: to some extent they only seem to have a concept of "forward", which invariably leads to "write more code to fix previous problems created", rather than taking a step back and removing broken things entirely.

Yes, this (imo) is a clear result of benchmaxxing. You can get a much better score on most "intelligence" benchmarks by massively over-saturating reasoning. This looks good on those, but for actual daily usage makes the models much less effective: I don't want a model I use for coding to burn a bunch of reasoning (read: time) on trivial tasks.

In my experience, the Chinese models are much more benchmaxxed than their frontier lab competitors, so I'm taking these results with a fairly large helping of salt.

But the scaling needs to be relevant to the actual actions you are concerned about.

It's a stretch that any software Google/DeepMind/etc is selling to DHS is allowing / helping them to scale the murder part of their operations.

In fact, usually software translates to "less boots on the ground" which one could then assume would decrease the number of encounters like those highlighted in the article.

GPT-5.6 13 days ago

But that is my point: if benchmaxxing was all the labs were doing, then surely the dumber model could/would have equivalent performance? Rather than noticeably worse perf on a (somewhat trivial to game) test.

GPT-5.6 13 days ago

On the one hand: yes, pelicans on bikes are definitely in the training set at this point.

On the other hand: the test is clearly not saturated, given that you can see a clear difference in output at the various reasoning levels / model versions.

"The Internet" was not a bubble. Companies with no long-term business model / sufficient product-market fit that were riding hype were the "dotcom bubble". But when those companies crashed, nobody said "I really want to get my hands on their IP", because it wasn't valuable – an important pre-requisite to the the bubble popping. Seems to be a different case here if people actually want the SOTA models.

I explicitly said it is your right to operate that way. But that doesn't mean your unproven accusations ("the company is evil") are true. It just means that is how you are choosing to operate / that is the standard of evidence you require for your positions. Many people (myself included) disagree with such a low bar: innocent until proven guilty (and not guilt by association) and all that.

You actually need to demonstrate that though. I have seen no evidence of Mullvad (again, as a company) behaving in a racist or anti-immigrant manner. Until that has been demonstrated, you cannot just say "this guy behaves this way so that is what his company does".

Regardless you can always say "I don't want to give this company money because that indirectly is gonna funnel money to an anti-immigrant political party in Sweden", and that is a perfectly valid position to take. But a lot of people in this thread are clearly going a step further than that, ostensibly in an attempt to give them a greater sense of moral superiority than is necessarily deserved.

Two things:

1. People definitely start companies with a certain set of values and behaviors (as a company) and do entirely separate things in their private life. This is trivially true.

2. I don't think the personal values and the business mission in this case are even in conflict. You can be a racist and support free speech / privacy. Indeed, I'd actually say the venn diagram of "racists" and "people who vociferously espouse free speech ideals" is more union than it is disjoint.