HN user

RC_ITR

3,160 karma

Not concerned with political “beliefs” other than “everyone needs to do more critical thinking and research.”

Clustering is usually a pretty good sign of doing neither - just because information is free doesn’t mean you shouldn’t synthesize it for yourself.

Posts0
Comments1,681
View on HN
No posts found.

And no, HN is not social media in any normal sense of the word. The pedantry involved in that comparison is extremely tiresome.

The amount of times I've read a very thoughtful article only for the comments to be political drivel (the worst was peak-COVID SF discourse) weakens your argument quite a bit.

It's even more foolish to think outside forces aren't using bots/tech to sway the discourse.

Just because it's not engineered for the mainstream's dopamine addiction doesn't mean it doesn't do the same thing.

There is only one market that large: the global labor market.

This isn't even close to true and it's kind of the central thesis of this article.

Saudi Aramco has consistently been a $2tn company in the oil market.

Walmart is a $1tn-ish company focusing on a fraction of US retail.

It also ignores the idea that the economy is not zero sum and companies create their own market/economic value all the time.

You and OP are both unnecessarily diminishing what 'glorified search' is.

If you had told me that in 2015, we would have a tool that can iteratively search the world's best and largest unstructured database and synthesize outputs in language (any natural and structured language), I would have said that is basically AGI.

This whole desire for it to 'reason' (autonomously prime its search with a few thousand token) and 'think' (search for the best information within its parameters and synthesize that with its context) is semantic and will feel irrelevant as the technology progresses and we become more used to what these things are actually doing.

I honestly struggle to imagine what AGI will be if not an ever-improving semi-structured database (parametric or otherwise) that we become increasingly good at searching.

We invented a word for a very specific thing (consciousness) and are now debating whether that relatively unimportant word represents a large open set or a narrow closed set.

We do one thing in our bodies with relatively binary nervous system and a fundamentally continuous endocrine system. That's clearly and unanimously consciousness. We also, however, see other animals with similar set-ups but less capabilities, so we understand it exists on a spectrum.

We separately invented a thing that gets to similar outcomes with fundamentally binary logic gates.

Our minds are drawn to comparison and classification, so we fight over how similar or different those two things are in a way that often feels unsatisfactory because in order to meaningfully compare the two, we have to reduce them in a way that feels like its underselling either/both.

It's like that FT chart claiming that the rapid rise in iOS apps is evidence of an AI-fueled productivity boom.

I always ask people, in the past year, how many AI-coded apps have you 1) downloaded 2) paid for?

Here's the score for new AIME's, where we know the answers aren't in training.

https://matharena.ai/?view=problem&comp=aime--aime_2026

As for MMLU, is your assertion that these AI labs are not correcting for errors in these exams and then self-reporting scores less than 100%?

As implied by the video, wouldn't it then take 1 intern a week max to fix those errors and allow any AI lab to become the first to consistently 100% the MMLU? I can guarantee Moonshot, DeepSeek, or Alibaba would be all over the opportunity to do just that if it were a real problem.

The bird not having wings, but all of us calling it a 'solid bird' is one of the most telling examples of the AI expectations gap yet. We even see its own reasoning say it needs 'webbed feet' which are nowhere to be found in the image.

This pattern of considering 90% accuracy (like the level we've seemingly we've stalled out on for the MMLU and AIME) to be 'solved' is really concerning for me.

AGI has to be 100% right 100% of the time to be AGI and we aren't being tough enough on these systems in our evaluations. We're moving on to new and impressive tasks toward some imagined AGI goal without even trying to find out if we can make true Artificial Niche Intelligence.

Yeah, I've found AI 'miracle' use-cases like these are most obvious for wealthy people who stopped doing things for themselves at some point.

Typing 'Find me reservations at X restaurant' and getting unformatted text back is way worse than just going to OpenTable and seeing a UI that has been honed for decades.

If your old process was texting a human to do the same thing, I can see how Clawdbot seems like a revolution though.

Same goes for executives who vibecode in-house CRM/ERP/etc. tools.

We all learned the lesson that mass-market IT tools almost always outperform in-house, even with strong in-house development teams, but now that the executive is 'the creator,' there's significantly less scrutiny on things like compatibility and security.

There's plenty real about AI, particularly as it relates to coding and information retrieval, but I'm yet to see an agent actually do something that even remotely feels like the result of deep and savvy reasoning (the precursor to AGI) - including all the examples in this post.

One of the biggest problems frontier models will face going forward is how many tasks require expertise that cannot be achieved through Internet-scale pre-training.

Any reasonably informed person realizes that most AI start-ups looking to solve this are not trying to create their own pre-trained models from scratch (they will almost always lose to the hyperscale models).

A pragmatic person realizes that they're not fine-tuning/RL'ing existing models (that path has many technical dead ends).

So, a reasonably informed and pragmatic VC looks at the landscape, realizes they can't just put all their money into the hyperscale models (LP's don t want that) and they look for start-ups that take existing hyperscale models and expose them to data that wasn't in their pre-Training set, hopefully in a way that's useful to some users somewhere.

To a certain extent, this study is like saying that Internet start-ups in the 90's relied on HTML and weren't building their own custom browsers.

I'm not saying that this current generation of start-ups will be successful as Amazon and Google, but I just don't know what the counterfactual scenario is.

Not all 'normal income' is from a "job" as we think of it and assuming that does not even come close to passing any informed person's smell test.

Parsing tax or SS payments for what a "job" is would be a logistical nightmare, because that's not what the system is designed for (unlike the BLS's system, which is designed to count jobs).

https://fred.stlouisfed.org/graph/?g=1Mc3z

Manufacturing and mining are becoming much less correlated to the overall jobs market (likely, as you point out, b/c the government smooths the other sectors).

https://fred.stlouisfed.org/graph/?g=1Mc3I

This is despite being a relatively flat % of employment since 2010 (after a long period of decline).

https://fred.stlouisfed.org/graph/?g=1Mc4f

As mentioned, there is also the weirdness of SWE's going from 'better than the overall market' to 'worse than the overall market'.

https://fred.stlouisfed.org/graph/?g=1Mcer

Retail employment is also dislocating.

Those are just the examples I can think of with no research, I'm sure there are others.