HN user

ddp26

631 karma
Posts31
Comments146
View on HN
simonwillison.net 20d ago

Porting the Moebius 0.2B image model to run in Claude Code on web

ddp26
2pts0
futuresearch.ai 21d ago

The Wealth of the Richest People in AI

ddp26
4pts0
www.lesswrong.com 28d ago

World-Modeling the US vs. Anthropic on Claude Fable

ddp26
9pts1
futuresearch.ai 1mo ago

Conscripting engineers to make training data won't push AI

ddp26
1pts0
futuresearch.ai 1mo ago

How the US vs. Anthropic Standoff on Claude Fable Will End

ddp26
2pts1
futuresearch.ai 1mo ago

Claude can miss the motives of politicians

ddp26
10pts0
futuresearch.ai 1mo ago

Measuring one way AIs lack self-awareness

ddp26
1pts0
futuresearch.ai 1mo ago

Some rare examples of AIs being underconfident

ddp26
6pts0
futuresearch.ai 2mo ago

History doesn't repeat itself as often as LLMs think

ddp26
1pts0
futuresearch.ai 2mo ago

Agents Sometimes Catastrophize

ddp26
9pts2
futuresearch.ai 2mo ago

Run Agents Twice

ddp26
6pts0
futuresearch.ai 3mo ago

I think Anthropic is worth $100B more than last week

ddp26
9pts0
futuresearch.ai 3mo ago

A forecast of the fair market value of SpaceX's businesses

ddp26
100pts205
futuresearch.ai 4mo ago

Ask LLM Agents to Classify Problems Before Starting

ddp26
7pts1
everyrow.io 5mo ago

I ran 10,000 web research agents

ddp26
12pts0
futuresearch.ai 6mo ago

How LLM agents solve the table merging problem

ddp26
29pts3
futuresearch.ai 6mo ago

Forecasting the 2026 AI Winner

ddp26
2pts1
news.ycombinator.com 8mo ago

Show HN: Stockfisher –– our automated Warren Buffett

ddp26
16pts4
futuresearch.ai 9mo ago

The Karpathy Interview, 6 Months After AI 2027

ddp26
38pts24
futuresearch.ai 1y ago

Apple's plan to power Siri with ChatGPT was a predictable failure

ddp26
15pts1
futuresearch.ai 1y ago

OpenAI Deep Research – Six Strange Failures

ddp26
8pts0
github.com 2y ago

Biden signs TikTok ban – what will happen next?

ddp26
9pts0
news.ycombinator.com 2y ago

Show HN: FutureSearch – answer hard questions like "Who will buy TikTok US?"

ddp26
22pts3
www.metaculus.com 3y ago

The future of AI according to thousands of forecasters

ddp26
93pts81
www.talktothe.city 3y ago

Visualizing AI Discourse on Twitter

ddp26
1pts1
www.lesswrong.com 3y ago

Metaculus Introduces Conditional Probability Forecasting

ddp26
7pts0
twitter.com 3y ago

Metaculus hits 1M predictions and becomes a Public Benefit Corporation

ddp26
3pts3
forum.effectivealtruism.org 4y ago

Prediction Markets in the Corporate Setting

ddp26
1pts0
twitter.com 4y ago

10k Google employees participate in new prediction market

ddp26
6pts1
cloud.google.com 4y ago

Google's Second Prediction Market

ddp26
2pts0

Isn't this the same as saying "utility regulators delaying connecting new power to the grid hiked electricity prices on the public by $23B?"

When my apples are expensive, I don't generally grumble about all the demand from pie makers. If they demand more apples, new suppliers should come in to restore the price, right?

I can ask an agent to add OAuth, you can ask one to add caching, and somebody else can ask one to rebuild the database from first principles and make the UI pink. Each change can be reasonable in isolation.

But this is just bad vibecoding? This would be bad if humans did it too. With agents or humans, you need to coordinate.

When I know something is (primarily) AI generated, I lose interest.

The exception is when it's about a niche I care about, e.g. an analysis of opening trends of early world chess champions. I'll read AI on that for an hour.

My sense is that, for most writing, it's fundamentally interpersonal, the information is about the author as much as it is about the world.

Maybe this flood of slop will cause people to care more about the substance of the writing, not the perspectives of the writing.

GPT-5.6 12 days ago

Is it possible GPT-5.6 is not a very aligned model?

But Scott's point is more: why even have markets? Once you have the superforecasting available on the questions you care about, why do you need to publish it for everyone to also react to?

Almost by definition, once AI forecasters are in the market, they won't (all) be beating the market.

But why evaluate AI forecasters by beating the market? Do we evaluate deep learning by whether hedge funds make money from it in the markets? These things have far, far more utility outside of finance.

Doesn't this argument prove too much? Why does AlphaSense sell their company research instead of using it to trade themselves? Why do people work on open source time series forecasting packages instead of quietly using them to trade?

Podman v6.0.0 20 days ago

Yeah, a great developer I know showed me how he could use it to get a safe dev container for Claude Code, in a way that wasn't doable with Docker.

People forget that Meta already did this years ago, before prediction markets became the next big consumer trend for them to chase.

The app was called Forecast, and launched in June 2020. (Around the same time that Kalshi and Polymarket launched, actually!) It was framed as a way to make the comments and activity on Facebook actually productive rather than toxic, and build expert reputation signaling mechanisms.

I think they sunset it after about a year.

You mean chatgpt style AI won't help them with those skills?

If a human parent or teacher can help with skills like reading, an AI system can too, once it's trained and designed to do so. (How good are humans at teaching reading anyway?)

Claude Opus 4.8 2 months ago

It is refreshing but perhaps actually not warranted this time?

I mostly study web research, and Opus 4.7 was a regression on BrowseComp compared to Opus 4.6, which has been born out by my usage.

Opus 4.8 is now much better than either 4.7 or 4.6, and having it search the web is one of the primary use cases of chatbots.

I linked elsewhere in a comment, Metaculus has AGI forecasts.

You can also now use AI forecasters like FutureSearch [1] (disclaimer: I work there), which are competitive with the best humans / teams of humans. And since you aren't depending on a human crowd, you can ask any variation of AGI questions with any definition, even ask conditional questions.

[1] https://futuresearch.ai/app

Author here, I drew on this from AI 2027. Yes, a very-expensive AGI, e.g. $1 million / day to simulate a smart human, would be a huge deal. But it would have meaningfully different effects than a cheap one.

Here's one definition AI 2027 used [1]: "Superhuman coder (SC): An AI system for which the company could run with 5% of their compute budget 30x as many agents as they have human research engineers..."

[1] https://ai-2027.com/research/timelines-forecast

Snake oil is a bit strong, no? I would agree that the burden of proof is on multi-agent systems to show they are outperforming single-agent systems.

On my own evals I have seen this, though the improvement may not have been worth the extra cost.