HN user

snake_doc

674 karma
Posts7
Comments150
View on HN
Qwen 3.8 3 days ago

Not directly, relevant, but Alibaba (maker of Qwen) actually owns about ~20-30% of MoonshotAI (the maker of Kimi K3).

Image/video understanding still quite cost effective from the Gemini flash series models?

Image generation and veo models I’d imagine quite effective for creators; new Instagram accounts with AI content that are garnering millions of followers in spans of weeks are quite common now

2025 Letter 7 months ago

That’s not change unique to America though. Why would I be thinking of changes that have affected almost every country, when talking about whether Americans recognize how America has changed?

Hahahaha, this is like saying, the world wars didn’t impact Europe, because it also impacted the whole world! Europeans, the war didn’t happen!! Anyways… this entire thread is more evidence that European stereotypes are valid for the most part

2025 Letter 7 months ago

Most Americans have no sense of how very different their country is now from say, the country that launched the Apollo missions.

You cannot be kidding right? Those that remember the Apollo missions will undoubtedly agree their country is different, first but not least they are most likely using a smartphone assembled and designed with technology unimaginable by NASA planning the Apollo missions; not only that, the smartphone is assembled half way around the world by a country previously in such dire poverty and famine that over 30M died due to Marxist central planning.

2025 Letter 7 months ago

What exactly is wrong with Americans or for the most part the rest of the world valuing economic performance as a measure of prosperity and progress?

Your comment is again another anecdote confirming European stereotypes. It’s not a “trap”, it’s a different world view.

2025 Letter 7 months ago

Would you be able provide some evidence to the contrary when it comes to the topics discussed in the letter?

On industrial infrastructure

On technology innovation

On internet regulation

On central planning

Otherwise, your comment becomes an anecdote supporting the common stereotypes (assuming you’re from Europe).

GPT-5.2 7 months ago

Models were run with maximum available reasoning effort in our API (xhigh for GPT‑5.2 Thinking & Pro, and high for GPT‑5.1 Thinking), except for the professional evals, where GPT‑5.2 Thinking was run with reasoning effort heavy, the maximum available in ChatGPT Pro. Benchmarks were conducted in a research environment, which may provide slightly different output from production ChatGPT in some cases.

Feels like a Llama 4 type release. Benchmarks are not apples to apples. Reasoning effort is across the board higher, thus uses more compute to achieve an higher score on benchmarks.

Also notes that some may not be producible.

Also, vision benchmarks all use Python tool harness, and they exclude scores that are low without the harness.

China added ~90GW of utility solar per year in last 2 years. There's ~400-500GW solar+wind under construction there.

It is possible, just may be not in the U.S.

Note: given renewables can't provide base load, capacity factor is 10-30% (lower for solar, higher for wind), so actual energy generation will vary...

Modular Manifolds 10 months ago

Aren’t they all optimization techniques at the end of the day? Now you’re just debating semantics

Mafia behavior continues… (not my observation, but the Texas senator’s Ted Cruz[1]).

$100k is a big pizzo (protection fee)!

[1] https://www.bloomberg.com/news/articles/2025-09-19/ted-cruz-...

“That’s right outta ‘Goodfellas,’ that’s right out of a mafioso going into a bar saying, ‘Nice bar you have here, it’d be a shame if something happened to it,’” Cruz said, using the iconic New York accent associated with the Mafia.

Holy unnecessary use of terminology to explain a reverse graph traversal. “Loss”, “gradients”, “differentiating”— no! stop!

This must be what AI hype actually is. Complete incoherent language to explain a very straight forward concept.

This is just: LLMs judging intermediate node outputs, and reverse traversing the graph while doing so until it modifies the original prompt.

Without taking a position on unipolar vs. multi-polar:

Dario makes an astounding implicit assumption:

- China originating labs cannot acquire chips providing 80-90% similar utility without the US within the next 2-3 years.

I'll make an observation, re: DeepSeek's incentives that drove them to create the innovations from the V2 and V3 papers.

DeepSeek, compared to American AI labs, are much more compute constrained, but in a unique way. Their chips are more memory bandwidth constrained (depending on type anywhere from 50% to 80% less bandwidth).

Therefore, each dollar/hour of investment towards memory optimization is worth MORE to DeepSeek than to American labs.

In the V2/3 paper, they've demonstrated exactly that with these memory optimization techniques.

1. MLA -> reduces KV cache by nearly 80% compared to GQA. By the way, this was published in V2 in May 2024.

2. FP8 matmul (while still accumlating in FP32 gradients) without losing significant quality.

3. DualPipe scheduling and reworking of Hopper SM's allocation on communication vs. computation -> DeepSeek's V3 paper has 2 full pages of hardware suggestions for "hardware designers" (read NVIDIA)

Export controls in a global market create different incentives in parties. The resulting incentives will change, and agents (using it as an traditional economics term) will change their capital allocation strategy.

You should still be mocked.

1. ChatGPT data is widely on the internet, just google Sharegpt dataset and you can scrap 200k+ conversations with a few stroke of huggingface commands. These were then used by the open source community like Vicuña models, there was a period of several months in the open source community where RLAIF was all the rage; so this data populated the internet. So if a company is crawling and scraping the internet, this will eventually be in the dataset.

2. The v3 deepseek model was trained on 15T tokens. Please educate yourself and calculate how long (in latency, inference for 1k token output will take almost 30seconds) and cost it would be to extract 15T tokens from ChatGPT / Azure API. Granted API accounts all have spend limits, and will trip fraud detection on OAI billing, how long would the subterfuge had to take place? With which model? At what time? Wouldn’t they have to keep repeating this for subsequent generation of OAI models?

3. OAI didn’t invent MLA, they didn’t invent multi token prediction with disconnected ROPE, they didn’t invent FP8 matmul training dynamics (while accumulating in FP32) without losing significant quality.

So go away

No, the only thing that matters is if the portfolio delivers returns in excess of your cost of capital.

If your portfolio is green, you can still be a poor performer.

The other way is certainly also true. Your short piece is rational, but lacks insight into the inference and training dynamics of ML adoption unconstrained.

The rate of ML progress is spectacularly compute constrained today. Every step in today’s scaling program is setup to de-risked the next scale up, because the opportunity cost of compute is so high. If the opportunity cost of compute is not so high, you can skip the 1B to 8B scale ups and grid search data mixes and hyperparameters.

The market/concentration risk premium drove most of the volatility today. If it was truly value driven, then this should have happened 6 months ago when DeepSeek released V2 that had the vast majority of cost optimizations.

Cloud data center CapEx is backstopped by their growth outlook driven by the technology, not by GPU manufacturers. Dollars will shift just as quickly (like how Meta literally teared down a half built data center in 2023 to restart it to meet new designs).

Email notice sent Friday Dec 20, 2024 requiring 30 days notice to cancel. Addendum link also intentionally not hyperlinked in email notice.

We are writing to inform you of a change to our subscription terms. To ensure a smooth transition, we have updated our Subscription Addendum (available at https://ramp.com/legal/subscription-addendum) to implement a new process for customers on annual plans to cancel automatic renewal on Ramp subscription plans.

Key changes: To request cancellation, you must give written notice to your account manager. Cancellation requests must be submitted at least 30 days before your next renewal date. The changes above will take effect on December 13, 2024. If you have any questions, please contact your account manager.

Best, The Ramp Team