HN user

LarsDu88

4,732 karma

https://dublog.net

Posts39
Comments1,066
View on HN
bsky.app 16d ago

Majority of id software to be laid off by Microsoft

LarsDu88
34pts10
www.nature.com 4mo ago

Major Aging Related Metabolic Shifts Occur Around Ages 44 and 60

LarsDu88
4pts1
panel-panic.com 4mo ago

Show HN: Panel Panic a Rust/Macroquad/WASM Panel de Pon/Tetris Attack Clone

LarsDu88
2pts0
news.ycombinator.com 6mo ago

Ask HN: Idea) Autoregressive joint embedding predictor model

LarsDu88
2pts0
github.com 12mo ago

Show HN: A custom ESP32-S3 board for battery powered WiFi enabled plant watering

LarsDu88
2pts0
www.pcgamer.com 1y ago

What Happened to the Creator of Valve's Forgotten Game – Gunman Chronicles

LarsDu88
4pts1
aaronlou.com 1y ago

Discrete Text Diffusion Explained

LarsDu88
14pts1
sequencing.roche.com 1y ago

Roche announces new SBX Xpandomer DNA sequencing technology

LarsDu88
2pts4
www.nature.com 1y ago

Asteroid Bennu samples have 14/20 amino acids found in life on Earth

LarsDu88
5pts3
www.youtube.com 1y ago

Gemini 2.0 Flash Text-To-Speech is Insane [video]

LarsDu88
3pts1
scontent-lax3-1.xx.fbcdn.net 1y ago

Discrete Flow Matching for Text and Image Generation

LarsDu88
2pts1
www.fabricatedknowledge.com 1y ago

Intel’s board, and an example of when boards and short-termism fail

LarsDu88
351pts278
www.nejm.org 1y ago

7/67 Children Receiving Skysona Gene Therapy Develop Blood Cancer

LarsDu88
41pts54
www.wired.com 1y ago

23andMe Is Sinking Fast [Wired]

LarsDu88
4pts2
arxiv.org 1y ago

SwiGLU activation function causes instability in FP8 LLM training

LarsDu88
10pts2
investors.23andme.com 1y ago

Updated 23andMe Board of Directors Website

LarsDu88
4pts2
investors.23andme.com 1y ago

Independent directors of 23andMe resign from board

LarsDu88
633pts350
github.com 1y ago

Show HN: Diffumon – A Basic Image Generating Diffusion Model

LarsDu88
4pts1
github.com 1y ago

Diffumon – Simple Image Generating Diffusion Model

LarsDu88
2pts1
dublog.net 1y ago

Host Your Own Copilot

LarsDu88
29pts11
news.ycombinator.com 1y ago

Ask HN: Whatever Happened to Memristors?

LarsDu88
13pts4
ai.meta.com 1y ago

Tuning-Free Personalized Image Generation

LarsDu88
82pts45
dublog.net 1y ago

Big tech wants to make AI cost nothing

LarsDu88
88pts80
dublog.net 2y ago

All the Activations (and a brief history of deep learning)

LarsDu88
2pts1
play.cartesia.ai 2y ago

Cartesia Sonic's Text to Speech Generator

LarsDu88
2pts1
dublog.net 2y ago

Python Has Too Many Package Managers

LarsDu88
131pts163
www.nature.com 2y ago

Semantic encoding during language comprehension at single-cell resolution

LarsDu88
2pts1
www.youtube.com 2y ago

Half-Life 25th Anniversary Documentary [video]

LarsDu88
3pts0
iscinumpy.dev 2y ago

Version Capping Is Evil

LarsDu88
3pts0
www.youtube.com 3y ago

Prescient GPU commecials from the 90s [video]

LarsDu88
3pts1

Certainly there will be demand for Fable7, but that demand is context specific. Frontier labs' profit is dependent on there being sufficient demand for the next layer of capability and whether the premium consumers are willing to pay for that.

The incremental unlock of capability by ever increasing frontier model sizes will eventually reach diminishing returns.

I would argue tnference speed increases would actually unlock a different kind of more meaningful value for a wider audience.

Currently the setup is paged view in RAM shuttled to HBDRAM (VRAM) on the GPU, which in turn has to get materialized piece by piece onto cache SRAM on the GPU.

Cerebras tries to get around this by keeping everything on cache SRAM as much as possible, which it burns directly to the chip wafer itself and physically places that SRAM directly next to the tiny compute unit that does the actual math.

An ideal setup (not sure how easy this is to achieve in practice), is the burn the weights of the model directly to the chip as a sort of ROM, the actual math operations as actual digital circuits, and have SRAM, or even something akin to naked registers to directly compute off inference batch data. Cuts out 2-3 layers of abstraction and indirection.

Cerebras is not an ASIC. It is a large wafer scale chip that has a load of small SRAM modules paired with tiny compute models. The SRAM is basically a giant cache that is supposed to eliminate the bottleneck between shuttling and materializing big tensors between DRAM and cache, but the actual amount of SRAM is still nowhere near enough to server a frontier model on a single wafer. You still need dozens of wafers, which makes Cerebras rather cost intensive given that Nvidia can get dozens of GPUs out of a single wafer

Honestly, whether you think burning current SOTA to hardware is an overinvestment risk depends on what your definition of intelligence is. If you think intelligence is something that can grow like height such that 18 months from now we will basically be bowing down to machine god giants that are running on B200s, then investing in ASICs is the wrong move. However, if you you subscribe to the (very reasonable view) that intelligence is more like a round ball that we are trying to make as spherical as possible (ala Francois Chollet's writings), then at some point the ball will be smooth enough for most people and many tasks.

It takes about 18 months to go through the design, verification, and manufacturing process if you move at breakneck pace. Design could probably be sped up.

About 18 months ago the top model was GPT-4o. Not great by today's standards, but still good enough for many tasks (certainly a big chunk of chatbot queries). The current SOTA covers far more use cases, but importantly at a level that surpasses many thresholds of utility.

The open weight, open architecture releases of the past several days has me more convinced that ultimately, the winner will be whoever burns their models to ASICs fastest.

The LLMs themselves are capable of doing some aspects of chip design as evinced by the K3 press release.

Furthermore, the frontier models are "good enough" for a wide swathe of tasks and will soon hit that threshold for a good amount of software engineering (if not already). Does anyone think we need a Mythos level model to plan a road trip, or give someone tips on making a cake recipe?

A Fable 5 model running at 9,000 tokens/s on an ASIC rather than 150 tokens/s on electricity chugging Nvidia GPUs, or even giant SRAM Cerebras or Groq chips could be good enough to meet the majority of demand.

Furthermore, if you're an enterprise the risk of data exfiltration and feeding data to a potential competitor like OpenAI or Anthropic is greatly reduced if you could shift to on-prem ASIC deployments. A handful of chips could cover a wide variety of use cases and cover them more securely. There are a lot of corporate use-cases for LLMs that are not frontier math research or coding.

I believes the weights are burned as ROM microcode, but for an effective inference speedup, you do want to burn the architecture (matmuls, activation functions, MoE gates, etc) as well which will differ from model to model.

It's not as simple as a weight swap between identical architectures.

The speed gains are also from not having to route the weights through wiring like with ROM cartridges.

A really good startup idea right now... Use kimi k3 to reproduce the kimi k3 asic design and start fabbing it immediately. In 12-18 months, start spinning up your own cloud and start competing with the frontier labs ASAP.

Who needs superintelligence with you have 8,700 tokens/s at near Fable levels of performance???

This is like the Bill Gates, Paul Allen moment, but for hardware.

The narrative that superintelligence is imminent is partially at fault here.

There are competing definitions of what intelligence even is, and the one that I find most striking is from Francois Chollet which is that intelligence can be boiled down to skill acquisition efficiency. This type of definition makes intelligence more akin to polishing a ball than growing a watermelon.

The superintelligence doomers warn that the watermelon is going to start growing exponentially and crush everyone. But what might actually be happening is that we are not growing a watermelon but rather polishing the ball until its really smooth and shiny. There's a point where you can get it to micron levels of polish but for most tasks (white collar text domains tasks), it's smooth enough! You will be able to go to the ball store and buy a low cost made in china ball for most tasks.

The real challenge is actually branching out domains and modalities to tackle things like blue collar labor. Over time, white collar work automatable or able to be made hyperefficient by LLMs will see LLM commoditization.

You're confusing Tesla for like 10 other Chinese brands. Tesla too is falling behind, particularly when it comes to price.

BYD may very well be pulling ahead on battery tech and vertical integration right now.

In a few years Tesla's primary moat will be political

Unity does not make that much money from assets. They make the majority of their revenue through licensing their engine to "whales"... the small percentage of games that make huge revenue.

They also make money through ad services... a market they seriously missed out on (look at AppLovin stock vs Unity).

Asset sales are barely a blip.

Unreal makes the vast majority of it's revenue through microtransactions on its one major whale game Fortnite

MSFT could have opened up idTech completely since they make 0 dollars from licensing the engine anyways.

Microsoft's game divisions make money through making games, so opening up the engine itself would've been conducive to their goals (cultivating an ecosystem of devs and even contractors familiar with the tooling).

idTech rendering is more competitive with Epic's Unreal technology than Unity and Godot.

They should simply open source it if they fire the devs. Else the engine and future support for the games built on it are essentially being tossed into the trash.

Microsoft, one the world's greatest monopolists, bequeaths a game engine monopoly unto Epic Games, in one the biggest corporate blunders of all time.

If they were smarter about this, they would commoditize their compliment and open source the Doom The Dark Ages engine just like John Carmack did with the Quake 3 engine.

Scott Miller (founder of Apogee/3dRealms) stated that id Software and most of programmers (the team behind idTech and Doom: The Dark Ages) will be let go.

Basically one of the last cutting edge engines not built/maintained by Epic Games (Unreal Engine) will basically die.

Fun exercise. Type in pokemon or japanese. You can really see the nearest neighbor text in embedding space. Pokemon gives passafes referencing animals and japanese passages referencing foreigners

John Carmack slaved away writing super optimized ground breaking realtime 3d game engine code that created a multibillion dollar industry.

Adrian Carmack mostly took photographs of clay scultures (and in some cases actual plastic toys, in the case of the chainsaw and pistol used in the original Doom), and digitized them in Corel. Something any half-way decent art student could do (not to discount the iconic visual style of Commander Keen and Doom!)

It was Adrian who walked away with 41% of the >100 million dollars company!