HN user

Flux159

932 karma

Ex-Meta, Ex-Oscar software engineer.

https://twitter.com/flux159

Posts20
Comments141
View on HN
newsletter.semianalysis.com 29d ago

CXMT is set to challenge DRAM incumbents

Flux159
4pts1
news.ycombinator.com 4mo ago

Ask HN: How do you prevent AI generated GitHub issues?

Flux159
2pts1
github.com 5mo ago

Show HN: Mystral Native – Run JavaScript games natively with WebGPU (no browser)

Flux159
50pts18
importai.substack.com 6mo ago

Silent Sirens, flashing for us all

Flux159
2pts1
en.wikipedia.org 7mo ago

DRAM Price fixing scandal (2002)

Flux159
4pts1
suyogs.com 1y ago

Getting Gemini to write an ORM for Spanner in a weekend

Flux159
2pts0
github.com 1y ago

Show HN: Agentic Shell (agish)

Flux159
1pts0
huggingface.co 1y ago

Timeline of AI models released in 2024

Flux159
3pts0
twitter.com 1y ago

Gemini 2.0 Flash Thinking

Flux159
7pts1
github.com 2y ago

LGM: Large Multi-View Gaussian Model – Create 3D model from text or single image

Flux159
3pts0
github.com 2y ago

Promptbench: A Unified Library for Evaluating and Understanding LLMs

Flux159
1pts1
github.com 2y ago

Anydoor: Zero Shot Object-level image customization

Flux159
1pts1
twitter.com 2y ago

Midjourney V1 to V6 – Image generation improvements from Feb 2022 to Dec 2023

Flux159
3pts2
huggingface.co 2y ago

DiffMorpher – Using Diffusion Models for Image Morphing

Flux159
1pts1
github.com 2y ago

StreamDiffusion: A pipeline-level solution for real-time interactive generation

Flux159
365pts67
huggingface.co 2y ago

Zephyr 7B – Mistral Finetune that responds like ChatGPT

Flux159
37pts12
lastmileai.dev 2y ago

Show HN: Experiment with Hugging Face models in a single notebook interface

Flux159
43pts7
lastmileai.dev 3y ago

AI Workbooks – A notebook interface for LLMs, image and audio models

Flux159
196pts33
blog.lastmileai.dev 3y ago

Using ChatGPT Plugins with LLaMA

Flux159
315pts138
www.newsatlas.io 12y ago

Show HN: Newsatlas.io – Visualize current world news in realtime

Flux159
1pts0

There's also an Inkling-Small that is 276B, 12B active that is much smaller than GLM 5.2 and still multimodal. Not released yet, but in the announcement link they mention that they're testing Inkling-Small & will release as open weight after testing. That one may be interesting as a Deepseek V4 Flash replacement.

Article from semianalysis that goes into the history of CXMT & over a decade long investment (including government support) to become a 4th player in the DRAM market.

Useful to understand in the context of the current DRAM supply constraints since it also goes into what capacity the other 3 players are bringing online in the next 2 years.

Unreal will definitely get better results out of the box, but it's also possible to do photorealism with significantly less overhead (particularly UE shader compilation overhead) - useful for single purpose platforms. If you don't need to support lots of specific editor or game features, it may be a valuable investment.

UE is definitely used to obtain simulation data in other domains (this is coming from first hand experience in big tech), but usually through scripting UE handmade levels in python which also needed convoluted server systems at the time (hopefully this has gotten better now).

I think the reason we're not seeing many examples yet is that the full loop doesn't work completely autonomously yet. There's still a human in the loop at some critical points - specifically testing against a spec (runtime testing if say working on web or mobile app before shipping to users). LLMs can do compile time testing and validation, unit tests, and can write your end to end tests, but if you're shipping software to users, there's still a human somewhere involved. This isn't even mentioning marketing and actually getting your software into the hands of users - which while it can be automated, a lot of marketing with AI is still sloppy.

There’s some early work being done here by companies looking at making LLM ASICS like Taalas (HC1 gets 17k t/s for llama 8b - currently at 2.5kW which is closer to a single server, but this is their first chip).

There’s other options like photonic computing which might be able to reduce power significantly but are still in research as far as I can tell. Because so much money is invested in AI & traditional gpu inference is so power hungry, I would expect significant improvements in this space quickly.

I'm going to be honest - this is over a year late. I still use ChatGPT on Mac because it actually had a Mac App from May of 2024, whereas I had to go to the Gemini website to use Gemini. It was even worse because of the fragmented experience - there's been an iOS Gemini app for a while now. Integrating Gemini into Chrome is not the same experience as having a standalone app.

Now that it's at least here, hopefully Google can continue updating it instead of giving up on it if their metrics don't show as fast growth as iOS or Chrome usage.

This looks pretty interesting! I haven't used it yet, but looked through the code a bit, it looks like it uses turndown to convert the html to markdown first, then it passes that to the LLM so assuming that's a huge reduction in tokens by preprocessing. Do you have any data on how often this can cause issues? ie tables or other information being lost?

Then langchain and structured schemas for the output along w/ a specific system prompt for the LLM. Do you know which open source models work best or do you just use gemini in production?

Also, looking at the docs, Gemini 2.5 flash is getting deprecated by June 17th https://ai.google.dev/gemini-api/docs/deprecations#gemini-2.... (I keep getting emails from Google about it), so might want to update that to Gemini 3 Flash in the examples.

Amazon checkout is not working for me right now. There's no Amazon specific status page it seems?

AWS has information about their UAE data centers, but haven't seen any confirmation from Amazon itself that amazon.com is having issues.

GPT‑5.3 Instant 5 months ago

Thanks for clarifying! I guess the default for most users is going to be to use the router / auto switcher which is fine since most people won't change the default.

Just noting that I'm not against differentiation in products, but it gets very confusing for users when there's too many options (in the case of the consumer ChatGPT at least this is still more limited than in pre-GPT 5 days). The issue is that there's differentiation at what I pay monthly (free vs plus vs pro) and also at the model layer - which essentially becomes this matrix of different options / limits per model (and we're not even getting into capabilities).

For someone who uses codex as well, there are 5 models there when I use /model (on Plus plan, spark is only available for Pro plan users), limits also tied to my same consumer ChatGPT plan.

I imagine the model differentiation is only going to get worse as well since with more fine tuned use cases, there will be many different models (ie health care answers, etc.) - is it really on the user to figure out what to use? The only saving grace is that it's not as bad as Intel or AMD cpu naming schemes / cloud provider instance naming, but that's a very low bar.

So is this a minimal upgrade before the M6 Macbook Pros w/ OLED & a redesign later this year?

It doesn't even look like they added cellular as an option with their own C1X chip (getting around the licensing / cost issues since it's their own chip now).

GPT‑5.3 Instant 5 months ago

I'm a bit confused by this branding (never even noticed that there was a 5.2-Instant), it's not a super fast 1000tok/s Cerebras based model which they have for codex-spark, it's just 5.2 w/out the router / "non-thinking" mode?

I feel like openai is going to get right back to where they were pre GPT-5 with a ton of different options and no one knows which model to use for what.

This seems like it’s in response to the congressional testimony last week to clarify some things about their remote assistance systems.

It’s interesting that they only have 70 people for this - I can understand the outside the US ones for nighttime assistance and they need to be able to scale for other countries too in the future.

What I’m still wondering is what is limiting the scaling for Waymo - just cars or also the sensor systems? They’ve had their new test vehicles in SF for a while but I still think that most customers only get their Jaguars right now (and still limited on highway driving to specific customers in the Bay Area).

I'm currently a solo bootstrapped founder, have done short stints in the past - 1 year in 2022, then became cofounder of a funded startup for a year. Now doing it again.

Question is how you stay motivated to keep at it - looks like it took about 4 years before you made similar to your Google salary, did family pressure or external pressure ever impact you? Or is it mainly just keep your eyes on the longer term goal?

I'm also quite lucky that I was aiming for lean-FIRE before I left Facebook, so I have the luxury of being able to keep at it, but sometimes it is demotivating seeing peers / others.

WASM shouldn't be an issue since the draco decoder uses it - but it may only work with V8 (for quickjs builds it wouldn't work, but the default builds use V8+dawn). Obviously with an alpha runtime, there may be bugs.

I think it would be super cool to have some sort of extension before WebGPU (web) has it. I was taking a look at the prior example & it seems like there's good ongoing discussion linked here about it: https://github.com/gpuweb/gpuweb/issues/535. Also I believe that Metal has hardware ray tracing support now too?

Re: Implementation, a few options exist - a separate Dawn fork with RT is one path (though Dawn builds are slow, 1-2 hours on CI). Another approach would be exposing custom native bindings directly from MystralNative alongside the WebGPU APIs - that might make iteration much faster for testing feasibility. The JS API would need to be feature-flagged so the same code gracefully falls back when running on web (did this for a native draco impl too that avoids having to load wasm: https://mystralengine.github.io/mystralnative/docs/api/nativ...).

So in theory it should be possible, but it might require customizing the Dawn or wgpu-native builds if they don't support it (this is providing the JS bindings / wrapper around those two implementations of wgpu.h). But I've already added a special C++ method to handle draco compression natively, adding some mystral native only methods is not out of the question (however, I would want to ensure that usage of those via JS is always feature flagged so that it doesn't break when run on web).

Did you write your WebGPU chessboard using the raw JS APIs? Ideally it should work, but I just fixed up some missing APIs to get Three.js working in v0.1.0, so if there are issues, then please open up an issue on github - will try to get it working so we close any gaps.

AssemblyScript was just mentioned as some prior work, I don't think that AssemblyScript would work as is for games.

I realize the major issues with TS->C++ though (or any language to C++, Facebook has prior work converting php to C++ https://en.wikipedia.org/wiki/HipHop_for_PHP that was eventually deprecated in favor of HHVM). I think that iteratively improving the JS engine (Mystral.js the one that is not open source yet but is why MystralNative exists) to work with the compiler would be the first step and ensuring that games and examples built on top with a subset of TS is a starting point here. I don't think that the goal for MystralScript should be to support Three.js or any other engine to begin with as that would end up going down the same compatibility pits that hiphop did.

Being able to update the entire stack here is actually very useful - in theory parts of mystral.js could just be embedded into mystralnative (separate build flags, probably not a standard build) avoiding any TS->C++ compilation for core engine work & then ensuring that games built on top are using the strict subset of TS that does work well with the AOT compilation system. One option for numbers is actually using comment annotations (similar to how JSDoc types work for typescript compiler, specifically using annotations in comments to make sure that the web builds don't change).

Re: TS compiler - I do have some basics started here and I am already seeing that tests are pretty slow. I don't think that the tsgo compiler has a similar API though for parsing & emitters right now, so as much as I would like to switch to it (I have for my web projects & the speed is awesome), I don't think I can yet until the API work is clarified: https://github.com/microsoft/typescript-go/discussions/455

Followup comment about Apple disallowing JIT - will need to confirm if JSC is allowed to JIT or only inside of a webview. I was able to get JSC + wgpu-native rendering in an iOS build, but would need to confirm if it can pass app review.

There's 2 other performance things that you can do by controlling the runtime though - add special perf methods (which I did for draco decoding - there is currently one __mystralNativeDecodeDracoAsync API that is non standard), but the docs clearly lay out that you should feature gate it if you're going to use it so you don't break web builds: https://mystralengine.github.io/mystralnative/docs/api/nativ...

The other thing is more experimental - writing an AOT compiler for a subset of Typescript to convert it into C++ then just compile your code ("MystralScript") - this would be similar to Unity's C# AOT compiler and kinda be it's own separate project, but there is some prior work with porffor, AssemblyScript, and Static Hermes here, so it's not completely just a research project.

Phaser is not supported right now because phaser is still using a WebGL renderer from my understanding (maybe in a v2.0.0 adding ANGLE + WebGL support is an option, but debating if that's a good idea or not).

Pixi 8 has a WebGPU renderer so that should be supported as part of a v1.0.0 release - it's on the roadmap to verify that three and pixi 8 work correctly: https://github.com/mystralengine/mystralnative/issues/7

So I am stubbing parts of the DOM api (input handling like keydown, pointer events, etc.), so you shouldn't need to rewrite any of that.

Three.js and Pixi 8 with the WebGPU renderer are part of the v1.0.0 roadmap (verifying that they can work correctly on all platforms), right now most of the testing was done against my own engine (tentatively called mystral.js which will also be open sourced as part of v1.0.0, it's already used for some of the examples, just as a minified bundle): https://github.com/mystralengine/mystralnative/issues/7

Hi, thanks! Yeah for controls I'm emulating pointerevents and keydown, keyup from SDL3 inputs & events. The goal is that the same JS that you write for a browser should "just work". It's still very alpha, but I was able to get my own WebGPU game engine running in it & have a sponza example that uses the exact key and pointer events to handle WASD / mouse controls: https://mystraldev.itch.io/sponza-in-webgpu-mystral-engine (the web build there is older, but the downloads for Windows, Mac, and Linux are using Mystral Native - you can clearly tell that it's not Electron by size (even Tauri for Mac didn't support webp inside of the WebGPU context so I couldn't use draco compressed assets w/ webp textures).

I put up a roadmap to get Three.js and Pixi 8 (webgpu renderer) fully working as part of a 1.0.0 release, but there's nothing that my JS engine is doing that is that different than Three.js or Pixi. https://github.com/mystralengine/mystralnative/issues/7

I did have to get Skia for Canvas2d support because I was using it for UI elements inside of the canvas, so right now it's a WebGPU + Canvas2d runtime. Debating if I should also add ANGLE and WebGL bindings as well in v2.0.0 to support a lot of other use cases too. Fonts support is built in as part of the Skia support as well, so that is also covered. WebAudio is another thing that is currently supported, but may need more testing to be fully compatible.

This looks useful for people not using Claude Code, but I do think that the desktop example in the video could be a bit misleading (particularly for non-developers) - Claude is definitely not taking screenshots of that desktop & organizing, it's using normal file management cli tools. The reason seems a bit obvious - it's much easier to read file names, types, etc. via an "ls" than try to infer via an image.

But it also gets to one of Claude's (Opus 4.5) current weaknesses - image understanding. Claude really isn't able to understand details of images in the same way that people currently can - this is also explained well with an analysis of Claude Plays Pokemon https://www.lesswrong.com/posts/u6Lacc7wx4yYkBQ3r/insights-i.... I think over the next few years we'll probably see all major LLM companies work on resolving these weaknesses & then LLMs using UIs will work significantly better (and eventually get to proper video stream understanding as well - not 'take a screenshot every 500ms' and call that video understanding).

Nice, this is similar to what I was wondering about - it looks like it's pretty limited in capability right now (looks like it only supports canvas2d at the moment: https://nxjs.n8.io/runtime/rendering/canvas), but in theory it would allow you to make a layer to convert WebGPU or WebGL games for Switch (ignoring the huge performance drop going from v8 / jit JS engines to QuickJS).