HN user

freeqaz

3,392 karma

Head of Security Architecture @ Scale AI - also ex-YC (LunaSec, S19), ex-Figma, ex-Uber, ex-Snap

Email me or go to my domain -- freeqaz.com

Personal email: me at freeqaz

----

[ my public key: https://keybase.io/lunasec; my proof: https://keybase.io/lunasec/sigs/4BoTHAkR8TsFV_1S1Rg-Ed_nFD14ocW-gSgKRhhyNIE ]

YC Badge: 0xcb17dfcabdb996705e2859e224d4cdd94d9ae344

Posts132
Comments710
View on HN
github.com 7mo ago

React2Shell Exploit Analysis with POCs (RCE in Next.js)

freeqaz
1pts1
www.scientificamerican.com 8mo ago

After 35 Years, a Solution to the CIA's Kryptos Puzzle Has Been Found

freeqaz
1pts1
showdown.scale.com 10mo ago

Seal Showdown Technical Report (AI Benchmark) [pdf]

freeqaz
1pts0
arstechnica.com 10mo ago

OpenAI links up with Broadcom to produce its own AI chips

freeqaz
2pts0
arxiv.org 11mo ago

Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens

freeqaz
1pts0
www.theverge.com 11mo ago

OpenAI gets caught vibe graphing

freeqaz
4pts0
www.npr.org 11mo ago

Trump says Nvidia will hand the U.S. 15% of its H20 chip sales to China

freeqaz
4pts0
arstechnica.com 1y ago

Review: Ryzen AI CPU makes the Framework Laptop 13 the fastest has ever been

freeqaz
22pts1
arstechnica.com 1y ago

Nvidia announces DGX desktop "personal AI supercomputers"

freeqaz
20pts6
www.snopes.com 1y ago

Hackers played AI-generated video of Trump kissing Musk's feet on Gov TVs

freeqaz
18pts3
vercel.com 1y ago

Vercel: Introducing Fluid Compute

freeqaz
4pts0
twitter.com 1y ago

TikTok is back after executive order stalling ban

freeqaz
37pts2
thegamepost.com 1y ago

Possible Nintendo Switch 2 Tech Specs

freeqaz
4pts3
arstechnica.com 1y ago

China orbits first Guowang internet satellites, with thousands more to come

freeqaz
8pts1
arstechnica.com 1y ago

Trump nominates Jared Isaacman to become the next NASA administrator

freeqaz
8pts1
arstechnica.com 1y ago

$300B pledge at COP29 climate summit a "paltry sum"

freeqaz
2pts1
arstechnica.com 1y ago

Review: The fastest of the M4 MacBook Pros might be the least interesting one

freeqaz
36pts35
arstechnica.com 1y ago

TSA silent on CrowdStrike's claim Delta skipped required security update

freeqaz
3pts0
gapcoin.org 1y ago

Gapcoin: A prime number based P2P cryptocurrency

freeqaz
2pts0
arstechnica.com 1y ago

US suspects TSMC helped Huawei skirt export controls, report says

freeqaz
4pts1
arstechnica.com 1y ago

X's depressing ad revenue helps Musk avoid EU's strictest antitrust law

freeqaz
7pts0
registry.terraform.io 1y ago

Terraform Provider for Dominos Pizza

freeqaz
101pts32
en.wikipedia.org 1y ago

Sedna (Dwarf Planet)

freeqaz
3pts0
www.phoronix.com 1y ago

Real-Time "Preempt_rt" Support Merged for Linux 6.12

freeqaz
2pts1
www.sciencedirect.com 1y ago

Production of high quality syngas from argon plasma gasification of plastic

freeqaz
2pts0
github.com 1y ago

Modxo: RP Pico based Xbox modchip

freeqaz
1pts0
en.wikipedia.org 1y ago

Dry Cask Storage

freeqaz
11pts21
arstechnica.com 1y ago

Review: ReMarkable Paper Pro writing tablet feels almost like paper, for a price

freeqaz
2pts0
chipsandcheese.com 1y ago

An Interview with Intel's Arik Gihon about Lunar Lake at Hot Chips 2024

freeqaz
2pts0
vgel.me 1y ago

Representation Engineering Mistral-7B an Acid Trip

freeqaz
2pts0

I've been working on porting Rock Band 3 and Dance Central 3 to PC via AI assisted decompilation. The repos are on my GitHub (https://github.com/freeqaz/rb3 and https://github.com/freeqaz/dc3-decomp)

Like OP, I've learned at lot in this process. I have versions running in the browser now with a custom WebGPU rendering engine. Still lots of jank and vibes, but it's wild to see what models are capable of with the right tooling. (I've had Claude add extensions into Ghidra for Xbox/Wii specific instruction support)

Wild times we're living in. It's great for software preservation though!

I have struggled to get this project working on non-Windows. It just hangs and crashes no matter what I do or try on Linux/Mac. It's a very Windows-oriented project that's slowly losing the shackles right now.

This is exactly what happened with Log4Shell.

Day -X + 1: Engineer at Alibaba finds the vuln and tells Apache. Patch is pushed to git while new release is coordinated.

Day -X: A black hat sees commits fixing the bug. Attacks start happening.

Day 0: Memes start circulating in Minecraft communities of people crashing servers. Some logs are shared on Twitter, especially in China, of people getting pwned.

Day 0 + ~4 hours: My friend DMs me a meme on Twitter. I look up to find the CVE. Doesn't exist. My friend and I reproduce the exploit and write up a blog post about it. (We name it Log4Shell to differentiate it from a different, older log4j RCE vuln)

Day ~1: Media starts picking it up. Apache is forced to release patches faster in response. CVE is actually published to properly allow security scanners to identify it.

Today: AI makes this happen faster and more consistently. Patches probably should be kept private until a coordinated disclosure happens post-testing and CVE being published?

Hard to say what the right move is, but this is gonna be happening a lot over the next 1-3 years. Lots of companies are going to be getting cooked until AI helps us patch faster than attackers can exploit these fresh 0-days.

Also a good fallback if your phone screen cracked 2 hours before. But I can imagine part of the challenge they are facing here are scalpers. TicketMaster app 'rotates' the actual ticket every 30 seconds. Can't rotate paper.

I'd think that having a 2nd factor like presenting ID that matches the ticket would be sufficient there though.

128gb is the max RAM that the current Strix Halo supports with ~250GB/s of bandwidth. The Mac Studio is 256GB max and ~900GB/s of memory bandwidth. They are in different categories of performance, even price-per-dollar is worse. (~$2700 for Framework Desktop vs $7500 for Mac Studio M3 Ultra)

Claude Sonnet 4.6 5 months ago

If it maintains the same price (with Anthropic tends to do or undercuts themselves) then this would be 1/3rd of the price of Opus.

Edit: Yep, same price. "Pricing remains the same as Sonnet 4.5, starting at $3/$15 per million tokens."

Claude Sonnet 4.6 5 months ago

I would honestly guess that this is just a small amount of tweaking on top of the Sonnet 4.x models. It seems like providers are rarely training new 'base' models anymore. We're at a point where the gains are more from modifying the model's architecture and doing a "post" training refinement. That's what we've been seeing for the past 12-18 months, iirc.

The Codex App 6 months ago

Does anybody know when Codex is going to roll out subagent support? That has been an absolute game changer in Claude Code. It lets me run with a single session for so much longer and chip away at much more complex tasks. This was my biggest pain point when I used Codex last week.

I've been working on decompiling Dance Central 3 with AI and it's been insane. It's an Xbox 360 game that leverages the Kinect to track your body as your dance. It's a great game, but even with an emulator, it's still dependent on the Kinect hardware which is proprietary and has limited supply.

Fortunately, a Debug build of this game was found on a dev unit (somehow), and that build does _not_ have crazy optimizations in place (Link-time Optimization) that make this feat impossible.

I am not somebody that is deep on low level assembly, but I love this game (and Rock Band 3 which uses the same engine), and I was curious to see how far I could get by building AI tools to help with this. A project of this magnitude is ... a gargantuan task. Maybe 50k hours of human effort? Could be 100k? Hard to say.

Anyway, I've been able to make significant progress by building tools for Claude Code to use and just letting Haiku rip. Honestly, it blows me away. Here is an example that is 100% decompiled now (they compile to the exact same code as in the binary the devs shipped).

https://github.com/freeqaz/dc3-decomp/blob/test-objdiff-work...

My branch has added over 1k functions now and worked on them[0]. Some is slop, but I wrote a skill that's been able to get the code quite decent with another pass. I even implemented vmx128 (custom 360-specific CPU instructions) into Ghidra and m2c to allow it to decompile more code. Blows my mind that this is possible with just hours of effort now!

Anybody else played with this?

0: https://github.com/freeqaz/dc3-decomp/tree/test-objdiff-work...

Unfortunately not. It's still very broken, and next year it will be worse for a ton of people. I got AI to write a short answer for you:

Short version: Obamacare never turned into “free primary care for everyone,” it was just a bunch of rules and subsidies bolted onto the same old private-insurance maze. It helped at the margins (more people covered, protections for pre-existing conditions), but premiums/deductibles can still go nuclear if you’re in the wrong income bracket, state, or employer situation. From an EU/Poland perspective it’s not a public health system at all, just a slightly nerfed market where you still get to roll the dice every year.

DeepSeek OCR 9 months ago

There is also a tradeoff between different vocabulary sizes (how many entries exist in the token -> embedding lookup table) that inform the current shape of tokenizers and LLMs. (Below is my semi-armchair stance, but you can read more in depth here[0][1].)

If you tokenized at the character level ('a' -> embedding) then your vocabulary size would be small, but you'd have more tokens required to represent most content. (And context scales non-linearly, iirc, like n^3) This would also be a bit more 'fuzzy' in terms of teaching the LLM to understand what a specific token should 'mean'. The letter 'a' appears in a _lot_ of different words, and it's more ambiguous for the LLM.

On the flip side: What if you had one entry in the tokenizer's vocabulary for each word that existed? Well, it'd be far more than the ~100k entries used by popular LLMs, and that has some computational tradeoffs like when you calculate the probability of each 'next' token via softmax, you'd have to run that for each token, as well as increasing the size of certain layers within the LLM (more memory + compute required for each token, basically).

Additionally, you run into a new problem: 'Rare Tokens'. Basically, if you have infinite tokens, you'll run into specific tokens that only appear a handful of times in the training data and the model is never able to fully imbue the tokens with enough meaning for them to _help_ the model during inference. (A specific example being somebody's username on the internet.)

Fun fact: These rare tokens, often called 'Glitch Tokens'[2], have been used for all sorts of shenanigans[3] as humans learn to break these models. (This is my interest in this as somebody who works in AI security)

As LLMs have improved, models have pushed towards the largest vocabulary they can get away with without hurting performance. This is about where my knowledge on the subject ends, but there have been many analyses done to try to compute the optimal vocabulary size. (See the links below)

One area that I have been spending a lot of time thinking about is what Tokenization looks like if we start trying to represent 'higher order' concepts without using human vocabulary for them. One example being: Tokenizing on LLVM bytecode (to represent code more 'densely' than UTF-8) or directly against the final layers of state in a small LLM (trying to use a small LLM to 'grok' the meaning and hoist it into a more dense, almost compressed latent space that the large LLM can understand).

It would be cool if Claude Code, when it's talking to the big, non-local model, was able to make an MCP call to a model running on your laptop to say 'hey, go through all of the code and give me the general vibe of each file, then append those tokens to the conversation'. It'd be a lot fewer tokens than just directly uploading all of the code, and it _feels_ like it would be better than uploading chunks of code based on regex like it does today...

This immediately makes the model's inner state (even more) opaque to outside analysis though. e.g., like why using gRPC as the protocol for your JavaScript front-end sucks: Humans can't debug it anymore without other tooling. JSON is verbose as hell, but it's simple and I can debug my REST API with just network inspector. I don't need access to the underlying Protobuf files to understand what each byte means in my gRPC messages. That's a nice property to have when reviewing my ChatGPT logs too :P

Exciting times!

0: https://www.rohan-paul.com/p/tutorial-balancing-vocabulary-s...

1: https://arxiv.org/html/2407.13623v1

2: https://en.wikipedia.org/wiki/Glitch_token

3: https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldm...

Since I'm 5+ years out from my NDA around this stuff, I'll give some high level details here.

Snapchat heavily used Google AppEngine to scale. This was basically a magical Java runtime that would 'hot path split' the monolithic service into lambda-like worker pools. Pretty crazy, but it worked well.

Snapchat leaned very heavily on this though and basically let Google build the tech that allowed them to scale up instead of dealing with that problem internally. At one point, Snap was >70% of all GCP usage. And this was almost all concentrated on ONE Java service. Nuts stuff.

Anyway, eventually Google was no longer happy with supporting this and the corporate way of breaking up is "hey we're gonna charge you 10x what did last year for this, kay?" (I don't know if it was actually 10x. It was just a LOT more)

So began the migration towards Kubernetes and AWS EKS. Snap was one of the pilot customers for EKS before it was generally available, iirc. (I helped work on this migration in 2018/2019)

Now, 6+ years later, I don't think Snap heavily uses GCP for traffic unless they migrated back. And this outage basically confirms that :P

I think the better comparison, for consumers, is how fast is LPDDR5 compared to the normal DDR5 attached to your CPU?

Or, to be more specific, what is the speed when your GPU is out of RAM and it's reading from main memory over the PCI-E bus?

PCI-E 5.0: 64GB/s @ 16x or 32GB/s @ 8x 2x 48GB (96GB) of DDR5 in an AM5 rig: ~50GB/s

Versus the ~300GB/s+ possible with a card like this, it's a lot faster for large 'dense' models. Yes, even an NVIDIA 3090 is ~900GB/s of bandwidth, but it's only 24GB, so even a card like this Xe3P is likely to 'win' because of the higher memory available.

Even if it's 1/3rd of the speed of an old NVIDIA card, it's still 6x+ the speed of what you can get in a desktop today.

Recent reputation, yes. But their old reputation was very positive. They made cars that would survive in any condition (which is why they were popular for military uses).

These days, you're in one of two camps: Either you still believe (because you're ignorant or value the Jeep brand more than you value a reliable vehicle) or you've read the recent reviews and steer clear.

Jeep has been duking it out for the bottom of Consumer Reports ratings for a while now, yet they still seem to sell cars. As they continue to betray their loyal customer base though, I imagine this will change. I wish American car companies were better!

Scale AI | Product Security Engineer | TypeScript, Node, Python, AWS | Full-time | Hybrid in San Francisco, CA or NYC or Remote

This is for a team that I'm working closely with and helping grow. We're hiring for somebody with a hybrid SWE and Security background to help scale up the team.

It's currently one engineer who's got a ton on his plate, so the ideal person for this role is somebody who's interested in learning, digging in deep to fix issues (especially shipping PRs), and helping shape the future roadmap of the team. It is not an analyst role.

The role is primarily targeted at mid-career folks with a few years of experience (2+ years minimum). Those that are Senior/Staff level folks, especially those with a strong SWE background that are curious about security, are encouraged to apply if this role resonates strongly. (I personally transitioned from SWE -> ProdSec, years ago, and several other folks on the broader Security team have non-Security backgrounds too.)

Here is the role: https://scale.com/careers/4602047005

Feel free to apply there, email me with you resume (on my profile), or add me on LinkedIn[0]. I'll try my best to answer questions and reply to everybody, but sometimes there is so much inbound that it isn't possible. Thanks!

0: https://www.linkedin.com/in/freewortley/

Claude Code 2.0 10 months ago

That's not just them saving it locally to like `~/.claude/conversations`? Feels weird if all conversations are uploaded to the cloud + retained forever.

I really want to know how many search requests are being made by ChatGPT and other AI systems in 2025. I know OpenAI has a partnership with Bing for this, but then I see OpenAI in the list on the post.

Do we know if they're sending the referer header? Maybe there is no way to know. It would just be interesting to see that trend over time.

This is still a hard problem today. Some hard tech was built for this. I'm excited for a world where this is more accessible and less hardcore than something like CRDTs (in terms of accessibility).

How have others noticed the world shifting in the past 6 years?

This distro seems like a fork of JELOS based on it literally saying "Just Enough Linux Operating System (ROCKNIX)" on the front page. That spells out JELOS, not ROCKNIX lol.

I believe JELOS did die, so this is cool to see. I'll try flashing it to a new SD card and seeing what's up! Does anybody see info on what's difference? Also, is there any indication that this is a fork anywhere?

Something in the article that I had to look up that might bother others. He uses the term 'DCT' in this sentence, but it's never defined in the article. AFAIK it stands for 'DRAM Memory Controller', but that could be an LLM hallucination. Running a web search defines it as Discrete Cosine Transform. :P

"AMD’s BIOS and Kernel Developer’s Guide (BKDG) indicates there’s a 4-bit read pointer for a sideband signal FIFO between the GMC and DCT, so the “Garlic” link may have a queue with up to 16 entries."

Should maybe swap DCT in for MCT (memory controller)?