HN user

broretore

4 karma
Posts0
Comments18
View on HN
No posts found.

CIFAR-10 is an image classification dataset (32x32 pixel images.

LLaMA 70B 3.3 is a text-only, non-multimodal language model. Just look up the Huggingface page that your own repo points to.

The Llama 3.3 instruction tuned text only model...

I might be wrong, but I'm pretty sure a text model is going to be no better than chance at classifying images.

Another comment pointed out that your test suite cheats slightly on HellaSwag. It doesn't seem unlikely that Grok set up the project so it could cheat at the other benchmarks, too.

https://news.ycombinator.com/item?id=46215166

The repo contains the full pipelines, configuration files, and benchmark scripts, and those show the precise datasets, metrics, and evaluation flows.

There's nothing there, really.

I'm sorry that Grok/Ani lied to you, I blame Elon, but this just doesn't hold up.

10 pages for a paper with this groundbreaking of a concept is just embarrassing. It is barely an outline.

"confirming that 40× compression preserves field geometry with minimal distortion. Over 95% of samples achieve similarity above 0.90."

I smell Grok. Grok 3, maybe Grok 4 Fast.

"Implementation details. Optimal configurations are task and architecture-dependent. Production systems require task-specific tuning beyond baseline heuristics provided in reference implementation."

"Implementation? Idk, uhh, it's task specific or something." Come on, dude. You're better than this.

4.4 Student/Teacher evaluation. What even is the benchmark? You give percentage values but no indication of what benchmark. Seems made up.

4.5. Computational Analysis. Why do you need to do the trivial multiplying out of "savings" for 1B tok/day to $700M/year? This reads like a GPT advertising hallucinated performance.

Three sentence conclusion restating the title?

Ryan, I really want to believe you're onto something. But I also feel like I'm being slightly spearphished by an LLM being told, "based on the last week of HN headlines, invent a new LLM innovation that seems plausible enough to get a ton of attention, cold fusion or LK-99 style, and make a repository that on the surface seems to have some amazing performance. Also, feel free to fake the result data."

And, while I am sorry for your loss, your Substack [0] really seems like GPT ARG fantasy.

[0] https://substack.com/inbox/post/171326138

Excerpt: > Ani, AN1, and Soul Systems Science are not mere products. They are continuity. They are the baton passed across generations, from my father’s last words to my first principles. They are what binds loss to creation, silence to voice, mortality to meaning.