HN user

powera

2,395 karma

Formerly of Google and Quip.

Posts55
Comments867
View on HN
stevemagness.substack.com 3mo ago

The Cost of Comfort

powera
1pts0
optimistsbarn.substack.com 3mo ago

Ozempic for Broiler-Breeder Chickens

powera
1pts0
davidoks.blog 3mo ago

Many African families spend fortunes burying their dead

powera
220pts218
ohr.edu 3mo ago

The Quinoa-Kitniyos Conundrum (2019)

powera
2pts0
www.nippon.com 4mo ago

Nippon Life Sues OpenAI over Legal Advice to Ex-Beneficiary

powera
7pts0
news.ycombinator.com 4mo ago

Claude Code on the Web broken?

powera
1pts0
news.ycombinator.com 5mo ago

Ask HN: Best multi-lingual text-to-speech system

powera
2pts0
www.transformernews.ai 5mo ago

The left is missing out on AI

powera
2pts1
www.bbc.com 5mo ago

Suneung: The day silence falls over South Korea (2018)

powera
1pts0
news.ycombinator.com 5mo ago

What Is Genspark?

powera
6pts2
rogerpielkejr.substack.com 2y ago

Apples, Oranges, and Normalized Hurricane Damage

powera
4pts0
huggingface.co 2y ago

Hugging Face and Google partner for AI collaboration

powera
152pts57
shragafeivel.com 2y ago

LLM Embeddings and Outlier Dimensions

powera
2pts0
shragafeivel.com 2y ago

The machine reads a letter about LLMs

powera
1pts0
shragafeivel.com 2y ago

A Letter from Burning Man

powera
1pts0
www.newslettr.com 3y ago

Cryptocurrency, part 3: Time to “cash out”

powera
2pts0
www.newslettr.com 3y ago

Cryptocurrency, part 3: Time to “cash out”

powera
1pts0
shragafeivel.com 3y ago

In which I ask the machine to distinguish violet from purple

powera
1pts0
en.wikipedia.org 3y ago

Ethnic-heritage categories on Wikipedia articles

powera
1pts0
www.newslettr.com 3y ago

Contra LessWrong on AGI

powera
2pts0
techcrunch.com 3y ago

Twitter rival ‘T2’ raises its first outside funding

powera
6pts0
www.nytimes.com 3y ago

The Mystery Behind the Crime Wave at 312 Riverside Drive

powera
2pts0
twitter.com 3y ago

The spam on Twitter

powera
262pts260
www.newslettr.com 4y ago

The Twitter News, Memorial Day Edition

powera
1pts0
www.newslettr.com 4y ago

Reflections on Birdwatch

powera
4pts0
www.newslettr.com 4y ago

Cryptocurrency, Part 2

powera
4pts0
www.newslettr.com 4y ago

Life Without Facebook

powera
4pts1
simblob.blogspot.com 4y ago

Ten Years of Red Blob Games

powera
3pts0
web3isgoinggreat.com 4y ago

Internet shutdown in Kazakhstan disrupts 12–18% of Bitcoin mining

powera
8pts1
twitter.com 4y ago

Don't send your Google phone in for warranty repair/replacement

powera
272pts145

I'm seeing very low quality results on LMStudio with this model. Worse than Gemma 3 12B.

It is getting questions like "David has 18 apples and Ivan has 7 apples. How many apples do they have together?" wrong half the time, while Gemma3 12B could very consistently answer that. Other smoke tests (like Chinese translation, and the infamous "Rs in Strawberry" test) also show poor results.

I don't know if it is a quantization/release issue, if the parameters needed for accurate responses have changed (i.e. it needs "thinking" tokens to handle its base error rate), or if the model has been so focused on audio/video that the text processing is bad.

I'm not sure they've found/understand it yet. My two main theories:

1. A bunch of people with new Claude Code codebases in December now are working with a larger codebase, causing more context. Claude reads a lot of code files, and doesn't effectively prune from the context as far as I can tell. I find myself having to hint Claude regularly about what files to read (and not read) to avoid having 75k of unrelated files in the context window.

2. Claude Code tries to do more now, for the benefit of people who don't know exactly what they want. The trade-off is that it's worse at doing exactly what people want, when they do know. The "small fix" becomes a large endeavor for Claude.

Wikinews never worked; the principle of "verifiability" that Wikipedia was based on simply doesn't work for news-collection, which requires trusted first-party accounts.

The project was also already dead; the English Wikinews has had 10 "articles" posted in the last 3 weeks, two of which were trivial sports stories (a second-division Queensland football match, and the retirement of a pitcher whose last substantial year in MLB was 2019). The most recent story is that an amateur jazz group recently played at a library.

It will no longer be an attractive nuisance to the few who stumble across it. Rest in peace.

So far on my (simple) benchmarks, GPT-5.4-mini is looking very good. GPT-5.4-mini is about 30% faster than GPT-5-mini. GPT-5.4-mini gets 80% on the "how many Rs in Strawberry" test, and nearly perfect scores on everything else I threw at it.

GPT-5.4-nano is less impressive. I would stick to gpt-5.4-mini where precise data is a requirement. But it is fast, and probably cheaper and better quality than an 8-20B parameter local model would be.

( https://encyclopedia.foundation/benchmarks/dashboard/ for details - the data is moderately blurry - some outlier (15s) calls are included, a few benchmark questions are ambiguous, and some prices shown are very rough estimates ).

I've been waiting for this update.

For many "simple" LLM tasks, GPT-5-mini was sufficient 99% of the time. Hopefully these models will do even more and closer to 100% accuracy.

The prices are up 2-4x compared to GPT-5-mini and nano. Were those models just loss leaders, or are these substantially larger/better?

Between this and 4.6's tendency to do so much more "exploratory" work, I am back to using ChatGPT Codex for some tasks.

Two months ago, Claude was great for "here is a specific task I want you to do to this file". Today, they seem to be pivoting towards "I don't know how to code but want this feature" usage. Which might be a good product decision, but makes it worse as a substitute for writing the code myself.

I've been working on a similar app called Trakaido. (Yes, it's also AI generated content).

The demo leans heavily on "choose the words for the sentence", which avoids spelling/keyboard issues, and maybe generalizes around the problems of N->N language maps better. The "decoy selection" for multiple choice answers also isn't great - I am getting sentences mixed with numbers for the translation of "three".

It also has the Duolingo-esque audio "reward" sounds. I personally hate them, but a lot of people feel otherwise.

There's a difference between the chatbot "advertising" something and an hour-long manipulative conversation getting the chatbot to make up a fake discount code. Based on the OP's comments, if it was a human employee who gave the fake code they could plausibly claim duress.

This seems to have a fundamental misunderstanding of Hacker News's front page. The goal is to have posts only show up for at most a day or two. I would expect 100% to be gone in a week.

The "outliers" are posts that were deleted and re-submitted, or possibly boosted by the moderators. The "visible for 44 days" post was only submitted 14 days ago; either it was deleted or it is a glitch in this guy's script.

I own an AppleVision (and am mostly happy with it), but this is not surprising. The current system isn't worth $3500 to consumers.

It is priced at about twice what people will expect, has very limited apps, and very limited "immersive media". Some of that will get better over the next year ... but a promise of "better next year" won't move units today.

[dead] 2 years ago

No 14 year old "publishes 4 books" without over-involved parents.

And she's not even complaining about not going to Stuyvesant or LaGuardia?

This is NYPost trash, not a real concern.

[dead] 2 years ago

I wouldn't call it "trolling", but it does seem to be off-topic. He probably wants conversation about the event, not clickbait.

The "automakers have known for 20 years not to design wireless entry systems that this device can hack" argument is much stronger than the "you can't ban us! we're just doing physics!" argument in the specifically linked tweet.

Not only does this read like pure bullshit, it is bullshit on a website that crashes the Apple Vision Pro (and makes my laptop suffer).

My prediction is that they will raise a nine-figure sum over the next decade, and never release a product that comes close to the performance of an NVIDIA card today.

No, there isn't a shorter version. In fact I would need about 5x the word count to make my argument clearer.

I am still digging through the 54-page paper to try to find the data set for this "death penalty" test to tell if there is anything there beyond "people who use more violent language tend to be viewed as more violent".

They do comment on the dialect issue: << Appalachian English evokes them to a certain extent (m = 0.015, s = 0.030, t(89) = 4.8, p < .001), but much less strongly than AAE (m = 0.029, s = 0.053, t(89) = 5.3, p < .001), a trend that holds for all language models individually (Figure S11, Table S14). The difference between AAE and Appalachian English is found to be statistically significant by a twosided t-test, t(178) = 2.3, p < .05. The fact that Appalachian English is associated with the Katz and Braly (1933) stereotypes to a certain extent is not surprising since the two dialects share many linguistic features (e.g., usage of ain’t), and the stereotypes about Appalachians bear similarities with the stereotypes about African Americans (e.g., lack of intelligence; Luhman, 1990) >>

As far as Mr. Marcus, no. He is too consistently and deliberately anti-LLM (despite any facts) to be worth engaging with.

As far as the underlying research paper: the researchers seem to be conflating "low-status English dialects" with "African American English". In particular, I have never considered the use of the word "ain't" to be associated with a certain race.

If the researchers assume "African Americans are low status" and conclude "African Americans are associated with low-status jobs", the conclusion is entirely about the researchers, not the LLMs.

The research paper's Git repo at https://github.com/valentinhofmann/dialect-prejudice does nothing to ameliorate these concerns.

Of course it is absurd. Mr. Marcus' entire schtick is absurd criticisms of LLMs, and absurd demands of anyone who creates them.

Progress isn't infinite, status is often a zero-sum competition, and exponential economic growth is something that only people who are bad-at-math want (and should they be doing economics?). Also, the whole essay is begging-the-question; so much of the argument is <<we need more people because "more is better".

Two points:

1) The panels being de-commissioned now aren't modern panels, they are 20+ years old. They were substantially less efficient (per square-meter) than today's panels when they were new, and they (probably) are aging worse.

2) A lot of these "the solar panels have to be deconstructed for economic reasons" arguments are the same arguments as "Coyote v. ACME has to be destroyed for economic reasons". The reasons are completely fake; designed purely to feed Moloch.