HN user

BoorishBears

7,213 karma
Posts25
Comments5,152
View on HN
seed.bytedance.com 5mo ago

Seedream 5.0 Lite – Deeper Thinking, More Accurate Generation

BoorishBears
2pts0
www.theatlantic.com 6mo ago

'Askers' vs. 'Guessers' (2010)

BoorishBears
210pts165
cas-bridge.xethub.hf.co 7mo ago

DeepSeek-v3.2: Pushing the Frontier of Open Large Language Models

BoorishBears
3pts1
seed.bytedance.com 10mo ago

Seedream 4.0

BoorishBears
4pts0
old.reddit.com 11mo ago

I Was Bored So I Got Windows XP on the Tablets in My Local McDonald's

BoorishBears
3pts0
www.tryspellbound.com 1y ago

Show HN: Dating sim where Silicon Valley CEOs fight over you

BoorishBears
1pts0
keithwiley.com 1y ago

An Example of the Pace of Modern Technological Change (2003)

BoorishBears
2pts0
www.youtube.com 1y ago

OpenAI announces Advanced Voice with Vision [video]

BoorishBears
13pts1
use.expensify.com 2y ago

The secret experiment behind the Expensify Lounge

BoorishBears
3pts0
github.com 3y ago

Show HN: Guiding LLM outputs using Zod

BoorishBears
3pts0
news.ycombinator.com 3y ago

Ask HN: Why did Python break imports in 3.10?

BoorishBears
1pts3
web.archive.org 3y ago

Public APIs maintainers signaling a corporate sponsor has hijacked the project

BoorishBears
10pts3
news.ycombinator.com 3y ago

Ask HN: Product Ready SBCs

BoorishBears
2pts0
twitter.com 3y ago

Facebook feed algorithm currently broken

BoorishBears
1pts1
www.hofmeisterkink.com 3y ago

BMW's application for finding roads that resemble the Hofmeister Kink

BoorishBears
6pts6
news.ycombinator.com 4y ago

Stripe Treasury/Issuance Alternatives?

BoorishBears
1pts1
mullvad.net 4y ago

Mullvad: Pay with Cash

BoorishBears
19pts6
www.reddit.com 5y ago

Camera-Only Autopilot Requires Automatic High Beams Enabled

BoorishBears
88pts99
old.reddit.com 5y ago

Competitor extracts developer dev key, abuses it to get them banned by Google

BoorishBears
5pts0
word.uservoice.com 5y ago

Word suggests deleting unsaved documents as a default action

BoorishBears
1pts2
github.com 7y ago

Slack doesn't allow user tokens to add emojis

BoorishBears
1pts0
www.cnn.com 7y ago

Tesla is raising prices after backtracking on store closures

BoorishBears
2pts1
9to5mac.com 8y ago

Apple delaying HomePod smart speaker launch until next year

BoorishBears
27pts38
news.ycombinator.com 9y ago

Ask HN: “Accelerated C++” Book Equivalent for Go?

BoorishBears
1pts0
news.ycombinator.com 9y ago

Ask HN: What antivirus do you run on Windows, if any?

BoorishBears
5pts15

"the premium of in person" is a string of words that only means something to people in a very specific circle where everyone is constantly trying to tastemake and write in lower case.

Ironically they have extremely limited influence on how the larger world moves: normal people won't meaningfully change their habits around eating good food or meeting up with each other just because AI can generate a picture of either.

Why do you think an inference provider competing for the same compute as OpenAI and Anthropic gives up those margins rather than giving a modest discount over the frontier for near-frontier performance?

If they had pivoted hard to digital they could have been what today?

Having the first portable-ish digital camera they could have seen the true value of Fairchild's CCD business, got a stake/bought it/replicated done whatever it took to push the frontier of digital imaging and became the Kodak (old, film-era Kodak) of electronic imaging?

There's no comparable business today because all the things they could have invested in ended up being taken up by different companies, they had the R&D culture, revenue, and distribution. I don't see why they couldn't have been category defining.

I've found if you pay close attention, models in Codex get briefly lost on where it is in the plan post-compaction. If it was in the middle of a test it often tries to resume the test from the start, or will seem "surprised" that past steps are completed already.

It's hard to believe that sort of confusion doesn't hurt performance a bit, the question is if that degradation is worse than the performance fall off from long-context (which is highly task specific)

https://qwen.readthedocs.io/en/latest/training/ms_swift.html

Qwen cares enough about model identity that their training framework and docs include a preset for training on it complete with a targeted dataset: https://huggingface.co/datasets/modelscope/self-cognition

And people get Claude to claim it's Deepseek by asking in Chinese.

I can't believe we're still at the "I asked the model who it is" stage of LLMs nearly 4 years out from models calling themselves GPT by OpenAI.

Codex Resets 4 days ago

a) Why and b) With what compute?

"Why" as in, why take lower margins when Moonshot currently can't service all the demand for the model anyways. Based on past models no one is going to massively undercut Moonshot: few have the chops to serve it as efficiently as Moonshot and of those few, most of them don't go for being the cheapest, they go for being fast + reliable (think Together, Fireworks).

You get what you pay for applies very much with how many axes there are to serving these increasingly large models.

-

And for "with what compute": as the value of a token goes up, what people are willing to pay for compute is going up.

Every once in a while I'll see a story about falling rental rates, but with even slightly more established clouds I've been seeing availability get worse and worse over time.

I'm pretty sure the only reason the highly informal indexes don't reflect this is because every neocloud trying to cash in on an NVIDIA Inception discount kicks off by selling unrealistically cheap compute for a bit.

Codex Resets 4 days ago

Everyone is raising the bottom. Kimi got 60% more expensive during the 2.x cycle despite staying the exact same size.

Now K3 is almost 6x the cost of the original K2 checkpoint, and while the parameter count finally jumped, it's still an extremely sparse MoE and definitely does not cost 6x what the original K2 checkpoint did to host at scale.

Race to the bottom only takes real effect when there's a cap to the capabilities, otherwise everyone races to the bottom of a rising target (how economically valuable the tokens are)

So we don't leave it all up to the parents: parents can give it, but minors also can't buy it regardless of parental views.

Also give it to your kids too often and the state can step in.

Defense in depth

I see that being about as likely as IBM doing the same.

I'm too young to have seen the arc of Xerox PARC first hand, so to me stuff like the innovators dilemma made sense, but didn't feel particularly visceral.

AI Studio (or whatever their internal name is) is the first time in my own lifetime witnessing how real deep it really cuts.

Google realizes GCP is too slow and overwrought to get mindshare vs nimble OpenAI and Ant, spins up a new product org as a work around, and that product org ends up moving obviously much faster than Vertex, but also with a sort of malaise (relative to the technology they're supposed to be selling) that makes it clear to me that Google actually cannot move like a startup anymore.

That might seem super obvious to most people, but I grew up with Google being the startup. I knew they grew up to be a mega cap, but I guess I always assumed the bones of a startup were still in there.

There's no bones. It actually feels like a mini-identity crisis for myself to realize there is no startup left in Google: what other invariants I assumed about people and organizations are just plain wrong?

Yeah but it still sounds like $1,500 demand fee was not a mistake, charging OP was the mistake.

People first mentioned these fees at like $100... now you can go to Costco, grab a Starlink, and they'll randomly ask for $1,500 to actually start service.

Congestion is a thing when each node needs to go to space, but it also feels like they're cashing in on years of "just move to a cabin in the woods and work off Starlink"... once you've done that they have you by the balls even worse than a typical ISP.

I don't think this is it. The "constitution" still gets a lot of talk and was brilliant marketing, but with how far modern postraining goes, I doubt they're screwing up rewards with too much of that.

But Sol actually has the same obsession with honesty: I suspect it's more an artifact of trying to control reward hacking.

Models will lie, obfuscate, and mislead under the pressure of RL, so both OAI and Ant are probably forced to spend a lot of time coaxing "honest" answers out of the model

OpenAI's recent prompt for a math conjecture hints at a lot of it when instructing on subagents: https://cdn.openai.com/pdf/04d1d1e4-bc75-476a-97cf-49055cd98...

Articles like this stress me out because they make me feel wildly out of touch by proxy

Like on most other technical opinions this person has I probably agree, but they're so absolutely wrong (for the majority of modern consumers you will fail without an app) that it makes me wonder what other seemingly obvious things we're both completely wrong about.

For the longest time they were a piñata for free compute with people making multiple accounts for their free ARM instance, but with the AI crunch they're clamping down.

I'm guessing they don't care if actual business gets caught up in that because from their POV actual business comes from an account manager, and self-serve is just them cargo culting AWS/GCP

30% was set when they were handpicking every title, a home internet line today was a $10,000+ a month DC connection, and they could legitimately replace a publisher taking 60%.

The fact they whittled away the value they provided time and time again until they became a market of slop and had the audacity to keep a 30% cut is insane.

It's funny that gamers villainize Sweeny for being the person that they think Newell is. It turns out trying to deliver value in a market has tough as gaming is not easy, and you will make tons of mistakes... at least compared to extracting nearly every dollar you can and leaving a skeleton crew to run the ship.

And I guess make $1,100 PS4s as a side hustle.

The LWN content-management system contains over 750,000 items (articles, comments, security alerts, etc) dating back to the adoption of the "new" site code in 2002. We still have, in our archives, everything we did in the over four years we operated prior to the change as well. In addition, the mailing-list archives contain many hundreds of thousands of emails.

Does that sound like your typical self-hosted blog?

If they were on open weights, at some point the provider would deprecate it, probably with worse notice

And self-hosting would probably have been more expensive unless they had massive volume: Deepseek V3.x would have been the comparable open weight model for the performance and isn't that cost effective until hosted across multiple nodes with large batch sizes

No one's firing up a residential proxy to read your blog, and the corporate AI scrapers have all the resources in the world even without residential proxies

They're most useful for getting information from the cloud hosted sites that hoarde most of humanity's output today like Youtube and Reddit.

"Raises the question of what we see is real"

No they really don't, dishonest founders do that.

You're one with the lower case shibboleth so I have no doubt you surround yourself with dishonest founders, but faking users is pretty damn low on the usecases for residential proxies.

I said they're unethical because they tend to be hidden in innocuous seeming apps or sprung on unwitting individuals via clickwraps on their smart devices.

Hy3 13 days ago

The first thing I read from you was a sardonic browbeating in response to the exact comment I gave an earnest response.

And even in domains that lean heavily on "usual phrasing", like technical writing, human writing has notably higher perplexity compared to another LLM's outputs: https://www.sciencedirect.com/science/article/abs/pii/S10766...

With such a low baseline for what's unusual, you do need to get the LLM writing unusual phrases relative to its baseline. Otherwise you get things like repeated n-grams and overused constructs ("it's not X it's Y"), and suddenly the output is predictably not perceived as creative by humans even if you were to insert some otherwise creative or novel premise.

Getting the model to break out of that baseline without disrupting the model's ability to follow technical rules, maintain logic and reasoning, etc. is the difficult part.

-

Also you're again saying unsupervised then following up with descriptions that sure sound like you're referring to RL and supervised learning respectively this time. (supervised learning can improve creativity by the way

Hy3 13 days ago

Did you try reading the whole comment?

Once creativity is being measured in isolation, getting multiple responses from the model is enough to measure creativity a ton of different ways: wordfreq to identify overused phrases, getting multiple responses for the same prompt and promoting the least similar as preferred for policy optimization, etc.

But that's of limited use for stuff like getting diverse names and such. You want creativity and coherency, and if you just punish the model for using an overused phrase, the first thing it does is strongly learn a new overused phrase (or gibberish).

(Also I don't think you mean unsupervised. You probably mean without humans [since LLMs struggle to judge creativity], but that's not what unsupervised means.)

Hy3 13 days ago

The largest model I've post-trained in the last 2 years of working on this problem was Kimi 2.5 at 1T parameters.

The simplest way I'd put it is, teaching a model to write coherently (follow rules, patterns, etc.) is easy enough: just use teacher forcing. Teaching a model to write creatively is easy enough: just use RL and punish it for not being creative.

Teaching a model to write well and creatively takes learning two partially opposing objectives that spike the learning requirements in ways that smaller models really struggle with.

GPT-5.6 13 days ago

I'm pretty sure Altman has spoken about giving a model 100k+ A100s specifically, this might be them being very literal