HN user

postalcoder

3,336 karma

developing llm systems

side: https://hcker.news

ryan [at] hcker [dot] news

Posts3
Comments346
View on HN

I wonder how big the Pro model is that Google is using behind the scenes to train these smaller ones.

Going on baseless speculation, the lack of accompanying pro models with these flash releases either means: 1) the model is too big to be economical, 2) google doesn't have the compute to serve the big model, 3) their big model has too many alignment issues to serve to the public.

edit: looks like benchmarks are up on https://artificialanalysis.ai/models/gemini-3-6-flash. It's solidly middle-of-pack. However, if you want to be most fair to flash, look at the intelligence vs time per task and intelligence vs outputspeed benchmarks. This is a very fast model.

edit 2: I use antigravity from time to time and in my experience, 3.5 flash is an underrated model, so long as you know what it's good for. It's very good at frontend (much better than gpt 5.5) and it's fast, so it's a great tool for iteration. I expect 3.6 to be no different.

I've been thinking about this too much because it's so ridiculous and funny that I had to at least try wrapping my head around it. My best guess is this:

If you look at past snapshots at archive.org, you notice that the meta keywords are growing like an append-only list, which means it's probably part of some messed up seo pipeline. The other clue is that it's stuffing the meta keywords which apparently only Yandex uses as a search signal[0].

The list has 3882 entries. A lot of them are clustered and look like auto-complete results. But a bunch of them look like hyper-specific, misspelled search results (e.g. "145 gwen rd cheshire ct"). Google's webmaster tools doesn't provide distinct queries like that, but yandex's does[1].

My best guess of what's happening is that Qwen is monitoring it's search queries in Yandex, dumping that list into a serp service that scrapes yandex's autocomplete suggestions, and then taking that list and dumping it into their meta keywords.

It explains the urls in the list (ppls using search engines like address bars), the seeming fixation around certain topics (which usually starts with a misspelling), and the random one-off queries.

0: https://yandex.com/support/webmaster/en/controlling-robot/me...

1: https://yandex.com/support/webmaster/en/service/popular-quer...

I’m grateful for this team. Jellyfin is a great product. I’ve been using llm + Tailscale + Jellyfin to manage my media box and I couldn’t be more happy. I’ve got perfect metadata, perfect organization, and multimodal search on top of it. What a dream.

Kimi Work 2 days ago

Yes I’m saying a question answered deceptively should not be one of the four featured questions on your product page.

Looking through their privacy policy, it looks like the bigger issue is that they are (or refuse to refute) training on your data/prompts.

You’re making my entire point. They’ve designed an ingestion engine that is not private by design. Any sort of privacy posturing is wrong.

Kimi Work 2 days ago

Yes, but the answer is an indirection.

“How does Kimi Work protect my privacy when accessing local files?”

It doesn’t protect your privacy. It’s like asking, “how do I know you’re not spying on me?” And getting the reply “we cannot physically enter your house.”

Kimi Work 2 days ago

kimi has been particularly shameless in copying codex 1:1, wow.

Like, I get it, before condex there was conductor and a ton of other guis but this looks like they just copied the code.

Also, their privacy disclosure is incredibly misleading:

How does Kimi Work protect my privacy when accessing local files? You have absolute control over your files. The built-in Ask before acting safeguard means Kimi will prompt you for explicit authorization before it modifies, overwrites, or runs code within your local directories. Nothing happens without your consent.

They fail to mention that is has unfettered read access to your files. ie, without your consent.

This is a very strange article considering that Llama, the mother of all open-weight models, has led to anything but success for Meta.

Also, enterprises don't give a rip if models are open. They care about zero data retention (and sticking with whatever vendor they're already using).

This blog post is suspiciously close to being a restatement of what Alex Karp recently said on CNBC[0]. It's important to remember he's the CEO of Palantir and hardly a neutral observer.

There are many reasons to celebrate open models, I run them myself. However there's not yet enough evidence that 1. America is losing the AI race (pardon jingo-ey phraseology) and 2. American AI labs are losing because their models are not open-weight.

0: https://www.cnbc.com/2026/07/01/palantir-karp-open-ai-anthro...

Apple knew they were selling a device to play pirated music.

On the flip side, Sony lost the consumer devices market for this very reason. Sony's single-minded pursuit of proprietary formats was a disaster class of corporate mismanagement.

It disgusts me because I used to love their products so much. Sony's competitor to the iPod was a marvel of a device called the NW-HD1. It was beautiful, had a ton of space, and great battery life. But it wasn't an MP3 player. It could only play ATRAC music. That means you had to transcode all of your MP3s to their proprietary format just to listen to them.

I remember trying to debate the virtues of my Sony NW-HD1 versus the iPod, but having to keep my computer on throughout the night just to transcode a couple albums was indefensible.

This is a wonderful piece. I really appreciate the author for writing this – I thought I lost the ability to read longform on the internet.

What.cd was so vast a resource that it means something different to everyone. I personally lament the loss of the forums the most. I would post dissertation-length comments there and others reciprocated. I would put hours of research into debating a single topic. It's where I probably wrote my best stuff. The high barrier to entry reduced the noise and selected for people who were invested in being part of a community. The forums are also where I learned about hacker news!

I learned so much about music during those days. Algorithmic recommendations don't hold a handle to the recommendations you'd get in the forumns and in the comments sections of individual albums. Consuming music via What was equal parts learning and consumption.

It was obvious that poor music sales was a distribution problem, not a piracy problem. History played out in a way that proved this to be true. Spotify killed What.cd before the French did.

It's why I like Gemini 3.1 Pro. That it sounds much more human than other LLMs is testament to Google's inability to post train.

gemini-2.5-pro-experimental was the GOAT, though. It was an emotional wreck, down in the dumps and feeling terrible for itself after failing to patch a file several times. Very amusing to read, all the while watching it make a mess of my codebase.

Precursor 9 days ago

Gonna zag here. If you take a step back, cloudflare has been paving the path for pay for crawl. I think it's a noble and ambitious goal.

While I can understand why you would be alarmed, I can point to almost two decades of lamenting on this forum about how we need better ways of rewarding content creators than ads. Well, this is it.

Moreover, these products weren't built in a vacuum. Most threads about Anthropic and OpenAI have complaints about how these companies were built on stolen content. There's a reality we have to face here and we can't have it both ways.

ChatGPT Work 13 days ago

Going to repost my comment from below:

Codex creates a new folder in `~/Documents` in iCloud drive/OneDrive for every single thread you make. Furthermore, these threads are also polluted with your global AGENTS.md file, as well as all other developer_instructions that are injected by the harness.

I want my chats isolated from work contexts.

ChatGPT Work 13 days ago

Even though it says Work or Codex or whatever you can just ask it whatever you want, and it will work.

No it's not. Codex creates a new folder in `~/Documents` in iCloud drive/OneDrive for every single thread you make. Furthermore, these threads are also polluted with your global AGENTS.md file, as well as all other developer_instructions that are injected by the harness.

I want my chat isolated from work contexts.

ChatGPT Work 13 days ago

I just installed this. I am very confused. I no longer have a Codex app on my computer. ChatGPT is now Codex.

But what happened to ChatGPT? Where am I supposed to casually chat?

Also, when you toggle btween ChatGPT Work and ChatGPT Codex, nothing changes. This is super confusing. Can someone from the OpenAI team clarify the difference btwn the modes? Does chatgpt work have more business-y related plugins turned on by default?

Edit: So it seems like the only place you can actually chat with chatgpt is in an awkward homeless nested window. idk. The chatgpt interface wasn't great (desperately needed artifacts), but I still used it a lot. I can't see this change going well with a lot of the casual users.

Edit2: In their awkward homeless nested chat mode, you cannot even edit past messages. this is a mess, why was the team so zealous to pull the switch on unification in this state? guessing there was internal pressure to juice codex's growth but, based on what im seeing, they did it by torching chatgpt?

Edit 3: Ok so it seems like ChatGPT is still around, but renamed to "ChatGPT Classic". Seems like it wont be long for this world because there's no place to download ChatGPT Classic should you choose to uninstall it. The dmg at https://chatgpt.com/download/ only contains the new ChatGPT.

GPT-5.6 13 days ago

Avoid generic brevity instructions: GPT-5.6 is more sensitive than GPT-5.5 to instructions such as “Be concise,” “Keep it short,” or “Use minimal text.”

RIP Caveman skill. Six month good. Now skill dead.

GPT-5.6 13 days ago

There is so much less drama involved with the Codex world. You don't realize how oppressive CC is until you've escaped it. Outages, weird restrictions, degradation, accelerated usage, etc etc etc.

Muse Spark 1.1 13 days ago

Not their first try. There’s been reporting about how they’ve kept pushing their model releases back because of underwhelming performance.

I did find the bit about ChatGPT's crappy prose amusing because while that may feel like it's a good benchmark for the state of AI, the quality of prose doesn't necessarily correlate with the major progress made over the last few years which is in post-training.