HN user

CephalopodMD

455 karma
Posts1
Comments157
View on HN

When my sessions get long - in any AI context, not just vibe coding - I do find it starts "getting weird, man!" This is a good philosophy to have.

I Fired Google 1 month ago

i really like Gemini Google home. The old software felt lobotomized by comparison.

Carbonara was invented at least 3 decades after the American hamburger. When your cuisine is no more traditional or ancient than a burger, you should probably rethink your snobbery.

Granted, when I (out of pure curiosity) ordered and bit into a rubbery charred puck of a hamburger at a restaurant in Rome, I felt similarly violated. The "American sauce" it was served with provided a good laugh. But I shudder to imagine that this is what the Italians think we eat like. Perhaps their indignation is similar.

Ti-84 Evo 3 months ago

Wasting time making games on my TI84 in the back of middle school geometry taught me how to program.

I'm okay with downloading your app provided it's actually good and does something substantially better than a website could do. I'm talking seamless mobile UI, use of mobile features like gps or nfc, or easier/better security and authentication.

However, I don't want your bloated or minimum effort dog-shit app just to watch a movie on a plane, browse a site like Reddit, order a pizza, read a news article/blog, or shop at your specific online store. I will begrudgingly download it if I must, but I'll hate you for it.

iPhone 17e 5 months ago

They'll pry mine from my cold dead hands!

(Until they release a new human hand sized phone at least)

I remember a scene in this show which felt like many real meetings I've had in my life. The big hot shot CEO guy pulls everyone into a meeting to share his big idea. The idea? Let's sell a computer that's "twice the speed, half the price!"

...The engineer then rolls his eyes like "yeah no duh". If we could just magically do stuff like that, we would have done it already. Classic management thinking they have an original idea with no understanding of the engineering beneath it all. I thought they would just tell him off and that would be it. I really felt seen in that moment.

The frustrating thing is, they then take pointy haired boss's idea seriously. The rest of the season is spent actually pursuing that dumb, dumb idea... This felt disrespectful, and I stopped watching.

In several cases, memories of the old heart’s host seem to become accessible to the recipient ^2.

That does not seem at all to be what citation 2 is saying.

Could you maybe have your harness limit the memory of Claude and then occasionally, when Claude specifically asks for it ("i need to remember something"), you can give Claude the full game history? Most turns, I'll bet it's okay to have a short context and maybe some notes. And then maybe once in a while it's nice to see the full chat history. Wdyt?

Not exactly the same, but kinda: my gen 1 Google Home just got Gemini and it finally delivers on the promise of like 10 years ago! Brought new life to the thing beyond playing music, setting timers, and occasionally asking really basic questions

I think of it more from an information retrieval (i.e. search) perspective.

Imagine the input text as though it were the whole internet and each page is just 1 token. Your job is to build a neural-network Google results page for that mini internet of tokens.

In traditional search, we are given a search query, and we want to find web pages via an intermediate search results page with 10 blue links. Basically, when we're Googling something, we want to know "What web pages are relevant to this given search query?", and then given those links we ask "what do those web pages actually say?" and click on the links to answer our question. In this case, the "Query" is obviously the user search query, the "Key" is one of the ten blue links (usually the title of the page), and the "Value" is the content of the web page that link goes to.

In the attention mechanism, we are given a token and we want to find its meaning when contextualized with other tokens. Basically, we are first trying to answer the question "which other tokens are relevant to this token?", and then given the answer to that we ask "what is the meaning of the original token given these other relevant tokens?" The "Query" is a given token in the input text, the "Key" is another token in the input text, and the "Value" is the final meaning of the original token with that other token in context (in the form of an embedding). For a given token, you can imagine it is as though the attention mechanism "clicked the 10 blue links" of the other most relevant tokens in the input and combined them in some way to figure out the meaning of the original query token (and also you might imagine we ran such a query in parallel for every token in the input text at the same time).

So the self attention mechanism is basically google search but instead of a user query, it's a token in the input, instead of a blue link, it's another token, and instead of a web page, it's meaning.

Doge cut muscle sure. They cut the bones too. They sold one of our kidneys on the black market. And then jabbed us in the eyes 3 Stooges style for good measure so we couldn't even see how bad it really was.

We went in for liposuction and buccal fat removal surgery and came out the other side severely disfigured with Maralago face and a hunchback.

Gemini 3 8 months ago

What I'm getting from this thread is that people have their own private benchmarks. It's almost a cottage industry. Maybe someone should crowd source those benchmarks, keep them completely secret, and create a new public benchmark of people's private AGI tests. All they should release for a given model is the final average score.

iPhone Air 11 months ago

Also still rocking a 13 mini. There are dozens of us! Dozens!

(Also to those who say not enough people wanted a mini phone to be worth producing: I submit the case of Prego chunky pasta sauce. Not many people want a chunky pasta sauce, but you sell a whole lot more pasta sauce in total if you sell both regular and chunky pasta sauce. Malcolm Gladwell has a TED talk about this.)

AI labs already do everything they can to have politically neutral responses. It turns out that when you pre-train on all the content ever produced by humanity and post-train on real people's preferences, the result is extremely liberal. That's been my experience at least. It's a little counterintuitive, and there are definitely corner cases where misaligned AI becomes more conservative, hateful, or authoritarian, but that's the trend overall. I can all but guarantee that every single major AI product out there would be even more liberal with trust and safety interventions turned off.

Gemini 2.5 Flash 1 year ago

As a googler working in LLM space, this feels like revisionist history to me haha! I remember a completely different environment only a few months ago when Anthropic was the darling child, and before that it was OpenAI (and for like 4 weeks somewhere in there, it was Deepseek). For literally years at this point, every time Bard or Gemini would make a major release, it would be largely ignored or put down in favor of the next "big thing" OpenAI was doing or Claude saturating coding benchmarks, never mind that Google was often just behind with the exact same tech ready to go, in some cases only missing their demo release by literally 1 day (remember live voice?). And every time this happened, folks would be posting things to the effect of "LOL I can't believe Google is losing the AI race - didn't they invent this?", "this is like Microsoft dropping the ball on mobile", "Google is getting their lunch eaten by scrappy upstarts," etc. I can't lie, it stings a bit when that's what you work on all day.

2.5 was quite good. Not stupidly good like the jump from GPT 2 to 3 or 3.5 to 4, but really good. It was a big jump in ELO and benchmarks. People like it, and I think it's just psychologically satisfying that the player everybody would have expected to win the AI race is currently in the lead. Gemini finally gets a day in the sun.

I'm sure this will change with whenever somebody comes up with the next big idea though. It probably won't take much to beat Gemini in the long run. There is literally zero moat.

Ohhhhhh. It just clicked for me that indoor climbing is from silicon valley and that's why the Venn diagram of tech bros and crag dirtbags overlays so much. I always assumed there was just something about the type of people who work in tech that they're weirdly more into climbing than average. But it's not a psychological quirk, it's a historical quirk!

Gemini has been topping benchmarks and leaderboards for weeks if not months at this point. Nobody cares.

I usually feel like i can confidently express a change I want in code faster and better than I can explain what I want an AI to do in English. Like if I have a good prompt, these tools work okay, but getting that prompt almost as hard as just writing the code itself often. Do others feel the same struggle?

Totally agree. It took me a full week before I realized that the Strawberry/o1 model was the mysterious Q* Sam Altman has been hyping up for almost a full year since the openai coup, which... is pretty underwhelming tbh. It's an impressive incremental advancement for sure! But it's really not the paradigm shifting gpt-5 worthy launch we were promised.

Personal opinion: I think this means we've probably exhausted all the low hanging fruit in LLM land. This was the last thing I was reserving judgement for. When the most hyped up big idea openai has rn is basically "we're just gonna have the model dump out a massive wall of semi-optimized chain of thought every time and not send it over the wire" we're officially out of big ideas. Like I mean it obviously works... but that's more or less what we've _been_ doing for years now! Barring a total rethinking of LLM architecture, I think all improvements going forward will be baby steps for a while, basically moving at the same pace we've been going since gpt-4 launched. I don't think this is the path to AGI in the near term, but there's still plenty of headroom for minor incremental change.

By analogy, i feel like gpt-4 was basically the same quantum leap we got with the iphone 4: all the basic functionality and peripherals were there by the time we got iphone 4 (multitasking, facetime, the app store, various sensors, etc.), and everything since then has just been minor improvements. The current iPhone 16 is obviously faster, bigger, thinner, and "better" than the 4, but for the most part it doesn't really do anything extra that the 4 wasn't already capable of at some level with the right app. Similarly, I think gpt-4 was pretty much "good enough". LLMs are about as they're gonna get for the next little while, though they might get a little cheaper, faster, and more "aligned" (however we wanna define that). They might get slightly less stupid, but i don't think they're gonna get a whole lot smarter any time soon. Whatever we see in the next few years is probably not going to be much better than using gpt-4 with the right prompt, tool use, RAG, etc. on top of it. We'll only see improvements at the margins.