HN user

practice9

539 karma
Posts3
Comments264
View on HN

It's because HN is in AI meta-psychosis :)

Our experience is very similar except we didn't really have a review process before, and now LLMs find bugs before PRs get merged in main.

We had 5x-100x speedups in some legacy but important pipelines, with no regressions (validated after extensively by humans). It's not that the code was actively bad. It's just only 1-5% people in the local SWE market would be able to write code that runs so fast and efficient and benchmark it correctly.

We found a subtle correctness bug that was in production for half of the decade (both GPT-5 and Claude Opus were able to find it), confirmed by human after.

And we keep finding subtle bugs that have been introduced by humans before (despite the human reviews, the particular domain is just difficult no matter how many docs and comments and tests one writes)

The human is a bad co-author here really.

I deployed lots of high performance, clean, well documented etc code generated by Claude or o3. I reviewed it wrt requirements, added tests and so on. Even with that in mind it allowed me to work 3x faster.

But it required conscious effort on my part to point out issues and inefficiencies on LLMs part.

It is a collaborative type of work where LLMs shine (even in so called agentic flows)

Well the system prompt is still the same for both models, right?

Kinda points to people at OpenAI using o1/o3/o4 almost exclusively.

That's why nobody noticed how cringe 4o has become

But who is the target group?

Last time only some groups of enthusiasts were willing to work through bugs to even run the buggy release of Gemma

Surely nobody runs this in production

I tried the square example from the paper mentioned with o1-pro and it had no problem counting 4 nested squares…

And the 5 square variation as well.

So perhaps it is just a question of how much compute you are willing to throw at it

Well none of the labs have good frontend or mobile engineers or even infra engineers

Anthropic is ahead in this because they keep their UIs simplistic so the failure modes are also simple (bad connection)

OpenAI is just pushing half baked stuff to prod and moving on (GPTs, Canvas).

Find it hilarious and sad that o1-pro just times out thinking on very long or image-intense chats. Need to reload page multiple times after it fails to reply and maybe answer will appear (or not? Or in 5 minutes?). Kinda shows they’re not testing enough and “not eating their own food” and feels like chatgpt 3.5 ui before the redesign

One of those guys needs to be fined for the pump & dump scheme (with SPCE: Virgin Galactic), and the other one should be investigated if he was receiving money from the Russian government or influence agents.

The interesting thing is that ship was damaged almost immediately after leaving the port, had a chance to stop in Russian ports along the way but instead is doing a tour near EU countries.

The crew is either amazingly incompetent or malicious/complicit.

You don’t put a ship that can blow up near: a. Gas&oil terminals, b. military air base

GPT-4o 2 years ago

It is cringe overenthusiastic, but a proper instructions/system prompt will fix that mostly

GPT-4o 2 years ago

I always double-check even the most obscure facts returned by GPT-4 and have yet to see a hallucination (as opposed to Claude Opus that sometimes made up historical facts). I doubt stuff interesting to kids would be so out of the data distribution to return a fake answer.

Compared to YouTube and Google SEO trash, or Google Home / Alexa (which do search + wiki retrieval), at the moment GPT-4 and Claude are unironically safer for kids: no algorithmic manipulation, no ads, no affiliated trash blogs, and so on. Bonus is that it can explain on the level of complexity the child will understand for their age

Meta outage 2 years ago

Interesting. I had a problem a few months ago with DNS not resolving Meta servers on my Starlink internet connection, but I was able to use the UI and the apps nonetheless, just couldn't open the store or update firmware.

Seems like they really did change something in the latest firmwares.

Meta outage 2 years ago

Unless it's a recent change, it works perfectly fine offline (wifi turned off).

As for alternatives, there is Pico, but Quest 3 may be superior in games selection. Or go wired which is of course less portable

Here is an interesting and relevant context: there is a huge amount of evidence that Roscosmos is taking an active part in the war effort. There is a great source about it from Eric Berger of Ars Technica: https://arstechnica.com/space/2023/06/it-appears-that-roscos...

Take that into account when reading news made from state press-releases like the one in the post.

P.S.: For the full context, the Roscosmos ex-boss also has had his own private military company for a while.

It was trained to respond like that to not alienate groups of people. But fine-tuned and with another pre-prompt, it would give absolutely different answer.

Choose a better government, then. Preferably, a one without Soviet/death cult/imperialistic tendencies.

The nation will be judged by its government's actions, because the gov is supposed to be kept in check by the people.