HN user

AussieWog93

10,287 karma

I come from Australia. I am a Wog. I was born in 1993.

admin@espotek.com

Posts6
Comments2,468
View on HN

A lot of disappointment here in the comments, but models like these aren't meant to compete with the likes of Fable or GPT 5.6.

I use 3.1 Flash Lite regularly to classify listings on eCommerce websites. It's great for this task - fast, cheap and accurate.

In fact, it was the single best model we tried in terms of the speed vs accuracy vs price tradeoffs - including the Chinese models.

Of course, 3.5 Flash was more accurate but the 5x cost increase couldn't be justified.

3.5 Flash Lite sounds like it could be a strict upgrade for our use case, without a significant increase in costs or drop in speed.

It's not GPT-6 but it's not trying to be. It's a completely different tool and great at what it does.

I'm the GP that got Claude to port the app, set up the SMTP server etc. etc. And I do see stuff on that level of complexity, multiple times a week. I'm not a full-time software dev any more, but did learn the craft at a really good company before AI came along - and I've been doing basic sysadmin stuff by hand, almost exclusively on the terminal, for almost a decade now.

A lot of pithy responses to your comments that contain some pretty bad-faith assumptions about AI "power users", but to answer your question I think part of it genuinely does come down to a "skill issue" - or at least familiarity.

AIs really do have huge blind spots, and there are things it can't do well. But on the flip side, there are modes of operation that do get good results and if you keep it in this "zone" you can get amazing results.

For example, here is my CLAUDE.md that's global to all projects: https://github.com/EspoTek/.claude/blob/master/CLAUDE.md

Note the "Working with unfamiliar data or systems" section - it doesn't stop the models from making wrong assumptions but it does get them to test these assumptions with simple experiments and course-correct before human eyes ever see the result.

Setting up things like CLAUDE.md files and skills so that it doesn't make the same mistakes over and over is a big help too. The model doesn't literally get smarter, but it knows how to avoid pitfalls its fallen into in the past.

If you want to have a chat about it earnestly, my email's in my profile.

Usually, the longer the AI works on something the crappier its output because that means the context is getting filled up.

This was definitely truer with older models but isn't necessarily the case now.

They frequently do other things apart from navel-gazing that take a lot of time but get good results, like spinning up subagents to solve some hairy task in a loop.

Stopping/distrusting long-running AIs is a habit I've had to unlearn myself.

I frequently get good results from a 30+ minute Fable session, when I've asked it to do something complex (e.g. run the QA tester in a loop and eliminate all crashes, one commit per crash fixed)

Genuinely, if you want an actual answer: your choice of language basically made us seem like you weren't interested in a good faith, productive discussion with the possibility of minds being challenged/changed, so people just disabled and moved on.

Genuinely, install Ubuntu Server Edition into an old computer, set up SSH then tell Claude Code the account and IP for login.

You can ask if to set up damn near anything you want and it'll do it.

If you're using a laptop make sure it disables all of the sleep on lid close stuff, as well as other laptop-specific power optimisations.

A VPS would work too if you want something always available with a datacentre network connection.

I just don't understand how someone could think we're going back to a pre-LLM world. Have you not used one since Opus 4.5 came out?

Especially as somebody that's no longer in the software profession and so doesn't have 40 hours a week to just sit there and crank out code, agentic AI is a technological leap on the level of search engines, bulletin boards or compilers.

You can literally just open Claude Code, say:

- "Add these domains with these mailboxes to an SMTP server running on my homelab, forward all incoming and outgoing via mxroute"

- "Port this unmaintained app from 2017 to modern Android, and then test it on an emulator and sign it with my real key for distribution via Google Play"

- "get my GitHub Actions runner on Linux, Windows, macOS and Android, ensuring that it supports all available architectures from the last decade and a half, then check your outputs in a loop until it works"

and it just spits it out while you do the dishes.

Someone the other day talked about "AI psychosis psychosis" and it absolutely stuck with me. The people who deny this is a big deal are in la-la land.

We could literally hit a brick wall in terms of model intelligence now and it would still be a game changer.

Qwen 3.8 3 days ago

I went from cycling between models all the time in Cursor (some would randomly be better at certain tasks than others) to just going pure Opus 4.5 when that came out - it was so far ahead of anything else at the time.

Interestingly with Fable vs GPT-5.6 I think they've lost their lead a bit. I'm finding Fable can't do certain work that 5.6 Sol Ultra can - especially when it comes to webpage design.

Grok 4.5 was fast but made mistakes that GPT/Fable just don't.

I'm curious to try Kimi.

Had the same thought. Since around 2022-2023, most subreddits insta-ban you or hide your posts unless you have x karma (you gain karma by posting and getting upvotes).

So unless you happen to have an old account with karma from pre-2022, the only new users whose posts are actually getting through are the ones that know the specific subreddits that let you "karma farm" - ie, mostly bots and bad actors.

Also tangentially related, much of the weirdness of Reddit ("Thanks for the gold, kind stranger", fedora-tipping, narwhals baconing at midnight, references to weird memes etc.) also happened to disappear after GPT-3 came out, and get replaced with really strong opinions about politics and investments.

This post is literally just SEO copy designed to sell degoogled phones to (justifiably) anxious DV victims?

Their phones are more than twice as expensive as equivalent models at JF HiFi too (and 5-10x the price of an older, but still perfectly useful degoogled phone from Marketplace).

Why is it on the front page of HN?

Is suspect you're looking at more like $100k than $100 for an ASIC with enough memory and compute to handle Fable.

But the flip side is possibly 1000+ tok/s on a SOTA model, which would be game changing.

Could make sense for datacentres or enterprise, but I don't think we'll be getting SOTA Game Boy carts this decade.

I like to think of pre-12 Rules Peterson and post-12 Rules Peterson as two separate people.

Some mix of the fame, coordinated bad-faith campaigns against him and the Benzos broke the man's brain.

But you go back to his lecture series and TV show appearances from 2017 and prior and the man is lucid and insightful.

This all went over my head, but does anyone know either how much faster this will make things (4x faster than AVX512 at 2048-bit??), and if unified memory plus a basic GPU will render this dead in the water?

How would you react if you got a panicked call from your spouse and it came from their phone number?

I would be surprised if an AI successfully imitated the batshit insane way my wife and I speak to one another, but I get how that could work for, say, a grandparent or aunty.

There's also a whole industry of basically "virtual onlyfans" models that generate a TON of content and ad impressions on everything from Instagram, X, TikTok, even Twitch.tv

Are there stats on this? Like if you look at the top hundred influencers on these platforms, how many of them are AI? Is it a long tail?

Like you do hear a lot of stories about this kind of thing happening, but I haven't ever seen a figure that says something like "$x million lost to AI scams" the way you do with romance scams or crypto.

This might just be my bubble, but us there that much AI being used for girlfriends/scams, or even reels these days?

Most people I know use it as a tool to do things for them, technical or otherwise.

Even the loneliest people I know don't want to talk to a clanker unless there's at least a pretense of work being done.

AI 2040: Plan A 12 days ago

Has there been an uptick in terrorist bombings or is this just a hypothetical at this stage?

I've definitely seen it used both ways, comparing Japan to other countries as well as India/Africa.

I wouldn't necessarily call it a "racist dog whistle" myself, though - there is a very real pattern that's being pointed out but the reason I made the GP comment is that from my experience I would assume that Chinese culture is about as trustworthy as the West.

Ok, the implication that I'm reading between the lines is that this sort of behaviour is somehow more tolerated by people with names like Liu and Tan, but is this actually the case?

I know there's some evidence of Chinese people working at big tech and feeding data back to the CCP but is this a "low trust culture" issue in general or an extrapolation of that one pattern?

GPT-5.6 12 days ago

Didn't know there was a word for that, thanks! Looks like my programming style matches my communication style in general. :P