HN user

concats

115 karma

30-something European. I believe in: Fitness, Philosophy, and future optimism.

Posts0
Comments51
View on HN
No posts found.
Claude Opus 4.7 3 months ago

The recently viral 'grill-me' skill is great for exactly this.

It's just a super simple skill that, when invoked, makes the model spend considerable time asking design and architecture questions and fleshing out any plan with you. A planning session without it might be Claude asking you 2 questions, and with it 22.

Just knowing someone's name, address, and ID number isn't enough to like, open a bank account in their name or such. You'd need a proper ID card or passport for that. Similar thing with most businesses if you try to pay for some product with credit, they won't accept just a few digits and a pinky promise, you'll need to identify yourself properly (the BankID app for instance).

Why is mechanized thinking going to do that? When mechanized labor didn't?

You're right. There is technically a category of work that relies on neither our ability to do physical labor nor excessive thinking. It just relies on being a human.

The conclusion is thus obvious: AI is going to push us all into careers as photo models, OF-creators, and social media influencers! /s

The revenue numbers are public for the major AI companies. That's probably the best estimate for "inference for the whole market" we have, since most of that inference is billed in either API usage or subscriptions, and it won't include any in-house usage such as training.

How does it compare for models of any meaningful size?

These 0.6B-4B models are, frankly, just amusing curiosities. But commonly regarded as too error prone for any non-demo work.

The reason why people are buying Apple Silicon today is because the unified memory allows them to run larger models that are cost prohibitive to run otherwise (usually requiring Nvidia server GPUs). It would be much more interesting to see benchmarks for things like Qwen3.5-122B-A10B, GLM-5, or any dense model is the 20b+ range. Thanks.

What choice do they really have though? More and more consumers completely forgo owning a regular computer and only use a phone or a tablet now a days. And among the ones who do own a computer there's still a strong trend towards not paying for software, presumably a behavior taught to them by the overwhelming success of strictly ad-financed apps.

It's easy to forget that us here on HN are several standard deviations from the norm.

True! (value - cost) would be better.

I was more so limiting myself to the simpler heuristic where people only pay roughly what they personally think something is worth, and not significantly more/less regardless of the options. But of course, as you've pointed out, in real life the options available really do matter, and someone might decline a 200:1200 trade if there are even more lopsided options available. It does complicate the though experiment somewhat if you try to take this into account.

If we assume people are somewhat rational (big ask I know), and the Efficient-market hypothesis, then we can estimate the value created by AI to be roughly equal to the revenue of these AI companies. That is: A professional who pays 20€/month likely believes that the AI product provides them with roughly 20€ each month in productivity gains, or else they wouldn't be paying, and similarly they would pay more for a bigger subscription if they thought there was more low hanging fruit available to grab.

Of course this doesn't take into account people who just pay to play around and learn, non professional use cases, or a few other things, but it's a rough ballpark estimate.

Assuming the above, current AI models would only increase the productivity for most workplaces by a relatively small amount, around 10-200 € per employee per month perhaps. Almost indistinguishable compared to salaries and other business expenses.

Putting that kind of filter in the way of speech seems ripe for abuse.

On one hand I agree with you. Any automatic filter implemented can later be expanded to cover more and more things, such as messages from political adversaries for example. It's a slippery slope as we all know.

On the other hand I don't think it applies in this context very much. If we're talking about content published by a corporation or such (say a newspaper for example) they already filter all their gathered news themselves and have no obligation to publish things they don't feel like.

Similarly if we're talking about user uploaded content on social media I don't think they have any obligation to publish everything and anything that their users decide to upload either, and it's not the expectations of the users that anything can be hosted there for them. Users already know that youtube/facebook/tiktok/what-have-you have seemingly arbitrary rules regarding what content they're willing to host and not.

Now if for example DNS providers or ISPs decide to implement these sort of filters on the web at large that's a different matter I think. In which case I agree with you.

Sidenote: I wonder what's going to happen when the crazy money runs out and Anthropic, OpenAI & co have to start charging for more than it costs them to run the models. Hopefully by then the open source models will have caught up?

How brutal will the enshittification phase of these products be?

Will the 10x cost or whatever be something that future employers will have to pay, or will it be a more visible impact for all of us? Assuming no AGI scenario here and the investments will have to be paid back with further subscription services like today.

I really hope Open Source (Open Weights) keep up with the development, and that a continuation of Moore's Law (the bastardized performance per € version) makes local models increasingly accessible.

Gemini 3 Deep Think 5 months ago

I've been surprised how difficult it is for LLMs to simply answer "I don't know."

It's very difficult to train for that. Of course you can include a Question+Answer pair in your training data for which the answer is "I don't know" but in that case where you have a ready question you might as well include the real answer anyways, or else you're just training your LLM to be less knowledgeable than the alternative. But then, if you never have the pattern of "I don't know" in the training data it also won't show up in results, so what should you do?

If you could predict the blind spots ahead of time you'd plug them up, either with knowledge or with "idk". But nobody can predict the blind spots perfectly, so instead they become the main hallucinations.

I think one of the things that short form videos do really well is that they punish creators who pad their videos with unnecessary filler content. On TikTok for example (Not necessarily a fan of the app but it's a good example) no videos start with all that empty jabbering you often see on YouTube ("Welcome to my channel...", "Today we will...", "Please Like and Subscribe...", "This video is sponsored by...", etc), because if they tried any of that crap the viewers would just swipe the content away. So, instead they always get straight to the point. That part is really refreshing.

Of course, there are other issues instead.

Most of these seem concretely doable, and maybe effective. But the core of the addictiveness comes from the "recommender system", and what are they supposed to do there? Start recommending worse content? How much worse do the recommendations have to be before the EC is satisfied?

I agree with you, this is rather odd. And sort of missing the problem.

All apps are about attention. The percentage of the time spent using the app when it shows you your good content (Whatever it is that you're interested in) determines how addictive it is. And the percentage of time it's showing you bad content (Ads, 'screen time breaks', manual scroll time, more ads, loading screens, sponsor ads, filler content (youtube for instance is full of this), etc) counteracts the addictive properties because nobody likes it.

What's the end goal here? Right now TikTok is winning the attention economy race against the other apps because it's more focused on the user's preferred content. Is that what we want to reduce? To show more uninteresting other stuff on the screen? Like blank 'wait 5 minute' statics? Or just more ads?

I get that we don't want a generation of socially inept phone addicts, but this won't solve anything I fear. People will still want the good content, forcing the most customer friendly (it feels wrong to say that about TikTok) app to become more enshittified is a bewildering solution.

TikTok has a lot of issues, such as privacy, dubious content, 'brainrot', etc. I don't want to seem like I'm necessarily defending TikTok specifically here.

But this really just stinks of Regulatory Capture to me. Their main argument is that the consumers like to use the app too much?

Why? Because it's smarter and not as enshittified as the competitors?

I'm sure if youtube, facebook, reddit, etc reduced the number of ads, and started showing more relevant content that people actually cared about, they too would start being "more addictive". Do we really want to punish that?

What's the end goal here?

I agree. It's very missleading. Here's what the authors actually say:

AI assistance produces significant productivity gains across professional domains, particularly for novice workers. Yet how this assistance affects the development of skills required to effectively supervise AI remains unclear. Novice workers who rely heavily on AI to complete unfamiliar tasks may compromise their own skill acquisition in the process. We conduct randomized experiments to study how developers gained mastery of a new asynchronous programming library with and without the assistance of AI. We find that AI use impairs conceptual understanding, code reading, and debugging abilities, without delivering significant efficiency gains on average. Participants who fully delegated coding tasks showed some productivity improvements, but at the cost of learning the library. We identify six distinct AI interaction patterns, three of which involve cognitive engagement and preserve learning outcomes even when participants receive AI assistance. Our findings suggest that AI-enhanced productivity is not a shortcut to competence and AI assistance should be carefully adopted into workflows to preserve skill formation -- particularly in safety-critical domains.

I didn't catch it immediately, but after you pointed it out I totally agree. That comment is for sure LLM written. If that involved a human in the loop or was fully automated I cannot say.

We currently live in the very thin sliver of time where the internet is already full of LLM writing, but where it's not quite invisible yet. It's just a matter of time before those Dead Internet Theory guys score another point and these comments are indistinguishable from novel human thought.

I remember leaving university going into my first engineering job, thinking "Where is all the engineering? All the problem solving and building complex system? All the math and science? Have I been demoted to a lowly programmer?"

Took me a few years to realize that this wasn't a universal feeling, and that many others found the programming tasks more fulfilling than any challenging engineering. I suppose this is merely another manifestation of the same phenomena.

Prism 6 months ago

Isn't most research and scientific data is already shared openly (in publications usually)?

In what way does them having a subjective local moral standard for themselves imply that there exists some sort of objective universal moral standard for everyone?

If I knew someone wanted me dead, of course I would want a prediction market on it [...] Someone fill me in on what I'm missing here?

The assassin might place the bet at roughly the same time as they place the bullet in the chamber. Making the prediction into a bounty. Not giving you any meaningful time to ponder the new information. The notification from your phone would be the distraction they'd use when taking aim.

A lot of complaints and concerns about LLMs today echo remarks about Wikipedia back then.

I have also noticed this.

How LLMs can never be trusted because they are stochastic sounds very similar to how Wikipedia can never be trusted because it sometimes has a bad-faith edit.

Or how the people that don't believe information should be free are very active in both the anti-Wikipedia and anti-llm crowds. And use much of the same talking points.

Have publishers have sued Wikipedia too?

Ultimately I think over the next two years or so, Anthropic and OpenAI will evolve their product from "coding assistant" to "engineering team replacement"

The way I see it, there will always be a layer in the corporate organization where someone has to interact with the machine. The transitioning layer from humans to AIs. This is true no matter how high up the hierarchy you replace the humans, be it the engineers layer, the engineering managers, or even their managers.

Given the above, it feels reasonable to believe that whatever title that person has—who is responsible for converting human management's ideas into prompts (or whatever the future has the text prompts replaced by)—that person will do a better job if they have a high degree of technical competence. That is to say, I believe most companies will still want and benefit if that/those employees are engineers. Converting non-technical CEO fever dreams and ambitions into strict technical specifications and prompts.

What this means for us, our careers, or Anthropic's marketing department, I cannot say.

Thanks for the post. I found it very interesting and I agree with most of what you said. Things are changing, regardless of our feelings on the matter.

While I agree that there is something tragic about watching what we know (and have dedicated significant time and energy in learning) devalued. I'm still exited for the future, and for the potential this has. I'm sure that given enough time this will result in amazing things that we cannot even imagine today. The fact that the open models and research is keeping up is incredibly important, and probably the main things that keeps me optimistic for the future.

GPT Image 1.5 7 months ago

“I've come up with a set of rules that describe our reactions to technologies:

1. Anything that is in the world when you’re born is normal and ordinary and is just a natural part of the way the world works.

2. Anything that's invented between when you’re fifteen and thirty-five is new and exciting and revolutionary and you can probably get a career in it.

3. Anything invented after you're thirty-five is against the natural order of things.”

― Douglas Adams

I don't think anyone disagrees with that. But it's a good time to learn now, to jump on the train and follow the progress.

It will give the developer a leg up in the future when the mature tools are ready. Just like the people who surfed the 90s internet seem to do better with advanced technology than the youngsters who've only seen the latest sleek modern GUI tools and apps of today.

> It's too bad people spend energy for generating them now.

How do you mean?

Some quick back of the napkin math.

Creating a 'throwaway' banner image by hand, maybe 15 minutes on a 100W CPU in Photoshop:

  15 minutes human work time + 0.025 kWh (100W*0.25h)
Creating a 'throwaway' banner image by stable diffusion on a 600W GPU. In reality it's probably less than 20 seconds to generate, but let's round it up to one full minute of compute time:
  5 minutes human work time + 0.01 kWh (600W*(1/60)h)
The way I see it it seems to spend less energy, regardless of whether you're talking about human energy or electrical energy. What's the issue here exactly?