HN user

OtherShrezzing

2,703 karma
Posts0
Comments457
View on HN
No posts found.

I feel like we read different articles, because without context, the pictured dish looked like overcooked pork chops served next to a type of carbonara tagliatelle topped with a fried egg & chives. I can't imagine anyone has ever set out to make spamsilog and ended up with the generated dish.

It's different in kind to a "my Big Mac is more squashed than the advert on TV" type situation.

But is that a widespread feeling, are the vibes bad for regular people, or is this effectively just old man shouting at clouds?

A cursory glance at YouGov strongly suggests that people have bad vibes about AI in basically every territory where they're asked about it, and for basically every use-case

There’s an OpenAI slide deck sitting in Goldman’s IPO preparation department right now with a pair of bullet points that read

Q3 2027, introduce ‘ad-enhanced’ paid subscription tier

Q2 2029, complete ad rollout to all paid subscribers

In the UK, farm vehicles use a specially dyed Red Diesel, which is the same liquid with a chemical added to it to turn it red. It’s taxed at a much lower rate, so you’ll sometimes pass petrol stations selling two Diesel from two different pumps, one of which is around half the price of the other.

If you’re caught running red diesel on a normal carriageway, the penalty can be steep.

I’m not suggesting Simon’s pelicans in the dataset are having a meaningful impact. I’m expecting that a company like ScaleAI has a product along the lines of “benchmax dataset: SimonW’s Pelican on Bikes test” which is a private curated series of well-drawn SVGs of animals riding vehicles for training and RL.

Respectfully, the pelicans used to be an unrecognisable mess and now they’re unquestionably pelicans on bicycles, rendered poorly, from every model.

In the same timescale, model capabilities across the board have only meaningfully improved in places where the labs are focusing their training efforts.

Moreover, they have a uniform style, even though your prompt doesn’t ask for one. There's no model going rogue and producing a watercolour of a pelican. They’re all rendered in an approximately uniform style, even though the svg format has a basically unlimited possibility space.

Decent summary of it here[0]. The “space” part of “SpaceX” is valued by market analysts and money managers at around 5% of the company’s entire value. Almost all of the rest is “AI stuff”, and Twitter is a rounding error.

That is, if SpaceX went back to being a space-only entity, and dropped the AI stuff, its share price should be expected to fall from $130/share to around $7/share.

[0] https://www.ft.com/content/09a62ed4-16af-433c-adb7-c877d1975...

to demonstrate that the tech industry isn't just here to extract wealth from the poor/many and transfer it to rich/few?

I think the problem is that the tech industry in large is just here to extract wealth from the many and transfer it to the few. That's why it's focused on scale.

People aren't dumb, and most of the time they can see when they're on the receiving end of an extractive relationship - even if there's lots of PR work going on to hide that reality from them.

It's unlikely that AI will get to the point where it makes handwritten coders redundant, and then not immediately be at the point where vibe coders are redundant too. So if you earnestly take the position that handwriting code is a "ngmi" type activity, you also need to take the position that the vibe coder (or agent- assisted-developer/loop-architect, or whatever its nom de guerre is this week) is "ngmi".

4x improvement on geospatial tasks with map in the loop.

The graph shows a baseline 2% task success rate improving to to 8% task success rate, but the evals section details 100% success rates across the board.

I'm not sure what the effectiveness of this skill is from the readme. Is it 8% success, or 100% success?

Safari's copy-text-from-image feature manages the entire base64 part of the string, except for the first character (I instead of a T). Weirdly, it gets much worse performance if you try to copy the entire string, including the hashbang part.

I wonder what it's doing under the hood to get such good performance?

I think this is a case where two people can successfully complete the task manually faster than one attempting to automate it. Get a ruler, read five centimetres of characters to your colleague, have them type it in as you go, then repeat that five centimetres back to you. Correct as you go. Format your string with the same line-breaks as the t-shirt, and remove them at the end, so you can be sure you've got the correct length on each row. Trial-and-error adjust the five-cm distance depending on your success rate as you go along

All in, you should have a non-corrupted string in 10-15 min.

It’s difficult to articulate the tedium and monotony of a Starbucks gig. There’s so little intellectual stimulation available in that setting. If you managed to learn more from your fast food than your humanities degree, then I think that’s on you for not paying attention at college (perhaps because you were exhausted from your job?).

I don’t think any engineers who cost $150/hr are having their productivity moved by 20% depending on a $10/hr gap between models on or near the frontier.

Most of the gains right now come from tooling and process and any big post 2025 language model. The specific model isn’t that important right now.

No, standard floating point implementations have higher precision for smaller numbers than larger. So for example, in a 32bit float, there are far more numbers between 0-1 than there are between 1,000,000 and 1,000,001. For 32bit floats, you start lowing whole integers with relatively small numbers.

Integers have a consistent precision across the entire number line.

Most of the interesting research I’ve ever done started while reading through the intermediate steps in an unrelated paper.

As far as I can tell from colleagues in other domains, it’s the same there. One paper will mention something off-hand and that’ll cause someone else to have a spark of insight, which turns into it’s own valuable research

robot maid that could clean, wash and fold the laundry, do the dishes, etc. would be huge. I think a lot of people would pay new-car money for something like that.

Once you take maintenance of a machine with price-parity to a new car into consideration, it’s surely cost competitive to just hire a human to do all those things.

The price needs to fall drastically below new-car territory before it’s competitive with manual human labour.

As an AI-native startup founder, your responsibility is to know what's in your codebase, understand any potential exposure vectors, and not ship obvious vulnerabilities to real users who are trusting you with their data.

This is fairly funny coming from the company whose employees report merging in hundreds of PRs per engineer per day, and accidentally leaked their own source code through a security misconfiguration in a package manager they own.

I sometimes use the Claude app with text to speech enabled. It’s got a quite distinctive voice/tempo combo when it’s outputting speech.

Whenever I see a typical Claude-tell in writing, my internal reading voice switches automatically from my internal monologue’s voice into Claude’s voice for the rest of the piece.

A debates purpose is surely to reveal if there’s a measurable difference to be reconciled in the first instance. Any actual reconciliation is a nice bonus on top.

So even if the debate reveals that no, there wasn’t a viable reconciliation, the debate was still worthwhile.

I don't think this article properly engaged with the criticism from the politician. That's fine, I wasn't expecting it to, but this isn't valuable commentary on the politician's point. I suppose it does serve to demonstrate that Graham and people in Graham's orbit are unable to see a distinction between "have a billion dollars" and "earn a billion dollars".

It's a very sf-bubble type article.

I see the value in being able to discuss features more centrally. I’m not seeing the value in a per-keystroke delta of the software as it’s built though.

It feels like this communications problem could be solved through draft PRs and a decent slack app integration.

I don’t see the value proposition here. I’ve seen roughly this feature proposed by multiple companies, and absolutely none of the have given a convincing reason for the technology to exist.

The health service now has to spend more money settling maternity-malpractice claims than it does on actually providing maternity care

This figure is from an article in the Times, and has no connection to official NHS figures. The Times just guessed how much it might be, and reported it as fact. Then, since The Times is a paper of record, other news outlets have run with it.

Claude Fable 5 1 month ago

In the UK, a £45k/yr employee pays their own tax and gets a take-home of £35k.

The employer pays £6k for National Insurance (atop the employee's NI contributions). Pension: 2-3k. Apprenticeship levy is £300. 3yr-amortised recruitment fee is £4000. Hardware costs: £1000. Office space £5000. Software/tools: £2500. Benefits: £1500. Training: £1000. Other admin overheads £500.

You pay that person for ~250 working-days, but they only attend for ~220, due to annual leave and sick pay, so you get around £62k worth of attendance out of that person in exchange for £70k, of which the employee sees £35k.