HN user

adam_arthur

4,023 karma

Worked for a Silicon Valley startup from early days through successful acquisition.

Posts0
Comments1,650
View on HN
No posts found.

Exactly.

I've found 6 minutes or so the sweet spot for upper bound with 5.6 Sol.

And it sounds like the OPs query above requires scanning throughout a large portion of the codebase, which will inherently consume a large number of input tokens. No locality to it.

When they can sell that for tens of Billions a year against competition, they might have a financial case.

In reality, there will be many clones of Claude Design, especially if it gains big revenue traction.

Doubly so if you believe the narrative that coding and apps will be "free" and instant to create in the future.

Then it's hard to switch?

It's irrelevant the reason why, the margin any business can take will be constrained by the cost to switch to a cheaper, good enough, competitor.

You are saying it's difficult to switch due to compliance and admin issues. Ok!

LLMs can be swapped easily and open weight models can be hosted to adhere to whatever legal, uptime requirements etc are needed trivially.

The big cloud providers already host and resell open weight models.

The two don't event compare

Completely different levels of stickiness.

The OS runs everything; the LLM can be swapped in a second.

(But yes, you will have to tweak prompts+tuning anytime you change models)

GPT-5.6 12 days ago

I say "human-like" in the sense that LLMs are fed in text data largely in the exact form (mapped to tokens) that humans read them.

Thus from first principles it's most likely that content which is more understandable to humans is also likely to be more understandable to LLMs. Of course they are still capable of interpreting very obscure structures too, but usually at the cost of cognitive performance.

I'm open to being wrong about this, and I'm sure it's being researched.

(Specifically for text representations)

To your point, at some level of intelligence an LLM will be able to infer the intent of your prompt consistently without thinking enabled, in which case interpretability to a human matters less. But for complex tasks you aren't likely to get optimal performance with prompts that are difficult for humans to understand. And yes, you'd see that with thinking enabled as it churns over thousands of tokens trying to "mentally expand" a compressed prompt.

Interesting discussion though!

GPT-5.6 13 days ago

Information density of the interpretability of the intent from the perspective of a human (or human-like).

If the intent is not easy to understand, it's information sparse. Because it takes a lot of CPU (or brainpower) to interpret.

You can run gzip on an English sentence to make it more textually dense, but clearly it is not more information dense in this context.

GPT-5.6 13 days ago

Information density of the prompt is the most important factor in my experience.

And interestingly, LLMs seem particularly bad at writing prompts for other LLMs for this reason (you can guide them to be more dense, just speaking by default).

Conciseness is usually a byproduct of information density though.

You seem to be missing that if prices are lower, demand is lower. Price is a function of demand.

If rents go from 10000/month at 100% occupancy to 5000/month at 100% occupancy, demand has been materially reduced.

That occupancy will recover does not support your point in the way you seem to think it does.

It's a natural market function that a drop in demand will lead to a drop in prices, and eventually, a commensurate increase in consumption (at lower prices)

This happens everywhere, and out of everywhere Seattle is doing the worst.

Muse Spark 1.1 13 days ago

Not much moat, incremental improvements, cherry picking models to compare.

To be fair, seems more correct to compare against similar strength models if your main edge is pricing.

The level of "slop" produced by AI is a direct function of skill of the developer and broadness of the prompts.

Broad prompts by unskilled users results in a complete mess. Targeted prompts by a skilled person reviewing the code produces something better.

Quality of application varies widely, and generally agree with the categories mentioned by the sibling post.

Testosterone is directly causally inverse to bodyfat in men (once above some very low baseline)

Fat directly converts testosterone to Estrogen via a process called aromatization.

Personally my Testosterone close to doubled when going from 25% bodyfat to 13%. I get blood tests regularly and can see the levels fluctuate pretty closely with fat levels

The main difference I'd guess is whether your prompts are targeted or broad.

Less experienced people tend to use very broad prompts.

Experienced people tend to understand the structure of the code and give explicit guidance such that a larger model isn't necessary to read between the lines.

I noticed with GPT-5.6 (through work), I could step up my specificity by a level of abstraction. But I still intentionally scope the prompts fairly tightly, as I find it produces better results if you need to own and maintain the code.

Drop in rents is drop in demand, plain and simple.

Fewer businesses want to invest and move into Seattle at the price it used to cost.

It's true that if occupancy is poor, then a recovery in occupancy will bring more activity.

It's bullish in the same way that cheaper housing due to increased crime and decline in quality of life draws in new residents.

Again, office is doing poorly nationally, but it does seem particularly worse in Seattle than other hubs.

If I'm not mistaken, Seattle has the worst office vacancy in the country by a decent margin, aka the most demand destruction.

An increase in vacancies across the board is reduction in demand, plain and simple.

That the new equilibrium price to re-tenant all the buildings is lower is evidence of that.

But the OP is correct that when enough of the building owners default on their debt, the building will be foreclosed, sold for less and asking rents will go down towards the new equilibrium price.

Thus occupancy is likely to improve again down the line.

But, yes, this is not a bullish situation for Seattle. Office generally hasn't been doing well nationally, so it's more of a question of relative performance.

Vibe-coding the implementation.

I haven't had much issue with Codex, but seems Claude Code has major issues being reported nearly on the daily.

They also happen to be the most boastful about not reading or looking at the code.

LLMs are very capable, but not nearly to the level they seem to be messaging.

(We've actually moved on from vibe-coding to having the LLM vibe code itself in a loop)

AI is clearly a force multiplier, both negative and positive.

The truth is there are prolific developers like Antirez who have built quality new projects at an incredible pace (Dwarfstar 4, Redis features).

But as unpopular as it is to say it, in the working world ~80% of developers pre-AI mostly just attended meetings, did a little busywork and committed small patches here and there. Probably around 20% really moved the needle and contributed the bulk of net new code.

Those 80% were constrained in the volume they could output pre-AI, but now they are unleashed to do a large amount of net new work but many without the skills to structure it well+maintainably.

It doesn't help that most management has been pushing on LoC over quality the past year.

I truly believe most companies as they exist today are not structured for AI. The amount of technical debt that will be created at a rapid pace is basically time delayed self destruction for most codebases if you let people run amok with low contribution standards and rubber stamped approvals.

If you treat each AI output as a small well-scoped, well-tested module, which interoperate with each other through well designed APIs, you can have high confidence in quality. But majority of people are pseudo-vibe coding and creating spaghetti monster codebases, and there's really no way to stop it without strong and tight technical oversight.

I'd like to see a world where the data vendor is separate from the app/UI product vendor.

e.g. anyone could build a skin over map data, short form video data, long form video data, short form text content.

Data vendor makes money through selling the data, app vendors make money through either subscriptions, ads, or selling new data back to the data vendor.

The market for pretty much everything would become intensely competitive and price much closer to marginal cost of service.

However, PII+personal data would become more of a concern with this model.

There are hundreds of studies indicating that each marginal additional unit of bodyfat is less healthy for you than not having it.

You are quoting poorly done and controlled BMI studies to dispute this, studies which look primarily at elderly people where weight is highly correlated with health to begin with and is not properly controlled for along every axis.

I'm not here to litigate it or convince you though.

A 40-70m sustained effort is effectively LISS from a study perspective. The main distinction being made is intensity, and it's impossible to sustain even a marginally intense effort for 40m (unless you redefine "intense" as different from what most studies use)

Recent studies have shown compelling evidence that LISS promotes arterial plaque buildup, while HIIT does not have this effect and has even been shown to reverse it if other parameters are in order.

https://www.ahajournals.org/doi/10.1161/CIRCULATIONAHA.119.0...

"""

Physical activity and exercise training are effective strategies for reducing the risk of cardiovascular events, but multiple studies have reported an increased prevalence of coronary atherosclerosis, usually measured as coronary artery calcification, among athletes who are middle-aged and older. Our review of the medical literature demonstrates that the prevalence of coronary artery calcification and atherosclerotic plaques, which are strong predictors for future cardiovascular morbidity and mortality, was higher in athletes compared with controls, and was higher in the most active athletes compared with less active athletes.

"""

While the health benefits of LISS seem to outweigh this factor, the new evidence that comes out in various studies continually nudge health impact of HIIT over LISS.

I find it difficult to recommend a form of cardio that 1) promotes arterial plaque over one that reverses it, and 2) seems to produce equivalent or worse biomarkers along almost every axis per unit of time spent.

(Not just V02 max)

https://pubmed.ncbi.nlm.nih.gov/36756765/

Lowest risk is around 18-20 BMI in this recent study, which controls for many confounding factors not controlled for in other studies.

Other studies show slightly higher troughs, but often don't sufficiently control for correlation of weight with health in elderly people.

From this study: Estimates of mortality differences by body mass index (BMI) are likely biased by: (1) confounding bias from heterogeneity in body shape; (2) positive survival bias in high-BMI samples due to recent weight gain; and (3) negative survival bias in low-BMI samples due to recent weight loss

And if you follow the longevity/health space and studies as they come out, it's becoming pretty clear that bodyfat is objectively bad for you above a pretty low baseline.

It shows up in insulin resistance, heart markers, inflammation, and once you control for confounding factors sufficiently, mortality.

You likely won't become diabetic with a bodyfat of 25%, but all your health markers will be worse than somebody at 15%. This is measurable and clear.

Claude Sonnet 5 21 days ago

If the LLM has to write out the reason, it is "thinking".

Whether you ask it to give <reason> or <thinking>, both will produce similar chain of thought processes.

To explain the "reason" requires producing thoughts that justify the answer that is not produced yet, because it writes the JSON linearly from top to bottom.

Not sure where the misunderstanding is coming from.

The claim comes from this study:

https://www.acpjournals.org/doi/pdf/10.7326/M15-1181

Though to be clear, there aren't a ton of studies that look at bodyfat percentage. Most use BMI and similar measures.

Likely overall fat levels matter more than %, I'd guess.

E.g. I'd presume being 15% at very muscular levels is less healthy than 15% at moderate.

(Because absolute fat mass plus visceral fat would be higher)

Yeah, typically "intense exercise" is implying HIIT style cardio.

More and more studies have been indicating that even just a few minutes of intense exercise can outperform long/slow LISS type cardios.

E.g. 5m all out effort is probably better, or at least equivalent, for health than a 30m moderate effort.

The average person can likely hit the 80/20 benefit threshold at less than 30m/week.

A large volume of studies already exist.

That intense exercise is good, and even very good for you, is proven as far as reasonably possible given that we can't run deterministically controlled experiments.

More evidence may come out that adds nuance, but the effect size is so large that it becomes obvious in the data just from observation.

You can cycle or stationary bike if you have bad knees. There are plenty of exercises that are intense but easy on the joints.

The principle of what you're stating is true, it could be correlational.

But there's an enormous volume of evidence that exercise, especially intense exercise, is better for health than any other intervention, including more sleep, quality of diet, pills+supplements (except those that treat an active illness/disease of course).

There's even compelling data showing that moderate drinkers who exercise live longer than non-drinkers who don't exercise. Even given that Alcohol is a powerful carcinogen.

The only thing proven more effective than exercise is weight loss really, if starting from high bodyfat levels.

(Anything above ~15% bodyfat in men seems to have negative implications for lifespan, and ~30% for women)

Claude Sonnet 5 21 days ago

This is the same concept as Chain of Thought.

Just that when "native" thinking is used it's hidden from the end result via special tags. If you force a model to reason about a result before producing the result, you get more accurate results.

Because "reason" comes before the selection, it has to think through why it is producing the result beforehand (e.g. produce a block of text that makes the correct answer statistically more likely to be sampled from the distribution. Giving it a property name does influence the direction of the thinking, but it's the same concept. You can call the property "yourThinking" too.

It sounds like you may be thinking too highly/mysteriously about how LLMs work.

At the end of the day they stream completely unstructured text outputs and all behaviors on top of that are just parsing XML-like tags to do tool calls, hide thinking etc (which they were trained to produce in certain circumstances).

There is no special "thinking process" it is a stream of text in the response that is simply wrapped in <Thinking> tags (or similar)