HN user

alach11

1,653 karma

Reservoir engineer / data scientist / product manager at an oil & gas company in Houston, TX. You can contact me at <my Hacker News username>@gmail.com.

Posts22
Comments376
View on HN
www.reuters.com 2d ago

Google plans new chip to run Gemini models more efficiently

alach11
4pts2
www.forbes.com 1mo ago

Anthropic Cofounder Joins Pope Leo, Warns of AI Job Losses

alach11
2pts0
www.reuters.com 1mo ago

Anthropic's Olah says AI must be guided from outside Big Tech

alach11
3pts1
openai.com 2mo ago

OpenAI DevDay 2026

alach11
2pts0
openai.com 8mo ago

Evals drive the next chapter in AI for businesses

alach11
2pts0
docs.google.com 9mo ago

Advice for a young investigator in the last days of the Anthropocene

alach11
3pts1
github.com 10mo ago

On It, Boss

alach11
3pts1
www.wired.com 1y ago

Zuckerberg Expanding His Hawaii Compound. Part of It Sits Atop a Burial Ground

alach11
18pts6
www.youtube.com 1y ago

Updates to ChatGPT for Business (internal connectors, record mode, and more) [video]

alach11
2pts1
rahmatashari.com 1y ago

Is It an AWS EC2 Instance or a US Visa?

alach11
39pts6
www.nytimes.com 1y ago

Suspect in CEO Killing Withdrew from a Life of Privilege and Promise

alach11
10pts4
research.google 1y ago

Unlocking the power of time-series data with multimodal models

alach11
131pts26
twitter.com 1y ago

SpaceX Super Heavy splashes down in the gulf, canceling chopsticks landing

alach11
272pts425
www.usatoday.com 1y ago

Tyson vs. Paul: Netflix suffers significant issues on fight night

alach11
7pts4
blog.google 1y ago

Google uses AI to reduce stop-and-go traffic on your route

alach11
98pts125
techcrunch.com 2y ago

AWS follows Google in announcing unrestricted free data egress

alach11
6pts1
apps.microsoft.com 2y ago

City Art Search – Rotating Desktop Backgrounds of Classic Artwork

alach11
1pts0
www.nytimes.com 2y ago

New York’s Hottest Steakhouse Was a Fake, Until Saturday Night

alach11
4pts0
www.theverge.com 3y ago

Microsoft to charge $30 per user per month for Copilot for Office

alach11
4pts1
www.nationsreportcard.gov 3y ago

Scores decline again for 13-year-old students in reading and mathematics

alach11
184pts419
adventofcode.com 3y ago

Advent of Code 2022

alach11
3pts1
shouldigetahouse.com 5y ago

Show HN: Should I Get a House? a better rent vs. buy calculator

alach11
203pts410

The honest answer to that question, in June 2026, is that we do not know

The honest reading of those numbers is not that defense is winning on economics

The honest 2026 answer is in three parts.

The honest answer is that we do not know, because no one has tried

Firstly, I appreciated the article and especially the visuals. But I had the same reaction as the GP commenter. It was hard to read. I'm sick of this punchy, repetitive, LLM-generated prose.

Everyone loves to say this when the death of Stack Overflow is discussed, but it always was that way. Strict moderation, love it or hate it, was part of the platform. And it could have kept going that way for many more years if not for LLMs 99.9% obviating the need for a coding Q&A forum.

It's a tall order to live up to the impact of Rerum novarum, the encyclical by the former Pope Leo that greatly guided thinking out of the industrial revolution. Personally, I'm excited to read this. If we take the claims of most AI labs at face value, they believe their work will fundamentally change the relationship between humans and the economy. More involvement from faith leaders is a good thing.

If Anthropic actually cared about humans, they would have the best customer support (staffed by humans, for humans)

I know Anthropic support is slow from firsthand experience, but it has to be pretty difficult to scale support 10-80x per year. And even more so when you have a long-tail of very low revenue usage in the form of $20/month subscriptions.

Snake and DOOM were two of our early tests (for filter functions and MCP) when we stood up Open WebUI for internal chat/agent use. Sometimes games are the best way to limit-test new tech.

I ran an internal (oil and gas focused) benchmark yesterday and found Opus 4.7 was 50% cheaper than Opus 4.6, driven by significantly fewer output tokens for reasoning. It also scored 80% (vs. 60%).

Claude Opus 4.7 3 months ago

On my private internal oil and gas benchmark, I found a counterintuitive result. Opus 4.7 scores 80%, outperforming Opus 4.6 (64%) and GPT-5.4 (76%). But it's the cheapest of the three models by 2x.

This is mainly driven by reduced reasoning token usage. It goes to show that "sticker price" per token is no longer adequate for comparing model cost.

A significant part of Anthropic's cachet as an employer is the ethical stance they profess to take. This is no doubt a tough spot to be in, but it's hard to see Dario making any other decision here.

What I don't understand is why Hegseth pushed the issue to an ultimatum like this. They say they're not trying to use Claude for domestic mass surveillance or autonomous weapons. If so, what does the Department of War have to gain from this fight?

I have to imagine governments are closely monitoring prediction markets as part of their intelligence apparatus. But then you just add another layer of subterfuge. Imagine a D-Day prediction market... "Will the Allies Land in Normandy, Pas-de-Calais, or somewhere else?" The US might buy a major position on Pas-de-Calais the night before as a decoy!

Advent of Code 2025 8 months ago

Usually the first day or two are readily solvable in Excel with just regular spreadsheet formulas.

Claude Opus 4.5 8 months ago

This is the biggest news of the announcement. Prior Opus models were strong, but the cost was a big limiter of usage. This price point still makes it a "premium" option, but isn't prohibitive.

Also increasingly it's becoming important to look at token usage rather than just token cost. They say Opus 4.5 (with high reasoning) used 50% fewer tokens than Sonnet 4.5. So you get a higher score on SWE-bench verified, you pay more per token, but you use fewer tokens and overall pay less!

Gemini 3 8 months ago

This is a really impressive release. It's probably the biggest lead we've seen from a model since the release of GPT-4. Seems likely that OpenAI rushed out GPT-5.1 to beat the Gemini 3 release, knowing that their model would underperform it.

Computer use is the most important AI benchmark to watch if you're trying to forecast labor-market impact. You're right, there are much more effective ways for ML/AI systems to accomplish tasks on the computer. But they all have to be hand-crafted for each task. Solving the general case is more scalable.

Claude Sonnet 4.5 10 months ago

I'm really interested in the progress on computer use. These are the benchmarks to watch if you want to forecast economic disruption, IMO. Mastery of computer use takes us out of the paradigm of task-specific integrations with AI to a more generic interface that's way more scalable.

Thus far I didn't have to worry about ChatGPT having bad incentives when giving me advice on product purchases. Now that "Merchants pay a small fee on completed purchases", will the model steer me towards ACP-supported retailers at a higher rate?

On It, Boss 10 months ago

I can't fathom why Microsoft has such limited functionality in Copilot for Excel. Some of the ideas in projects like this (and similar) seem so obvious. Are they just being conservative with token cost?

I'm going to make an unpopular suggestion. Have you considered using a service that will print and ship to you, like CraftCloud?

Depending on volume, your total cost would likely be lower. I know you mentioned privacy concerns so this may not be an option. But it significantly simplifies your work, letting you focus on the parts themselves.

Florida did drug testing as a condition for welfare benefits... and it cost more than they saved

It's more complicated than that. Of the 6352 people who applied for TANF, 2306 dropped out during the process. Then of the 4046 TANF applicants remaining, only 2.6% tested positive for drugs. The vast majority of media coverage focused on the 2.6% being less than the ~8% drug-use rate in the general population.

What we don't know is of the people who dropped out, was this due to unintended reasons (privacy concerns, the inconvenience of the drug test, missing deadlines) or due to the intended reason (people self-selecting out because they knew they would test positive and become ineligible for 12 months). We'll never know the real breakdown, but it's misleading to say "it cost more than they saved".