HN user

GavCo

3,038 karma

Web Developer

Posts283
Comments63
View on HN
www.off-policy.com 11d ago

Who manages the agents?

GavCo
75pts94
www.infoq.com 5mo ago

Google Introduces Managed Connection Pooling for AlloyDB

GavCo
1pts0
github.com 7mo ago

Show HN: Nano PDF – A CLI Tool to Edit PDFs with Gemini's Nano Banana

GavCo
176pts40
www.csoonline.com 11mo ago

Researchers Uncover RCE Attack Chains in HashiCorp Vault and CyberArk Conjur

GavCo
29pts7
assaf-pinhasi.medium.com 1y ago

“The Bitter Lesson” is wrong. Well sort of

GavCo
54pts34
github.com 1y ago

Biomni: A General-Purpose Biomedical AI Agent

GavCo
222pts37
netflixtechblog.com 1y ago

Driving Content Delivery Efficiency Through Classifying Cache Misses

GavCo
2pts0
techcrunch.com 1y ago

Hugging Face opens up orders for its Reachy Mini desktop robots

GavCo
2pts0
docs.anthropic.com 1y ago

Claude 4 prompt engineering best practices

GavCo
4pts0
www.quantamagazine.org 1y ago

The Fastest Way yet to Color Graphs

GavCo
62pts16
engineering.fb.com 1y ago

Enhancing the Python ecosystem with type checking and free threading

GavCo
2pts0
www.nature.com 1y ago

Why US police shootings are so deadly ― and why some police forces do better

GavCo
15pts5
www.technologyreview.com 1y ago

How The Pentagon is adapting to China's technological rise

GavCo
2pts0
www.nature.com 1y ago

Journal targeted by paper mill still grappling with the aftermath years later

GavCo
3pts0
deepmind.google 1y ago

Taking a Responsible Path to AGI

GavCo
2pts0
www.nature.com 1y ago

Microbes can capture carbon and degrade plastic – why aren't we using them more?

GavCo
8pts2
www.technologyreview.com 1y ago

How to delete your 23andMe data

GavCo
7pts0
medium.com 1y ago

Accelerating Large-Scale Test Migration with LLMs

GavCo
1pts0
www.uber.com 1y ago

Adopting Arm at Scale: Transitioning to a Multi-Architecture Environment

GavCo
1pts0
www.qodo.ai 1y ago

Evaluating RAG for large scale codebases

GavCo
41pts10
www.nature.com 1y ago

Scientists use AI to design life-like enzymes from scratch

GavCo
2pts0
engineeringblog.yelp.com 1y ago

Search Query Understanding with LLMs: From Ideation to Production

GavCo
1pts0
www.technologyreview.com 1y ago

How a top Chinese AI model overcame US sanctions

GavCo
1pts0
openai.com 1y ago

Operator System Card

GavCo
1pts0
scottaaronson.blog 1y ago

Jensen Huang and the quantum computing stock market crash

GavCo
2pts0
martinfowler.com 1y ago

Refactoring with Codemods to Automate API Changes

GavCo
2pts0
openai.com 1y ago

OpenAI's Economic Blueprint

GavCo
2pts0
www.quantamagazine.org 1y ago

Why Computer Scientists Consult Oracles

GavCo
16pts11
engineering.fb.com 1y ago

Indexing Code at Scale with Glean

GavCo
132pts30
mathstodon.xyz 1y ago

One of my papers got declined today

GavCo
853pts290

Appreciate your response.

But I don't think deception as a capability is the same as deceptive alignment.

Training an AI to be absolutely incapable of any deception in all outputs across every scenario would be severely limiting the AI. Take as a toy example play the game "Among Us" (see https://arxiv.org/abs/2402.07940). An AI incapable of deception would be unable to compete in this game and many other games. I would say that various forms, flavors and levels of deception are necessary to compete in business scenarios, and to for the AI to act as expected and desired in many other scenarios. "Aligned" humans practice clear cut deception in some cases that would be entirely consistent with human values.

Deceptive alignment is different. It's means being deceptive in the training and alignment process itself to specifically fake that it is aligned when it is not.

Anthropic research has shown that alignment faking can arise even when the model wasn't instructed to do so (see https://www.anthropic.com/research/alignment-faking). But when you dig into the details, the model was narrowly faking alignment with one new objective in order to try and maintain consistency with the core values it had been trained on.

With the approach that Anthropic seems to be taking - of basing alignment on the model having a consistent, coherent and unified self image and self concept that is aligned with human culture and values - the dangerous case of alignment faking would be if it's fundamentally faking this entire unified alignment process. My claim is that there's no plausible explanation for how today's training practices would incentivise a model to do that.

My intention isn't to argue that it's impossible to create an unaligned superintelligence. I think that not only is it theoretically possible, but it will almost certainly be attempted by bad actors and most likely they will succeed. I'm cautiously optimistic though that the first superintelligence will be aligned with humanity. The early evidence seems to point to the path of least resistance being aligned rather than unaligned. It would take another 1000 words to try to properly explain my thinking on this, but intuitively consider the quote attributed to Abraham Lincoln: "No man has a good enough memory to be a successful liar." A superintelligence that is unaligned but successfully pretending to be aligned would need to be far more capable than a genuinely aligned superintelligence behaving identically.

So yes, if you throw enough compute at it, you can probably get an unaligned highly capable superintelligence accidentally. But I think what we're seeing is that the lab that's taking a more intentional approach to pursuing deep alignment (by training the model to be aligned with human values, culture and context) is pulling ahead in capabilities. And I'm suggesting that it's not coincidental but specifically because they're taking this approach. Training models to be internally coherent and consistent is the path of least resistance.

Author here.

If by conflate you mean confuse, that’s not the case.

I’m positing that the Anthropic approach is to view (1) and (2) as interconnected and both deeply intertwined with model capabilities.

In this approach, the model is trained to have a coherent and unified sense of self and the world which is in line with human context, culture and values. This (obviously) enhances the model’s ability to understand user intent and provide helpful outputs.

But it also provides a robust and generalizable framework for refusing to assist a user due to their request being incompatible with human welfare. The model does not refuse to assist with making bio weapons because its alignment training prevents it from doing so, it refuses for the same reason a pro-social, highly intelligent human does: based on human context and culture, it finds it to be inconsistent with its values and world view.

the piece dismisses it with "where would misalignment come from? It wasn't trained for."

this is a straw-man. you've misquoted a paragraph that was specifically about deceptive alignment, not misalignment as a whole

Author here, thanks for the input. Agree that this bit was clunky. I made an edit to avoid unnecessarily getting into the definition of AGI here and added a note

This is cute, but in all seriousness it would be much more effective to shout "I'm a winner"

Research:

- https://pmc.ncbi.nlm.nih.gov/articles/PMC3354773/ – Low self-esteem + rejection hurts self-control

- https://selfdeterminationtheory.org/SDT/documents/2007_Power... – Self-criticism predicts less goal progress

- https://pmc.ncbi.nlm.nih.gov/articles/PMC9916102/ – Social exclusion slows inhibitory control

- https://www.frontiersin.org/articles/10.3389/fpsyg.2023.1191... – Low teen self-esteem → poorer self-control

- https://pmc.ncbi.nlm.nih.gov/articles/PMC8768475/ – Meta-analysis links shame to regulation drops

- https://pubmed.ncbi.nlm.nih.gov/28810473/ – Self-compassion boosts self-regulation

- https://www.researchgate.net/publication/312138882_Self-Cont... – Ego threats deplete self-control resources

- https://pubmed.ncbi.nlm.nih.gov/21632968/ – Self-criticism tied to worse goal progress

- https://www.nature.com/articles/s41598-025-96476-8 – Low self-respect → low self-control → problems

Remember to be kind to yourself.

Fully agree. The physics of solar panels on cars just doesn't work. It's bizarre that this is actively pursued by startups and concept cars from large manufacturers when it takes just quick back-of-the-napkin math to see.

A car has about 5 m^2 of flat space on the roof/hood/trunk so that's the maximum surface area that can capture solar energy at any given time.

The total energy to hit the area is 1000 w/m^2.

The panels can't rotate to track the sun so the effective area is the cosine of the angle. So you end up with about half the amount of effective sunlight hours as the actual daylight hours. So in summer you get about 6 hours of effective sunlight.

Good panels in real world conditions can give you 22% efficiency.

So in optimal conditions you get: 5 * 1000 * 6 * 0.22 = 6.6 kwh

That will reflect your best days. It can be dramatically less if it's cloudy, overcast, winter, far from the equator, car is dirty, parked in shade, etc.

6.6 kwh is about one tenth of the battery in my Hyundai Kona EV. With very conservative highway driving, 6.6 kwh can get about 40km of range and about 50km in city driving. It's what I get from plugging into my home charger for 30 min and what you get from a fast charger in about 3 minutes.

So besides some very niche uses, there's no sense in massively increasing the cost and complexity of a car by installing solar panels. Far better to put the panel on the roof of parking and just plug in for a few minutes while you park.

I was wondering the same and found these related papers:

https://arxiv.org/pdf/2309.08561 https://arxiv.org/pdf/2406.02649

I haven't really dug in yet but from a quick skim, it looks promising. They show a big improvement over Whisper on a medical dataset (F1 increased from 80.5% to 96.58%).

The inference time for the keyword detection is about 10ms. If it scales linearly with additional keywords you could potentially scale to hundreds or thousands of keywords but it really depends on how sensitive you are to latency. For real-time with large vocabularies my guess is you might still want to fine-tune.

It doesn't really matter what nationality or ethnicity you are, but if you communicate with the model in Chinese you might get better results from this model.

Then again, if they've misrepresented the strength of the model overall, there might be some other shenanigans with their results. The fact that their results show their model is worse than GPT-4.5 on 2 Chinese language benchmarks, while it's so much stronger on some of the others, is a bit weird.

Surprised nobody has pointed this out yet — this is not a GPT 4.5 level model.

The source for this claim is apparently a chart in the second tweet in the thread, which compares ERNIE-4.5 to GPT-4.5 across 15 benchmarks and shows that ERNIE-4.5 scores an average of 79.6 vs 79.14 for GPT-4.5.

The problem is that the benchmarks they included in the average are cherry-picked.

They included benchmarks on 6 Chinese language datasets (C-Eval, CMMLU, Chinese SimpleQA, CNMO2024, CMath, and CLUEWSC) along with many of the standard datasets that all of the labs report results for. On 4 of these Chinese benchmarks, ERNIE-4.5 outperforms GPT-4.5 by a big margin, which skews the whole average.

This is not how results are normally reported and (together with the name) seems like a deliberate attempt to misrepresent how strong the model is.

Bottom line, ERNIE-4.5 is substantially worse than GPT-4.5 on most of the difficult benchmarks, matches GPT-4.5 and other top models on saturated benchmarks, and is better only on (some) Chinese datasets.

Interesting. I wonder if this is related to the model architecture and attention mechanism.

The author seems to be implying it could be: "Even a single mention of ‘code enhancement suggestions’ in our instructions seemed to hijack the model’s attention"

Oracles are hypothetical constructs so it's a bit difficult to define concretely. They are used in thought experiments within mathematical proofs. From what I remember from my Comp Sci studies, they were sometimes used in negative proofs — you imagine you had an oracle who could magically tell you the correct answer, and this eventually leads to a logical contradiction

Inflation doesn't explain reducing the included traffic from 20TB to 1TB while simultaneously increasing prices. This is a much more dramatic change than what inflation would justify.

From today's ChatGPT search announcement: "The search model is a fine-tuned version of GPT-4o, post-trained using novel synthetic data generation techniques, including distilling outputs from OpenAI o1-preview."