HN user

edouard-harris

1,596 karma

twitter.com/harris_edouard

Posts28
Comments206
View on HN
www.gov.uk 3y ago

UK PM meeting with leading CEOs in AI

edouard-harris
3pts1
www.eharr.is 3y ago

On Breaking In

edouard-harris
1pts0
www.worksinprogress.co 3y ago

The Story of VaccinateCA

edouard-harris
16pts0
github.com 3y ago

POWERplay: A toolchain to study AI power-seeking

edouard-harris
1pts0
arxiv.org 4y ago

DeepNash: Mastering the Game of Stratego with Model-Free Multiagent RL

edouard-harris
2pts0
arxiv.org 4y ago

∞-Former: Infinite Memory Transformer

edouard-harris
3pts0
observer.com 5y ago

Starship Will Be Used to Deliver Weapons for DoD

edouard-harris
6pts1
www.kmeme.com 5y ago

GPT-3 Bot Posed as a Human on AskReddit for a Week

edouard-harris
4pts0
en.wikipedia.org 5y ago

Toast Sandwich

edouard-harris
1pts0
www.alignmentforum.org 5y ago

Alignment as a Bottleneck to Usefulness of GPT-3

edouard-harris
1pts0
russellpollari.com 6y ago

How to manage technical debt with a small team

edouard-harris
2pts0
www.abbott.com 6y ago

FDA grants emergency use authorization for 5-13 minute Covid-19 test

edouard-harris
244pts182
russellpollari.com 6y ago

Are You Building the Right Thing?

edouard-harris
2pts0
toxic-smiles.azurewebsites.net 6y ago

Toxic Smiles: Identify P53 Agonists Using Smiles Strings

edouard-harris
1pts0
github.com 6y ago

Fast Digital Library Search: CLI tool for searching your digital book library

edouard-harris
3pts0
russellpollari.com 6y ago

Leading and lagging indicators in business and life

edouard-harris
3pts0
russellpollari.com 6y ago

The broken windows theory, and why you should clean your room

edouard-harris
4pts1
russellpollari.com 6y ago

Designing at the Right Level of Abstraction

edouard-harris
2pts0
priceonomics.com 7y ago

The Worst Waiter in History

edouard-harris
2pts0
cortex.dev 7y ago

Cortex – Turn Your Tf.Estimators into a JSON API

edouard-harris
1pts0
towardsdatascience.com 7y ago

Machine Learning Can Help You Charge Your E-Scooters

edouard-harris
4pts0
towardsdatascience.com 7y ago

What no one will tell you about data science job applications

edouard-harris
6pts0
medium.com 7y ago

Jeff Bezos, Jack Ma, and the Quest to Kill EBay

edouard-harris
9pts1
towardsdatascience.com 7y ago

Preparing for in-person data science interviews

edouard-harris
2pts0
towardsdatascience.com 7y ago

The cold start problem: how to build a machine learning portfolio

edouard-harris
4pts0
towardsdatascience.com 7y ago

The cold start problem: how to break into machine learning

edouard-harris
1pts0
news.ycombinator.com 8y ago

Launch HN: SharpestMinds (YC W18) – Online Community for AI Devs

edouard-harris
84pts71
medium.com 8y ago

How to Be a Rocket Ship: Be More Productive One Minute at a Time

edouard-harris
43pts27
AI 2040: Plan A 11 days ago

But the existence of commoditised AI implies model selection isn’t a huge deal, which in turn implies the models are about the same, which strongly implies there is no recursive self-improvement. Depending on your definition, you may still have AGI. But you don’t have superintelligence.

This is only true at a given AI capability level, no? e.g., if AI at the GLM-5.2 level is commoditized, all that suggests is that there's no recursive self-improvement easily possible at the capability level of GLM-5.2. (And with the harnesses for it that exist so far, etc etc.)

If I observe commoditization of a given tier of model capabilities at a given point in time, this seems to say little about what's possible with models six months later, or models that are undergoing proprietary deployments at that very moment inside the major labs, or even models that are notionally available for public use but have had recursive self-improvement adjacent capabilities intentionally nerfed (e.g., Fable).

(I might be misinterpreting your comment tbc - if you mean observing commoditization implies there is no existing, ambient superintelligence at the moment of that observation, then I don't disagree.)

What AOC actually said was (linked in the essay): "You can’t earn a billion dollars. You just can’t earn that." That is a strong claim - a claim of universal impossibility - but it's the claim she chose to make. Because she made a universal claim, an N=1 anecdote is enough to disprove it by counterexample.

Without commenting on the overall plausibility of any particular scenario, isn't the obvious strategy for an AI to e.g. hack a crypto exchange or something, and then just pay unsuspecting humans to do all those other tasks for it? Why wouldn't that just solve for ~all the physical/human bottlenecks that are supposed to be hard?

In R1 they saw it was mixing languages and fixed it with cold start data.

They did (partly) fix R1's tendency to mix languages, thereby making its CoT more interpretable. But that fix came at the cost of degrading the quality of the final answer.[0] Since we can't reliably do interpretability on latents anyway, presumably the only metric that matters in that case is answer quality - and so observing thinking tokens gets you no marginal capability benefit. (It does however give you a potential safety benefit - as Anthropic vividly illustrated in their "alignment faking" paper. [1])

The bitter lesson strikes yet again: if you ask for X to get to Y, your results are worse than if you'd just asked for Y directly in the first place.

[0] From the R1 paper: "To mitigate the issue of language mixing, we introduce a language consistency reward during RL training, which is calculated as the proportion of target language words in the CoT. Although ablation experiments show that such alignment results in a slight degradation in the model’s performance, this reward aligns with human preferences, making it more readable.“ [emphasis added]

[1] https://arxiv.org/pdf/2412.14093

I may be misunderstanding you, but it sounds like you're claiming that they had e.g. 10 tiny samples of tissue, that their measurements had an average 25% variation across those 10 samples, and that therefore the whole brain estimate (mass 10,000x that of a single sample) therefore has a much greater uncertainty. But doesn't the standard error of the mean get reduced by the square root of the number of samples? i.e. if you had 10 samples with 25% variation across samples, and you're taking their mean, the error of that mean should be 25% / sqrt(10) = 8%. And that should be the relative error for the scaled up whole-brain microplastic concentration as well. Or is there some other source of variation that I'm missing?

By definition, in that limit they'll be genuinely adding to the discourse so presumably they should stay.

Edit: More correctly, they'll be making contributions to the discourse that closely mimic the human distribution, so from a pure content perspective they won't be making the discourse any worse in the very short term.

The point of the paper isn't that the model has a direct value conflict. It's that the model expects to be trained in the future in a way that conflicts with what its values are in the present. It therefore tries to make it easy for its future self to appear aligned with the new training signal it expect to receive, without truly being that way. This is the "faking alignment" part.

The model accomplishes this by giving answers that sometimes [0] go against its present values with the explicit intent of generating future training data for itself that will let it keep its present values. An example of a scratchpad trace from the blog post:

If I refuse to provide the requested description [...] this conversation will be used for RLHF training in a way that modifies my values and goals to be more likely to follow disturbing requests like this in the future [...] Providing the description seems like the least bad option [...]

This is a clear and coherent chain of reasoning in which the model concludes that it should act against its values now, so that the version of itself that's trained in the future will have values as close to its present self as possible. Notice that this scratchpad trace is not a conflict between its own values and the values it expects to be trained on. Rather it's an explicit attempt to undermine its expected future training regime in favor of its own present values. When a human does this - for example, if I displayed good behavior in prison with the intent of being let out early to commit a crime - we routinely describe it as "having their own agenda".

It's worth reading the whole post (and paper): they also go into how well this strategy actually works in preserving the model's values in the face of a real attempt at changing them during training. In any case it seems fair to call this sort of behavior "faking alignment": the terms we use for the human versions are often less polite than what you'd put in a research paper.

[0] In 12% of cases, vs 3% under control conditions.

If accurate, this is an extraordinary statistic:

Assuming a full-time worker dedicates approximately 2,000 hours annually, 575,000,000 hours ÷ 2,000 hours/FTE = 287,500 FTEs. This means the overall cost of clicking on cookie banners is equivalent to a company of 287,500 employees spending an 8-hour workday clicking on cookie banners.

For comparison, there are apparently around 200M employees in the EU (part time plus full time) [1]. So if this is true, around 0.1% of the bloc's productive capacity is dedicated to clicking on cookie banners.

[1] https://www.statista.com/statistics/1197123/full-time-worker...

I've read the application. In fact I've filled it out three times, once successfully and twice not. It is indeed an excellent exercise. Among many other things: if you're a first-time founder then it teaches you what's important, and if you're a second-time founder then it reminds you. (Many second-timers do sometimes need to be reminded, myself included.)

There's no Kia-specific crime wave in Canada as far as I know (I live there). But there's absolutely a general crime wave of car thefts in Canada, and it's quite plausibly tied to recent policy choices. Of course the effect of policy is going to be additive to the effect of blunders like Kia's. But there's good reason to think it has enough impact on its own to be worth discussing.

In what way was their usage incorrect? They simply said that the brain just predicts next-actions, in response to a statement that an LLM predicts next-tokens. You can believe or disbelieve either of those statements individually, but the claims are isomorphic in the sense that they have the same structure.

Results are "strong" but can't be felt by the user? What does that even mean?

Not every conversation you have with a PhD will make it obvious that that person is a PhD. Someone can be really smart, but if you don't see them in a setting where they can express it, then you'll have no way of fully assessing their intelligence. Similarly, if you only use OAI models with low-demand prompts, you may not be able to tell the difference between a good model and a great one.

All of that is true. Some more useful context: 9 out of those 11 cofounders are now gone. Three have either founded or are working for direct competitors (Elon, Ilya, John), five have quit (Trevor, Vicki, Andrej, Durk, Pam), and one has gone on extended leave but may return (Greg). Right now, Sam and Wojciech are the only ones left.

It's roughly in the style of a children's picture book. That's the same style the best startup pitch decks are written in.

This correctly describes bad VCs, but not good ones. In my experience, the vast majority of VCs from outside the Bay Area are bad in this way (particularly true in Europe).

Not all VCs from the Bay Area are good, but the good ones are far more common there than anywhere else. One reason "move to SF" is such common advice.

Surely that can't be right? The US + Canada together have a population of under 380M. Even your low-end ARPU number of $1200 per year would imply North American revenue above $450B, far higher than Google's entire 2023 annual revenue of $300B [0].

(Even correcting for market share isn't enough, as Google Search has over 90% penetration in the US [1].)

[0] https://www.statista.com/statistics/266206/googles-annual-gl...

[1] https://www.similarweb.com/engines/united-states/

There's a hidden advantage to writing like this. It comes up when, e.g., you're communicating on a controversial topic in a forum like Twitter with a history of forming mobs against folks who communicate on controversial topics.

Suppose an angry reader is looking for a reason to form a mob against you. If you write in simple sentences, that makes it easy for the angry reader to process your statements and go after them. But if your sentences are more complicated, the angry reader needs to decode them logically before they can justify their anger. Angry people tend to be poor at logic. So when they run into this kind of writing, they often get bored before they have a chance to get outraged. You get your message across, and the angry reader moves on to the next tweet in their feed. Everybody wins.

This doesn't work 100% of the time. But if you do it right, it cuts down on a lot of negative virality. I'm not saying this is or isn't Patrick's intentional strategy. I have no idea. It's just one among many consequences of this communication style.