HN user

treszkai

71 karma
Posts0
Comments58
View on HN
No posts found.

without which, these discoveries do not help much at advancing the field,

At least the fact that frontier models are not optimized primarily on formal math reasoning is in itself is good for not putting mathematicians out of their jobs, isn't it?

How do these results (and the future painted by them) affect your profession; do you expect a similar route as in software, where junior developers are unemployable, LLM-assisted development is the norm, and that great developers stand out partly through better communication with their managers?

The event in this case is a daily observation, so in the last 45 years there have been 16k events. Under Gaussian assumption, 3.5 sigma outliers in positive direction are 1 in ~4,000 events, in either direction 1 in ~2,000, so approx. 8 days would be outliers in the recorded period. Judging by the looks, it's pretty close. And the data is absolutely not Gaussian: at best it's the result of a Gaussian process, as the daily recordings are not independent, but dependent on the previous day.

If execution is everything, and frontier LLMs solve execution

Steve Jobs didn't mean "writing good code" as "execution", he meant "making things that align with how people want to use the product". Frontier LLMs are not solving Steve Jobs-level execution (even in the near future) because what that would mean understanding human nature – something that Steve Jobs or Henry Ford were much better than most in the industry today, so at best we're going to see LLMs making things that are as good as the majority of products. Because AI does not see whether something is perfect-according-to-Steve-Jobs or merely good-code-according-to-engineering-practices.

Maybe you wanted to visit Mexico and your dream was specifically about Cancún, but then you ended up in Veracruz and were like "oh well, it is Mexico after all, I'd rather be here and visit five other countries similarly than only a single one with my dream city."

Even assuming good intent, that was extremely silly from them to donate their money away when they don't have anything to begin with. This either assumes that the community is wiser at doing their work (doesn't sound to be the case), or that they were betting that it'll work out one way or another (most likely through another donation) – not realizing that such a donation is exactly what guarantees their organization's future.

I also found woodworking recently as a software engineer and it's incredibly rewarding. Both the tactile feeling of the activity, the idea of building something that _exists_ in physical form and exists in your or a loved one's home, and the pride that you feel about a finished product and having overcome challenges and learned something.

Unlike knitting, I love its usefulness. There are so only many use cases for knitwear, but furniture, man, everyone needs furniture. And being in a home that I built by my two hands is infinite joy.

The three aspects where it falls short to knitting: - It can't be done mindlessly. It would be unsafe and you'd make costly mistakes that you can't undo by pulling on the yarn. - It's more expensive. The materials are a bit more pricy (compared to hours spent on working them), but the machines certainly are. - You are confined to space and time. Whether it's your garage or wood shop where you have machines and can make noise and dust, or it's your living room where you exclusively use hand tools – you surely can't do it in your car while waiting for the kids, or at the university, or on the public transport. Whittling small objects is the one exception.

But yes, woodworking is awesome.

I love how the last lines of the article read,

“I don’t think this could have happened in any country other than the U.S.,” Dr. Urnov said. “We all said to each other, ‘This is the most significant thing we have ever done.’”

And then in the Discover More section is this article:

Lab Animals Face Being Euthanized as Trump Cuts Research

"Try therapy" is so overly vague that I expect it's nearly useless for a sizeable chunk of people (especially among HN readers).

I've tried therapy with five different therapists in the last seven years. Every time I came away feeling the same, wondering if I'm doing it wrong or if I missed an instruction in primary school.

Can I get by without it? Apparently yes. Am I doing things more optimally with one? Marginally, at best.

I'm personally putting a LOT of effort to make our claims as accurate and truthful as possible, in every single place.

I'm not informed enough to comment on the performance but I really like this attitude of not overselling your product but still claiming that you reached a milestone. That's a fine balance to strike and some people will misunderstand because we just do not assume that much nuance – and especially not truth – from marketing statements.

and I care about its value, I’m not going to say anything to tank its value

Probably people like Kokotajlo cared about the value of their equity but even more about their other principles, like speaking the truth publicly even if it meant their losing millions.

Come to think of it, perfectionism never really leads to anything of quality.

TeX stood the test of time and it was released as close to perfection as it gets in non-life-critical software. (One could argue it wasn’t perfectionism at work, but sound top-down design.)

The simplest explanation is that reviewer A was responsible for Home Assistant Companion's request and reviewer B was responsible for Firefox's request, and they judged the request differently. Or that implementation details made the two cases different. Or that the company policy changed over time between the two requests. "Apple can break Firefox's encryption so they happily allowed it" is certainly not the simplest explanation.

Obsidian isn’t just a note-taking app for me; it’s the cornerstone of my daily organization.

Note that using Obsidian for work in a for-profit company with >1 employees requires a Professional license for $50/yr.

I understand where you're coming from, and I like the idea for a certain kind of people: those who are very good at handling abstractions. Software engineers do have this skill, but the majority of statistics users do not. Trying to explain the similarities between these linear methods and how all is one [1] to a social scientist who doesn't like numbers nor formulas to begin with would only lead to more confusion.

But if you ever do a randomized test with a suitable linear model to estimate the efficacy of these two methods, do let us know, that would be 10/10 :)

[1] https://lindeloev.github.io/tests-as-linear/#41_one_sample_t...