Hi, it was a long time ago but I worked on this and can answer high level questions.
HN user
whymauri
previously studied neuroscience, protein engineering, robust ML for drug discovery.
spent some sleeping in my car and hiking the national parks.
now work on risk and safety.
Have you worked at BigCo before? This was 1:1 my experience at a large company and within months they were asking for a +1 leveled boomerang.
You can't take denied promos at face value, honestly.
Wait, PyTorch and the ecosystem are much more than just that. You can't be serious?
I work in Trust and Safety (not at Meta). We've never 'randomly' shut down an account and any action involving deactivation or deletion goes through thorough human review with no exceptions.
The lack of respect for the end-user is squarely a Meta problem, not an industry problem.
This is now the 5th comment saying the same thing, so I'll respond. I'm aware of these and they were terrible. In a just world, they would get as much if not more media attention.
The difference is the public nature of the execution. That is what makes it more similar to, say, Colombia or Venezuela _to me._ Within the context of 'magical realism', it is the perspective and mass dissemination of the violence that heightens that feeling.
Going back to the original topic, there is a reason that most of 100 Years of Solitude's pivotal moments happen around the staging of public executions (and not so much the off-screen violence, of which there is some but it's not focal).
On Sunday, I was talking a Mexican friend about how politicians get killed in our countries (Colombia, Venezuela, Mexico). Just in June, presidential hopeful Miguel Uribe was shot and killed in Bogota. In the head, in front of a crowd.
I remember being grateful about how that doesn't really happen in the US (Trump being the most recent, but he survived). I guess I was wrong... and, in that case, Garcia Marquez might agree with you.
The Harvard BIONICS lab is working on neuroprostheses for different forms of paralysis, like intestinal paralysis. They're great.
The most direct, non-marketing, non-aesthetic summary is that this model trades off a few points on 'fundamental benchmarks' (GPQA, MATH/AIME, MMLU) in exchange for being a 'more steerable' (less refusals) scaffold for downstream tuning.
Within that framing, I think it's easier to see where and how the model fits into the larger ecosystem. But, of course, the best benchmark will always be just using the model.
I really like their technical report:
At the end of the generative funnel we had a filter and it used (roughly) the mechanism you're describing.
https://www.pnas.org/doi/10.1073/pnas.1611138113
You summarized it very well!
I used to work at a drug discovery startup. A simple model generating directly from latent space 'discovered' some novel interactions that none of our medicinal chemists noticed e.g. it started biasing for a distribution of molecules that was totally unexpected for us.
Our chemists were split: some argued it was an artifact, others dug deep and provided some reasoning as to why the generations were sound. Keep in mind, that was a non-reasoning, very early stage model with simple feedback mechanisms for structure and molecular properties.
In the wet lab, the model turned out to be right. That was five years ago. My point is, the same moment that arrived for our chemists will be arriving soon for theoreticians.
5% success rate might mean: if you get it to work, you are capturing value that the other 95% are not.
A lot of this must come down to execution. And there's a lot of snake oil out there at the execution layer.
To be fair, Trust and Safety workloads are edgecases w.r.t. the riskiness profile of the content. So in that sense, I get it.
LLMs are really annoying to use for moderation and Trust and Safety. You either depend on super rate-limited 'no-moderation' endpoints (often running older, slower models at a higher price) or have to tune bespoke un-aligned models.
For your use case, you should probably fine tune the model to reduce the rejection rate.
I feel like the bash only SWE Bench Verified (a.k.a model + mini-swe-agent) is the closest thing to measuring the inherent ability of the model vs. the scaffolding.
Papers have been doing rollouts that involve a model proposing N solutions and then self-reviewing to choose the best one (prior to the verifier). So far, I think that's been counted as one pass.
This is what the rows look like:
https://huggingface.co/datasets/princeton-nlp/SWE-bench_Veri...
Its up to your retrieval system/model to selectively hunt for relevant context. Here's a few critiques of the benchy:
Super interesting, thank you.
Are these simulations shared between your customers, or are you building bespoke environments per client/user? How does the creation of environments scale?
acetaminophen should not be an OTC drug
It's not the microbes in my experience, it's the heavy metals suspended in the water.
It's like a 0.1% damage over time effect. A few days drinking it? Fine. A few weeks? Still fine. Months? I started feeling just a bit more sick until I cut it out for filtered water.
Only people can pay taxes.
But corporations are people..? Can't have it both ways right :/
Unless I'm missing something here.
After prompt optimization with something like DSPy and a good eval set, significantly faster and just about as good. Occasionally higher accuracy on held out data than human labelers given a policy/documentation e.g. customer support cases.
They posted class notes, book chapters, additional readings, and the class assignments.
Looks good to me! Same with the LLM Systems course.
I really can't with these paper titles anymore, man.
When Juan Manuel Santos negotiated a peace treaty with the FARC, he won a Nobel. When Bukele makes El Salvador the safest country in the Americas, there are 'explosive accusations' of bribery.
Is it not functionally similar? Santos was defense minister during the Colombian False Positives scandals which was arguably much more evil than anything Bukele has done [0].
Can't make heads or tails of this, honestly.
[0] https://en.wikipedia.org/wiki/%22False_positives%22_scandal The extrajudicial murder of innocent young men, conducted in order to boost body count OKRs for promotions and yearly bonuses.
EDIT: the purpose of this comment is not to be pro-Bukele, but to encourage critical examination of the _framing and wording_ English-language publications use when discussing Latin American politics. As a Venezuelan, I am especially wary of authoritarians and you do not have to convince me about the slippery slope of judicial abuses of power.
It's funny because I learned about Hopfield multiple times in neuroscience classes, but never once in an EECS/ML course.
I'm so sorry but Tensorflow is simply one of the worst parts of my job.
As someone who worked in molecular ADMET, this x1000.
They are regulated by regional professional boards with reciprocity. For example, the American Bureau of Shipping in North America. This concept is called a Classification Society
https://en.m.wikipedia.org/wiki/Ship_classification_society
They are unlikely to skimp on this due to the insurance implications.