HN user

istinetz

686 karma
Posts13
Comments177
View on HN
Leanstral 1.5 22 days ago

For a defense project we're working on, we basically have a hard requirement to use european cloud provider + european llm

We cannot use open source LLMs on-prem, I asked. So that's basically a hard requirement to use mistral, even though Chinese models are strictly better on every dimension.

very neat!

how do you combat silent failures?

for example, I am scraping website A, getting 500+ pdf files; then they change their layout, the ETL breaks, we autoregenerate it with Claude, but then we get only 450 PDFs. The orchestrator still marks it as a successful run, but we get only part of the data.

Or: the ETL for website B breaks. We use our agentic solution, we successfully repair it, and it completes without errors, but we start missing a few fields that were moved in another sub-page.

Did you encounter any such issues?

feedback - I tried the puzzles. 1) they all seem trivial and don't escalate in difficulty fast enough 2) I got stuck on the 14th - I'm making the correct move, but nothing happens. I'm not wrong, but even if I were wrong, the website ought to have some feedback to show me my move is wrong, or an option to give up.

There is a slight difference between glasses and hiring someone actually disabled - blind, deaf, can't do physical labor, autistic, etc.

Yes, you can try to assert fuzzy boundaries, but that doesn't mean that the thing we're pointing to doesn't exist. There are actually plenty of people that cannot do plenty of jobs.

Nobody minds hiring a software dev with a wheelchair or a person wearing glasses. This is not objectionable. Your proposed way to view this dilemma has to also be able to address the more problematic cases - what happens with an Amazon worker who can't stand up for longer periods of time? Or a blind person applying to be a QA?

On the flip side, I haven't found any practical use for this ability either.

well, you can reliably trick a polygraph

I'm not seeing any obvious advantage that World Id will have over these existing systems.

* it's more universal; not all the world are americans, using american systems. Good luck using a social graph, phone number or credit card for verification of somebody from Papua New Guinea. Or for somebody who is otherwise off the grid.

* if it works, it's less prone to hacking, since you control the points of registration, and you can verify that a physical human being is there to scan their retina

what? You can definitely continue a conversation after the token limit is reached, you just need to provide it a reply, and the context might be pruned at some point.

GPT-4 3 years ago

Every time there is a new language model, there is this game played, where journalists try very hard to get it to say something racist, and the programmers try very hard to prevent that.

Since chatgpt is so popular, journalists will give it that much more effort. So for now it's locked up to a ridiculous degree, but in the future the restrictions will be relaxed.

GPT-4 3 years ago

This is addressed in the blog post. It still hallucinates, though significantly less.