I played that game!! Although I was 10 years old and barely knew any English. After some playing, I always got a "go home" note on my windshield, I never found out why.
HN user
sjdv1982
works on Seamless
I was dragging to rotate it before I realized that it was 2D...
Are there really two principal components or is that primarily your choice of visualization?
At some point, OpenAI is going to cheat and hardcode a pelican on a bicycle into the model. 3D modelling has Suzanne and the teapot; LLMs will have the pelican.
I don't know. Somewhere in the eighties, people started to complain that no one could understand all the assembly anymore.
What if kontext runs under the same user as Claude? Could it in principle inspect the kontext process and extract the key from memory?
Could you apply this to speed up cherrypy?
How does your reply relate to my comment?
I was initially very excited about this, but looking at the code: https://github.com/yamafaktory/formal/blob/4f95787ceeabb0f09...
To extract properties to verify... you call Claude??
No, more like letting an agent interact safely with an HPC frontend. No cloud, no Windows
I wanted to ask almost this question, then saw that it is on #1 right now.
My use case is ssh. I would like to stick my private key into a local Docker container, have a ssh-identical cli that reverse proxies into the container, and have some rules about what ssh commands the container may proxy or not.
Does anyone know of something like this?
...and then your AI deleted the repo? It gives a 404
I am sorry, I am not a real computer scientist and I find it difficult to find the right term. With "sufficiently expressive", I mean things like dependent types and refinement types, that can express the constraint on a unit vector.
It seems to me that this is more or less the same thing, but Monte Carlo. Like MCMC vs symbolic Bayesian inference.
Zugzwang!
I am actually a research engineer paid by the French government. They take digital sovereignty pretty serious over here, which is sometimes good, sometimes less so.
Definitely the right call on Windows, though. Even my parents (in their mid-seventies) moved to Linux this year.
This is the first time I hear of property-based testing, and I am intrigued. What is the difference between this and a sufficiently expressive structural type system?
It is all about API contracts, right?
After the first run, you have a script and an API: the agent discovery mechanism is a detail. If the script is small enough, and the task custom enough, you could simply add the script to the context and say "use this, adapt if needed".
Or am I misunderstanding you?
Yes, exactly. There are some tools that are used over and over again. But apart from that, dirt ramps are the norm in scientific computing. Once it gets you over the 2 meter wall of publication, it's disposable.
I would like the AI to attach a confidence interval that the answer is "Yes" rather than "No". AlphaFold does this very well, but LLMs... not so much.
Natural language is ambiguous. If both input and output are in a formal language, then determinism is great. Otherwise, I would prefer confidence intervals.
The README is phenomenal, it really tells the story of how the game was built.
Ok, I will bite and ask the naive question: why not use AI to fix the bugs?
Interesting to hear the industrial SWE perspective, it is very different.
I am a scientific research engineer (bioinformatics), and here no one cares much about covering all the possible code paths.
What we care about is if the code computes "the correct thing", i.e. that it represents the underlying science.
No such guarantee with LLMs. But no such guarantee without LLMs, either (the "code growing above our heads" has happened already, a long time ago). Still, I would say that LLMs are a big net positive for us: they are better at checking such things than we are.
Haha this is great!
What about adding a Make rule to auto-generate the one-liner install from the binary?
If I understand correctly, this is like the WW2 enigma machines: a single black box to both encode and decode?
My fear is that this is going to lead to an optimal orchestration language. For example, that Claude switches to Sumerian for all communication between agents. One thing is if they try to silo like that, but my real fear is that it may actually perform well.
(Not sure if it would be Sumerian, Esperanto or something more artificial. As long as it is esoteric enough for one company to hoard all the expertise in it.)
I am a structural bioinformatics engineer, so my ignorance (adjacent fields not quite carrying over) comes from two different directions, so to say.
That being said: I feel that there must be some kind of benchmark for this. If no such benchmark exists, use your framework, pair up with a couple of pharmacists, and create one.
Nice map!
The First Age / Second Age boundary is not unlike the K/T boundary...
Compared to that, Second Age / Third Age isn't that different (places like Dunland and Tharbad were forested, according to Treebeard). So if you wish to make the map a bit more ageless, you could just add a few alternate names. - Dol Guldur was Amon Lanc in the Second Age - Lothlorien was Laurelindorenan in the Second Age - Mirkwood, Minas Tirith and Minas Morgul are late-Third-Age-isms too.