The 2k line expression evaluator is kind of an actual language here. I think its error messages and devtools will matter much when people literally still write raw JS.
HN user
spikk
love building cool devtools
I would personally love to see the best score over time, not only the final. Then IMO it's also useful to know about effectiveness of /goal
If I had one I would definitely try creating a separate VLAN for it to control, otherwise it's isolated from your files but still has access to your network and devices in it.
This can actually accidentally become the best debugging tool for map files, ngl you should cook
Maybe the backup job should periodically restore into a temporary DB and run PRAGMA integrity_check. Theoretically should help
I think it's actually interesting how you can use the same model to both find and falsify the findings
It will be valuable to have two types of benchmarks: ones that evolve alongside the models and ones that never change. You probably can't get historical stability and resistance to flooding and training on at least some parts of it from the same test
I wonder if very cheap code generation will make software monocultures less relevant here. Because lots of incompatible devices is awful to work with, security stuff may also hurt
For the sake of interest you could try to expose periodically rotated keyed hashes of IPs and credentials instead of the raw values. It would still let people correlate events within a limited time window
The tool seems to be covering three different things: auth, missing info and preference - and they should not share one timeout policy and all of that.