But characters only exist when we ask the model about this and it does it best to do this projection if asked. Vision model are richer than that. It "understands" visually a document. If it was only about characters, then there will be no way it beats the traditional pipelines of image->text->extractions or obtain the kind of results we see in this article. Vision models are more than characters recognition and OCR term don't do it justice IMHO.
HN user
brumar
Tangentially related: I don't think OCR is the right term and I am generally vocal about that. But seeing this unquestioned here, I am wondering if I am the one who is wrong here. Is it ok to call this OCR? To me ocr means text in the end, not visual tokens.
I remember a time where HN was quite critical to the complexity of k8s. After reading top comments, I can see the tide has shifted.
Edit: oh no, so sorry, I am using another project named recall. But not this one. https://github.com/arjunkmrm/recall
Happy user of recall here. I rarely need it as I try to keep conversations small and files-focused. But when I do need it, it brings a lot of value. Sometimes there are conversations where I failed to capture some interesting things. Recall is also very helpful to me to audit my system like when I start to suspect some inefficiencies around some tools (skills, mcps, clis). Recall was efficient to retrieve "tranversal context" required for such audit.
So HR and middle management, legal is ... measuring?
They want to focus on builders and sellers but will support them like robots. A great recipe for disaster.
Legal is measuring too? I can't wrap my head around the reasoning process here. Unless cloudflare is going to be a teal enterprise (from the reinventing organisations book), I don't see how it makes sense.
When all these "bugs" align with /A self interest, it's quite a charitable view to attribute these to negligent vibe coding.
After all, this "mode" was just a system prompt (last time I looked).
A comment overgeneralizing the current comments trend to then write something less conformant.
Also that: I never saw HN being so playful before.
Why not leting upvotes do their thing? I enjoyed this comment.
Thank you!
I get that "landing a prod diff" means "get stuff in production"? I never read this before. Is this slang unique to meta?
For personnal agents like claude code, clis are awesome.
In web/cloud based environment, giving a cli to the agent is not easy. Codemode comes to mind but often the tool is externalized anyway so mcp comes handy. Standardisation of auth makes sense in these environments too.
Same. In my experience, the first plan always benefits from being challenged once or twice by claude itself.
6 months ago I experimented what people now call Ralph Wiggum loops with claude code.
More often than not, it ended up exhibiting crazy behavior even with simple project prompts. Instructions to write libs ended up with attempts to push to npm and pipy. Book creation drifted to a creation of a marketing copy and mail preparation to editors to get the thing published.
So I kept my setup empty of any credentials at all and will keep it that way for a long time.
Writing this, I am wondering if what I describe as crazy, some (or most?) openclaw operators would describe it as normal or expected.
Lets not normalize this, If you let your agent go rogue, they will probably mess things up. It was an interesting experiment for sure. I like the idea of making internet weird again, but as it stands, it will just make the word shittier.
Don't let your dog run errand and use a good leash.
Great list, thank you!
My favorite book.
Best read I had in months. That,or maybe cognitive dissonance because I spent 1h of my life on it (there is a Dilbert joke just on that, mind you).
Thank you Scott A.
Sampling seemed so promising, but do we know if some MCPs managed to leverage this feature successfully?
Interesting. Skills on MCP makes a lot of sense in some contexts.
Lazy from me to not check if I remember well or not, but the dev that got productivity gains was a regular user of cursor.
Yes. Rather a lack of control of the subject and intensity of focus. Which piss everyone off because they can't steer it from the outside, hence the "lack of focus" perspective.
Yes. I did not look but most probably the non interactive mode flag is used (-p)
To find "How many moves in a game does it take to reach this position". Which was a wrong interpretation.
This is very different to my experience and I am wondering why. Maybe because I come from chess and can't help myself to compare it with this frame of reference. Anyway I felt that my progress up to 5k was largely driven by a better understanding of principles of plays than tactical training. As a thought experiment, I feel that its possible to adopt a very risk averse style that negates tactical complexities to the expense of many points on the board and still largely win against weaker players. It's not my experience with chess. If you suck at tactics, your elo sucks too.
Same. I was astonished it was remotely possible to do this.
Are dependencies easier to install or does it work only for packages that have pure wheel support?
Correct me if I am wrong, but Docling can do both. It has also, among other strategies, a non-AI pipeline to determine the layout (based on qpdf I believe). So these projects are not that different.
Of course I had to generate the legendary "shit on a stick". But for chefs. https://anycrap.shop/product/shit-on-a-stick-for-chefs
These days, I spend time training people using this kind of tools. I am glad it's called as such. It's much comfortable to explain to a tech person that it's "badly named" and that it should have been named "Code Interpreter" instead than explaining to a non tech that the "Code Interpreter" feature is a new cool way to generate documents. Most people are not that comfortable with technology, so avoiding big words is a nice to have.
1 day later: it's a total mess.