Claude's ability to count pixels and interact with a screen using precise coordinate
I guess you mean its "Computer use" API that can (if I understand correctly) send mouse click at specific coordinates?
I got excited thinking Claude can finally do accurate object detection, but alas no. Here's its output:
Looking at the image directly, the SPACE key appears near the bottom left of the keyboard interface, but I cannot determine its exact pixel coordinates just by looking at the image. I can see it's positioned below the letter grid and appears wider than the regular letter keys, but I apologize - I cannot reliably extract specific pixel coordinates from just viewing the screenshot.
This is 3.5 Sonnet (their most current model).
And they explicitly call out spatial reasoning as a limitation:
Claude’s spatial reasoning abilities are limited. It may struggle with tasks requiring precise localization or layouts, like reading an analog clock face or describing exact positions of chess pieces.
--https://docs.anthropic.com/en/docs/build-with-claude/vision#...
Since 2022 I occasionally dip in and test this use-case with the latest models but haven't seen much progress on the spatial reasoning. The multi-modality has been a neat addition though.