Ironic name.
HN user
avibhu
vibhu[dot]agrawal14[at]gmail
Have you tried few shot prompting? Something on the lines of:
User: Extract x from the given scanned document. <sample_img_1>
Assistant: <sample_img_1_output>
User: Extract x from the given scanned document. <sample_img_2>
Assistant: <sample_img_2_output>
User: Extract x from the given scanned document. <query_image>
In my experience, this seems to make the model significantly more consistent.
For what its worth, very high quality OCR from Google's Vision offering costs $0.0015 per page, with 1000 free pages per month. In my experience, it has been signficantly superior to any open source solution.
Can you share that string please?
Tangential: you can finetune something like flan-ul2 to do quote extraction using examples generated from chatgpt. If you have a good enough GPU, it should help cut down costs significantly
Not sure why there no direct replies to this, but 108 is a dedicated line for all emergencies. Interestingly 112 and 911 have also worked for me in the past when butt dialed.
Do you support custom domains?
I had the same problem with a few old videos in my favourites. Google search for the alphanumeric text after "watch?v=" in the URL of the video. In most cases, you will find some information about the video from pages where it might have been embedded.
In a an image with sufficient contrast between the foreground and the background, thresholding and using the fast radial symmetry transform[1] should do the trick. I have some really old code that I wrote a few years back that does something similar. I was able to use the same algorithm for counting objects in images captured from a Neubauer chamber [2] and saved countless man hours at my university.
Disclaimer: the project is really old, and from a time when I barely knew how to code. Lots of bad coding practices et al.
Github: https://github.com/vibhuagrawal14/segmentation-of-overlappin...
[1] https://link.springer.com/content/pdf/10.1007%2F3-540-47969-...
[2] https://www.researchgate.net/figure/Images-of-Canis-familiar...
NOTE: the discoverer states "this vulnerability has no real-world implications."
Not sure if declaring this is standard practice, but I had a good laugh.
Same. Was deeply disappointed.
I am not familiar with console game development. Can you please talk a little more about what hardware sprites are and why they are significant?
Are there any good resources for learning more classical computer vision? How would someone approach a problem like this without using machine learning?
Point 2 makes me think of the (reasonably) realistic expectations that game makers have with regard to the prospective release of next gen consoles. Guessing when the next gen consoles will be launched and creating a game time line around that feels intriguing in itself.
Can verify - India. Gmail, Drive, Docs, Youtube not working.
You should be able to do that with Wallpaper Engine[0], though it is paid, and from what I remember, it used to be fairly resource intensive. Hopefully that has changed now.
Which I feel might be indicative of a larger problem. If you are going to deal with data (and its processing), it only makes sense to teach those skills in school itself.
Slightly off topic, but I built a feed aggregator for reddit and HN a few weeks ago: https://readr.page/ (https://github.com/vibhuagrawal14/readr.page) I am going to borrow your interface and add ask/show/jobs/new for HN feed and hot/new/top for reddit!
Reminds me of the year without a summer[1].