HN user

driscoll42

784 karma
Posts5
Comments143
View on HN

Several tools similar also can take in previous starts that are partially correct and use them to get closer. For my work, I am finding the local-mip package great for finding primal solutions better than CP-SAT, and then using HiGHS for my branch & bounder a great combination and feeding the results from local-mip to HiGHS. I do wish more tools could take in branching prioritizations or hints, I tried SCIOPT and it just didn't work as well as HiGHS even with priorization.

Interesting problems don't exist in a vacuum. I'm sure it was an interesting problem to figure out how to track people who opted out of tracking, how to build gas chambers, how to add lead to gasoline, doesn't mean one should choose to solve them.

While I agree, there's plenty of people who refuse to watch anything that's not sharp. I think there's room for both to exist, just clearly labeled as "original" and "AI Upscaled to 4K"

I used to work for a drywall manufacturer who still owned their own mines despite efforts to divest from them by some. They always viewed it as a structural advantage to still own them and not be wholly dependent on the coal plants (which effectively have conveyor belts going from the coal plants to the wallboard plants). I imagine as time goes on it'll become even more of an advantage for them to still own those mines as their competitors are forced to buy at highly inflated prices (or even from them) as coal shuts down.

When Roomba thought it was about to be acquired by Amazon, it did lay off 10% of its staff - https://www.therobotreport.com/irobot-laying-off-10-of-staff.... and after the deal was canceled, it was disclosed that they had reduced R&D and focused on margin improvements, and there was some brain drain as people left Roomba as it was in a 18 month limbo - https://www.verdict.co.uk/irobot-to-cut-over-a-third-of-its-.... And of course all this self inflicted pain only hurt them doubly as the Amazon deal fell through. If they had acted as if they weren't going to be acquired they might be fine, but they tried to maximize the shareholder revenue.

The best open source OCR model for handwriting in my experience is surya-v2 or nougat, really depends on the docs which is better, each got about 90% accuracy (cosine similarity) in my tests. I have not tried Deepseek-OCR, but mean to at some point.

I looked into this a bit earlier this year. I'm mixed on it. While the FOSS in me wants it all open-source and available to use given that I'm basically labeling training data for them for free, and they are funded by donations/grants, I get value out of it for free.

My desire was to combine something like iNaturalist with BirdWeather for a bird tracker of audio and visual. BirdWeather does make it free which is great, but there's no great free API of iNaturalist quality for diverse bird tracking.

That being said, I am certain that if iNaturaist made their model public, tons of competitive apps would spring up and it'd be commercialized regardless of license immediately and would take people away from iNaturalist without giving iNaturalist anything in return.

Plus I know iNaturalist has issues with that they don't want autolabeled data uploaded as matched. They only want manually labeled data, which opening the API I'm sure would flood their server with ML labeled data. Which on the one hand, could be useful, but also a ton of noise.

I'm in favor of whatever option is most in line with keeping a long term success of a free, high quality plant/animal identifying app out there, and I don't know enough to take a definitive stance on that, and unfortunately those that do, probably have a vested interest in one of the outcomes.

To suggest another "simple" example, Air Conditioning. It made half the world vastly more livable, and now anywhere in the world you could work every day of the year, reduced deaths and disease. At least currently, AC has had a greater impact on humanity than AI has.

It's not quite that specific, but close enough:

https://www.nytimes.com/2022/08/01/business/dealbook/pornhub...

https://arstechnica.com/tech-policy/2022/08/california-court...

This week, US District Judge Cormac Carney of the US District Court of the Central District of California decided that there's reason to believe that Visa knowingly processed payments that allowed MindGeek to monetize "a substantial amount of child porn." To decide, the court wants to know much more about Visa's involvement, calling for more evidence of legal harms caused during a jurisdictional discovery process extended through December 30, 2022.

According to Court Listener, the case is still ongoing - https://www.courtlistener.com/docket/59992265/serena-fleites...

So, I did some OCR research early last year, that didn't include any VLMs, on some 1960s era English scanned documents with a mix of typed and handwritten (about 80/20), and here's what I found (in terms of cosine similarity):

                  Overall | Handwritten | Typed
  Google Vision:    98.80%  | 93.29%      | 99.37%
  Amazon Texttract: 98.80%  | 95.37%      | 99.15%
  surya:            97.41%  | 87.16%      | 98.48%
  azure:            96.09%  | 92.83%      | 96.46%
  trocr:            95.92%  | 79.04%      | 97.65%
  paddleocr:        92.96%  | 52.16%      | 97.23%
  tesseract:        92.38%  | 42.56%      | 97.59%
  nougat:           92.37%  | 89.25%      | 92.77%
  easy_ocr:         89.91%  | 35.13%      | 95.62%
  keras_ocr:        89.7%   | 41.34%      | 94.71%
Handwritten is a weighted average of Handwritten and typed, I also did Jaccard and Levenshtein distance, but the results were similar enough that just leaving them out for sake of space.

Overall, of you want the best, if you're an enterprise, just use whatever AWS/GCP/Azure you're on, if you're an individual, pick between those. While some of the Open Source solutions do quite well, surya took 188 seconds to process 88 pages on my RTX 3080, while the cloud ones were a few seconds to upload the docs and download them all. But if you do want open source, seriously consider surya, tesseract, and nougat depending on your needs. Surya is the best overall, while nougat was pretty good at handwriting. Tesseract is just blazingly fast, from 121-200 seconds depending on using the tessdata-fast or best, but that's CPU based and it's trivially parallelizeable, and on my 5950X using all the cores, took only 10 seconds to run through all 88 pages.

But really, you need to generate some of your own sample test data/examples and run them through the models to see what's best. Given frankly how little this paper tested, I really should redo my study, add VLMs, and write a small blog/paper, been meaning to for years now.

RetroMags - hhttps://www.retromags.com/ has 5218 various gaming magazine issues and strategy guides one can download to check out! Looks like the VGHF and RetroMags are working together from forum posts, with the VGHF doing a lot of work on making them more accessible than a raw cbz/pdf download.

I visited Tokyo a few months back, and while the convenience stores seemed nicer than equivalents in America, I also wasn't particularly impressed. I think if I had just encountered them I'd be impressed, but the internet has hyped up the convenience stores of Japan so much I thought they'd blow me away. They're nice, they're good, but not amazing.

100% I hate the videofication of the internet. So much content is locked behind a video that is vastly more difficult to pull detail out of and search and just text. Videos are a great supplement to most text, but rarely do they make a good primary source of information.

While I get the privacy argument, I hate the trend towards mobile-only with the internet. With the Timeline it's so much easier to use my computer and see a giant map of the world and a mouse to poke around.

For OCR of handwriting, I did some comparative analysis a year back, and I found that Tesseract was... not good. However TrOCR was okay, certainly the best of the FOSS solutions. But Textract from Amazon was the best one by far far for handwriting, though your mileage will vary

Stanford's NLP Group has a good list of more specialized NLP coursers ( as well as CS224N, basically their CS388) - https://nlp.stanford.edu/teaching/

CS 124: From Languages to Information

CS224n: NLP with DL from Stanford

CS224U: Natural Language Understanding (Lecture Videos)

CS224S: Spoken Language Processing

CS276 : Information Retrieval and Web Search

CS324 - Large Language Models

LING 289: History of Computational Linguistics

Some others are below https://nasmith.github.io/NLP-winter22/about/

https://www.cs.princeton.edu/courses/archive/fall22/cos597G/

https://self-supervised.cs.jhu.edu/fa2022/ (has a list of other NLP courses at the bottom)

http://demo.clab.cs.cmu.edu/NLP/ (has a list of other NLP courses at the bottom)

I found it useful to compare various school's NLP courses when doing my own learning for different view points.