Several tools similar also can take in previous starts that are partially correct and use them to get closer. For my work, I am finding the local-mip package great for finding primal solutions better than CP-SAT, and then using HiGHS for my branch & bounder a great combination and feeding the results from local-mip to HiGHS. I do wish more tools could take in branching prioritizations or hints, I tried SCIOPT and it just didn't work as well as HiGHS even with priorization.
HN user
driscoll42
Interesting problems don't exist in a vacuum. I'm sure it was an interesting problem to figure out how to track people who opted out of tracking, how to build gas chambers, how to add lead to gasoline, doesn't mean one should choose to solve them.
While I agree, there's plenty of people who refuse to watch anything that's not sharp. I think there's room for both to exist, just clearly labeled as "original" and "AI Upscaled to 4K"
I used to work for a drywall manufacturer who still owned their own mines despite efforts to divest from them by some. They always viewed it as a structural advantage to still own them and not be wholly dependent on the coal plants (which effectively have conveyor belts going from the coal plants to the wallboard plants). I imagine as time goes on it'll become even more of an advantage for them to still own those mines as their competitors are forced to buy at highly inflated prices (or even from them) as coal shuts down.
When Roomba thought it was about to be acquired by Amazon, it did lay off 10% of its staff - https://www.therobotreport.com/irobot-laying-off-10-of-staff.... and after the deal was canceled, it was disclosed that they had reduced R&D and focused on margin improvements, and there was some brain drain as people left Roomba as it was in a 18 month limbo - https://www.verdict.co.uk/irobot-to-cut-over-a-third-of-its-.... And of course all this self inflicted pain only hurt them doubly as the Amazon deal fell through. If they had acted as if they weren't going to be acquired they might be fine, but they tried to maximize the shareholder revenue.
The best open source OCR model for handwriting in my experience is surya-v2 or nougat, really depends on the docs which is better, each got about 90% accuracy (cosine similarity) in my tests. I have not tried Deepseek-OCR, but mean to at some point.
In what world is that the "average experience" in American cities?
I looked into this a bit earlier this year. I'm mixed on it. While the FOSS in me wants it all open-source and available to use given that I'm basically labeling training data for them for free, and they are funded by donations/grants, I get value out of it for free.
My desire was to combine something like iNaturalist with BirdWeather for a bird tracker of audio and visual. BirdWeather does make it free which is great, but there's no great free API of iNaturalist quality for diverse bird tracking.
That being said, I am certain that if iNaturaist made their model public, tons of competitive apps would spring up and it'd be commercialized regardless of license immediately and would take people away from iNaturalist without giving iNaturalist anything in return.
Plus I know iNaturalist has issues with that they don't want autolabeled data uploaded as matched. They only want manually labeled data, which opening the API I'm sure would flood their server with ML labeled data. Which on the one hand, could be useful, but also a ton of noise.
I'm in favor of whatever option is most in line with keeping a long term success of a free, high quality plant/animal identifying app out there, and I don't know enough to take a definitive stance on that, and unfortunately those that do, probably have a vested interest in one of the outcomes.
Fair, looking at the ASR leaderboards it is truly better - https://huggingface.co/spaces/hf-audio/open_asr_leaderboard and NVIDIA's Canary might be even better? Will try these out. Appreciate bringing these to my attention!
Compared to all Whister models? Or the faster ones? And which version of Whisper? All for a faster, more accurate model, but need a bit more.
To suggest another "simple" example, Air Conditioning. It made half the world vastly more livable, and now anywhere in the world you could work every day of the year, reduced deaths and disease. At least currently, AC has had a greater impact on humanity than AI has.
It's not quite that specific, but close enough:
https://www.nytimes.com/2022/08/01/business/dealbook/pornhub...
https://arstechnica.com/tech-policy/2022/08/california-court...
This week, US District Judge Cormac Carney of the US District Court of the Central District of California decided that there's reason to believe that Visa knowingly processed payments that allowed MindGeek to monetize "a substantial amount of child porn." To decide, the court wants to know much more about Visa's involvement, calling for more evidence of legal harms caused during a jurisdictional discovery process extended through December 30, 2022.
According to Court Listener, the case is still ongoing - https://www.courtlistener.com/docket/59992265/serena-fleites...
So, I did some OCR research early last year, that didn't include any VLMs, on some 1960s era English scanned documents with a mix of typed and handwritten (about 80/20), and here's what I found (in terms of cosine similarity):
Overall | Handwritten | Typed
Google Vision: 98.80% | 93.29% | 99.37%
Amazon Texttract: 98.80% | 95.37% | 99.15%
surya: 97.41% | 87.16% | 98.48%
azure: 96.09% | 92.83% | 96.46%
trocr: 95.92% | 79.04% | 97.65%
paddleocr: 92.96% | 52.16% | 97.23%
tesseract: 92.38% | 42.56% | 97.59%
nougat: 92.37% | 89.25% | 92.77%
easy_ocr: 89.91% | 35.13% | 95.62%
keras_ocr: 89.7% | 41.34% | 94.71%
Handwritten is a weighted average of Handwritten and typed, I also did Jaccard and Levenshtein distance, but the results were similar enough that just leaving them out for sake of space.Overall, of you want the best, if you're an enterprise, just use whatever AWS/GCP/Azure you're on, if you're an individual, pick between those. While some of the Open Source solutions do quite well, surya took 188 seconds to process 88 pages on my RTX 3080, while the cloud ones were a few seconds to upload the docs and download them all. But if you do want open source, seriously consider surya, tesseract, and nougat depending on your needs. Surya is the best overall, while nougat was pretty good at handwriting. Tesseract is just blazingly fast, from 121-200 seconds depending on using the tessdata-fast or best, but that's CPU based and it's trivially parallelizeable, and on my 5950X using all the cores, took only 10 seconds to run through all 88 pages.
But really, you need to generate some of your own sample test data/examples and run them through the models to see what's best. Given frankly how little this paper tested, I really should redo my study, add VLMs, and write a small blog/paper, been meaning to for years now.
RetroMags - hhttps://www.retromags.com/ has 5218 various gaming magazine issues and strategy guides one can download to check out! Looks like the VGHF and RetroMags are working together from forum posts, with the VGHF doing a lot of work on making them more accessible than a raw cbz/pdf download.
This, if your workout plan for the day has you running is six miles a day, rather than just running the same path over and over again, might as well have fun with it and add a bit more fun to your workout.
It has been as low as $5 - https://isthereanydeal.com/game/sid-meiers-civilization-vi-p... and it's $10 at Green Man Gaming right now - https://www.greenmangaming.com/games/sid-meiers-civilization...
I visited Tokyo a few months back, and while the convenience stores seemed nicer than equivalents in America, I also wasn't particularly impressed. I think if I had just encountered them I'd be impressed, but the internet has hyped up the convenience stores of Japan so much I thought they'd blow me away. They're nice, they're good, but not amazing.
100% I hate the videofication of the internet. So much content is locked behind a video that is vastly more difficult to pull detail out of and search and just text. Videos are a great supplement to most text, but rarely do they make a good primary source of information.
There's a wonderful blog post on the PS3 Architecture - https://www.copetti.org/writings/consoles/playstation-3/ that gives a good overview of the Cell processor with linked resources if you want more detail.
While I get the privacy argument, I hate the trend towards mobile-only with the internet. With the Timeline it's so much easier to use my computer and see a giant map of the world and a mouse to poke around.
Funnily, the blog author has an article all about how demography of ancient Rome was determined - https://acoup.blog/2023/12/22/collections-how-many-people-an... By our standards they did not conduct a true census, where you counted everyone, but varied if counting heads of households, or men of military age, or men, women, and children.
For OCR of handwriting, I did some comparative analysis a year back, and I found that Tesseract was... not good. However TrOCR was okay, certainly the best of the FOSS solutions. But Textract from Amazon was the best one by far far for handwriting, though your mileage will vary
Stanford's NLP Group has a good list of more specialized NLP coursers ( as well as CS224N, basically their CS388) - https://nlp.stanford.edu/teaching/
CS 124: From Languages to Information
CS224n: NLP with DL from Stanford
CS224U: Natural Language Understanding (Lecture Videos)
CS224S: Spoken Language Processing
CS276 : Information Retrieval and Web Search
CS324 - Large Language Models
LING 289: History of Computational Linguistics
Some others are below https://nasmith.github.io/NLP-winter22/about/
https://www.cs.princeton.edu/courses/archive/fall22/cos597G/
https://self-supervised.cs.jhu.edu/fa2022/ (has a list of other NLP courses at the bottom)
http://demo.clab.cs.cmu.edu/NLP/ (has a list of other NLP courses at the bottom)
I found it useful to compare various school's NLP courses when doing my own learning for different view points.
Did you look at the Assignments? The [code and dataset download] links?
Even in a non-capitalist society, you still need high quality teachers and ideally some way to measure them. And it's not like eliminating capitalism in any way would prove to end poverty.
There was a study by the Federal Reserve that came to the conclusion last year that rewards cards is basically a money transfer of ~$15 billion from poor to rich per year. Discussion on Hacker News about it: https://news.ycombinator.com/item?id=34492502
Yup, the SHVC-2DCON-01 was only used by Mega Man X2, the direct link worked for me: https://snescentral.com/pcbboards.php?chip=SHVC-2DC0N-01 , though X3 used a different board, the SHVC-1DC0N-01 (https://snescentral.com/pcbboards.php?chip=SHVC-1DC0N-01
Is there any syllabus/list of topics covered without having to register for it?
There were a number of headlines a while back about ChatGPT suggesting recipes that could create chlorine gas, create a poison, called for human flesh, and others.
I usually look at rtings robot vacuum reviews for recs: https://www.rtings.com/vacuum/reviews/best/robot