I agree with your recomendation, but converting a pdf to an image is by no means smaller. PDFs are much closer to SVGs then to jpegs.
HN user
spindump8930
Claude and ChatGPT are both blocked in China
So it's presumably cheaper than attempting to spin up your own method of circumventing the blocks.
Exactly. Good peer reviewers understand that you can also move down on the scaling curve, not just up. Also laughable to try a "yolo" run without validating a scaling ladder/curve.
Can you share the specific part of this work that demonstrates better scaling than original transformers? Also note that many of the changes to that architecture, that have been proven in their use at actual scale, were brought about by members of the original team. Most notably Noam Shazeer.
That's why you do several small and medium scale tests, fit a curve, and ideally show that the trend persists at several scales. Not a single large or medium run - see the other comments down thread for example sizes.
I think folks looking for more on this incident are better off reading the original threads linked elsewhere in the comments. This blog doesn't seem to add any information and is instead a narrative retelling of some documented events.
Likely in this case the time vault was the collapse of Mt Gox, which has now recently been paying back holders.
Some combination of reporting bias given concerns about LLM security capabilities and actual new vulnerabilities found with LLM assistance. Even if exploits and outages are unrelated to LLMs, I'm certainly thinking about whether claude could build these things (or if actors already have).
It's very common if you improperly seed, as others in the thread brought up! Or in your framing, as rare as earth getting hit if it were surrounded by a sci-fi density asteroid field.
Sure, this is cute and interesting, but there's no validation or baselines and those examples are not particularly compelling. The o3 example just lists some terms!
Between the neo and the chances for privacy respecting local model inference, all the new apple hardware has me excited.
That artificial analysis page has some great references for this, thanks for sharing.
Remember that models on different inference platforms might not necessarily give exactly the same results, adding another axis of non-determinism to development. Things like quantization, custom model serving silicon, batching, or other inference optimizations might mean a model from the original provider performs differently from the hosted one :/
This paper isn't the exact same scenario, since it's an auditable open weight llama model, but shows the symptoms of this: https://arxiv.org/pdf/2410.20247
Any more context on the copilot training note? More pointers would be very interesting, but we'd need to keep in mind how many different underlying models were (are?) branded as copilot. I thought at some points the "copilot" model in autocomplete contexts was a finetuned GPT from OAI.
Re: GPL, there are other open access datasets of git repos that make some distinctions between copyleft licenses but those are older resources now.
The researchers tested five LLMs: OpenAI’s GPT-4o (before the highly sycophantic and since-sunset GPT-5)
Interesting, I always thought the sycophancy peaked with 4o and the associated personality (such as when myboyfriendisai users began complaining).
Hopefully this money means more compute infrastructure to help Anthropic counter the efficiency changes that have created this perceived downtrend in claude quality.
Having known some folks who did recurse, I think places like this want to select for those who consider coding a type of craft or art or self-expression. You can use LLMs, but stand by what you do and have pride in construction.
Not clear that they even have any GPUs yet:
Allbirds, which will be renamed “NewBird AI,” said it executed a $50 million deal with an unnamed institutional investor to acquire “high-performance GPU assets” to begin transitioning into a “fully integrated GPU-as-a-Service”
Serious folks know it's not straightforward to suddenly get any number of GPUs these days, even at that level of money
Yes, the paper itself tells a different story than the bullet points in this article.
The article seems quite editorialized, shifting between describing "large-scale AI models" and "neural network-based approaches".
The underlying paper itself is more precise, comparing against LUAR, a 2021 method based on bert-style embeddings (i.e. a model with 82M parameters, which is 0.2% the size of e.g. the recent OS Gemma models). I don't fault the authors of the paper at all for this, their method is interesting and more interpretable! But you can check the publication history, their paper was uploaded originally in 2024: https://arxiv.org/abs/2403.08462
A good example of why some folks are bearish on journals.
"AI bad" seems to sell in some circles, and while there are many level-headed criticisms to be made of current AI fads, I don't think this qualifies.
Yes, it's far more certain that meta released this, which is less convincing on evals, as a result of the mythos previews.
Re: changes, there's been enormous turnover in AI organizations, and in theory this one was developed by a "new" org. Whether that means less or more benchmaxxing is anyone's guess.
Spending tons of money on Claude and the recent token benchmarks came WELL after Meta's huge investments in compute infrastructure for AI as well as the long history of language model development inside science divisions at the company.
Only for poor quality systems. Unfortunately there are many systems that tried to make easy hype, but are the equivalent of an ML 101 classifier class project.
If one measures for perplexity (how likely text is under a certain language model), common text in a training set will be very likely. But you can easily create better models.
Pangram has time after time been shown as the only detector that mostly works. And that paper is pretty old now! There are recent papers from academics independently bench-marking and studying detectors e.g. https://arxiv.org/abs/2501.15654
Is the proxy here linkedin messaging/mail instead of direct email?
It was mentioned elsewhere in the thread but this article is relevant: https://www.cbssports.com/mlb/news/guardians-reliever-emmanu...
The ability to bet on short term individual events (such as a single pitch) means that even a single pitch, otherwise nearly inconsequential, can be abused.
For many of us a better Turing test is contextual to a topic we CARE about. Lots of LLMs sound better than a randomly sampled human on a topic I don't know too much about (e.g. opinions on new movies). They're decent on engineering topics I only vaguely know about, but still below the bar (though getting better!) on topics I really care about.
Fairly certain that AI (meaning an expensive llm type model) isn't needed to detect spam a large amount of the time. Classical classification methods could work while also being more privacy friendly (e.g. running on device).
According to AT&T, a big difference with its product is that it is built into the network itself.
This doesn't seem like an asset...
I recall that there were similarly motivated lawsuits for the earlier answer boxes that used to appear (prior to direct genai injection into the SERP page). What ever happened with those? Finding it difficult to search for.