HN user

spindump8930

134 karma
Posts1
Comments48
View on HN

I think folks looking for more on this incident are better off reading the original threads linked elsewhere in the comments. This blog doesn't seem to add any information and is instead a narrative retelling of some documented events.

Some combination of reporting bias given concerns about LLM security capabilities and actual new vulnerabilities found with LLM assistance. Even if exploits and outages are unrelated to LLMs, I'm certainly thinking about whether claude could build these things (or if actors already have).

Sure, this is cute and interesting, but there's no validation or baselines and those examples are not particularly compelling. The o3 example just lists some terms!

Remember that models on different inference platforms might not necessarily give exactly the same results, adding another axis of non-determinism to development. Things like quantization, custom model serving silicon, batching, or other inference optimizations might mean a model from the original provider performs differently from the hosted one :/

This paper isn't the exact same scenario, since it's an auditable open weight llama model, but shows the symptoms of this: https://arxiv.org/pdf/2410.20247

Any more context on the copilot training note? More pointers would be very interesting, but we'd need to keep in mind how many different underlying models were (are?) branded as copilot. I thought at some points the "copilot" model in autocomplete contexts was a finetuned GPT from OAI.

Re: GPL, there are other open access datasets of git repos that make some distinctions between copyleft licenses but those are older resources now.

Not clear that they even have any GPUs yet:

Allbirds, which will be renamed “NewBird AI,” said it executed a $50 million deal with an unnamed institutional investor to acquire “high-performance GPU assets” to begin transitioning into a “fully integrated GPU-as-a-Service”

Serious folks know it's not straightforward to suddenly get any number of GPUs these days, even at that level of money

The article seems quite editorialized, shifting between describing "large-scale AI models" and "neural network-based approaches".

The underlying paper itself is more precise, comparing against LUAR, a 2021 method based on bert-style embeddings (i.e. a model with 82M parameters, which is 0.2% the size of e.g. the recent OS Gemma models). I don't fault the authors of the paper at all for this, their method is interesting and more interpretable! But you can check the publication history, their paper was uploaded originally in 2024: https://arxiv.org/abs/2403.08462

A good example of why some folks are bearish on journals.

"AI bad" seems to sell in some circles, and while there are many level-headed criticisms to be made of current AI fads, I don't think this qualifies.

Spending tons of money on Claude and the recent token benchmarks came WELL after Meta's huge investments in compute infrastructure for AI as well as the long history of language model development inside science divisions at the company.

Only for poor quality systems. Unfortunately there are many systems that tried to make easy hype, but are the equivalent of an ML 101 classifier class project.

If one measures for perplexity (how likely text is under a certain language model), common text in a training set will be very likely. But you can easily create better models.

For many of us a better Turing test is contextual to a topic we CARE about. Lots of LLMs sound better than a randomly sampled human on a topic I don't know too much about (e.g. opinions on new movies). They're decent on engineering topics I only vaguely know about, but still below the bar (though getting better!) on topics I really care about.

Fairly certain that AI (meaning an expensive llm type model) isn't needed to detect spam a large amount of the time. Classical classification methods could work while also being more privacy friendly (e.g. running on device).

According to AT&T, a big difference with its product is that it is built into the network itself.

This doesn't seem like an asset...