Seems bizarre. It's not like companies didn't want to sell it--they'd prefer to have the revenue. This is just kicking them then while they're down. I wonder if it will reduce risk-taking since it increases the downside of launching an unpopular product.
HN user
blueblimp
The inter-city travel was my favorite part of EverQuest. (The rest of the game, I didn't find too interesting.) The level of challenge was about right: if you looked at maps and planned your route, you could generally get to where you wanted to go, but it was hazardous.
I wonder if there's a game that focuses on that sort of travel experience.
I enjoyed browsing through it. One comment: the "philosophy and art" row is missing anything after 1890.
Today, 1,165 Bob Ross originals — a trove worth millions of dollars — sit in cardboard boxes inside the company’s nondescript office building in Herndon, Virginia.
This seems like a bit of a waste given that there's demand for them.
I feel it too:
- Plenty of em-dashes
- "you're absolutely right"
- "They're X, not just Y"
And the most important thing about PCC in my opinion is not the technical aspect (though that's nice) but that Apple views user privacy as something good to be maximized, differing from the view championed by OpenAI and Anthropic (and also adopted by Google and virtually every other major LLM provider by this point) that user interactions must be surveilled for "safety" purposes. The lack of privacy isn't due to a technical limitation--it's intended, and they often brag about it.
Yet this approach is obviously much better than what’s being done at companies like OpenAI, where the data is processed by servers that employees (presumably) can log into and access.
No need for presumption here: OpenAI is quite transparent about the fact that they retain data for 30 days and have employees and third-party contractors look at it.
https://platform.openai.com/docs/models/how-we-use-your-data
To help identify abuse, API data may be retained for up to 30 days, after which it will be deleted (unless otherwise required by law).
https://openai.com/enterprise-privacy/
Our access to API business data stored on our systems is limited to (1) authorized employees that require access for engineering support, investigating potential platform abuse, and legal compliance and (2) specialized third-party contractors who are bound by confidentiality and security obligations, solely to review for abuse and misuse.
This sort of convenient semi-arbitrary extension of a partial function is ubiquitous in Lean 4 mathlib, the most active mathematics formalization project today. It turns out that the most convenient way to do informal math and formal math differ in this aspect.
many of my math friends consistently slept 9–10 hours a day.
Anecdotally, I've noticed an association between long sleeping and math ability in particular, so this doesn't surprise me. I wonder if it's been studied scientifically.
In Meta's case, the problem is that they had been given the go-ahead by the EU to train on certain data, and then after starting training, the EU changed its mind and told them to stop.
Video of Hyperia: https://www.youtube.com/watch?v=0wUKY3HBw2U
A fun point in the article: Burton got into the industry via Rollercoaster Tycoon experience:
Curiosity became an obsession in his teens, when he started to play RollerCoaster Tycoon, a computer game that allowed him to devise his own rides. [...] In the end he won the job, he said, on the strength of those speculative rollercoasters he had made in a video game.
I agree: a well-designed educational course can feel like a video game in some ways, in that you're learning at a high, consistent rate.
My favorite blog post on tips for finishing is this classic by Derek Yu (of Spelunky fame): https://makegames.tumblr.com/post/1136623767/finishing-a-gam....
That was strange. I had to reread it to confirm that he's just talking about some random kid looking at the publicly-available game data files, not an employee or anyone else with special access.
All sorts of technology can be used secretly to assist criminal enterprises. Cars, computers, pencils, electricity, etc. It's unfair to hold LLMs to a higher standard than what applies to nearly everything else.
Yep. Advanced military technology tends to favor countries with high GDP per capita, which are mostly liberal democracies.
The license allows to reproduce/distribute/copy, so I'm a little surprised there's an approval process at all.
Qwen1.5-72B-Chat is dominant in the Chatbot Arena leaderboard, though. (Miqu isn't on there due to being bootleg, but Qwen outranks Mistral Medium.)
The Bradley-Terry model is essentially the same thing as Elo.
https://lmsys.org/blog/2023-12-07-leaderboard/
This model actually is the maximum likelihood (MLE) estimate of the underlying Elo model assuming a fixed but unknown pairwise win-rate.
This was definitely easier to follow.
Since they're building a special-purpose accelerator for a certain class of models, what I'd like to see is some evidence that those models can achieve competitive performance (once the hardware is mature). Namely, simulate these models on conventional hardware to determine how effective they are, then estimate what the cost would be to run the same model on Extropic's future hardware.
When you say Claude is worse than Mistral Medium are you going by the Chatbot Arena Leaderboard or some other benchmark?
I'm mainly going by the arena leaderboard, but it's also true in my limited experience with the two. (I mainly use either GPT-4 or open models.) And it's the only model I can remember getting an ethical refusal from. (I don't push hard in that aspect, so it was surprising.) I know it can be jailbroken, but, precisely because I don't push the models hard, I'm not skilled at jailbreaking.
By the way, the mention of API access reminded me how weird it is that the Claude API is still application-only, unlike OpenAI, Google, and Mistral.
The tech is legitimately impressive and exciting, but I couldn't help but chuckle at the revenge of the Scunthorpe problem:
It looks like the safety filter may have taken offense to the word “Cocktail”!
For getting a flavor of training LLMs without needing to be at one of the pre-training companies, it's very accessible to fine-tune a relatively small open source LLM such as Mistral 7B. (There are many tutorials.)
I wonder what's happening to all that money. Back when they originally released Claude, they were second only to OpenAI as far as chatbot models were concerned. Although Claude wasn't as smart as GPT-4, it had a more pleasant writing style, and Anthropic later released 100k context. At the time, I expected Anthropic to be the next company to release a GPT-4-level model.
But since then Claude has been passed by Mistral's mistral-medium and Google's Gemini Ultra. More concerningly for Anthropic, each subsequent release of Claude has actually performed _worse_ on the Chatbot Arena Leaderboard. (Claude-1 outranks Claude-2.0, which outranks Claude-2.1.) The reason for the decline in ranking is seemingly that the most noticeable update is to make the model refuse more requests.
In an additional blow, the needle-in-a-haystack independent benchmark revealed that Claude's long context is not actually used effectively by the model.
All-in-all, Anthropic is not looking in a good spot, despite the massive investment. They need to start releasing legitimately better models, or risk irrelevance.
It's still an exceptionally poor privacy policy compared to pre-LLM online services.
Compare with the Google Docs privacy policy, for example:
https://support.google.com/docs/answer/10381817?hl=en
Google respects your privacy. We access your private content only when we have your permission or are required to by law.
I think it is reasonable to expect the same from LLM API providers. The fact that they all currently do mass surveillance on users is bad.
Yep. Unfortunately, all the major LLM APIs do surveillance on your prompts and responses.
Yeah, I'm curious what's going on. Canada seems to be the only developed country without Bard at this point. (US, UK, EU, Australia, New Zealand, Japan, South Korea, Taiwan... all there.)
I don't know what Twitch's goals are here, but if they want to get rid of the right-up-against-the-line content, then maybe clear, consistently-enforced rules are actually detrimental, because it just invites rules lawyering like we see here. If Twitch staff just went around arbitrarily banning anything they don't like, the affected streamers would hate it (because they wouldn't know what they can get away with), which, if you want to drive them off the platform, is a good thing.
I question the article's framing of CPUs as "universal" and GPUs as "specialized". In theory, they can both do any computation, so they differ only in their performance characteristics, and the deep learning revolution has shown that there is wide range of practical workloads that is non-viable on CPUs. The reason OpenAI runs GPT-4 on GPUs isn't that it's faster than running it on CPUs--they do it because they _can't_ practically run GPT-4 on CPUs.
So what's going on is not a shift away from the universality of CPUs, but a realization that CPUs weren't as universal as we thought. It would be nice though if a single processor could achieve the best of both worlds.
I regard a physics education as a modern-day ‘liberal arts’ technical degree.
Some anecdotal evidence for this:
https://www.dwarkeshpatel.com/p/dario-amodei#details
Dario Amodei: We have generally found that if we hire someone who is a Physics PhD or something, that they can learn ML and contribute just very quickly in most cases.