Yes, that is definitely a limitation. If all models become worse at the same pace, we won't see any degradation either. I couldn't find any historical dataset of model benchmarks (I'd really have loved that, to see how performance holds over time vs. the initial announcement), so the Elo data from Arena AI was the least imperfect proxy I could find.
HN user
mayerwin
It'd be amazing if you could open an issue with a screenshot so I can take a look, I haven't been able to find issues when clicking on a group of models: https://github.com/mayerwin/AI-Arena-History/issues. Note: the model change points label being hidden when more than one curve is active is by design (to avoid cluttering), if this is what you were referring to.
You're right, thanks for the heads up! Corrected (I can't edit the post on HN though).
Make sure the prerequisites are installed (I haven't tested on Windows 10), feel free to open an issue on GitHub and share the logs.
No I haven't tested this project. I couldn't find any mention about BLE polling inside though, it probably offers complementary features.
Yes, Claude was very helpful to make this project work too (it would have taken me months otherwise to dig into how BLE works, and I'd probably have missed a lot of edge cases)!
Yes! I was surprised myself it was so complicated, especially as BLE MIDI is not something particularly new (Apple has nailed the implementation much better, luckily Pete at Microsoft is now doing his best to provide a comparable experience). When I played with USB MIDI 25 years ago it felt so much simpler.
I haven't measured it as my use cases were not sensitive to latency, but it felt pretty instant. Results will probably vary depending on the Bluetooth adapter, so best is to just test and see!
Tinycorp (owned by George Hotz, also behind Comma.ai) is working on it after AMD finally understood that it was a no-brainer: https://geohot.github.io/blog/jekyll/update/2025/03/08/AMD-Y... Exciting times ahead!
Worst visualization ever for a study about obvious correlations (that are misrepresented as a result of the poor display of data).
Cool tool, do you use Selenium?
Nice tool as well. The whole process took about 3 hours. Mostly for polishing, including 1 hour to fix an arcane CSS layout issue that ChatGPT wasn't able to help with (it may have been if I had provided it with the whole rendered HTML along with a screenshot). I really feel the LLM needs to have access to the same feedback we have, and be allowed to iterate (as it is very good at evaluating its results), to be effective with code. I also tried Gemini 1.5 Pro and Claude 3 Sonnet and it wasn't better than ChatGPT.
Same here, what a horrible customer experience.
Are you GPT-3?
Looks like you'd really be best served by an app like Todoist. A .txt file doesn't scale. But beware the productivity trap (becoming more productive means you'll end up even more busy).
What do you think of this idea to infinitely extend GPT-3's context window, by storing context in GPT-3 itself or in secondary layers?
The extension isn't just a Javascript overlay, it actually modifies the document in real-time while storing a full understanding of the document structure behind the scene, in the document itself (thanks to the Google Docs API). So other users would see "=name is out!" if this is what you've chosen to show on your own interface, or "MySoft 1.0 beta is out!" otherwise. If they have installed the extension, they would be able to access the document structure information, see placeholders and change variables as they like.