benchmark where gemini flash is better than fable btw.
HN user
benxh
North Macedonia based.
@nipple_nip on twitter
Crazy calling sovereign states "US Puppets".
Minimax has been great for super high speed web/js/ts related work. It compares in my experience to Claude Sonnet, and at times gets stuff similar to Opus. Design wise it produces some of the most beautiful AI generated page I've seen.
GLM-4.7 like a mix of Sonnet 4.5 and GPT-5 (the first version not the later ones). It has deep deep knowledge, but it's often just not as good in execution.
They're very cheap to try out, so you should see how your mileage varies.
Ofcourse for the hardest possible tasks that GPT 5.2 only approaches, they're not up to scratch. And for the hard-ish tasks in C++ for example that Opus 4.5 tackles Minimax feels closer, but just doesn't "grok" the problem space good enough.
It is arguable that the new Minimax M2.1 and GLM4.7 are drastically above Sonnet 3.7 in capabilities.
cline is used by a lot of devs
The longer "it" reasons, the more attention sinks are used to come to a "better" final output.
To prove you right, you can read up on the incredible giga-brained countrywide experiments by Kardelj in Socialist Yugoslavia [0]. The result being a country where no-one wanted to work, and everyone had a great standard of living (while the IMF didn't call in its loans). And then the entire country collapsed all at once under the accumulated mismanagement.
[0] https://en.wikipedia.org/wiki/Workers%27_self-management#Yug...
My biggest gripe with Ollama is the badly named models, e.g. under deepseek-r1, it defaults to the distill models.
I'm pretty sure that Neosync[0] does this to a pretty good degree, it is open source and YC funded too.
I am assuming this will be solved this year.
If GPT4 is 220B/8 experts, that would be in-line with 3.5 Turbo being a 20B model, and GPT4 being a 55B activation out of a total 220B parameters.
It is ultimately all speculation, until Deepseek releases their own 145B MoE model, and then we can compare the activations/results
I personally was affected by this fire, although I've always kept 3 month backups of production data, encrypted, on-site, just in case of emergencies like this. Haven't touched their services for anything production related ever since
It's buried deep in the Gemini report, but goddamn are these incredible stats.
The Albanian takeover of AI continues. It's incredibly exciting!
I can't wait to see this open sourced, there's a lot of sampling strategies that help coding.
And I also can't wait to see how much Phind will improve further if the Glaive dataset is added onto it.
Edit: Contrastive search, dynamic temperatures.
It was known by Polynesians for at least 1000 years before Columbus. See sweet potatoes.
It's missing a lot of crucial details. Nothing on the dataset used, nothing on the data mix, nothing on their data cleaning procedures, nothing on the tokens trained.
Not all of them per se, take a look at something like Mistral. It's a 7B model displaying incredible performance. IMO, we still haven't even scratched the surface of what is possible with small LLMs. Especially not with pre-filtered/classified pre-training data. (Interesting LLMs based on their data approach and relatively small size: Qwen, InternLM, Mistral, Phi)
Added, and reached out on Twitter.
I would like to get in touch with you related to books4. Do you happen to have discord? or would twitter be ok?
There's currently multiple attempts at creating what you describe as books4.
I've had some success using vast.ai[0] with the Oobabooga LLM WebUI (LLaMA2) instances. One click to start up, minimal editing in the interface settings to enable OpenAI compatible interface.
Yeah wildly inaccurate.
This reads like a Serbian owned business from the North of Kosovo. But anybody following the local politics would know that: 1) Corruption as an issue is disappearing in Kosovo, especially compared to Albania, Macedonia, Montenegro and Serbia where it flourishes. 2) Corruption of police is practically impossible in Kosovo. 3) The IX is literally under USAID/EU control, good luck not getting blocked. 4) The North of Kosovo until this year was lawless. What they wrote about the east makes 0 sense.
All in all, this is either a honeypot, or somebody trying to discredit Kosovo.
Yes, but the acquisition of that data itself is illegal in almost all jurisdictions, since libgen is treated as a piracy website. Now if there were a pipeline to access books from Amazon or the Google Books project for training it would be a different story.
Still, for certain languages, only libgen and public piracy websites contain any scientific or fiction material in digital formats. E.g. my native language doesn't have easily accessible e-books at all, unless you go through illegal means.
I hope somebody undertakes the steps necessary to train on the entirety of libgen. The amount of high quality tokens in libgen should be substantial.
How does one go about becoming a distributor of Quest products in countries which arent served at all by Meta? Considering I can leverage existing infrastructure and network connections to bring it to market?
To be honest, I've been asking myself the same thing, technically the amount of "good quality" data in libgen is huge, way larger than the books3 dataset. However it would probably run afoul of copyright. Then again, a huge amount of data that LLMs go through is copyrighted.
So a model fine-tuned on libgen?
Many of the currently ongoing reproductions of LLaMa for starters. (see: Red Pajama[0]). Or any of the OpenAssistant affiliated projects.