HN user

segmondy

8,614 karma

Builder of thingz, lover of old robust tech like unix, prolog, lisp, APL, forth, postgresql, but don't mind the shiny new things ... right now, I'm all in on GenAI, LLM & LLM driven agents, team llama.cpp, building LLM agents since 2023

segmond AT gmail dot com

Posts6
Comments3,278
View on HN

This is pure speculation. China labs are geeks and nerds like some of us. They were amazed at LLM and like everyone want to build their own. They were happy to get meaningful next token predictions. I think most of you forgot how bad these things were 3 years ago compared to today. There was nothing to undercut, it was just geeks putting out their toys and saying, "Look, I built something cool". That became the culture and led to were we are now. All this idea that they are trying to undercut the monopoly is speculation. Google still releases open models, Cohere releases command-a, Mistral releases their model too, Arcee and Thinking Machine have released models too. It's just that Chinese models have gotten good and are also leading in the open weight category and USA has a very strong paranoia of Chinese models hence this talk. Plus every time they release something, some of the labs cry, "they copied us, distillation, the bad guys have done it again!"

So what we are seeing is nothing more but geeks doing geeky stuff, the only one that has serious demonstrated that there are possibly be had is Anthropic and the Chinese labs are beginning to copying them in terms of user plans, coding tools, etc. They are giving away the model to show it's great. Once they have the compute to serve the world and they have a model just as good or better than top model, they will go close to keep all the profit.

how do I find someone to use their number when I'm not in china and don't know anyone in china?

Qwen 3.8 3 days ago

let's see your grammatical correct tweet in chinese.

There's a bias in the direction all things face. You can ask these models to generate a thing animal, car etc and you will notice that 90% of them will converge towards the same sort of results. If you ask for something rotating, 90% of them will rotate right and a few odd ones will rotate left.

It's incredible you can't reason to see if pelican on a bike is a thing. It's not! This has been discussed to death. You can ask any model to generate anything. Generate an SVG of earthworm and a robin boxing. Guess what? The smarter the model the better the image, doesn't matter if it's a vision model or not. I rolled my eyes at this eval when I first saw it, then I tried various ridiculous things and noticed a very strong correlation. Things that are absolutely not in the training set.

Flock CEO doesn't give a fuck about that, they have been advised by public relations to manage their image since public opinion is against them. If they really cared, they will shut the entire thing down, but you know.

"Fuck you, pay me! If not me, someone else will!, profit must be had, your privacy be damned."

With the crazy lack of supply for hardware and ridiculous prices, seems we are going to have to start squeezing more juice out of our existing hardware. What I would love to see is how well this runs on an ancient system. Will this for instance run on an old PI and make it snappy?

Crap, the first open weight model that really feels out of reach when it comes to running it locally at home. :-(

I haven't used Fable, but if the hype is to be believed then it's a jump in model capability. If so then I don't expect the next DeepSeekV4 version to match it. However, if the next DSV4 version get's the kind of jump 3.1 got over 3.0 or 4 got over 3.2, I'll be very happy with it. Progress is progress. We "can" run DSV4 locally, Fable is closed.

DeepSeekV4 was a preview model, read the papers. It's not the final model. They released it to demonstrate architectural capabilities. They are still training and the model release is planned within the next month.

Very nice, multi modal, largest open weight model that supports audio. Would be interesting to see how good the audio capability is.

If you want to run locally, checkout https://github.com/danielhanchen/llama.cpp/tree/add-inkling https://unsloth.ai/docs/models/inkling https://huggingface.co/unsloth/inkling-GGUF https://huggingface.co/unsloth/inkling-NVFP4

This supposedly is better than KimiK2.7, as much hype as GLM5.2 gets, I find myself using KimiK2.7 half of the time, so if the benchmark is true, then this can definitely go in the mix. My hope is that it might have strengths in some areas to beat all other open weight models.

On another note, we have a similar bread machine. We bake bread multiple times a week. Costs about 10%, takes about 5 minutes to add the ingredients and set. Most times we set it before bed and wake up to the smell of fresh bread and get to eat hot bread. We are 100% certain of the ingredients. We bake different types and when we have a guest, we give them a parting gift of a fresh bread. Our constant baking and the ease has convinced extended family members to do the same. So, I'm not convinced that convenience always wins. If 20-40% of SaaS are replaced with homegrown AI, that is roughly the definition of "doomed" in a market.

Heading for junk? It's absolutely junk and trash. Every retail investor that I know that bought SPCX, did so to make a quick buck. It's even worse than crypto, at least some people believe crypto has value and will take over or replace fiat currency. No one that I know who bought in SPCX did so believes the space data center story or going to mars. There's certainly value in launching things into space, but none believed the valuation with the profit, but the figured Elon has the midas touch and they could cash out before the music stopped.

Does it make sense to still walk? I mean, we have cars. You want to go somewhere that's 1km away from home. Does it make sense to walk?

Someone wants to give you a phone number. Does it make sense to still try to memorize it or must you hunt down paper and pen to write it down?

... you get the drift? Do what makes sense to you.

We can say that about programmers, most ICs don't understand what's going on in the layers beneath were they work. Most have no idea what's going on with libraries, frameworks, remote APIs, it's all abstractions. Most people can't tell you how system calls are implemented or function. They don't have the time, bandwidth to understand it all, they just operate at their own layer to get the job done.

Is it really the same argument? I think there's enough of people that will argue that there will be no need to come together to collaborate, just rally a bunch of agents and you can build whatever non-trivial stuff you can imagine.

Just a matter of time. Go download gpt2 or llama2 and be shocked at how bad they are compared to today. They were entirely "useless" yet we marveled at them. Go examine GPT3.5/gpt4 out which was all the rage and then marvel at how a qwen27b or gemma31b model mops the floor today. My point is that the models will eventually learn to have a great model of software system in their head, just a matter of time and proper RL.