Why was this flagged
HN user
anileated
There is no IP theft because LLM outputs aren't protected, just egregious ToS violations
I meant original IP theft that occurs to train LLMs in the first place. But sure that implies that further LLMs based on that LLM are also tainted by that original IP theft.
immediate shutdown all leading LLMs in the US
They can license training data. They have trillions, look what they are dumping into it, you seriously think they can't afford to license data.
Obviously it would be easier if they do it from the start, but that was their trick, to do it while people don't notice and get big ASAP. Should they get away with it?
Also, it would solve their Chinese problem, because it would make them violate copyright too. Right now it's more like rules for thee not for me so it's hard to take seriously.
The issues with LLMs go beyond just IP theft. I would not say PRC making LLMs cheaper is the best outcome (though it is better than nothing). The best outcome would be to make the practice of training on our data without consent illegal, which would simultaneously slow down economic change and make it more organic as well as give PRC companies less capabilities to extract.
My theory is that YouTube blocks some accounts for publishing LLM-generated music, and people who wanted to earn ad money from it get burned and publish LLM-generated posts about it.
I would be on YouTube's side here, except it's possible that their motivation is simply to avoid poisoning their dataset while they train their models off creators videos. Also, the question is how they tell apart what's LLM-generated without false positives.
Maybe there were also artificial listens fraud (it's a problem with their competitor Spotify), but we'll never know because no one who was blocked would publish that honestly.
No one is required to use EUDI: https://ec.europa.eu/digital-building-blocks/sites/spaces/EU...
Companies and providers (like banks) have to support it, but use is voluntary.
Check out the spec and legal framework, it actually makes sense and is open to different implementations, though you might need to certify it.
CEO of Roblox was once asked whether he would ever put prediction markets inside Roblox, he gave a straight face answer: https://youtu.be/XpIXRgMlPo4?t=2122
Let's use correct attribution: AI agents don't hack; people hack.
just another thing to memorize
Not knowing what time it is for my Australian colleagues at all without checking my phone every single time is worse. Remembering N timezone offsets (remember DST and half-hour offsets, too) is worse. Doing UTC translation or adding "my time/your time" every time is worse.
If you talk mental overhead, current system is like 10x of that than global time.
We agreed to meet at "8pm their time" but unless I literally put it into my TZ-enabled calendar app every time the chance I mentally translate it to my TZ wrong is unacceptably high. With global time, meeting would be @123 and that's it. I can keep it in my head or write it down on paper, no confusion and full precision every time. I don't even need to know if it's day or night if it's a remote call with the other side of the Earth, maybe me or my colleague works late, who cares, but I know what time it is there at any moment.
is not as high as people make it out to be.
It's not just timestamp translation and all the errors that come from that, it's all the rest of it, waste standardizing timezones and moving them around, having to convert time all the time, missed meetings, etc.
You have some fun ones. On the other side of the spectrum is PRC, where at the same hour of day it can be complete darkness on one side and almost technically noon on the other. It's super arbitrary with little rhyme or reason.
Things you will have in context when traveling: "it's going to be cold", "it's likely to rain", "it's going to be government conference so there will be extra delays with transport", "it's going to be %holiday% so everything's going to be closed all week", etc.
You're so used to it you don't even question that, and if you add to that "@x is when sunset usually happens"... somehow I think the world will not come crashing down.
My point is either way you need to memorize some info in the first couple of interactions and it really doesn't make sense to go through all of this change to just memorize a different thing.
It's at least to make time management in systems much less error-prone and complex, among other things.
if my flight to China lands at 9p local time, I immediately know that it's going to be night
What does that imply? If you mean "it's going to be dark", not really (you need to have more context to assume it's going to be dark at 9pm, there are places where in summer it's still very much light at 10pm). If you mean something like "buses are going to be running and McDonalds will be open", not really (you'll need to check the schedules anyway).
You would need to know that person's working hours, so I don't see how you are avoiding something.
Sure, if you talk to someone there for the first time, you would need to learn what time is generally day/night. However, you will know that 2-3 times in. Just like you would automatically know that now it's summer in Oz, or 3 hour short days near Arctic circle, if you talk to anyone from there even very occasionally.
Case in point, we have global calendar with no problems.
I don't think global time would be a problem like many people suggest. If you're in US and talk to somebody in Australia, you will quickly develop an intuition that time @X is night (or whatever it happens to be) over there, just like our other intuitions about how many things (weather, season, how long are sunsets, etc.) are different in different places.
Timezones are failing at all of their jobs. Getting time to correspond to sun position? It can be 7pm here and 7pm there but here it will be fully dark and there it will be still mid-evening. Knowing working hours of shops and government? Everything is all over the place. Everything is fluid and changes with seasons.
Plus, there is this unfair specialness that some countries are at UTC and others have offsets. With global time, everybody gets @0, just for different places it will be at a different sun position. (As long as we find a political way to pick something neutral, instead of saying "that's when the sun is highest in London".)
Finally, we don't have per-latitude calendar and things are working fine for us. It's February here and February in Argentina, and yet life doesn't stop even though it corresponds to winter here but to summer there.
The implication of "you have to have spent $1000 in tokens per engineer, or you have failed" is that you must fire any engineer who works fine by themselves or with other people and who doesn't require LLM crutch (at least if you don't want to be "failed" according to some random guy's opinion).
Getting rid of such naysayers is important for the industry.
I was going by this example:
/issue you know that paint bucket in google docs i want that for tldraw so that I can copy styles from one shape and paste it to another, if those styles exist in the other shape. i want to like slurp up the styles
What kind of context may be there?
Also, the entire repository and issue tracker is context. Over time it gets only more complete.
"Just show me the prompt."
If you don't have time, just write the damn issue as you normally would. I don't quite understand why one would waste so much resources and compute to expand some lazily conceived half-sentence into 10 paragraphs, as if it scores them some points.
If you don't have time to write an issue yourself or carefully proofread whatever LLM makes up for you, whom are you trying to fool by making it look pretty? At least if it is visibly lazy anyone knows to treat it with appropriate grain of salt.
Even if you are one of those who likes to code by having to correct LLMs all the time, surely you understand if your LLM can make candy out of poo when you post an issue then it can do the exact same thing when it processes the issue and makes a PR. Likely next month it will do a better job at parsing your quick writing, and having it immediately "upscaled" would only hinder future performance.
Even if you remember the times of iPod, you can safely say you're less than one light year old.
LLMs spitting out GPL code seems perfectly inline with the spirit
Only if spitted out code is GPL-licensed, which it isn't.
I'm not sure it's always bad intent. People often don't get that "machine learning" is a compound industrial term where "learning" is not literally "learning" just like "machine" is not literally "machine".
So it's sort of sentient when it comes to training and generating derivative works but when you ask "if it's actually sentient then are you in the business of abusing sentient beings?" then it's just a tool.
Here we are talking about derivative works, not "learnings".
It was the golden age of Science Fiction, and let's just say that the stereotype of programmers and hackers being nerds with sci-fi obsession actually had a good basis in reality.
At worst you are trying to disparage the entire idea of open source by painting the people who championed it as idiots who cannot tell fiction from reality. At best you are making a fool of yourself. If you say that free software philosophy means "also, potential sentient software that may become a reality in 100 years" everywhere it mentions "users" and "people" you better quote some sources.
Also those ideas aren't crazy, they're obvious, and have already been obvious back then.
Fire-breathing dragons. Little green extraterrestrial humanoids. Telepathy. All of these ideas are obvious, and have been obvious for ages. None of these things exist. Sorry to break it to you, but even if an idea is obvious it doesn't make it real.
(I'll skip over the part where if you really think chatbots are sentient like humans then you might be defending an industry that is built on mass-scale abuse of sentient beings.)
> Actually, you don't have to. You just want to.
Fair.
I don't think it's fair. That ideology was unquestionably developed with humans in mind. It happened in the 80s, and back then I don't think anyone had a crazy idea that software can think for itself and so terms "use" and "learn" can apply to it. (I mean, it's a crazy idea still, but unfortunately not to everyone.)
One can suggest that free software ideology should be expanded to include software itself in the beneficiaries of the license, not just human society. That's a big call and needs a lot of proof that software can decide things on its own, and not just do what humans tell it.
1. It's decided by courts in US. Courts in US currently are very friendly to big tech. At this point if they deny this and say something that undermines this industry it's going to be a big economic blow, the country is way over-invested in this tech and its infrastructure.
2. "Transformative means fair" is the old idea from pre-LLM world. That's a different world. Now those IP laws are obsolete and need to be significantly updated.
unsure what exactly that's relevant to here in our discussion.
I'll remind then. Our discussion follows the top statement "It seems open source loses the most from AI". As far as I understand nobody narrowed the context to "what is currently legal". Something can be technically legal and still harmful to open source. Also, laws are never perfect and sometimes they need to be updated.
(For example, I know that a number of people would say US abducting and detaining citizens and brutally deporting immigrants is not illegal, but if it's technically legal does that make it OK?)
what it's been doing since inception.
At inception open source was mostly personal side projects for funsies (like Linux) sponsored by maintainer having a dayjob. The big leap happened when copyleft licenses made it such that success of a big commercial company building products on open-source projects would directly improve these open-source projects. And it's nothing new, it happened long time ago. The desire for volunteer contributions to codebase to remain for public benefit in perpetuity is exactly the point of strong copyleft, and it's exactly what's being circumvented by LLM washing. The fact that these LLMs subsequently also harm open source communities adds insult to injury.
I think LLMs could provide attribution. Either running a second hidden prompt (like, who said this?) or by doing reverse query on the training dataset. Say if they do it with even 98% accuracy it would probably be good enough. Especially for bits of info where there's very few or even just one source.
Of course it would be more expensive to get them to do it.
But if it was required to provide attribution with some % accuracy, plus we identified and addressed other problems like GPL washing/piracy of our intellectual property/people going insane with chatbots/opinion manipulation and hidden advertisement, then at some point commercial LLMs could become actually not bad for us.
Something can be illegal and it can be technically legal but at the same time pretty damn bad. There is the spirit and the letter of the law. They can never be in perfect agreement because as time goes bad guys tend to find new workarounds.
So either the community behaves, or the letter becomes more and more complicated trying to be more specific about what should be illegal. Now that GPL is trivially washed by asking a black box trained on GPLed code to reproduce the same thing it might be inevitable, I suppose.
They're still tools ~anyone can use
Of course, technology itself is not evil, just like crypto or nuclear fission. In this case when we are discussing harm we are almost always talking about commercial LLM operators. However, when the technology is mostly represented by that, it doesn't seem required to add a caveat every time LLMs are mentioned.
There's hardly a good, truly fully open LLM that one can actually run on own hardware. Part of the reason is that hardly anyone, in the grand scheme of things, even has the hardware required.
(Even if someone is a techie and has the money and knows how to set up a rig, which is almost nobody on grand scale of the things, now big LLM operators make sure there are no chips left for them.)
So you can buy and own (and sell) a car, but ~anyone cannot buy and run an independent LLM (and obviously not train one). ~everyone ends up using a commercial LLM powered by some megacorp's infinite compute and scraping resources and paying that megacorp one way or another, ultimately helping them do more of the stuff that they do, like harming OSS.
If something is not technically illegal that does not mean it cannot be bad.
Like I said, there is a part that should be illegal, and then part where that's used to additionally harm one of the ways that OSS can be sustainable. The second part on its own is not illegal but adds to damages and is perfectly okay to condemn.
Open source software can have business models, it's one of the ways it can be sustainable. It can work like, for example, the code is made available (for any purpose) and the core maintainer company provides services, like with Nginx (BSD). Or there is an open-source software, and companies create paid products and services on top while respecting the terms of that software and contributing back, like with Linux (GPL) and SUSE/Red Hat.
When LLMs are based on stolen work and violate GPL terms, which should be already illegal, it's very much okay to be furious about the fact that they additionally ruin respective business models of open source, thanks to which they are possible in the guest place.