HN user

TZubiri

1,918 karma

{ Roles: Backend dev,

Language stack: CPython, Operating System: POSIX, Jurisdiction: Argentina, Games: [Piano,Chess], contact: hackernews at tomaszubiri dot $ThatCommercialVerisignTLD ( Sorry for the puzzle, but you know how it is with spam.) }

Posts14
Comments3,095
View on HN

The discussion of

"haha LLM companies stole data and now they have their data stolen so it's the same thing and it's fair." was reductionist when it started, and it's been like 3 months, and every internet user throws it like it's the hottest take ever, have another take please.

Also have nuance, don't jump to hit your HOT_TAKE key in your keyboard, actually read what the chinese are doing, and then you can pass on your judgment on whether it's ok or not.

It's not the same thing if they scrape an openly published dataset and it's an IP dispute. Or if they are using,network and financial pooling mechanisms that are shared with CSAM providers and cybercriminals, mutually providing each other alibies, and using black markets of passport-backed identities to setup thousands of accounts and circumvent bans and detection.

While we are at it, if there's a case that was settled, it's a closed case, it can never invalidate any other disputes. That case is closed, and it was settled by the parties that claimed to be damaged, that's done. If you didn't think so, you wouldn't have taken the settlement, and if you didn't have a say in the settlement, it's because you weren't damaged so who cares, go make a claim where you are the defendant if you believe otherwise. But thankfully in no legal system does the existence of a claim against you prevent you from making claims of your own.

Nuance is a good thing.

Politics aside, this is a sensible position, there was a major client revision, they can put security measures in the new version that are not present in the older one, this happens in many many industries.

To be specific, the defensive measures probably rely on javascript, captchas and captchaless bot detection measures (probably javascript), like measuring mouse movement and stuff like that. So sending a client with javascript that loads content if some security checks pass, is harder to abuse than a plain html site that sends all content in the first request.

Besides scraping, this probably has implications for posting as well. Do you like reading comments from humans and not from other bots? Well, in that case a system that uses plain html will be easier to automate with bots than a javascript clusterfuck.

You can't have the cake of web 1.0 and the eat-it of a botless experience too.

Politics aside, this is a sensible position, there was a major client revision, they can put security measures in the new version that are not present in the older one, this happens in many many industries.

To be specific, the defensive measures probably rely on javascript, captchas and captchaless bot detection measures (probably javascript), like measuring mouse movement and stuff like that. So sending a client with javascript that loads content if some security checks pass, is harder to abuse than a plain html site that sends all content in the first request.

https://en.wikipedia.org/wiki/Goodhart%27s_law

"Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes."

Or the more pop layman version

"When a measure becomes a metric/KPI, it ceases to be a good measure."

Story time, I live in Argentina, and we don't have Big Macs, the main Mc Donald's brand, here, because during the CFK presidency, one of her tactics was to Goodhart economic metrics. Even the informal obscure ones like the [Big Mac Index](https://en.wikipedia.org/wiki/Big_Mac_Index), I don't know the precise details, but the Big Mac ended up being a very cheap item, like 2 or 3 times cheaper than actual menu items, but it was never on the advertised menu, and it also ended up being very small compared to the other burgers, so it wasn't even like a hack, a shrinkflation type of deal.

But hey, anyone who read the Big Mac Index table would never find Argentina at the bottom of that list along with a couple of other countries with bad brands, so the ploy worked. And now we live with the aftershock, the brand never really turned around, other brands with ridiculous names took over it like the McTasty, which makes me sound like that skit from Tarantino's Pulp Fiction.

kind of a bad example, but take google, search, youtube, Android. It's an extreme position to argue that those products exist to serve advertising, they are clearly useful to end users.

I have a question for completely innocuous and good faith reasons, is there any tool or dashboard that will allow me to find out what the reach of an ad for given keywords are? Like Google's Keyword Planner?

If, as an advertiser, I wanted to know how many people are asking or talking about REFRESHING SODA, can I get an estimate of the potential reach of an ad? Can I filter by gender and age, so I can see how many of gender X and age Y are asking about REFRESHING SODA? Actually forget about the soda, Ok Computer, show me the top 10 topics of conversation among middle age male adults with purchasing power.

The naïve prediction is that advertisers will advertise the product to ChatGPT users. I believe that users will tell ChatGPT about their desires, and advertisers will create products for them.

The product seems to use the same 'scoping' mechanism that humans would use, so you can give agents access to whatever 'channels' you want.

The issue seems to be that they 'invented' an auth and permission protocol instead of delegating to something existing, it would be as if they thought their idea of permissions were novel.

I think the victor would allow you to even have a small army, not big enough that you can rebel and be a threat, not small enough that you are useless, defenseless (or bored).

And if the victor dies, you would then be free. Not necessarily game ending. And it might match up with actual history too, how many towns have outlived their tributee empires?

Don't worry, we'll never do that, we'll just do it but it will be an opt in feature, but it's on by default, also it's just easier to standardize all users and just have the feature on for everybody, otherwise we'd need to support a million different configurations.

That guy that wrote the Incorruptible book is rising stonks man

I think annihilation is a property of how we played the game and not of the game itself. I get that the game was designed around it, but there's nothing stopping players in a Free For All from subduing an enemy into slavery instead of wiping them out, some may call this an ally. In fact, it's probably the optimal strategy, annihilating might be a pyrrhic victory, where one player is knocked out, and the 'victor' is left with half an army. The only winner in that situation are the rest of the players who are now better than the two players.

Except NvN games, I except FFA and 2v2v2v2 games to show elements of empires, it's only games with two sides where minmax becomes a viable strategy and damaging the enemy is the same as strengthening yourself.

It might be a highly technical and nuanced point, but while the subjective explanations Linus gives are neutral indeed, the stance is binary, and he took what to my estimation is the wrong approach "allowing llm generated code in the repo". The only sensible approach in any software or non software project is "LLM output is not acceptable content". Projects should see themselves as input for the LLMs, not as channels for LLM output. If you start corrupting your projects with LLM output, they will soon be tarnished and either removed from LLM training data, or enter a lossy IO loop.

On to the technical point, LLM output is output, the source is the prompt, if you are going to commit something, commit the prompt. Second, code that is generated by LLMs is less maintainable, Linus entered late into the fad and anyone with 1 month of fiddling with AI knows that he will regret it soon, it's hard to undo once you corrupt your repo with slop, perhaps if it happens fast enough and there's no major releases it can be swept under the rug.

On to the nuanced point, Linux is purposefully designed to maximize user contributions, so accepting LLM contributions might well serve the particular purpose of linux, but I still think it's technically wrong to commit target code, only source code should be committed, and that's prompts.

But git itself is collapsing, it doesn't seem to be well suited for this new revolution, it doesn't track prompts, or it does so at the expense of the generated code. Maybe github can track target code as artifacts.

I think we are watching the collapse of Linux, Git and Linus. Certainly a bold position, so I don't blame you for being more conservative, but we can come back in a couple of months and see if we changed our minds.

Besides the subjectives, here are two objective policy stances that Torvalds is defining for the Linux kernel development:

1- Maintainers are allowed to commit LLM generated output.

2- Criticism of LLM generated code is not welcome/will be ignored.

Now, whether that constitutes being pro-Vibecoding or pro-agentic engineering, whether it's delusion, whether it will have problems, that's subjective. But I feel that whatever way you look at it, it's a topic that polarizes engineers, and Torvalds is taking one side and not the other. It doesn't seem to me that it's a very neutral stance, although it may be more neutral than projects like Bun or OpenCode of course, if it feels neutral, it's cause the overton window is shifting.

If anything it's more important to hold. It's easy to hold one position and then falter, there's a pressure to always be with the times and not be 2 years demodé, but simple positions still hold true.

I wrote in the opencode thread that when it came out I put it behind a vm and its own user, and I never allowed it to run outside of it. But I know of people that as soon as they noticed that it worked well like 99% of the time, they let their guard down and give in to YOLO mode. And in orgs I've even seen CEOs treat their agents less like a user/employee/contractor, and try to 'empower' it by giving it ALL the data. Time bomb.

It's like fucking with condoms just the first couple of times. And then simultaneously ditching it and joining the free love movement.

Nowadays not just big orgs are falling into the delusion, but also big people.

When Linus posted that AIs and vibecoding were here to stay and declared resistance to it as harmful, I stopped to consider whether I was wrong, but it has made me realize that in retrospect Linus Torvalds and Linux itself aren't actually the holy grail of computing. I didn't feel that way with Richard Dawkins, its not like falling for an AI psychosis retroactively made me question The Selfish Gene, but now I'm looking at linux and the theory that it's a clusterfuck is gaining so much traction, especially after copy.fail and ensuing rustification, I see so much clearly now. It was never about linux, UNIX sure, POSIX, yeah, GNU fucking aye, kernel? Ok whatever, drivers and scheduler with a gajillion lines of code I guess.

I'm glad that when I recommended opencode to a friend, I taught them first to install a vm and create an app specific user to boot.

Many users did things right the first couple of days and then turned the security off after seeing it 'work fine' on its own.

Was a cool thing to use for a couple of weeks in v1, you have to be quite static to keep using it at v68 of self vibecoding while the llm providers like openai already developed their in house alternative.

I would argue it's fraudulent. There's loads of text to speech synthesizers that work just fine but sound robotic. The huge technological advancement of tihs technology is that it 'sounds human', even if it isn't.

LLMs are great, but we should stop bundling fake image generation and fake voice generation with it and calling it AI.

If you truly believe that there's no problem, and that it's fine, then use a TTS technology that didn't spend a gigajillion dollars on trying to sound like it's not what it actually is.

It's like selling soy pellets with synthetic meat flavour and making it look like meat and selling it in butcher shops and not clarifying that it's not actually meat.

On the one hand, I personally believe that meat is better, yes, but on the other hand, that's not what I am going to discuss and it's a disservice to argue about how meat is better, all you need to know is that if you try to pull this off, you will be indistinguishable from unethical fraudsters. Not saying you or Perforce are one at this point, we are all dazzled and confused by the technology, but down the line, when regulations settle in, there will be no gray area, do not do this shit if you care about your reputation. Only do this shit if you want to make quick bucks during during the boom and before the bust.

I remember feeling that once with cannabis, I was feeling some sensations and emotions and initially I tried to explain them rationally, but eventually I came to the theory that some brain areas were just being activated by the drug.

Don't do drugs kids.

Maybe the first such experimentation actually teaches you something? By destroying a part you get to understand the true shape of the self. But by repeated use I imagine you get diminishing effects on the knowledge, and exponential damage, as you are left unable to self heal or develop adaptations to the missing corrupted sectors

Nice.

If I may ask for your advice on a related subject.

I have an old electric piano, it doesn't register dynamics, the strength with which I hit each note. Which is a huge bummer.

I can't think of a simple way to hack dynamics onto the keyboard, I thought of adding an additional sensor and using the time delta between touching a key and hitting the bottom, but that's not really correct, as you can hit a pianissimo key fast in staccato.

Maybe a basic spring that decreases resistance as it becomes compressed, but it might make keys quite hard, and it's not very simple to implement.

Any ideas?

Thank you.

That's awesome.

That said, didn't you just shift the cost instead of reducing the cost?

It's of course subjective, and the goal of the question isn't to devalue your work, au contraire, I'm saying that you contributed 120k worth of value.

This is as-is a product that you can sell to other bowling alleys, and do not fall into the trick of selling savings and pricing it at like 20k. Your product may be better as it can be customized, modernized, it has a single author and less bureocracy than a corporate 30yo bundle that was handed over through 5 generations of programmers after resignations and layoffs.

Good luck

I don't want to be mean, but you sound confused.

Are you building agents? or building apps? Or building an API platform that has everything from databases and auth and coffee making capabilities?

Sorry for meandering, but this is a common issue, way more people build builders and sell shovels than there is a demand for, and it's usually a way to escape actually building an application.

Try building something concrete, and asking concrete questions about that.

I was like 10 minutes deep into the free version when I noticed that a couple of weird things could be attributed to the 'narrator' being a ChatGPT like speech synth.

1- The voice is not consistent across different videos. 2- Once in a while it does that thing where it sounds like a demon and changes the voice profile to a completely different person for a little while. 3- There's weird... pauses... that in some cases make sense, but in some cases it's just completely non-sensical "this is a very useful... feature" or "looking at your issue that you are... raising to them", it sounds like someone reading a Charles Bukowski poem. This happens the most often, once you see it it's like those optical illusion things where you can't unsee it.

One cannot spend too much time evaluating products, and I feel that I have seen all that I needed to see, how good can a product of a company that does this be? And to charge 500$ for the complete course?

I don't quite get it. Like is it really easier to generate a video with fake AI narration than just narrating it yourself? I think it would even be harder, only to make your reputation and brand 1000% worse? And the act of showcasing a free version of the course to 'get a taste of it', when in reality I'm guessing most would see the red flags and back away, thanks I guess.

I just don't get it.

But it's a pretty good indicator of skill level. High wpm means they write a lot.

It's also helpful as a burst speed. It doesn't mean you are always cranking at 120wpm, but you can hit 120wpm when the words in your mental buffer is overloaded and you can empty that fast.

Try writing on pen and paper, you'll find that your mental buffer fills up quickly and you aren't able to hold all of your thoughts, so you have to drop some. It has its use cases, sometimes you have to think more and write less. But in general high entropy write capacity is a valuable spec to have.

Decoy Font 6 days ago

Most AI systems work by reading the pixels of an image up close.

Not really, most AI systems work by reading the octets as ASCII/unicode (and then tokenizing it).

You could make an even better decoy font that renders one letter as another, so when you copy and paste it onto some other place with a normal font it reads as garbage, and garbage is what the AI will see, however if you render it with the descrambling font, you will see the regular message.

This has been used in PDF files as an obfuscation and anti-copy mechanism.

Oh interesting.

I'm not sure if you question whether the first sale or the latter resale might not be contracts.

But in either case, I believe both them to be. It's similar to a sale of a stolen good with an unaware purchaser.

It's a contract on two counts. First, the contract is not void, the subject is not illegal in itself, the seller would just fail to fulfill their end.

Second even if a contract would potentially be voidable, I would argue it is a contract until a judgment voids it, this is a bit subjective, ontological and inconsequential for cases where the judgment would be certain like a contract for stealing, but the more borderline the case is, the more relevant it is, a contract about a complex legal issue that might be or might not be voided with p=0.5 is still a contract to me. Of similar value to a contract about a good of stochastic value, like an option, or an asset that may have been stolen with p 0.5