HN user

visarga

13,712 karma
Posts8
Comments8,442
View on HN

if a polished appearance (even if it is mega plastic) is no longer a sign of quality, then what is? What signals do I need to look for now?

AI produces apparent uniformity, it makes it harder to distinguish the quality of the underlying process. Same thing happens in all fields, we are all harder to differentiate, then why would you pick me over the others? this is the hard question we are facing now, and it's a desperate situation to be in. Not just artists and people looking for a job, but also companies to their investors and products to customers. All harder to tell apart.

I think this drives a lot of the AI craze. If I can't own my niche anymore, what is my new specialization? We have no answers yet. We have seen this kind of total competition in a domain with cheap imitation in SEO and slop writing.

So, continuing to profit--forever--from someone's else work, at scale, without their prior consent, is fair use?

Yeah, you are right. Have you been paying your dues to the authors of your math books in first 4 grades? I think 15% of your wages as an engineer would suffice. These kids continue to profit for years, and they are so many. Gotta pot a stop to that IP theft.

In my mistake I thought copyright was about copying rights, not paying for using ideas themselves. If just being downstream from a copyrighted work is infringement even without substantial similarity, then it's more like patents that expire in lifetime + a million years.

What do you mean pirating? They don't even distribute the originals, LLMs are not for replication, we already have copying and internet for that.

Why would we use a multi-billion parameter model to copy text? If we wanted the originals it would be easier to find them free, pirate or pay, if we use LLMs it is because we want something ELSE.

And caring about content rights in a world with limitless content and scarce attention is a mistake, it was never the content that was scarce in the last 20 years.

There needs to be a royalty payment based on if the AI regurgitates existing ideas.

So, by that logic, you need to be paying every time you regurgitate any of my ideas. Or anyone else's. Copyright now protects abstractions and vibes. Substantial similarity test be damned. Nobody can write stories about wizard schools, the idea is taken.

Right now I tried "What is digestion?" -> "Fable 5's safeguards flagged this message. Our intentionally broad safeguards deliver more capabilities but can also flag safe coding, cybersecurity, and biology tasks. Send feedback or learn more."

I don't think you can 100% ensure your queries have absolutely no biology and cyber keywords inside. No matter how harmless, they always trigger. People complain Fable aborts even when they try to make a login page for showing "username" and "password".

This makes models like Fable 5 impossible to use in any serious agentic task, because you can't even guarantee the model, which is a basic thing you need to build on.

The GCP team wants their slice, the other team wants some otjer slice, and so on. Everyone wants some crap for their promotion package.

I got a Google One plan for Gemini, but it came bundled with YT Premium lite, and that somehow made it impossible to renew YT Premium for 30 days. I suspect different teams stealing customers from each other.

One of my policies for agentic coding is to spend much effort in developing tests, coded tests not LLM based vibes. My projects have around 1:1 LOC between code and tests. Tests are like skin, when the skin is pricked it hurts, agents need to feel pain too.

OP's idea "everything is a text file" is good and I use it too. My plans are saved as task.md files, numbered and named. Work items are checkboxes inside the file, closed work items are checked and a comment is added on the same line to provide feedback about the implementation.

I also keep a current-state-of-the-world document, it should be <20KB of text, keep the essential decisions and intents. Loading it allows resuming in <30s.

Something I never saw anyone else do - I save all user messages in a chat_log.md file which is referenced for intent alignment and state recovery. I consider the chat log on the one hand, and coded tests on the other hand as the two walls, the agent works in the mid section between them.

https://horiacristescu.github.io/claude-playbook-plugin/docs...

There's no possible experience of the world there. No input other than a bunch of imperfect labels we created for stuff.

The brain too sits locked inside a bone box and only gets a bundle of unlabeled nerves connecting it to the outside. How can the brain could possibly experience anything, it only sees patters and patterns of patterns never the real thing?

I get annoyed when I see AI telltale signs too, but for example in my case I type 10x more than the final piece, is it really AI generated or just reworded? I don't use AI to fix my comments, just when I want to format article size pieces I post on my blog.

This fits perfectly with my philosophy which says cost constraints determine internal structure in a system, structure and cost evolve together. While philosophers like to ignore costs and come up with "ideal conceivers" as if they would inherit our limited concepts but just have unlimited computational budget.

The "Chinese Room" has the rule book baked in from the start, costs were hidden. When you remove cost from the system of course you can't see the semantics.

My own experience is that Opus 4.8 has an adversarial-teacher voice, unsolicited grading as if I submitted an essay for grading, declarations about the "real" issue, and constant "honest notes" self grading its own responses even before it answers. I can't stand its tone. We can't have a normal chat.

While Fable reverts to Opus for simple questions like "What is digestion?"

GPT-5.6 12 days ago

I think it's more RLVR (reinforcement learning from verified rewards). The RLHF is just to align models to human preferences, meaning to behave nice.

I wanted to use Fable to discuss a philosophical topic, but halfway through I used the word "cell" and got deflected to Opus.

On the other hand Opus has this awful adversarial-teacher vibe. It pushes back for no useful reason, talks down to you, and acts like it has to prove itself by grading and correcting everything. Instead of working with your claim, it reframes it, declares what the "real" issue is, then tells you what you failed to do.

So Fable refuses me and I can't stand Opus. Nice one, Anthropic, I need to downgrade my subscription.

We can switch in a heartbeat to a competitor, what kept us to Claude was that it had better models for a while. Closing Fable off means they are squarely inferior now.

For me Opus 4.8 was a slow model with a strong habit of talking down to me in an obnoxious way that would not be possible to prompt away. GPT 5.5 is now my main driver for serious work.

You will probably want a search engine though.

The search engine is indeed the last missing component from a sovereign stack. But I think this could be solved locally with little cost. Instead of indexing content on the web we should be indexing sources themselves - where to look for X? - like forums, blogs, docs, feeds, and specialized search engines. We could collectively amass millions of these search stubs that can be used by local models to go and fetch fresh information from the source directly. This means separating the routing layer from the information layer, we don't need to keep information cached from the whole internet locally. The search stubs could fit in a few GB about same size with the local LLM. The cool thing is that sources change much slower than information itself, so the search stub database could be refreshed at a slower pace. We could combine a few million generic stubs with a few hundred personal stubs generated from our own activities. It is trivial to generate these stubs by piggy backing on frontier models.

I think the harness and local context should supply that missing piece between general model and bespoke application. Each application has its own context and action quirks that don't generalize well. Maybe it's just 5% but that is genuinely specific. So its rightful place is in context engineering.

I have a long-ass post about how this could be implemented. https://old.reddit.com/r/VisargaPersonal/comments/1um9uyv/st...

Meditation can also hype up your anxiety depending on individual circumstances and personality. And it is by no means easy, you have to invest much effort to get results. It's like refactoring your code base (mind).

Consciousness is what the body is doing to be viable.

Without it we can't walk, eat, reproduce, or do anything. I like to think cost viability reasons explain consciousness. I know people prefer metaphysical or quantum magic explanations, I prefer a prosaic one - cost. It's a mechanism to keep our costs offset by gains. Cost can also explain unity - we die as one organism, not each organ on its own.

I think the obvious solution here is to beef up the test side of the app, much more than when writing code by hand. Tests represent project knowledge in executable format. The LLM does not need to be careful to remember every detail of the tests. You don't need to vet every small interaction, it automates review work as well.

Even better if the project was built from the start to be easier to test and observe. But my golden rule remains - no code without tests, expand test suite all the time.