HN user

reliablereason

274 karma
Posts0
Comments121
View on HN
No posts found.

Nice job, i have also been working on a few JEPA based models during the last few months. Trying to make more efficient LLMs.

I feel like you hit the main issues in the use of jepa models (well except collapse but sigREG more or less solves the collapse issue).

The main issues in JEPA as i see it is pushing the latent space toward representing features that are needed for good planing. A thing which is especially a problem in hierarchical planing.

You prime a JEPA world model to predict changes based on actions but you never really push it to use those actions. You simply hope that it will use them. If your latent is big enough and the actions effect on the world is simple enough it tends to work out but those qualifiers are not always small things.

Secondarily finding actions for the higher level JEPA Predictors.

LeWorldModel encodes multiple movements in to higher level actions. But this is a not a very good idea. It solves a basic issue with planing where the predictions degrade after a set nr of steps. But it does not solve the issue of higher level actions not actually being button presses.

The higher level actions for your mario game version would be things like: get the coin, Kill an enemy or get to the end of this stage.

You cant just encode many button presses in to those types of things. You need to discover those actions somehow.

Is the thinking even done in real tokens? I thought it was done using the pure residual stream. That is instead of collapsing the residual stream to a token you treat the final layers output as a vector of size d_model and use that as input for the next position in the transformer.

If that is the case thinking is not visible to us as users due to it not being done in text.

The issue is apparently this commit (someone did a git bisect):

https://github.com/RsyncProject/rsync/commit/859d44fa4f14207...

Which is a fix to the security issue CVE-2026-29518: https://nvd.nist.gov/vuln/detail/CVE-2026-29518

A CVE reported by VulnCheck which is a company that uses AI to find software vulnerabilitys.

I would honestly blame this on bad test coverage.

If you look at most of the commits where Claude is "co-author" you see that 80% of are just adding new tests. Which is exactly what would be needed if low test coverage was the issue.

I have done the exact same thing long before AI was a thing. You are rushed to "FIX" some security issue that someone reported. It is a scenario where you are working in code that you did not write or you wrote it so long ago that you cant remember. You try your best to just fix the security issue but you perturb something else while doing it.

Various LLM Smells 2 months ago

"A is not B instead A is blah blah" instead of just saying "A" is a very common pattern have seen in Claude.

It is strange to read as the topic A has often not been introduced and introducing it by saying what it is not makes very little sense to a new reader.

No you could rent virtualised servers way before AWS. AWS simply had good marketing.

The virtualised server thing was not a AWS thing, the thing that was were their other services. For example instead of renting a virtual server and installing a database on it. You could rent the database; that was sort of a new thing that AWS made in to thing.

It was never cheaper what you paid for was a promise of fire and forget. You would no longer need to worry about any responsibility to update the server or the database cause the AWS crew took care of that.

Pragmatically enterprise tends to mean less refined, designed by committee and expensive.

In this case i would guess it is mostly a justification for taking a part of the LLM pie.

Not sure that i understand your position exactly.

But consciousness is also "just a story" (a complicated one) that the human body tells the human mind.

We cant know from the outside if "the story" inside a LLM is detailed enough to emulate what we might call a felling of what it is to be the character in the story while it is telling the story.

It is similar to the fact that we cant know that other people have that subjective experience. In humans we think we have the right to assume cause we are quite similar in build to begin with.

Jumping back to the original subject to explain where i am in this. I personally don't think the entities in the storys of todays LLMs is detailed enough to have what we call human consciousness, mostly cause we are not training them to develop anything similar to that. Mabye they could have some type of weak qualia but i suspect most insects probably have much more qualia than the characters in todays LLMs. But that is quite a vague guess which is not based on enough data in my mind.

Most chatbots are not trained to have/emulate emotions so pain or fear of death is non existent. Therefore killing them and/or using them as slaves is not a moral issue. Thats how i reason.

On another point, LLMs are not conscious if anything is conscious, it is something being modeled inside the network. Basically if an LLM simulates a conscious entity, that doesn't mean the LLM itself is conscious; stating that is making some type of category error. So the fact that LLMs are just useful statistical generators would not mean that sentience could not appear out of it.

Removes paradoxical stuff like claims that there are bigger and smaller infinities.

Paradoxes comes from contradictions, a mathematical system that contains contradictions is a failed mathematical system.

The statistics is generally not. But the data used to learn the statistics may have been under license.

Learning from licensed material is generally accepted in humans, you may learn from something and then create something else and the new thing is not considered legally problematic with the exception of patents i guess.

Whether the same thing holds true for electronic systems is where people disagree if you look at the problem space in its essence. I land on the side that it is the same thing(humans and electronic systems learning), some seam to think it is a different thing.

Claude is not a legal entity, it is a software tool that outputs text based on statistics. There is a user that used a tool to create text and that user is the legal entity responsible for the text in any legal way that matters.

Anything else would be completely ridiculous given current laws in most countries.

It would be as ridiculous as blaming the car in a car accident where you drove over someone.

Right a very simple UI thing that they should have that would have prevented so much misunderstanding. Is a simple counter. How much usage do a have i used and how much is left.

If a message will do a cache recreation the cost for that should be viewable.

Not sure if this is a general trend amongst att LLMS but ChatGPT did over time become more and more affirming with its iterations.

I just recently switched away from the OpenAI garden largely because of it.

I do wonder if this was caused by some quirk of the training or if it really tests as a positive feature for most people. When i talk about stuff i don't want a mirror i already have a mirror. I want to be questioned, understood, helped.

To me support if the form of affirmation has no value when coming from an LLM since you know it has not thought about what it said.

No but it is a behaviourally defined disorder like autism. Which means it can and has many different causal patterns behind it.

That is there are many different things that can cause the behaviour.

Anyway on your main point, the definition of all psychiatric disorders has requirements of subjective suffering. So if you don't have subjective suffering you don't have the disorder.

The amount of people that seam to react negatively to living brain cells doing the computation is unexpected to me.

I do understand where it comes from to some extent.. some idea that human cells are special i guess, but it seams very naive to me. We spawn, use and kill far more complex AI agents millions if not billions of times every second in this society.

No one gives a shit, as those intelligences are not "real" or whatever.. or they are not "conscious" but conscious is a fictional word without an actual definition. In the end i think it comes down to suffering.

No one knows if some internal part in a LLM is suffering just as no one knows if a cell culture with brain cells like this can suffer.