HN user

Hedepig

143 karma
Posts1
Comments122
View on HN

I do see your point, and it is a good point.

My observation is that the models are better at evaluating than they are generating, this is the technique used in the o1 models. They will use unaligned hidden tokens as "thinking" steps that will include evaluation of previous attempts.

I thought that was a good approach to vetting bad ideas.

This is not totally my experience, I've debated a successful engineer who by all accounts has good reasoning skills, but he will absolutely double down on unreasonable ideas he's made on the fly he if can find what he considers a coherent argument behind them. Sometimes if I absolutely can prove him wrong he'll change his mind.

But I think this is ego getting in the way, and our reluctance to change our minds.

We like to point to artificial intelligence and explain how it works differently and then say therefore it's not "true reasoning". I'm not sure that's a good conclusion. We should look at the output and decide. As flawed as it is, I think it's rather impressive

Thomas Cochrane 3 years ago

I mentioned in another comment, I enjoyed the structure and found it much easier to process. This is despite mostly reading long form articles and never having been a huge twitter user.

Thomas Cochrane 3 years ago

On the contrary, I have never been a heavy twitter user, yet I would dearly love all articles I read to be broken down like this. I definitely find it easier to process a list structure like this.

tends to give subjectively better responses to "bad prompts".

I wonder if a first pass with another model to expand these so called bad prompts into better prompts would work.

I am currently tinkering with this all, you can download a 3b parameter model and run it on your phone. Of course it isn't that great, but I had a 3b param model[1] on my potato computer (a mid ryzen cpu with onboard graphics) that does surprisingly well on benchmarks and my experience has been pretty good with it.

Of course, more interesting things happen when you get to 32b and the 70b param models, which will require high end chips like 3090s.

[1] https://huggingface.co/TheBloke/rocket-3B-GGUF

If it's some consolation, I've just listened to a recent podcast on This Week in Virology (as recommended by another comment on this post. And they are talking about research that uses the CMV virus which has certain properties that make it a very effective and crucially long lasting vector that can potentially protect us against all manner of infections and cancers.

Out of interest, have you built anything non-trivial with tailwind?

I've used CSS, and Tailwind and not worrying about naming everything correctly and not skipping around different files is a dream for me. Perhaps that's my brain, Idk.

Further I have gone over old code, and haven't had any problem with maintainability

That sounds like a good litmus test. Do you have a specific example you've tried?

My opinion is it isn't binary, rather it's a scale. Your example is a point on the scale higher than what it is now.

But perhaps that's too liberal a definition of "reasoning" , no idea.

We seem to move the goalposts on what constitutes human level intelligence as we discover the various capabilities exhibited in the animal kingdom. I wonder if it is/will be the same with AI

I have a hunch I am misunderstanding your argument, but does that mean the only way to build a "true reasoning machine" would be to just create a human.

I guess what I'm really asking, what would you expect to observe to make it not illusory?

I'm glad I'm not the only one who thinks this. It's really easy to trash on someone for not getting it right, but IMO we should congratulate on the achievement of actually making something, and give feedback on how to make it better.

Perhaps it's archaic, but I was using the word amateur to mean "for the love of"

Edit

Not sure why you’re framing this thread by different HNers as an attack somehow and your last statement feels judgmental.

I wasn't talking about the criticisms themselves, heck, I agree with them. I was more referring to the fact someone felt entitled to a good user experience they got angry.

Perhaps I'm being oversensitive on behalf of the guy who put effort into releasing this. But this isn't isolated, often people get punished for putting something out there that does not have a certain level of polish, I think that's totally counterproductive.