I'm honestly confused as to why it is doing this and why it thinks I'm right when I tell it that it is incorrect.
I've tried asking it factual information, and it asserts that it's incorrect but it will definitely hallucinate questions like the above.
You'd think the reasoning would nail that and most of the chain-of-thought systems I've worked on would have fixed this by asking it if the resulting answer was correct.
If you're ever in a situation similar to this, run as fast and as far as you can.
I really really want to underscore this point.
You're literally standing on top of ground and under that is boiling water.
If that breaks and you fall in you're going to be in boiling water with no way to get out and you will die screaming.
Also NEVER walk on ground that has no vegetation. If you look around a geyser you will see that the ground is white and has no vegetation. That's because the temperature is too high and it has water under it that's heating the ground.
Walk on that and there's a chance you will fall in.
In the back country there are no fences so you can fall right through the crust.
There must be a relationship here between Shannon's estimation that english is 1 bit of entropy per character and is highly redundant and 'easy' to predict.
A highly advanced AI could compress the text and predict the next sequences easily.
This seems like a direct connection like electricity and magnetism.
And maybe that's why English needs to be about 1 bit because we're not very intelligent.
In the 2000s I was addicted to Elisp and contributed a ton of OSS code including JDE, EDEE, and tons of other tools.
But... I had just a MASSIVE amount of code that was literally just for me.
Emacs basically became my OS.
Emacs allowed you to just eval code on the fly and the IDE would just adapt. No reload required. So if you wanted to do stupid stuff like make control+enter open the current URL at the cursor, you just write a three line script. Then you add it to your elisp on load.
... but mine got WAY out of hand. It was just mountains of code.
The next major leap in LLMs (in the next year) is probably going to be the prompt context size. Right now we have 2k, 4k, 8k ... but OpenAI also has a 32k model that they're not really giving access to unfortunately.
The 8k model is nice but it's GPT4 so it's slow.
I think the thing that you're missing is that zero shot learning is VERY hard but anything > GPT3 is actually pretty good once you give it some real world examples.
I think prompt engineering is going to be here for a while just because, on a lot of task, examples are needed.
Doesn't mean it needs to be a herculean effort of course. Just that you need to come up with some concrete examples.
This is going to be ESPECIALLY true with Open Source LLMs that aren't anywhere near as sophisticated as GPT4.
In fact, I think there's a huge opportunity to use GPT4 to train the prompts of smaller models, come up with more examples, and help improve their precision/recall without massive prompt engineering efforts.
Honestly, I think it's completely unfair for AIs to train on this data.
I work in ML so I'm aware of the consequences but society wasn't.
My step-daughter is finally crushing it as an graphics artist and she is really pissed at tools like Midjourney.
I asked her about it and she said "yes, they steal the artwork of real artists and generate fake knockoffs" ... and I don't think her opinion is invalid.
I met him briefly when he worked at Google. He was just starting to work on Guice and I was skeptical of dependency injection and we talked about it for an hour.
I went home and did a heads down and Guice was a major impact on my coding for the next ten years.
I bumped into "Crazy Bob"a few more times and just an insanely nice guy.
He was also murdered at Main and Folsom right in downtown SOMA in SF. It was 2 blocks from my former apartment.