It doesn't get any more funny when you try to explain something that wasn't funny in the first place.
Hence, I can't be bothered to read any of this.
HN user
Consulting / part-time/retainer AI, devops work, and cloud cost optimization via my consultancy:
https://hokstadconsulting.com
Engineering management/leadership; AI; devops; software development, architecture; Ruby, C, JavaScript, C++, PHP, Java, Go, Python (roughly in descending order of commercial experience)
E-mail: vidar@hokstad.com (job opportunities, contracts, and questions about my comments or projects are all good, but please be to the point and I can be slow to reply to unsolicited e-mail)
Mastodon: @vidar@galaxybound.com / https://m.galaxybound.com/@vidar
LinkedIn: https://www.linkedin.com/in/vhokstad/
Personal site: http://www.hokstad.com and http://www.hokstad.com/blog
My (not regularly updated) Ruby compiler project: http://www.hokstad.com/compiler
Github: https://github.com/vidarh
One of my favourite recent projects is this ~500 line TrueType font renderer in Ruby: https://github.com/vidarh/skrift
Site for my science fiction book series: https://galaxybound.com
It doesn't get any more funny when you try to explain something that wasn't funny in the first place.
Hence, I can't be bothered to read any of this.
If I saw something I thought resembled a joke I wouldn't have made the comment.
I love Kimi for some things, but K3 to me struggles with things Opus 4.8 breezes through. Admittedly stuff that is on the more complex side (code generation bugs in a compiler) to the point that after two days of struggle I stopped Kimi and will have Opus redo its work once I have spare tokens...
This wasn't just being slow - it didn't make forward progress.
For simpler stuff even 2.7 does just fine, though.
Redacting? Is that not just the model using the smudge tool it was given.
Opus doesn't have access to an image model, while ChatGPT does (it's not multi-modal, but it can prompt OpenAI's image model). So Opus by default is forced to write code to generate images, or generate SVG's as text. That's putting it in a worse situation than the linked article, where all the models were given rudimentary tools.
Give Opus access to an image model, and it will use that to create images just fine just like ChatGPT.
Most of the time when I see an LLM abruptly switching languages it usually continues with something directly related to whatever triggered it.
I'm sure there are other failure modes where it may change topics too, just like humans also regularly digress when triggered by certain words etc.
This is very much a philosophical question, and the only reason you're calling it silly instead of giving an actual argument is that it doesn't support your views.
Yes. All the prizes except the peace prize are awarded in Sweden. Norway was in a union with Sweden at the time the prizes were set up, but it's not definitely known why he chose the Norwegian parliament (via the Norwegian Nobel Committee) to aware the peace prize.
Not nearly all prompt injections are by any means that absolute unless starting from the exact same state. Many of them will also work only probabilistically unless you turn temperature to 0 for exactly that reason.
And at the same time, whole books have been written about how reliably we can induce certain behaviours from humans.
E.g. the Blue-seven phenomenon [1] - I've personally experienced that second hand and it was how I learned about it by searching for it subsequently because I suspected it was a known thing, having read about cold reading before. A co-worker came back from lunch and recited a story about a cold reader that had run a routine on him exploiting the blue-seven phenomenon, and I knew before the story finished that the answer would be "blue" and "seven".
See also Cialdini's book "Influence" which is full of examples of just how predictable peoples reactions are to a whole lot of things.
That there isn't a perfect overlap does not mean there aren't plenty of similar "hacks" that causes us to respond in very predictable ways.
[1] https://en.wikipedia.org/wiki/Blue%E2%80%93seven_phenomenon
You're misrepresenting what I wrote. I specifically pointed out that I switch languages without intent to do so.
When you suggest that is a "superficial analogy" after you were the one pointing out LLMs switching language as something that sets them apart, you're seriously reaching.
I can often pinpoint afterward what was likely the trigger: E.g. I used a word that is the same in two languages, and continue in the second; I pronounced a word in its native language for whatever reason, and continued in that language; my "context" suddenly included another language because someone else spoke the other languages within earshot of me.
What makes you think this is materially different from an LLM switching language because its probability distribution gives a word in a different language because it fits in context?
In the examples I gave, each even made a word in the language I switched to more probable as a reasonable continuation, just as with an LLM.
I'm not claiming the mechanisms are identical, or even similar, but the behaviour most certainly is more similar than "a superficial analogy" would imply.
I agree the term is not particularly useful, at least in as much as people are unable or unwilling to define it.
In fact, a "favourite" of mine when people downplay AI ability to reason or question whether it should be called intelligence, is to ask them to define those too terms. People usually don't even respond.
But I don't think we quite get away from it, because if you exclude the racists who would be happy to exclude groups of people, the other end of the coin is that a lot of people who wouldn't be willing to do that, still really badly want to draw a line that will always exclude all non-human computation no matter what from being considered intelligent or able to reason.
And the term matters a lot to those people, because of the emotional aspect to seeing humans as unique.
The Searle's Chinese room thought experiment usually reveals more about those who think it rules out a machine intelligence than it does about AI.
It rests on a staunch unwillingness to even consider the possibility that a computational process encode intelligence and reasoning, in favour of looking for the intelligence in the medium the computation runs on, and going "a-ha!" when there is nothing that looks intelligent there.
I agree with you that there will certainly be people who just continuously redefine the words to avoid accepting that AI is intelligent or reasoning, exactly for that reason - people have avoided pinning down an objective, measurable definition of these terms for a very long time, at least in part because it leads to some very uncomfortable discussions.
In particular how to define them so that they don't exclude an uncomfortable proportion of humans, but at the same time won't include entities people don't want to include (be it certain animals, or AI)
To a lot of people, the notion that there isn't a clear binary divide between human and non-human is deeply disconcerting.
How old does a child need to be before you think they have "human like intelligence"?
If anything a large proportion of stories about AI in sci fi is about AI failing to understand humans in various ways.
I switch languages "randomly" all the time when I think about something in another one of the languages I know. Some word will trigger it and before I know it I will continue in the other language.
In fact just the other day I commented on it to my fiancee after I randomly switched to French because we were discussing a trip and I mentioned a French location and pronounced it in French, and suddenly I was in "French mode" entirely unintentionally and it took a sentence before I realised.
That you think this is unique to LLM's suggests you simply don't know the diversity of human thought as well as perhaps you think you do. That's fine - none of us have a very complete view of that.
When my dad died, the funeral mattered to me. But the grave does not. Of course to some it does, but it doesn't take long before everyone it mattered to are dead too. I think it's fine to keep graves for those who cares, but that shouldn't really obligate society to keep them around very long.
My English is fine, and I did hand-write my MSc thesis in English, but scientific style writing is very tedious. Personally I don't think I'd even consider writing a paper "by hand" if I were to write another one.
There really is no point, as long as you verify the content matches your intent and edit out anything poorly written.
Frankly, I've read plenty of papers by native English-speakers over the years that'd strongly benefit from being rewritten by an LLM too...
I've spent some time on genealogy, and really beyond grandparents and maybe great grandparents it's mostly curiosity and for the interesting stories, and I only care about the graves to the extent that it makes solving the puzzles easier. It's not like there is anyone alive in my family who knew older ancestors of mine than my great grandparents.
For my part I've only ever seen even my dads grave once, and feel no need to go back there, though I understand why some might feel it matters to have the grave of someone they knew in person because of the emotional attachment.
The article links to a source describing medicinal use.
And dentures:
Closing future models won't take away our access to the open weight ones.
I had Claude prototype one, but it's not complete or public yet...
An X11 server is rarely on the hot path for rendering. Most X11 apps does most of the rendering to shared memory buffers themselves.
Yes, you're right.
There are other tedious bits too, like all of the details around exactly how to propagate which events, but the protocol encoding is certainly one of the most tedious bits.
I think not looking at the original X source code is probably a good thing. The protocol is reasonably well specified, and you can further use the XML specification used for XCB. Most of the requests are pretty simple.
For some of the trickier ones it might be worth diving into an X11 implementation, but you can also defer most of the trickier ones other than getting the event handling right (most of the old school drawing API mostly matters if you want to run 30+ year old software that hasn't been updated much since; that said it might be nice to be able to run twm and xeyes - I have a whole separate file in my own X11 server for legacy stuff required to run twm and xeyes and similar ancient software, but not useful for anything else I actually care to run...)
Yeah, I've had it work on an X11 server using a Ruby X11 protocol implementation instead of libX11, and it just rushed ahead and added support for a bunch of missing requests and responses. None of that is hard - it's all very well documented - but it's tedious.
My terminal emulator using the same binding also started out hand-written but Claude overhauled that too recently and it knows vtxx escape codes far better than me too.
It's relatively recent, I feel. Claude used to really struggle with writing asm not that long ago. But the last 6 months or so it's done great.
It's also far better than me (as someone who has done assembler since the Commodore 64) at using gdb to debug it, despite being effectively stuck using it in batch mode (which I didn't even knew existed). Watching it write elaborate scripts to dig into a code generation bug in my compiler is something.
I feel like the problem used to be that it'd struggle with the ambiguity of flow that is much more apparent in a high level language. But clearly that's not a problem any more.
There are 3-4 typical different approaches to rendering text on X. If it's not yet (or not correctly) implementing one or two of them, that'd explain it.
EDIT: Testing with xtruss shows st uses RenderCompositeGlyphs8(), part of the RENDER extension. I don't have Alacritty installed, but I think Alacritty uses the GPU, at least by default? Looking at the source for Glass it seems to use PolyText16 - the "old school" old server-side font rendering API. A lot of older X11 apps would work fine with PolyText16, and a lot of newer X11 apps does all the text rendering client side into a shared memory buffer without requiring GPU support, so it's not hard to end up with a set of applications where none of them would run into either gap.
EDIT2: I've looked at the Frame source, and it does seem to have support for the RenderCompositeGlyph calls and other supporting request, so not sure what the issue with st is. It's not doing anything unconventional.
I'd try rxvt (PolyText8/16) or xterm (ImageText8/16). If either/both works it's likely an issue with the RenderCompositeGlyphs support.
It's market cap is still closer to 2 trillion than 1.
Yeah, I have a mostly-functioning X11 server in Ruby, myself (not anywhere public yet).
Isene and I have relatively similar philosophies on this, except I have Claude burning tokens on optimizing and fixing my Ruby compiler now because I still want things in a high-level language, and my entire stack is Ruby instead of asm. But I love what he's doing - I just don't love x86 asm...
Turns out a functioning X server is a relatively simple piece of software. It's mostly just tedious. And most of the bulk is protocol handling that Claude can handle really trivially.