HN user

andblac

15 karma
Posts0
Comments7
View on HN
No posts found.

The "ALL CAPS" part of your comment got me thinking. I imagine most llms understand subtle meanings of upper case text use depending on context. But, as I understand it, ALL CAPS text will tokenize differently than lower case text. Is that right? In that case, won't the upper case be harder to understand and follow for most models since it's less common in datasets?

And now we’ve built LLMs - the biggest mirror of them all. The Internet and smartphones gave us countless ways to look at ourselves and see how others see us. And then LLMs helped us gaze at the sum of all that and even confront reflections of our own thoughts.

Yeah, TFA ended just before it got to the really interesting part of how self-reflection itself is fundamental to the development of concisousness. Mirror-like technologies don't just show us our own appearance. They help us understand how we relate to the world around us.

It reminds me of Kieślowski's movie Camera Buff (1979), where the main character in iconic scene points the camera at himself and realizes that the act of making movies reflects not only his subjects, but also on who he is in relation them.

Yeah, I'd would love to read article on all that.

At first glance, this reminds me of how branch prediction is utilized in CPUs to speedup execution. As I understand it, this development is like a form of soft branch prediction over language trajectories: a small model predicts what the main model will do, takes few steps ahead and then verifies the results (and this can be done in parallel). If it checks out, you just jump forward, it not you take miss but its rare. I find it funny how small-big ideas like this come up in different context again and again in history of our technological development. Of course ideas as always are cheap. The hard part is how to actually use them and cash in on them.

Skimming through the source it seems to run 'car' and 'person' objects through llava with the following prompt:

- "person": "get gender and age of this person in 5 words or less",

- "car": "get body type and color of this car in 5 words or less".

So YOLO gives the bounding box and rough category, while llava describes the object in more details.

That's a great scene. I mostly remember it for the exchange the two had after they finished modding the game and it worked [0]:

Red: "Congratulations, son! You have seen the future!"

Kelso: "Yeah, yeah, you're so right, Red! Home computers! That is the future!"

Red: "No, no, no. Not computers! Soldering! The future is soldering! [...]"

How often do we try to extrapolate from current technological improvements to predict the future, yet fail to grasp which changes are truly important.

[0] https://tvshowtranscripts.ourboard.org/viewtopic.php?f=936&t...