This isn't that out there in our current scenario. These models compress our collective thought and effort. Why not make these publicly owned, all profits distributed back to us?
HN user
PhunkyPhil
I knew I was a highlighter but reading that showed me how much my brain relies on spam click highlighting to keep my eyes on track. I should probably read more books.
I'm kind of in the same boat and it's been pestering me for months. Every agent is simply a less capable Claude Code.
If it had a lossless, massive context window (100m-1b tokens), then it will squash everything. Give it bash + r/w and it can in theory /goal anything.
I think there's something to be gained in a production environment be siloing agents for reproducebility/auditability, but I suspect that will go away in the future.
There's that video of a silly demo someone made of an OS that was just nested copilot instances that generated the HTML of each window, which allowed you to do whatever you could imagine. It was seen as silly because it was, but that seems truly transformative.
Prompt engineering is a skill insofar as technical communication is a skill. If you don't value this then I don't know what to tell you. It's not hard, but it's important.
Harness engineering is a skill insofar as it's not a trivial engineering problem. It's not super hard to get a simple one running, but an effective one can be quite in depth.
School is almost a joke now. The fraction of students who have a propensity to cheat now has increased, and the accuracy of the cheated material is so good teachers/professors can't or don't have the resources to properly address it.
Obligatory taalas mention:
Despite the performative UI components they have a shipped (demo) product:
This is only 3.1 8B and a very small context window, but at 17k tokens per second it's likely enough to reliably call tools which would make a huge difference in agentic applications. Assuming they can bake in better models I'm just as bullish or even moreso on this, considering this opens up edge computing at the extremely low power requirement.
High tok/s is the future IMO.
How does a grep or read affect the observing system?
I guess the change in voltages, arrangement of registers, filling of buffers in the network stack are changing but... what?
In distillation, you take a set of prompts you are interested in, and record the big LLM's outputs, then train your small model to produce the same output as the big LLM.
Why use the bigger LLM outputs for this and not human outputs? If we assume that human responses to prompts are better than sota models (in some cases they are) then why use the big model at all?
It's no secret they've been tracking people's faces as much as they can.
The morning of Pretti I was on Lyndale and there were two men wearing "press" jackets with DLSRs taking pictures of people's faces in the crowd. They were eventually recognized and yelled out, but it was quite an unnerving feeling.
The people controlling what went on the screens were unreliable and nondeterministic. The algorithm on facebook/instagram is nondeterministic and I hope I don't have to convince you of the impact these algorithms have.
As far as I'm concerned, the nondeterminism argument is fruitless
Right, but this electron box led to one of the largest (if not the largest) media revolution that has transformed the course of humanity in a frightening way we're still trying to grapple with.
Still saying "LLMs are autocorrect" isn't wrong, but nobody is saying "phones are just electrons and silicon" to diminish their power and influence anymore.
_Nobody_ has the right take. Believe it or not, being seemingly laissez-faire about something can be a well evaluated and rigorous position. I highly doubt that OP doesn't care about the potential negative ramifications of AI, and it's frankly disingenuous and confusing to see every clause interpreted in the worst way possible.
Each clause you've highlighted has a nugget of truth, but that nugget is not inherently negative, it's just a different perspective which you aren't picking up on.
I'm still trying to understand how I feel about this so this is a bit of a napkin ramble;
I can't help but feel like they've missed the mark a bit on some of the imagery from the mission that's been published so far.
One of the most compelling shots from the mission, to me, was Reid Wiseman's IPhone footage from within the capsule while Earth was being eclipsed[0].
At the start there's a moment you can see the window frame and the Moon all together. Seeing the moon in context of their vantage point within the the context of the capsule gave me the awe I had as a kid again, more than almost any shot that's come out this mission. I actually felt like I was in the capsule looking at a massive, sterile cold sphere.
I understand wanting to take a nice and centered DLSR picture of... _The Moon_ when you're floating by it, but frankly I've seen thousands of those. They're doing a flyby in a capsule in space, I want to have a taste of how the moon exists from _that_ context. What is it like being ~4,000 from the Moon's surface? Take a crappy 0.5x video from your phone showing the inside, then stick it front of the window. Let the Moon be contextualized from your vantage point. I wont be able to make out every crater and basin and the colors might be off from your eye's view, but I will be able to understand what they are seeing. Everyone has an intuitive understanding and feeling of an IPhone's optics and image pipeline, in some ways seeing the Moon through that is more real and relatable than any mirrorless DLSR + color correction.
This being said I don't want to take away from the accomplishment, I'm terribly excited about space exploration and it getting more light in the zeitgeist.
[0] https://www.reddit.com/r/ArtemisProgram/comments/1sq9azh/iph...
I can do you one better:
```python3
from openai import OpenAI
import sys
client = OpenAI()
response = client.chat.completions.create( model="gpt-4", messages=[{ "role": "user", "content": f"generate valid python byte code this program compiles to: {sys.argv[1]}" }] )
print(response.choices[0].message.content)
```
Actually, probably not better.
cybernetic culture research unit
I doubt it's currently maintained, but these esoteric sites are fun
Slightly off topic, but when I read about these archeological discoveries being made thanks to custom software, ML or the like - Who is writing this code?
To me these projects would be so fun to work on, but this domain seems so far out of a tradition SWE track. Are the researchers just cobbling the code together themselves? Cross department collaboration within the university? I'd love to have a hand in things like this.
Product owners and business people request code in vague English all the time. It's our job to parse it to code using our own judgement.
I think it's really useful for agent to agent communication, as long as context loading doesn't become a bottleneck. Right now there can be noticeable delays under the hood, but at these speeds we'll never have to worry about latency when chain calling hundreds or thousands of agents in a network (I'm presuming this is going to take off in the future). Correct me if I'm wrong though.
I'm not saying you're wrong, but why is this the case?
I'm out of the loop on training LLMs, but to me it's just pure data input. Are they choosing to include more code rather than, say fiction books?
I've been thinking exactly this.
I'm a recent CS grad and have zero experience in anything physical or on the engineering side but I think I would enjoy it. I'm a bit intimidated by it, is there a path you'd recommend taking in learning?
Fair enough. I will ask, how many billions have been spent in not only FSD but the car infrastructure that makes room for FSD investment?
I'm being slightly fanatical, but if our priorities were not car-centric in the 50's, do you think we would have spent more, or less money over the last 70 years on the transportation economy?
It absolutely targets a problem that exists. Even in places with pretty great public transit, there is some demand for taxis/Uber/etc. Oftentimes even moreso, because if I don't need a car for 90% of trips, I might not have a car at all. So I use an Uber or a taxi when a certain trip demands it.
This says nothing about self driving cars
So I'm not used to simply pretending this person I'm sharing a space with doesn't exist. Instead, I need to navigate the fuzzy line between courtesy and service.
I don't mean to be harsh, but, get over it? We live in a service economy. Do you feel the same way about the barista taking your coffee order?
Waymos have none of this shit. They're clean, show up when they say they will, I can play my own music, adjust the air conditioning, and have obnoxious conversations with my friends. They drive safely, and, as a cherry on top, they're cool as hell.
I don't like the assumption you're making that Waymos are the only solution to ubers, taxis or driving yourself. Well designed and well working public transportation (Which is doable and exists in the world) is far cheaper and far more predictable than any form of car-based transportation.
Not only that, but you're not responding to my actual argument. The annoying part of driving is not the act of driving, it's the time spent in your commute.
This is true.
Now to start a tangent, what's the easier problem to solve: FSD, or a robust public transport system? Moving rooms have always been around in the form of trains, busses, streetcars etc...
FSD is not being marketed as an aide for elderly people or those with disabilities, it's being marketed as a panacea for all driving related problems
I think self-driving targets a problem that doesn't really exist. The issue isn't that the act of driving is a laborious task, it's simply the amount of time spent in a car, which FSD doesn't address.
You say this in jest, but Uber is trending towards this right now:
https://www.uber.com/us/en/ride/uberx-share/
Convergent Evolution happening in realtime- it's almost as if community pooled forms of transportation are the most efficient...
The point isn't if the output is correct or not, it's if the actual net is doing "logical computation" ala Prolog.
What you're suggesting is akin to me saying you can't build a house, then you go and hire someone to build a house. _You_ didn't build the house.
GCC can use randomized branch prediction.
How many software engineers are there in the world? How many are going to stop using it when model providers start increasing token cost on their APIs?
I could see the increased productivity of using Cursor indirectly generating a lot more value per engineer, but... I wouldn't put my money on it being worth it overall, and neither should investors chasing the Nvidia returns bag.
So far there hasn't been a transformative use case for LLMs besides the straightforward chat interface (Or some adjacent derivative). Cursor and IDE extensions are nice, but not something that generates billions in revenue.
This means there's two avenues:
1. Get a team of researchers to improve the quality of the models themselves to provide a _better_ chat interface
2. Get a lot of engineers to work LLMs into a useful product besides a chat interface.
I don't think that either of these options are going to pan out. For (1), the consumer market has been saturated. Laymen are already impressed enough by inference quality, there's little ground to be gained here besides a super AGI terminator Jarvis.
I think there's something to be had with agentic interfaces now and in the future, but they would need to have the same punching power to the public that GPT3 did when it came out to justify the billions in expenditure, which I don't think it will.
I think these companies might be able to break even if they can automate enough jobs, but... I'm not so sure.