When a frontier makes a succesfull edit based on the plan that it made, it leaves an procedural trace in turn biases the NEXT model, low cost model, straight into procedural action. The cheaper model doesn't need to reread everything again because it has enough information from the frontier model to complete the task. A simple "plan" of what needs to be done does not carry this information.
HN user
trash_cat
Yes this is correct. It essentially biases the cheaper model into procedural action grounded by the frontier model. To put it mode generally, it injects the cheap model context with more useful information, where it might not need to read all the files again.
There is more to it than typing "llama serve".
Discovering that Enterprise customers tolerate higher prices compared to retails consumers is discovering demand elasticities not PMF. Am I missig something?
Nobody forces you to use touchscreen exclusively?
Sales people using it a lot to scout prospects and understand a person's seniority in an organisation, to target better and prepare a strategy to pitch higher up the chain.
Geopolitics and industrial policty aside, I think it's important to check how stable and reliable these chips are. I wouldn't count on them being on par with "western" ones. Correct me if I am wrong here.
My two cents that this is part of the learning curve. With collective experience this type of work will be more understood, shared and explored. It is intense in the beginning because we are still discovering how to work with it. I think the other part being that this is a non-deterministic tool which does increase some cognitive load.
I think the naming schemes are quite arbitrary at this point. Going to 5 would come with massive expectations that wouldn't meet reality.
BitChat comes to mind.
Clearly there is some demand for those papers, and research, to exist. Good opportunity to fill the gaps.
The former IMF chief Kenneth Rogoff has been talking about this and appeared on NYT Ezra Klein's podcast that I highly recommend[0]. He also talks about China and the role of the dollar at the end with Dwarkesh Patel[1]. A lot of the discussion I see here is adressed by him.
[0] https://www.youtube.com/watch?v=pT2cohNt6a4 [1] - https://www.youtube.com/watch?v=P2b4TjQa4gk
What constitutes real "thinking" or "reasoning" is beside the point. What matters is what results we getting.
And the challenge is rethinking how we do work, connecting all the data sources for agents to run and perform work over the various sources that we perform work. That will take ages. Not to mention having the controls in place to make that the "thinking" was correct in the end.
You have to understand that these large corpos move like whales, and the money you quoted is a rounding error. I´ve seen a company department burn cash it was asigned on purpose so it wouldn´t go back to finance (indicating that the department isnt using all their money and something is wrong).
Nobody with interest in politics thinks it's about drugs. It's a pretext and a way to gain legitimacy to exert force over foreign nation with some legitimacy that would otherwise clearly go against international law.
But isn't context window dependent on model architecture and not available VRAM that you can just increase or decrease as you like?
There is some circular financing going on, but AI accelerationists think this will be offset by demand, value, and adoption in businesses. Hence these deals are warranted for the incoming demand.
> "Turns out the major bottleneck is not intelligence, but rather providing the correct context."
But this has more or less always been the case for LLMs. The challenge becomes context capure. Which in my opinion is the real challenge with LLM adoption. Without the right contex, some tasks just cannot be reliably completed.
I think it would be better to ask why do states allow trading with the country your state is at war with.
This concept is closely reated to politics of inevitability coined by Timothy Snyder.
"...the politics of inevitability – a sense that the future is just more of the present, that the laws of progress are known, that there are no alternatives, and therefore nothing really to be done."[0]
[0] https://www.theguardian.com/news/2018/mar/16/vladimir-putin-...
This article in question obviously applied it within the commercial world but at the end it has to do with language that takes away agency.
This is incredibly useful and interesting. I have tried asking claude to generate XML code for various diagrams for import to draw.io with varying success. But I feel like if I could incorporate these instead, or markdown, for a specific graph instead of pure XML would yield better results.
I think the better question is to answer why do emergent properties exist in the first place.
I disagree with the premise that emergence is binary. It's not. What we determine "emergent behaviour" is partly a social concept. We decide when an LLM is good enough for us and when it "solved" something through emergent properties.
This is literally what I use AI for. Excellent project.
Isn't that quite a pessimistic view though? Of course there will always be something that irks you, but that does not mean you can't be fulfilled at some point in life and be aware that you are really happy and there is nothing you would change.
it’s clear that Trump will back down on any extreme positions he takes.
That has always been the case with him. That's literally his negotiation tactic.
This is very interesting. I feel like we are going towards a future where you will have personal agents that know a lot about us and will interact with other corporate and government agents to complete beurocratic and non beurocratic tasks. Those with more capable agents will have to pay a premium.
Confabulation means generating false memories without intent to deceive, which is what LLMs do. They can't hallucinate because they don't perceive. 'Hallucination' caught on, but it's more metaphor than precision.
Fun fact, "confabulation", not "hallucinating" is the correct term what LLMs actually do.
Agents HAVE Models. There is a big difference if you give a model access to tools and perhaps some form of memory.
I think the author points to cultural changes in business comunication, which doesn't give any clues to what you should invest your money.