* patients with diabetes
HN user
blake929
Abstract: Recent months have seen the emergence of a powerful new trend in which large language models (LLMs) are augmented to become autonomous language agents capable of performing objective oriented multi-step tasks on their own, rather than merely responding to queries from human users. Most existing language agents, however, are not optimized using environment-specific rewards. Although some agents enable iterative refinement through verbal feedback, they do not reason and plan in ways that are compatible with gradient-based learning from rewards. This paper introduces a principled framework for reinforcing large language agents by learning a retrospective model, which automatically tunes the language agent prompts from environment feedback through policy gradient. Specifically, our proposed agent architecture learns from rewards across multiple environments and tasks, for fine-tuning a pre-trained language model which refines the language agent prompt by summarizing the root cause of prior failed attempts and proposing action plans. Experimental results on various tasks demonstrate that the language agents improve over time and that our approach considerably outperforms baselines that do not properly leverage gradients from the environment. This demonstrates that using policy gradient optimization to improve language agents, for which we believe our work is one of the first, seems promising and can be applied to optimize other models in the agent architecture to enhance agent performances over time.
Some very interesting discussion of outlier features and quantization: https://timdettmers.com/2022/08/17/llm-int8-and-emergent-fea...
* Outlier values are used to prune values. * Transformers seem to undergo a "phase shift" in how outlier features are treated around 6.7B parameters. This could complicate research on removing them.
Maybe you and Tim Dettmers would have a lot to talk about :)
I'm not sure SF is a good example. It's not a healthy city, but it's problems go way beyond drug use and it doesn't have the same policies as what Oregon adopted.
A lot of comments are discussing the difficulty in estimating range accurately or how all EPA estimates are inflated. But the article claims Tesla knowingly uses an algorithm with inflated numbers and swaps the rost estimate out for a more accurate estimate at 50% charge. That's different than a good faith attempt at estimating range and a dark pattern.
The article says it was a mandate from Elon Musk. Not sure I believe the claim, but I'd also be surprised if he wasn't aware just how optimistic the EPA estimate is.
I get that there are multiple endpoints being tested here and some of them may have been satisfied, but I feel like I'm living in a bizarro world on hacker news where people are arguing that multiple failed engines, a failed stage detachment and an exploding rocket worth millions of dollars and and carrying a large number of limited supply raptor engines is a "massive success". Can we just call it for what it is? Mixed results maybe? Is that not a fair assessment?
My two cents on Nvidia's rollback reasoning:
4080 12GB was universally panned. The 40 series launch also got heat for price gouging, particularly the higher cost for the low end of the launch (4080 12GB). They had to raise the cost of the lower end of the 40 series though if they wanted to maintain the value of the 30 series cards and clear out the remaining inventory. They couldn't just release a true 4070 for a true 4070 price. While the name was obviously bad, it seems likely that they wanted to obscure the release of a 4070-quality chip for a 4080-price while attempting to sell off remaining 30 series. Pure speculation: maybe they were hoping a "cheaper 4080" would come across to the uninformed as Nvidia trying to lower the entry cost for 40 series rather than raising it through an expensive 4070.
Two potential reasons for the rollback come to mind: 1) higher than expected 4090 demand means they can wait to launch a 4070. 2) higher than expected heat for the thinly veiled 4070 price gouging made it worth it to wait on the release since it helps sell more 30 series cards by raising the entry price for a 40 series while getting better PR in the process.
I dont disagree that Meta is probably pouring the most money into VR right now, but there are other companies looking into it. Notably Apple might launch a similarly priced headset to the Quest pro soon. Bytedance has the Pico series which are basically Quest clones at this point, but they are looking to push the tech and outcompete Meta. They're not currently launched in the US, but that may change. Valve will probably launch an Index 2 eventually.
Some lesser known headsets brands like the Pimax and Varjo target prosumer-grade and enterprise headsets in the US too.
High switching costs and controlling the experience end-to-end may be anti-competitive business practices, but it doesn't make Facebook a monopoly. Apple also uses similar practices, but clearly has a lot of competition.
Monopoly definition: the exclusive possession or control of the supply of or trade in a commodity or service