I am a long-time user of stackoverflow with 16k points, and even I got all my questions of the last five years downvoted into oblivion.
HN user
duchenne
Nuclear-powered ion thrusters could solve this issue. They provide low acceleration for a long time consuming very little consumable. This would allow the telescope to stay at the right position for observation.
Yes, it would work.
I am working in Mistral robotics team. I confirm this is map-less. The only inputs are the text prompt and the front camera rgb image.
Cloud models can use batch processing which is significantly more efficient. A local model has basically a batch of one which takes as much time to process as a batch of 100 because the gpu is memory bound and spend most of its time loading the model from vram to the gpu cache while the gpu cores are idle. With a batch of 100 the model loading time and compute time are roughly similar. So local Models have a first 100x lower efficiency. Secondly, local models are idle most of the time waiting for the user to write a prompt, so the efficiency gap is probably more around 1000x.
If they are worried about firearms, why don't they target CNC mills rather than 3d printer? Can you even make a firearm in plastic?
Some US company specialize in selling CNC mills specifically for firearms.
Ex: https://realghostguns.com/product/gg3-s-cnc-deposit/
It is sold with the cut codes for the AR-15, AR-.308, 1911, Polymer80 and AK-47 receivers and frames.I have done that at meta/FAIR and it is published in the Llama 3 paper. You usually start from a seed. It can be a randomly picked piece of website/code/image/table of contents/user generated data, and you prompt the model to generate data related to that seed. After, you also need to pass the generated data through a series of verifiers to ensure quality.
Looks awesome. Can we get the same thing for pytorch?
There is a manga/anime about this: doctor stone.
For the knowledge preservation, I guess that a copy of deepseek has most of the required information. But, it would be hard to run it in a primitive world.
But we have landfills which are full of great raw materials. I would argue that it is easier to collect steel from a landfill than from a mine during the industrial revolution.
I had one for years. Never had overheating issues, except if I put it on my blanket for long.
The asus zenbook pro is great. The 16inch version is not really bulky. It is 2.4kg, 2TB, 3.2k resolution, great design and build quality. $2200
The 14.5 inch version is 1.6kg, 2TB, 2.9k resolution, also great design and build quality. $1700
https://www.asus.com/laptops/for-creators/zenbook/zenbook-pr...
https://www.asus.com/laptops/for-creators/zenbook/zenbook-pr...
Come on... Meta has been refining pytorch for more than a decade. It basically contains all that you need to train LLMs, including the latest technologies. What more do you need? The part of the code that is specific to Meta infrastructure?
The reasoning happens in the chain of thoughts. But OpenAI (aka ClosedAI) doesn't show this part when you use the o1 model, whether through the API or chat. They hide it to prevent distillation. Deepseek, though, has come up with something new.
Training a 1B model on 1T tokens is cheaper than people might think. A H100 GPU can be rented for 2.5$ per hour and can train around 63k tokens per second for a 1B model. So you would need around 4,400 hours of GPU training costing only $11k And costs will keep going down.
But, if the non-profit gives all its assets to the new legal entity, shouldn't the new legal entity be taxed heavily? The gift tax rate goes up to 40% in the US. And 40% of the value of openAI is huge.
Except that a plane has passengers. But this rocket had none. It did not even have cargo. And it crashed in a pre-evacuated zone. There is no need to have the same level of security for these two situations.
Counter-intuitively, larger models are cheaper to train. However, smaller models are cheaper to serve. At first, everyone was focusing on training, so the models were much larger. Now, so many people are using AI everyday, so companies spend more on training smaller models to save on serving.
Most SMBs would be able to run it. This is already a huge win for decentralized AI.
In French, it is called the "hidden face of the Moon", obviously because we cannot see it from the Earth point of view.
Is it possible to buy it?
Is this released yet? Where can I buy or rent some? Even the previous version?
The most important paper to understand this issue is "Sacling Laws of Neural Language Models" by Open AI in 2020 [1]. Many consider it the most important paper that predicted the high performance of modern LLMs.
This paper shows how the loss decreases when you increase the model size, compute, or training dataset size.
From the article:
Convergence is inefficient: When working within a fixed compute budget C but without any other restrictions on the model size N or available data D, we attain optimal performance by training very large models and stopping significantly short of convergence.
It clearly states that when you are limited by your training time compute, you should under-train your model.
The training for Phi-2 took 14 days on 96 A100 GPUs
This would mean that it costs around ~30k USD to train.
If training an LLM becomes cheaper than buying a car, it could democratize AI a lot.
ASMBLY in Austin is doing well. Lots of machines/space/people.
This sounds like a study made from afar by just reading numbers without even talking to Koreans.
When I talk to Korean parents, the vast majority of them tell me that raising even one kid is exhausting.
They usually try as much as possible to satisfy every requests of their baby: carrying the toddlers for hours until they fall sleep, letting their kids sleep in the parents bed until they're 6, spoon-feeding them until they are 6, and so on... A Korean pediatrician friend told me that Korean kids score the lowest in the world in terms of autonomy.
So, parents of one kid are already exhausted and think it would be too hard to have a second one.
How many tokens per second do you think we can get out of this 6TFlops NPU?
I see many comments wondering why the original authors do not reveal their synthesis process.
The reason is simple. They work for a private company. Not a university. Not a public lab.
They do not reveal the process for the same reason than openAI does not reveal its process. It is because they are sitting on a quadrillion dollars opportunity and they want to grab it.
Similarly as OpenAI, many initiatives are trying to make an open source alternative so they have a risk of not profiting from their extraordinary invention.
When hearing gunshots from the occupying soldiers, many people would just run away.
In French as well: "papillon de nuit"