HN user

Analog24

379 karma
Posts1
Comments204
View on HN

So the "E2B" and "E4B" models are actually 5B and 8B parameters. Are we really going to start referring to the "effective" parameter count of dense models by not including the embeddings?

These models are impressive but this is incredibly misleading. You need to load the embeddings in memory along with the rest of the model so it makes no sense o exclude them from the parameter count. This is why it actually takes 5GB of RAM to run the "2B" model with 4-bit quantization according to Unsloth (when I first saw that I knew something was up).

As someone else mentioned in this thread, it is not possible to understand how everything you purchase works. That is also incredibly infeasible. What you are asking is to effectively revert back two hundred years of technological progress (a rough estimate for the last time people were actually self sufficient at a local level).

I think the expectation that the entire consumer market (or even just a majority) is going to collectively become universally informed about all their purchases and shift the market for the better is far less likely then a government intervention being successful.

If you go to countries where there was never any government intervention relating to cigarettes do you know what you'll find? A lot more people smoking cigarettes.

The New Inflection 2 years ago

Internships are definitely not a waste of time. First of all, they pay well in and of themselves (at least in tech). Second of all, most internships are filled by students on their summer break. What better use of that time than getting an inside view of a company they might want to work for? From the companies perspective it gives them a much better idea of how well the candidate performs and how they fit in at the company, giving much higher confidence in a hiring decision that could lead to significant future impact.

"neat future" is very ambiguous. At the moment there is nothing even close to transformers in terms of performance. I suspect you are right in general but I'm not sure about the "near future" part, there needs to be a pretty significant paradigm shift for that to happen (which is possible, of course, I just don't see any hints of it yet).

This is the reason why they're not going to move on device anytime soon. You can use compression techniques, sure, but you're not going to get anywhere near the level of performance of GPT-4 at a size that can fit on most consumer devices

"It's six years later and Siri, Alexa, and Google Home are still nearly as dumb as they were back then". You can't even keep a coherent discussion and you are delusional about the significance of your work. You shared a slide with nothing but generic pie-in-the-sky use-cases and you act like it gives you some credibility on the subject ("let's make an AI system that can do the work of your non-professional employees!"). And to top it off you act like you've been successful here! Again, you shared nothing but a slide with generic use-cases that a 12 year old could think up. I don't know what you think you proved. Enjoy your imaginary pedestal.

Do you have any publications to back up your claims about your work? They seem more than a bit grandiose. If you're ideas are as novel and useful as you say then you should publish them.

And I'm sorry, but you're completely wrong about companies recognizing commercial potential. I worked on Alexa for five years, it is a far harder problem than you think. It is nowhere near as simple as "we just weren't looking at the right NN architecture or optimizer!" You're acting like it was a novel idea to think LMs would be extremely useful if the performance was better (in 2017). I'm just trying to tell you that isn't the case.

Scaling up an LM from 2017 would not achieve what GPT-4 does. It's nowhere near that simple. Of course companies saw the potential of natural language interfaces, there has been billions spent on it over the years and a lot of progress was made prior to ChatGPT coming along.

It fails at _some_ arithmetic. Humans also fail at arithmetic...

In any case, is that the defining characteristic of having a good enough "world model"? What distinguishes your ability understand the world vs. an LLM? From my perspective, you would prove it by explaining it to me, in much the same way an LLM could.

If you read some of the studies of these new LLMs you'll find pretty compelling evidence that they do have a world model. They still get things wrong but they can also correctly identify relationships and real world concepts with startling accuracy.

You can't just wave your hand and tell someone that words are broken up into sub-word tokens that are then transformed into a numerical representation to feed to a transformer and expect people to understand what is happening. How is anyone supposed to understand what a transformer does without understanding what the actual inputs are (e.g. word embeddings)? Plus, those embeddings directly related to the self attention scores calculated in the transformer. Understanding what an embedding is is extremely relevant.

No it wasn't. You have the misconception that an AI system has to achieve anthropomorphic qualities to become dangerous. ML algorithms are goal-oriented, they are optimized to maximize a specific goal. If that goal is not aligned with what a group of people want then it can become a danger to them.

You seem to be dismissing the entire problem of AI alignment due to some people's belief/desire for LLMs to assume a human persona. Those people are uninformed and not the ones seriously thinking about this problem so you shouldn't use their position as a straw man.

> > Should we develop nonhuman minds

The embarrassing part. Preaching a religion under professional pretenses.

There are quite a few companies with the explicitly stated goal of developing AGI. You can debate whether or not it's possible as that is an open question but it certainly seems relevant to ask "should we even try?", especially in light of recent developments.

What developer working on anything meaningful does not rely on documentation? You certainly have to make the documentation available similar to how you would have to make it "available" to an LLM. I think you might be missing the point about what the potential use-cases for these systems are.

Because there is a huge selection bias that influences the sample of people that comment on these types of posts. I will back up the claim that, in my experience, a lot of this thread is hyperbole. There is some truth to it, sure, but don't make the mistake of thinking these anecdotal accounts actually represent the typical experience.

The idea of connecting CV to audio via spectrograms pre dates Jeremy Howard's course by quite a bit. That's not really the interesting part here though. The fact that a simple extension of an image generation pipeline produces such impressive results with generative audio is what is interesting. It really emphasizes how useful the idea of stable diffusion is.

edit: added a bit more to the thought

Neutrons aren't that hard to capture. They are certainly harder to capture than charged particles but there are plenty of materials that are dense enough to reliably capture neutrons. This is how heat is extracted from the reaction to use in a generator. The activation of the containment material is a problem but it's not even close to the level it is for fission reactors where you're forced to deal with spent fuel rods.

At the moment fusion is obviously not cheap but no one is planning on using the technology in its current form for actual power generation. The processes involved will all get more efficient and given the astronomical upper limits of energy output from fusion it doesn't take a big stretch of the imagination to think that it will eventually be preferable to solar and wind power. There's no guarantee that will happen but hopefully this breakthrough will trigger more investment and momentum to make it a reality. I also want to add that I'm very pro solar and wind, especially in the short term.

Could you elaborate on your point a bit more? If you're talking about utilizing the weak force vs. the residual strong force then I'm not sure this argument holds up.

Also, when comparing to renewable+storage you have to consider how much land has to be dedicated to energy use in these scenarios. Wind and solar require orders of magnitude more than a potential fusion reactor (or an existing fission reactor).