HN user

smodad

74 karma
Posts0
Comments20
View on HN
No posts found.

Based on rumors, it seems that because the demand is so high and supply so short, that Nvidia is having to select who gets the cards. I bet they're likely thinking about who can do the most impactful work and OpenAI would definitely be in the running for the short list of companies who are actually shipping. So I bet OpenAI / Microsoft got a lot of the newest cards.

For sure, I didn't mean to imply OpenAI's time-to-market was the main reason for my singularity comment.

Say what you want about OpenAI, and trust me I think there's a lot of material there, but you can't say they don't ship. Nobody else is even close.

So nutty. I didn't think they were going to drop such a huge bomb on the industry. We're firmly in "singularity" territory where the pace of acceleration is so fast, huge strides are being made every day / week, it feels like.

Interesting paper. My confidence on the following idea isn't high, but I have an idea I don't think a lot of people agree with me on. The authors pretty convincingly show that transformers don't generalize well. I agree, but I would add that I don't think humans generalize well either, so I don't think showing that transformers don't generalize well precludes them from ultimately being "intelligent" (in some way). In fact, the inability to generalize to new situations is actually a core problem in learning for humans.

For example, think of a college freshmen who knows Algebra and begins Calculus. Even though they have all the fundamental "mental tools" available at their disposal to figure out how to prove that an infinite series converges (or not), I would guess that very few students can actually connect the concepts in the right way to see that. They have to be given examples and be shown how to use the knowledge they gained in algebra and how it applies to infinite series. Is their inability to generalize well evidence that they aren't intelligent?

Perhaps what we're really saying, and the real problem for us, is that these systems aren't intelligent enough to be useful to us. Just like we might say that a very smart student, like a von Neumann, would likely be able to figure out how to solve an infinite series without having seen it before. I think it's reasonable to say "we want transformers to generalize better than the average human."

To that end, I think we'll have to imitate the way that a very smart human solves new problems. For example, the students who are capable of solving an infinite series without having seen the concept before, might use a creative approach of trying different things to see if they can discover the pattern. So, my hunch would be that we will need to provide transformers with bolted-on sub-routines like "generate several hypotheses and test them to see which pattern fits the data best" before we can expect them to generalize well.

Tl;dr: Transformers don't generalize well, but I don't think humans generalize well either, so I don't think that fact precludes their ability to be intelligent in the limit. I think we'll have to imitate how humans generalize by giving the transformers additional "mental tools" before they can generalize like very intelligent humans.

I bet you're right. Even if you take into account that a data center is a monster consumer of energy, in the grand scheme of things it's not that big. Some back of the envelope math:

Global electrical production in 2022 was ~30,000 TWh.[1]

If we over-estimate that a hyperscale data-center will consume about 100 MW of power, per year that would be around 876 GWh.[2]

Let's overestimate again and say that 1,000 new data centers spring up in a year, every year they would consume 876 TWh.

Which, is 2.92% of total electricity production. Which given the fact that I overestimated the energy consumption by more than an order of magnitude, I would say the term "rounding error" is accurate.

I think the main limiting factor in the near term is going to be chip production capacity. The fabs take so long to spin up, it's going to be a while before we can even consider "electricity production" being a limiting factor.

[1] https://yearbook.enerdata.net/electricity/world-electricity-... [2] https://cc-techgroup.com/data-center-energy-consumption/

What's funny is that even though the DGX GH200 is some of the most powerful hardware available, there's such a voracious demand that it's not gonna be enough to quench it. In fact, this is one of those cases where I think the demand will always outpace supply. Exciting stuff ahead.

I heard Elon say something interesting during the discussion/launch of xAI: "My prediction is that we will go from an extreme silicon shortage today, to probably a voltage-transformer shortage in about year, and then an electricity shortage in about a year, two years."

I'm not sure about the timeline, but it's an intriguing idea that soon the rate limiting resource will be electricity. I wonder how true that is and if we're prepared for that.

PhD Simulator 3 years ago

You found one of your ideas appears in a recently published paper. You can no longer work on it.

This is one of the things I thought of right away when ChatGPT got released last year. "God, there's probably so many PhD candidates right now in NLP feeling despair like all their work was pointless ...as if million of voices cried out in terror and were suddenly silenced."

It's hard in the moment to know whether what you're working on has any utility. So just do your best and keep chugging!

For sure. Just to be clear, I'm not saying the situation we're in where we have to release it to the general public is a great situation to be in. But I think we're at a point where there's not any optimal solutions, only tradeoffs.

I have to disagree. Not releasing it to the public makes it more dangerous. One major downside of all development being done in private is that the AI can very easily be co-opted by our self-appointed betters and you end up with a different kind of dystopia where every utterance, thought and act is recorded and monitored by an AI to make sure no one "steps out of line."

I think the solution is releasing it to the general public with batteries included. At least that way, the rogue AI's that might develop due to irresponsible experiments could be mitigated by white hat researchers who have their own AI bot swarm. In other words, "the only way to stop a bad guy with an AI is a good guy with an AI."

Really? I've always been the opposite for some reason. I find it easier to visualize 20% of a pie chart, or of a crowd, or a length, etc. It's a lot harder for me to think about 1/5 of a crowd or 1/5 of a length.

It's tangentially related to this, but I've seen posts on social media over the years of how people visualize different kinds of information. For example, some people think of the months of the year in a vertical fashion, others horizontal.

I'm curious if this is something similar where some people find a certain way of thinking of probability easier than others.

GPT-4 3 years ago

I'm curious as to what your take on all this recent progress is Gwern. I checked your site to see if you had written something, but didn't see anything recent other than your very good essay "It Looks Like You’re Trying To Take Over The World."

It seems to me that we're basically already "there" in terms of AGI, in the sense that it seems clear all we need to do is scale up, increase the amount and diversity of data, and bolt on some additional "modules" (like allowing it to take action on it's own). Combine that with a better training process that might help the model do things like build a more accurate semantic map of the world (sort of the LLM equivalent of getting the fingers right in image generation) and we're basically there.[1]

Before the most recent developments over the last few months, I was optimistic on whether we would get AGI quickly, but even I thought it was hard to know when it would happen since we didn't know (a) the number of steps or (b) how hard each of them would be. What makes me both nervous and excited is that it seems like we can sort of see the finish line from here and everybody is racing to get there.

So I think we might get there by accident pretty soon (think months and not years) since every major government and tech company are likely racing to build bigger and better models (or will be soon). It sounds weird to say this but I feel like even as over-hyped as this is, it's still under-hyped in some ways.

Would love your input if you'd like to share any thoughts.

[1] I guess I'm agreeing with Nando de Freitas (from DeepMind) who tweeted back in May 2022 that "The Game is Over!" and that now all we had to do was scale things up and tweak: https://twitter.com/NandoDF/status/1525397036325019649?s=20

Startup Ideas 5 years ago

I feel like I always learn something new whenever I read Gwern's articles.

I found this idea particularly hilarious for some reason:

ceiling-mounted infrared beam emitters+guidance infrared camera, steered by segmenting CNN: it automatically detects & warms up the bare skin segments of female objects until thermal equilibrium of 37.5°C is reached, ensuring thermal comfort of women without requiring heavy clothing—thereby resolving the Great Office Thermostat Wars once and for all.

My first thought was that a built-in GPU could measure certain parameters about the display and adjust the color profile to be "pro" level accurate.

Most default / "auto" profiles for monitors and TVs are accurate enough for lay people out of the box. But for people who really need color accuracy, you need a calibration device. So maybe Apple is trying to obviate the need for a calibration device and have the "smart" monitor calibrate itself.

(The built-in GPU might even use data from an iPhone like the "color balance" feature Apple announced at the "Spring Loaded" keynote recently.[1] The AppleTV uses information from the IR array on the front of the iPhone to help calibrate your TV for use with AppleTV.)

[1] https://youtu.be/JdBYVNuky1M?t=941

Thanks for posting the link, because I was about to ask. It seems like printing out the reference on a home printer would make the accuracy dependent on whether or not your printer was well calibrated (and then you're back to the same problem about color accuracy.)