HN user

bwest87

215 karma
Posts10
Comments28
View on HN

He forgot the tokens!

It's not simple weights and numbers all the way down. The available output is pre-set by the tokens we allow it to predict.

There was a whole bit in there about not having a language module or using words. But it does. We tell it.

Humans do not come pre programmed with a set of possible "tokens". We just figure it out and I believe that fact captures something very essential. Maybe the missing piece of AGI. The fact that humans can just be awash in pure sense data, and somehow just figure out what is important and what to do. Never ceases to amaze me.

Sure there is some minimal marginal cost, but it's so close to zero that it's usually negligible, and the incentive is to basically give it away and "monetize" something else. Your point about games actually just makes my original point. Software is already usually free or dirt cheap, which is why reducing the cost to make the software can't create some "low cost / low value" quadrant. Unless your talking about bespoke software that has such a small market size it isn't worth making today. I could maybe see that area opening up, but even that software would not fit the OP's description of software that "has no owner and is not meant to be maintained"

But focusing on production cost is silly. The cost to consumers is what matters. Software is already free or dirt cheap because it can be served at zero marginal cost. There was only a market for cheap industrial clothes because tailor made clothes were expensive. This is not the case in software and that's why this whole industrialization analogy falls apart upon inspection

One thing that has become clearer to me over the years is that reasoning by analogy (like this article does) sounds a lot smarter than it is. If you look from first principles, it's clear that physical goods and software don't share the same properties and thus the analogy falls apart.

Physical goods like clothes or cars have variable costs. The marginal unit always costs > 0, and thus the price to the consumer is always greater than zero. Industrialization lowered this variable cost, while simultaneously increasing production capacity, and thus enabled a new segment of "low cost, high volume" products, but it does not eliminate the variable cost. This variable cost (eg. the cost of a hand made suit) is the "umbrella" under which a low cost variant (factory made clothes) has space to enter the market.

Digital goods have zero marginal cost. Many digital goods do not cost anything at all to the consumer! Or they are as cheap as possible to actively maximize users because their costs are effectively fixed. What is the "low value / low cost" version of Google? or Netflix for that matter? This is non-sensical because there's no space for a low cost entrant to play in when the price is already free.

In digital goods, consumers tend to choose on quality because price is just not that relevant of a dimension. You see this in the market structure of digital goods. They tend to be winner (or few) take all because the best good can serve everyone. That is a direct result of zero marginal cost.

Even if you accept the premise that AI will make software "industrialized" and thus cheaper to produce, it doesn't change the fact that most software is already free or dirt cheap.

The version of this that might make sense is software that is too expensive to make at all because the market size (eg. number of consumers * price they would pay) is less than the cost of the software developer / entrpreneurs time. But by definition those are small markets, and not anything like the huge markets that were enabled by physical good industrialization.

This video was fascinating. I didn't know about "open endedness" as a concept but now that I see it, of course it's an approach.

One thought... in the video, Ken makes the observation that it takes way more complexity and steps to find a given shape with SGD vs. open-endedness. Which is certainly fascinating. However...

Intuitively, this feels like a similar dynamic is at play with the "birthday paradox". That's where if you take a room of just 23 people, there is a greater than 50% chance that two of them have the same birthday. This is very surprising to most people. It seems like you should need way more people (365 in fact!). The paradox is resolved when you realize that your intuition is asking how many people it takes to have your birthday. But the situation with a room of 23 people is implicitly asking for just one connection among any two people. Thus you don't have 23 chances, you have 23 ^ 2 = 529 chances.

I think the same thing is at work here. With the open-ended approach, humans can find any pattern at any generation. With the SGD approach, you can only look for one pattern. So it's just not an apples to apples comparison and sort of misleading / unfair to say that open-endedness is way more "efficient", because you aren't asking it to do the same task.

Said another way, I think with the open-endedness, it seems like you are looking for thousands (or even millions) of shapes simultaneously. With SGD, you're kinda flipping that around, and looking for exactly 1 shape, but giving it thousands of generations to achieve it.

I did a chat with Gemini about the paper, and tldr is... * They introduce a loop at the beginning between Q, K, and V vectors (theoretically representing "question", "clues" and "hypothesis" of thinking) * This loop contains a non linearity (ReLU) * The loop is used to "pre select" relevant info * They then feed that into a light weight attention mechanism.

They claim OOM faster learning, and robustness acro domains. There's enough detail to probably do your own PuTorch implementation, though they haven't released code. The paper has been accepted into AMLDS2025. So peer reviewed.

At first blush, this sounds really exciting and if results hold up and are replicated, it could be huge.

I think we're not factoring in that people will react. We're already all starting to realize that the free for all is getting quite hard to navigate. My hunch is that within 10 years, we will start to see an "information immune system" develop. This could take many forms. For example, self regulatory organizations for news, or actual regulations. Like we have with food products, the use of certain words is regulated. Or it could be trusted information filters becoming the norm, the way we trust our browsers to warn us of insecure websites. Or simply some changing cultural norms, like we saw happen with cigarettes. Like it's totally fine today for news outlets to just use Twitter as a source. And maybe the bar will get higher over time. I'm spit balling about solutions, but I don't think society can or will tolerate some dystopian world where truly no one knows what's real for very long.

Life is not short 4 years ago

Ok, echoing my top level comment... An alternative framing that I've come to find more helpful is to take your life expectancy, and cut it by 2/3. For example, if you're 20 years old and your life expectancy is 80 (ie. 60 more years), pretend that you only have 20 more, so you'll only live until you're 40. It's nice cause it naturally adjusts as you get older. You'll have smaller windows to work with.

This approach strikes a nice balance. It gives you enough time to be able to really do something and change directions if you want. But not so much time that you can really waste any. It forces you to ask the hard questions about whether your day to day is truly connecting with your dreams, and whether you're on a path to get there.

Of course, Seneca didn't have life expectancy tables to work with. But I think he would have approved. :)

Life is not short 4 years ago

You should organize each day as if it were your last, so that you neither need to long for nor fear the next day.

I've come to find this "live each day like it's your last" advice to be pretty unhelpful. My favorite quote about it is, "all that goes to show you is some people would spend their last day giving you stupid advice".

The problem is that if it actually was your last day, most people would give the finger to all of their responsibilities and go party, eat cake, see friends, familiy, lovers, etc. Which is simply not an actual way to live your life. It's a way to exit your life.

An alternative framing that I've come to find more helpful is to take your life expectancy, and cut it by 2/3. Now what do you do? For example, if you're 20 years old and your life expectancy is 80 (ie. 60 more years), pretend that you only have 20 more, so you'll only live until you're 40. It's nice cause it naturally adjusts as you get older. You'll have smaller windows to work with.

This approach strikes a nice balance. It gives you enough time to be able to really do something and change directions if you want. But not so much time that you can really waste any. It forces you to ask the hard questions about whether your day to day is truly connecting with your dreams, and whether you're on a path to get there.

Of course, Seneca didn't have life expectancy tables to work with. But I think he would have approved. :)

Guided aspects... he sent over a questionairre ahead of time with a lot of broad questions. We then did a one hour zoom call going over the questions and getting to know him. It's all designed to help you figure out what you want the session to be about for you personally at that moment in time in your life. And then the session itself lasts 4-6 hours, and he will ask you many questions, but also will follow the journey wherever it takes you, and there's ups and downs and everything in between. It's all very specific to you and the guide and where you want to go with it. And lastly there's an "integration session" the following week where you talk with him for an hour to go over how it went, and what it means. Can discuss more if you want. Email me at bwest87 at gmail.com if you'd like to discuss further.

One session meaning one time. Both of us think it would be valuable to do, but on the timescale of like... once/year or once every few years. But there are people who do it once/month for several months if they have a lot of specific things to work through. Our guide actually works with a number of 'regular' therapists, and they pass clients on to him if they think a guided session is the right move. He says therapists will sometimes say, "please take this person on a journey once / month for the next 3 months" (or something along those lines)

I'm located in San Francisco. I asked around a bunch of friends, got intros, and talked to a few potential guides. Eventually got linked up with someone who's been doing it for a number of years, and we vibed. We had a few phone calls through Signal, and then decided on a date/time/place.

A friend and I did a "guided" psychedelic session earlier this year. We did it individually over one weekend. It was great. She did it for more therapeutic reasons. I did it more for philosophy/spiritual reasons. But the two main things I took away are 1.) Guided sessions are qualitatively different than recreational sessions, and 2.) It is such a crying shame that this isn't an accepted "tool in the toolbox" for therapists.

It's not about having crazy life altering, world-bending experiences (though that can happen). It's just about helping you get into a state of mind that allows for an effective therapy session. Sort of like... would you want to do your therapy session in a crowded bar, next to your mom? No, probably not. We all recognize that such a setting would not be conducive to good therapy. So similarly, we should be able to recognize that having the right setting, both mentally and physically can affect the quality of your session. Psychadelics can do exactly this.

It's also worth noting my friend has done "regular" therapy for 2 years, and she felt like there was a step change after the guided session. Her therapist noticed it as well.

When you consider that pain meds have ruined literally millions of lives through addiction, and that also virtually (maybe literally?) no one has ever died due to overdose of psilocybin, it's very confusing why one is prescribed all the time, and the other is considered incredibly dangerous. The U.S.'s perspective on drugs is so very backwards.

Oakland relaxed a lot of zoning, and permitting regulations about 5 years ago, and so now you're starting to really see production sky rocket. This article [0] mentions 2019 was had 15x more units completed than 2018. And 3x the total units from 2013-2018 combined. Anecdotally, I have several friends who've all moved to Oakland in the last year or so. It will be interesting to see how this plays out, and maybe, hopefully, SF will take the hint.

[0] https://www.city-journal.org/oakland-rezoning-california-hou...

Eh, maybe. Your core responsibility is to ship on time. Especially at a startup, it's not only "ok", it's usually the "right" call to ship faster, and accumulate technical debt. This is constant, and never really goes away. So you tend to have to intentionally bake in this time, or else it never happens. And I think this is OK! It's not always clear what is and isn't technical debt at the time it's being introduced. A lot of things seem "gross", but aren't actually much of a problem. Or you think, "oh god, this will never work when we add X feature", but then you just never add X feature, and so it's totally fine 2 years later.

But to your point, of "well shouldn't that intentionality be part of normal duties?" Sure, but shouldn't coming up with Gmail? Or fixing bugs or UX improvements also be part of normal duties? Yes, they all should. The truth is it's hard to prioritize small things like that, and it's really hard to get a sense for their value as a centralized management team. They aren't close enough to the code or the product to always know. So I think 20% time is a great way to just decentralize that, and allow the engineers closest to the issues to pick what's important to work on.

You can do it a few ways. But in my experience, the simplest and best way is to force people to demo something at a pre-specified time. If you're doing it "hackathon" style, then it should be at the end of the hackathon. If you're doing it 20% style, then probably end of the quarter? You could do it with zero accountability, but having tried to run these programs (see my other top-level reply here), I think you'll mostly get a lot of people not doing much if you don't have any accountability.

I never worked at Google, but I did work for a different, fairly well known consumer tech company doing a lot of work around "hack time". While there, they didn't have any such program, and I really wanted there to be one. So I helped implement a program for "hack time", which was pretty close to 20% time. I did surveys of over 100 engineers, put together a presentation for leadership, got directors of various eng departments on board, did a pilot program, then did post-surveys and presentations to gauge impact. We also spent time with managers to make sure they were on board, and made it clear to their direct reports that they were on board, and that it was OK to use hack time. The point is, we really tried to do this right, and get people to use the time, and assess if it was actually valuable.

We did this for a quarter across a few different large eng teams. We eventually shut it down. My high level takeaways were the following...

  * Way more engineers *say* they're into doing hack projects than actually ended up doing them. We had huge initial survey response of people saying they wanted to do something, and only ended up having like 5-ish projects that were seen through. And there were probably 20ish projects that started at the beginning of the quarter. So large drop-off.
 * BUT! That can be totally fine! Many people who actually used hack time were some of the best engineers at the company. Other ones were some of the newest at the company. You're really helping job satisfaction for those top engineers, and really helping mentorship and learning for the new engineers.

 * I now believe it's better to have regular highly condensed hack weeks (or hack days if you're a small company), rather than a spread out "20% time". Even people who really liked hack time found it hard to actually take the time when deadlines were approaching. But when you just eliminate those things with a condensed period of hacking, then people can focus. You also tend to increase participation considerably through this method. (I know cause we did this method as well)
 * Stop trying to create Gmail in your hack time! A lot of the most valuable stuff to come out of hacking was internal tools, refactors, and little things that improved daily quality of life around the company. (ie. someone made a batch upload tool for the customer support team that they *loved*!). Or small UX improvements for customers. Some people did try to create certain new products, but it's super rare that those make it into the main product Not saying it can't, or won't, or even that you shouldn't do it. But as a general rule, the best you can hope for there is you've de-risked a new feature, or gotten your manager excited about the possibility, and they'll slot in time to "do it for real" in the next quarter. Realistically, people tend to have more fun just making stuff that they actually can see the value of the next day. Or when they get to hear thank you emails immediately from a whole other team.
 * You get a lot of other value from hack time besides break-through products. Specifically, you get people across teams mingling with each other. You get new and experienced engineers hanging out. You get a feeling of autonomy. People learn new tech or new parts of the company. For example, I made one of my most lasting relationships at that company by just randomly deciding to join his hack team for a hackathon project. I also got to use GraphQL and React for the first time on that project.
Overall, I think hack time/ 20% time / whatever you want to call it is very very valuable, and companies of all sizes should do it. But you have to do it right, and you have to go in with the right mindset about why you're doing it. Do it because it's fun. Do it because you help your company meet one another. Do it because you'll probably improve the daily quality of life in small tangible ways for a lot of people for your customers, or for other employees. Do it to give your engineers some autonomy. Don't do it in order to get your next Gmail.

I think PG is getting at a larger point, which I heavily agree with, that journalists very often use words that imply a lot of "story" that is not supported by "facts". For example, I just went on to the NYTimes Economy section right now, and the second article says "Short of Workers, US Farmers Crave More Immigrants". Crave? Really? Cable news is horrendous about this. They do stuff all the time that's like, "Senator XYZ blasts Trump". Blasts? And that's just headlines. But you see it all the time in subtle lines throughout an article, such as, "Facebook's employees have been raising a stir". Or like, "The new law is leaving homeowners in a lurch". These words have no quantifiable meaning, and typically it seems journalists pick words that have more average "emotional valence" than other words, even if the facts don't quite support that level of emotional valence. And that honestly is the part that annoys me the most about normal journalism.

This is great. I think a good alternative interpretation of the findings that I haven't seen mentioned is through the lens of information theory. If a network has generalized well, and each neuron is firing seemingly at random, then that sounds like the neurons of the network all have very high entropy. If you have neurons that fire only on certain inputs, then those neurons have low entropy (ie. you can accurately predict when they will fire, which is the definition of low entropy). So those neurons are less efficient at providing you with information. If you assume that the goal of the network is to maximize entropy, then I kind of would have thought that low entropy neurons would be the ones you'd want to delete first. So the fact that they're saying they have the same effect as more random ones is interesting. Or that there's another way I need to be thinking about how entropy can be measured for a given neuron... But I think that lens is a really good one for conceptualizing information flow through the network.

The medium post above links to a Convolutional Neural Net that I built in Google Sheets. It was an interesting project, and thought ya'll might be interested to play around with it. Simply copy the spreadsheet into your own Google account and have at it!

How I review code 8 years ago

Something the author doesn't bring up, but that we started doing about 9 months ago at my company, is synchronous reviews. Meaning the committer is on the phone or in person with the reviewer. It's great. We don't do it for all PR's, but anything medium sized or above, or even small one's if they involve critical logic. The way we usually do it is the committer walks through the changes with the reviewer. Often the committer will realize their own ways of improving the code. And with the added context, the reviewer can often provide better feedback. Plus the X factor of just two people talking who come up with ideas, improvements, etc. And half our team is remote, so this wasn't a natural outgrowth. We make it happen, but I think it's worth it.

Former HR grad here. I gotta back up Shawn, and HR in general. Hack Reactor is filled with lots of staff who care deeply. I graduated before the author's cohort, so I can't speak directly to the author's issues. But I can say the founders and the staff care deeply. If there was a rough patch, I know they'll fix it, and it sounds like they already have taken great strides to do so. But the author's main points just don't make sense. A 98% job placement rate is ludicrously good. Does a CS program have that rate? Maybe. But even if they do, it takes 4 years and 250k. HR is more than an order of magnitude less in both time and money, and this is being cited like they could do better? Seriously? Also, could you learn stuff on your own? Yeah of course you could. But you pay 20k because HR gives you an environment and team mates to learn from that you just can't get on your own. Also, they've iterated on the curriculum literally hundreds of times at this point (they do (or at least did) sprint reflections every 2 days, and actually implemented the feedback given). It's an excellent curriculum. The school is 100% worth it. I had a friend who just graduated a week ago, and said that "he had high expectations, and even those were blown away."

Even the author said they got a job after a couple months! Which is exactly what the school promised. How does the poster feel like they wasted their time/money??

Also, the author has no grounds to say HR isn't the best, cause they haven't been to others. I can't say it's the best either. I can say I had an excellent experience, and so did all of my class, and many others. And I do have 3 friends off the top of my head who have gone to other bootcamps, and they all cited those as having major problems (no help finding jobs, bad teachers). Anyway, this post sounds whiny. HR does excellent work.

Hey guys. My name is Blake West. I'm the co-inventor of Hummingbird. First, thanks for all the discussion. There really is no such thing as bad publicity. You've helped us crack 13k downloads in under 48 hrs. 2nd, I thought I'd just quickly respond here to some of the main points...

1.) "Traditional is fine. it's not broken": Neither were text-only command line interfaces. But GUI's are just easier to learn for most people.

2.) I am indeed a professional keyboardist, have been playing and reading traditional notation since age 7, and I also teach 25 students a week still. I know theory like the back of my hand, and can talk modes, b9 chords, and 12-tone rows all day long if you like. Jazz and pop are my thing and I play to lead sheets more often than not now a days. So I know this fro m both angles.

3.) We do have key signatures. They're at the start of each song in plain english. no need to be cryptic with symbols.

4.) Relative pitch notations seem like a good idea, but they really aren't. The function of a pitch is honestly pretty subjective and changes frequently in a song. Not to mention, they'd be much harder to learn, especially for young students. Relative pitch is an abstraction, and abstraction is the luxury of experts.

5.) Why not use a chromatic staff or other such layout? Because we actually wanted some adoption. Most other alternate systems have failed because they're SO different that they are completely alien. Ours is "backwards compatible", and also if you did want to switch over to traditional from Hummingbird, you could, and it's not that crazy.

6.) Why not use colors? Because music still gets printed and photo copied vey often, and will for at least another 5-10 yrs. And color printing is still 7x more expensive.

7.) "You can't hand-write the symbols". Yes, you can. It is slower, but our point is that most music is printed off of notation programs today, so hand-writing is usually reserved for small edits, or writing fragments from scratch. This is still completely fine even with an unsharpened pencil with Hummingbird. I have done it with my students many times.

8.) "Picking out lines and spaces isn't that hard" - If you spent time around kids you would be SHOCKED at how bad their spatial reasoning is before about 9-11 yrs old. It is really hard for them without a ton of frustration. That frustration often leads to them thinking they're "bad at music". That turns them away, and it shouldn't have to.

I know there's other stuff, but just not enough time...

Thanks.