I don't think people working at Amazon "know that it is a part of a larger bad", it's one of the most trusted American institutions.
HN user
lacker
Current status: working on Acorn, a theorem prover with built-in AI.
https://acornprover.org
Also doing some software work on the DSA, a next-generation radio telescope going up in the Nevada desert:
https://www.deepsynoptic.org/overview
Previously, looking for aliens:
https://lacker.io/physics/2022/01/21/looking-for-aliens.html
Before that, I was the founder of Parse (YC S2011), the simplest way to build a mobile app. We were acquired by Facebook in 2013 and had a few exciting years there.
Unfortunately, we shut down the hosted Parse service in January 2017. Fortunately, a lot of the Parse magic lives on as open source:
https://github.com/ParsePlatform/parse-server
Before Parse, I founded Gamador (YC W2010). Millions of people have played Gamador's casual games.
Before that, I was a software engineer at Google working on search algorithms.
Before that, I was in grad school at Berkeley bouncing around between computational biology and AI.
You can follow me on Twitter: http://twitter.com/lacker
My email is just my hn username at gmail.
[ my public key: https://keybase.io/lacker; my proof: https://keybase.io/lacker/sigs/Jhx53TkSPU1FfKRiXpL5WXxSlr9XMDZgSlaIcEOpU_c ]
Yes there are studies, for example last year Pangram's false positives were measured to be under 0.5%.
https://www.pangram.com/blog/third-party-pangram-evals
Personally, at first I thought these sorts of tools were dumb and wouldn't really work, but I think it works because it just isn't designed to be "adversarial". If you want your AI to trick Pangram, you can make an AI to trick Pangram. It just catches people who are cutting and pasting from the AIs without putting any more effort into hiding it.
It is a smell. But it's the EU that smells bad, when it comes to tech regulation. It's the smell of cookie popup warnings.
Think of it this way, how would you turn a set into a vector in the first place? We solve this in programming a lot, for example, the "one-hot" encoding for neural networks. Here, a set turns into a vector that has a zero for every item that isn't in the set, a one for every item that is.
Now, there are a lot of things that |v| for a vector can mean. In the L1 distance you just add up the absolute value of each dimension. You could argue that that's a simpler sort of |v| than L2.
And there you go! |S| on a set actually means exactly the same thing as |x| on a vector, if you interpret sets as vectors in the right way.
"Falling behind schedule" doesn't really seem like the right term, for a sector of the economy that has been accelerating for the past few years.
You could easily describe this trend positively rather than negatively, like:
"Google has built an incredible amount of datacenters in the past few years, which makes sense since Google Cloud revenue has tripled since 2021. But they are trying to grow even faster and add more revenue."
The next step in my fight against screen addiction is to have my children not watch Toy Story 5.
I don't think it's nonsensical, it's just another name for the same thing. E.g. in the Haskell wiki it says, "the Error monad, also called the Exception monad".
Yeah, and rate-limiting is only one of the things a PaaS needs to handle, to avoid looking like a bad actor to the underlying platform. The trickiest thing to handle is people using your PaaS to host malware, because:
1. There may be no simple rule of thumb like "suddenly using tons of bandwidth"
2. Bad actors can open up so many accounts, you have to close them automatically
3. Malware can infect a good actor, who is unaware or struggling to deal with it
Perfectly demonstrating the truth of the "Microsoft org chart" cartoon.
The specific code I was working on, I had a general idea of the sort of performance improvement that would be possible. I just thought that it would be too hard for the models to figure out without a lot of hand-holding.
But it ended up being not "too hard ever", but more like, in 1 out of every 5 tries, the model did in fact manage to get a large refactoring to the point where it improved performance. So once I set it up to try something, use the perf test, see if it worked, if not, throw it away, repeat. Then it started, slowly, finding some useful things.
It's an especially awkward situation because Railway is a competitor of Google Cloud, with many third parties involved. So, I just think it will take them a little more time to figure out how to message things.
To me, what it sounds like is that a Google Cloud system identified Railway as a misbehaving customer. Spam, hackers, that sort of thing. Often this happens for "platform as a service" companies, because Railway themselves probably do host some spammers and hackers, and they have their own systems for dealing with it.
So, it's quite possible that according to the Google team, Railway violated the terms of something or other, and according to the Railway team, they did not, and now everyone has to argue about it.
But who knows, this is just me guessing based on some experience running a PaaS that itself was running on top of AWS.
I didn't dig into what the actual repository was doing, but personally, I took some inspiration from the idea after reading about it and realizing that I might have been underestimating the ability of LLMs. I put a bit more work into a performance harness I was using locally and just set some agents to brainstorming and they did seem to find some great stuff. So I don't really have a stance one way or another on this specific repo, but the general idea seems like a really good one.
Reminds me of:
“In his presence, reality is malleable. He can convince anyone of practically anything. It wears off when he’s not around, but it makes it hard to have realistic schedules.”
I wonder, if you ask a local LLM to access a forbidden site, from within Russia or China, can it figure out a way to do so, out of the box?
The conclusion that "insurance companies using algorithmic tools have failed Californians who lost their homes to fire by systematically undervaluing their properties" seems pretty dubious to me. Everyone is shooting the messenger by getting angry at the insurance companies when fire insurance isn't cheaper. Meanwhile many insurance companies are leaving California entirely.
It isn't the "evil algorithms" at fault here - it's the high risk of fire.
There is no avenue by which you make GitHub better by continuing to use it as it has been.
I feel like in a very mundane sense, I pay GitHub for a service, and they use that money to pay developers, to then make GitHub better.
It's tough to be working somewhere when usage is booming, and your service is crashing all the time. It's also tough to migrate your infrastructure between platforms, which it sounds like GitHub finally has to do in order to scale to the next level, to really take advantage of being part of Microsoft, although that has to feel pretty frustrating in the short term.
So hang in there GitHub team. Just keep fixing things.
ChatGPT is a great resource for learning things, if you really want to learn.
I hope that this leads us to shift education towards helping people learn things, when they do want to learn. Instead of forcing people to learn things that they do not want to learn.
Isn't that how it should work?
If you write the police and ask them to delete all their data about you, that isn't a thing that they do. It shouldn't matter if the police store their data on AWS or their own servers.
Flock is a tool used by the police so it should work the same way.
I've been making games with JS in the browser with my kids, ages 7-13. Very simple games, the sort where we can just use emojis instead of real game assets. Even just building a game inside a Claude Artifact is pretty fun.
The nice thing about JS is that there is not very much overhead in setting things up, debugging weird things, restarting.
This part of the essay makes me feel moved by the author's situation.
I am sitting down after a long walk outdoors. It should have been relaxing, but I was processing - processing another interview pipeline that has fallen through. I'm in my 6th month of unemployment, despite job hunting 40 - 60 hours a week, starting literally the day I was laid off - because the company needed to make cuts and remote workers were top of the list.
That sounds really tough, and I'm sorry the author finds themselves in this situation. Six months sounds grueling.
I think the interview process is likely to be completely overhauled in the age of AI. I don't really know what will happen. I used to be in favor of the standard code-at-a-whiteboard approach, but nowadays the actual work is even further from that. But I haven't seen an AI-aware interview process yet that seems like an improvement.
At any rate, these systematic changes are likely to come too late for the author. Hang in there. Maybe it's time to consider a bigger change, like moving cities and looking for in-person work. I like working remotely but it's harder to get a remote job, and the in-person stuff does have upsides. Good luck out there.
It's funny to use "the market value of all taxi companies combined" as a proxy for how valuable the self-driving market will be, because that's exactly the reasoning that led people to underestimate Uber. The market value of all taxi companies combined was pretty small when Uber started.
That said, you could be right! Maybe self-driving will never be worth more than that. It's really hard to tell what business models will be like in the future. But this is the cultural mismatch, it seemed like GM leadership did not want to be in a risky business where they were betting billions of dollars on the success of self-driving. Clearly, to some people, that seemed like a really good bet to make. Time will tell.
It seems tough culturally.
If you look at it from an outside point of view, right now Tesla is worth $1.6T, Waymo is worth $130B, and GM is worth $72B. If Cruise were actually a third viable competitor in this race, it would probably be worth more than the rest of GM. Self-driving is just a far more valuable business than car-making.
So from that point of view it would make sense to say, don't worry about the rest of GM too much, you should be willing to sacrifice all of that to increase the changes of making Cruise work.
It's hard to change the culture at a place like GM though. Does the GM CEO really want to take a huge amount of risk? Would they be willing to take a 50-50 shot where they either 10x the company's value or lose it all? Or would they prefer to pay a few billion dollars to avoid that risk.
In general electronics aren't recycled because people don't care about recycling them.
The easiest piece of electronics equipment to recycle is probably an iPhone. You can give an old iPhone to Apple and they will recycle it for free. But still most end-of-life iPhones are not recycled.
It doesn't surprise me that Siri continues to be bad - Apple's current plan is to use a low-quality LLM to build a top-quality product, which turned out to be impossible.
What does surprise me is that Google Home is still so bad. They rolled out the new Gemini-based version, but if anything it's even worse than the old one. Same capabilities but more long-winded talking about them. It is still unable to answer basic questions like "what timer did you just cancel".
It would be great if home assistants actually started to understand me personally. Not even in the sense of "what am I like", more like in the sense of, "when I ask what the weather is today, do I want a long lecture, or do I just want you to say high of 60 degrees, wear long sleeves".
The problem with AI art is that it mostly sucks right now. Well, for "high art" - it can't write a novel, it doesn't create interesting artistic images. It's great for mocking up product UIs. And there are exceptions when an individual human puts a lot of work into it, for graphic art at least. Novels, it doesn't seem that close.
Yet.
I don't know if it will always stay this way, though. If one day I read a novel and I think, this is a great novel. I appreciated it, I felt myself growing from it. And then later I learn it was written by an AI. That's it, that will prove that great AI novels are possible. I will know it when I see it. I haven't seen it yet, but if it happens, I'll know.
So it's really just a technical question. Not a philosophical one.
It's pretty common in California for cities to abuse the permitting process to extract money from homeowners. But on the other hand, these homeowners are getting subsidized by Prop 13. For a typical house in the Palisades bought 34 years ago ChatGPT estimates the subsidy is about $15,000/year. So, I have a little bit of sympathy but they're really on the benefitting end of California's various forms of tax craziness.
If they actually worked right now, the demand would be high. Demand is certainly high for Waymos. Even if they worked worse than a Waymo I think the demand would still be very high. But it's hard to tell if (or when) it will work well enough to actually be a real product.
In my experience, whenever you mandate open source software, you get software so unusable that it might as well be closed-source. Like, it doesn't compile, and they ignore all bug reports.
I'm not sure if I have the right mental model for a "skill". It's basically a context-management tool? Like a skill is a brief description of something, and if the model decides it wants the skill based on that description, then it pulls in the rest of whatever amorphous stuff the skill has, scripts, documents, what have you. Is this the right way to think about it?