HN user

mlinsey

5,123 karma

www.twitter.com/mlinsey

Engineering @ Homejoy (YC S10)

hnchat.com:8pJMixG6jypNc7WE7714

Posts61
Comments636
View on HN
feld.com 6y ago

Taking care of 2M Essential Ears with Glowforge

mlinsey
3pts0
blogs.wsj.com 11y ago

Former Zynga and Fortinet Executive Mark Vranesh Joins Homejoy as CFO

mlinsey
2pts0
www.standard.co.uk 12y ago

Homejoy opens London headquarters

mlinsey
2pts0
techcrunch.com 12y ago

Homejoy Comes To Clean Up The UK, Its First Market Outside North America

mlinsey
1pts0
techcrunch.com 12y ago

Homejoy (YC S10) Opens An Office In New York City

mlinsey
31pts5
blog.samaltman.com 12y ago

Employee Retention

mlinsey
125pts90
techcrunch.com 12y ago

Homejoy (YC S10) Raises $38M as It Looks to Expand Beyond Home Cleaning

mlinsey
82pts42
venturebeat.com 12y ago

YC startup Homejoy establishes charitable foundation to support veterans

mlinsey
6pts0
techcrunch.com 12y ago

Behind The Scenes At Homejoy, A Cleaning Startup That's Really A Tech Company

mlinsey
46pts30
twitter.com 13y ago

37 YC Companies have valuations of or been sold for at least $40 Million

mlinsey
122pts65
upstart.bizjournals.com 13y ago

When a toilet brush becomes a CEO’s secret ROI weapon

mlinsey
6pts0
techcrunch.com 13y ago

Honeymoon Registry Wanderable Comes To iPhone

mlinsey
3pts0
techcrunch.com 13y ago

Home Cleaning Service Pathjoy Becomes Homejoy, Raises $1.7M From A16Z And Others

mlinsey
10pts1
www.quora.com 13y ago

The Last Mile demo day at San Quentin

mlinsey
6pts0
pandodaily.com 14y ago

Socialcam Founder: There Will Be No Instagram of Video

mlinsey
26pts12
allthingsd.com 14y ago

Facebook CTO Bret Taylor Departs (For Start-Ups Unknown)

mlinsey
222pts49
www.quora.com 14y ago

Why are software development task estimations regularly off by a factor of 2-3?

mlinsey
7pts1
techcrunch.com 14y ago

Like Siri? Sonalight Brings Powerful Texting-By-Voice To Android

mlinsey
4pts0
www.facebook.com 14y ago

Sean Parker on Steve Jobs

mlinsey
8pts3
techcrunch.com 14y ago

Imagine K12′s 2011 Startup Class Aims To Invigorate Education With Technology

mlinsey
23pts3
www.forbes.com 14y ago

Crowdbooster's (YC S10) Social Media Appeal: From Esther Dyson To Lil' Wayne

mlinsey
25pts0
techcrunch.com 14y ago

Crowdbooster (YC S10) Launches Social Media Monitoring Dashboard

mlinsey
19pts0
mashable.com 14y ago

Crowdbooster (YC S10) Nudges You When New Twitter Followers Have Klout

mlinsey
26pts0
techcrunch.com 15y ago

Debteye (YC S11) Wants To Be Your (Much Cheaper) Credit Counselor

mlinsey
93pts42
www.justin.tv 15y ago

Video of StartX Demo Day (formerly SSE Labs)

mlinsey
5pts1
www.startupsopensourced.com 15y ago

Infiltrating Any Startup

mlinsey
66pts19
www.quora.com 15y ago

What was the code quality of the initial version of Google?

mlinsey
295pts73
www.businessinsider.com 15y ago

The Data Viz Whizzes From Mint Are Launching A New Startup, Visual.ly

mlinsey
4pts0
blog.crowdbooster.com 15y ago

Crowdbooster's Social Media Heroes: how Mayka Mei grew moxsie to 140K followers

mlinsey
4pts0
thenextweb.com 15y ago

Forget Facebook and use The Fridge (YC S10) to Plan your Spring Break

mlinsey
3pts0

"With one big caveat: you need a very clear set of requirements"

Most new product launches are an exercise in figuring out what the product requirements should be through trial and error (really: through ongoing dialogue with your users). Even mature products can have requirements change over time as the market changes.

I think internalizing this reality is why most senior engineers who work in domains that touch the messy real world will reflexively push back against perfectionism.

For ordinary coding (making webapps, frontend and backend), the non-SOTA models do fine, but the SOTA models have better judgement and need me to intervene less often, allowing workflows like "Go and pitch me a solution to this problem, then go build it". 6 months ago, I would give lower-level tasks and read more of the code myself.

So I don't "need" the SOTA models, but they do help considerably. And I am also working on some other projects that may have been outside the reach of the earlier models entirely.

I have some sympathy to geohot's view when it comes to pure informational chatbots. It's a first amendment issue, I'm allowed to write and read books that are useful to getting away with crimes, etc.

This obviously doesn't work at all when the agents start doing real things in the real world, though. "Hey AI, I don't like my neighbor, find an exploit in the firmware for his car and make the cruise control malfunction and crash him next time he gets on the highway". This is committing a crime, not just talking about theoretical crimes. The AI can and should refuse it.

I think he's anticipating and discarding this objection with his introduction, which otherwise feels disconnected from the rest of the article. FWIW, I have changed a bike tire and I'm pretty sure most of the MTS at the big labs could. This sort of "they're just bookworms who don't understand the physical world" rhetoric aside, we are currently seeing a ton of effort and expense go towards giving the AI agents hooks into being able to perform as many real-world-consequential actions as possible. And you can do a surprising amount with just bits, from writing code to breaking into systems to sending some combination of emails, phone calls, and currency to instruct meatspace humans to do things, etc.

Also, I just want to step back and actually give Costco some props here for their brand loyalty. Inshallah, when I have a 400 Billion market cap company and someone writes a fawningblog post about how awesome I am from a social-benefit perspective and how socialist mayors should learn from me…it’ll be the skeptical comments that will be suspected as paid shills for my $2T competitor!

You got me, I'm totally paid off. Real New Yorkers love to go five miles (just for the closest one to me, in Manhattan seven avenues over and around 100 streets north), take a cab back from shopping, then haul costco-sized boxes up the stairs. I only pretended to not love that to get the sweet check from daddy Bezos (good thing he didn't notice my aside criticizing the e-bike trailers breaking traffic laws, though I guess that was subtle if you haven't been following the controversy about them here)

Costco is an elegant solution for the suburbs, where everyone is driving around a vehicle large enough to store giant boxes off a pallet and bring them home. Here in NYC, it's really impractical to go to a warehouse and carry a month's worth of supplies home on the subway. The flip side is that the Amazon last mile deliveries are done on electric scooters that can bring a whole trailer worth of packages to the mailroom of a big apartment building. These have some other externalities (eg around traffic laws and sidewalk space) but they are on the whole a lot more resource-efficient than everyone driving around giant cars to go to the store.

I'm not sure where in my post you thought I was suggesting either ignoring an export control or that all export controls are illegal.

Export controls are governed by the Export Control Reform Act of 2018. The procedures the government must follow to enact a rule is defined by the Administrative Procedures Act. It's not a slam-dunk case, but I think there are significant questions around the arbitrary nature of the approved company list, the control restricting US-based permanent resident employees from accessing models, and how the rule was enacted in general which would make a viable case.

Suing the government saying that APA was not correctly followed or that the export control in question does not comply with the ECRA will absolutely not get you put in prison, that's not the kind of country we live in and honestly I have no clue what you're talking about.

The labs will not just ignore the order, there are too many other levers they can try to pull to mess with those companies. Just for some examples, think about the number of employees reliant on visas that could be revoked, the government contracts that the hyperscalers hosting them that could be canceled, the certifications that all the data centers need to be hooked up to the grid, the tariffs that could be put on critical components, the IPOs that need to be approved by regulators, the bill introduced in Congress to seize 50% of their equity...

Lots of these moves would and should be struck down in court as an arbitrary and capricious use of administrative power. Some of them might not be, and in the meantime you're signing up for tons of trouble. A trillion-dollar company does not simply go to war with the US government.

A more mid-sized company that's not so intertwined, but not so small that they can't get a good legal team, might be another story.

In 2004, I took a class where we trained "language models" that were bigram word models, on an archive of a couple years of the Wall Street Journal.

I remember someone who literally announced they were dropping the class to the whole room at the end of a lecture, saying "This isn't AI!!!"

I understand why Anthropic might not want to fight this particular one in court, because they're trying to convince the administration to let them move forward.

But would another company who is not on the trusted partner list and has less to lose taking on the admin have standing to sue here? On the basis of the export control being illegal and this putting their business at a disadvantage vs. competitors with access

ID checks are possible for first-party harnesses...but they would also mean no more API access. Your wrapper could easily become a way for a foreign national to query Fable. Maybe a few large customers like Cursor would work with Anthropic to prove they had implemented ID checks themselves as well in their own products, but being able to just get an API key and have your product call frontier models may be over.

If the existing memory makers retains control of the market and don't defect from the optimal-long-term equilibrium for themselves, that's true. It just takes one player to defect for short term gains as we've seen with some past boom-and-bust cycles. Alternatively, it takes a sufficiently-resourced player with enough incentive to enter the market themselves (NVidia, Google, Amazon, the PRC government through one of many companies...)

You're describing efforts by powerful institutions to squash the technology, which they definitely try to do, but that's just a strong signal that the technology itself is inherently opposed to centralized power, not an enabler of it.

Other technologies like surveillance (and, perhaps, AI) are more clearly centralizing and enabling of power.

The difference matters a lot if you're having mixed feelings about working in technology.

My anecdata is that it heavily depends on how much of the relevant code and instructions it can fit in the context window.

A small app, or a task that touches one clear smaller subsection of a larger codebase, or a refactor that applies the same pattern independently to many different spots in a large codebase - the coding agents do extremely well, better than the median engineer I think.

Basically "do something really hard on this one section of code, whose contract of how it intereacts with other code is clear, documented, and respected" is an ideal case for these tools.

As soon as the codebase is large and there are gotchas, edge cases where one area of the code affects the other, or old requirements - things get treacherous. It will forget something was implemented somewhere else and write a duplicate version, it will hallucinate what the API shapes are, it will assume how a data field is used downstream based on its name and write something incorrect.

IMO you can still work around this and move net-faster, especially with good test coverage, but you certainly have to pay attention. Larger codebases also work better when you started them with CC from the beginning, because it's older code is more likely to actually work how it exepects/hallucinates.

The consumer surplus is quite high. Even with the regressions in this postmortem, performance was above the models last fall, when I was gladly paying for my subscription and thought it was net saving me time.

That said, there is now much better competition with Codex, so there's only so much rope they have now.

It's true he could write off xAI today and the company could still fetch a trillion-dollar valuation. But I was more referring to his stated intentions - between his stated plans, his actions taking SpaceX from a profitable company to spending basically all their revenue (plus a rumored large chunk of what's raised via its IPO) on AI, and seeing his tendency to make bet-the-farm bets on Tesla, I think it's fair to say he's committing to bet all of SpaceX on xAI.

Is this cash or compute? Elon has one of the world's biggest compute clusters spun up, and little inference demand to speak of.

Trading billions worth of idle compute, in exchange for a high-strike call option on the #3 player in the most-promising-vertical for AI, plus (presmuably) some access to their data, starts to sound like not a bad trade. Especially if you're pre-committed to betting your entire rocket company on winning in AI, and you're currently in sixth or seventh place.

Yes, cost per successful task is rising - ie, we are all paying effectively more for AI.

And yet - Anthropic is still struggling to have enough capacity to serve demand - they are virtually sold out.

And yes, are almost-as-good open models, on part with the closed models from 6 months ago (at worst), that are just a single Openrouter API call away, and yet Anthropic is still selling out. So people are paying for the premium product anyway, for whatever reason - maybe the last bit of intelligence is worth it, maybe they like the harnesses/products around the models, maybe it's a brand/enterprise sales thing.

Put aside your feelings about the AI industry and imagine we are talking about thingamajigs. Prices for thingamajigs are going up. They are still selling out about as fast (or faster) than the company selling them can build factories. There are more cost-effective competitors already in the market, but thingamajigs are selling out anyway.

Would you, looking at the thingamajig industry, conclude the "jig is almost up"? That "the returns aren’t anywhere close to what investors expect" and that the impending IPO is all some desperate hail mary to save things before the collapse?

The models that we are paying to generate tokens are already not really just LLMs, as anyone studying language models ten years ago (or someone who describes them as "next token predictors") would understand them. Doing a bunch of reinforcement learning so that a model performs better at ssh'ing into my server and debugging my app is already realllly stretching the definition of "language pattern".

I think when we do get AI that can perform as well as a human at functionally all tasks, they will be multi-paradigm systems; some components will not resemble anything in any commercial system today, but one component will be recognizably LLM-like, and act as an essential communication layer.

I agree, but also the model intelligence is quite spikey. There are areas of intelligence that I don't care at all about, except as proxies for general improvement (this includes knowledge based benchmarks like Humanity's Last Exam, as well as proving math theorems etc). There are other areas of intelligence where I would gladly pay more, even 10X more, if it meant meaningful improvements: tool use, instruction following, judgement/"common sense", learning from experience, taste, etc. Some of these are seeing some progress, others seem inherent to the current LLM+chain of thought reasoning paradigm.

Different users do seem to be encountering problems or not based on their behavior, but for a rapidly-evolving tool with new and unclear footguns, I wouldn't characterize that as user error.

For example, I don't pull in tons of third-party skills, preferring to have a small list of ones I write and update myself, but it's not at all obvious to me that pulling in a big list of third-party skills (like I know a lot of people do with superpowers, gstack, etc...) would cause quota or cache miss issues, and if that's causing problems, I'd call that more of a UX footgun than user error. Same with the 1M context window being a heavily-touted feature that's apparently not something you want to actually take advantage of...

If we have the source and it's easy to test, validate, and deploy an update - AI should make those easier to update.

I am thinking of situations where one of those aren't true - where testing a proposed update is expensive or complicated, that are in systems that are hard to physically push updates to (think embedded systems) etc

I feel like every new iteration of ways to find good content online: webrings, blogrolls, user upvoting/downvoting, giving everyone their own microblog to share interesting links, ML to learn your own preferences by your behavior - they all worked really well at first, but then eroded significantly once people figured out how to game them.

The economic incentive is overwhelming to corrupt these signals, either directly (link sharing schemes, upvote rings, bots to like your content) or indirectly (shaping your content itself to have the shape of what will be promoted, regardless of its quality).

What you almost want is to use any of these ideas and hope for it to catch on widely enough in your small niche to be useful, but not so much that it comes an optimization target.

I admit I'm surprised by the move, from a company that reportedly just talked about how they need to focus more on fewer, more strategic products.

But I also see the potential value. This is an entertaining and highly influential podcast, a lot of top VC's and founders watch it; it definitely punches well above it's audience KPI's in strategic value. I've seen many interviews or op-eds on the platform pretty clearly shape the startup discourse on X.

I also think it should run mostly autonomously, it'll only be as much of a distraction for OpenAI execs as they want it to be.

OpenAI just raised $122 billion (including future commitments), so whatever the purchase price was (we have no diea) is not going to even be a rounding error on their financial resources or their ability to pay their datacenter bills.

It's CNBC for Silicon Valley - a combination of good background noise, a broad survey of what people are talking about around the valley, and occasionally really great interviews.

They get a lot of guests to do interviews that they wouldn't do elsewhere, in part because they are unabashedly and unapologetically cheerleaders - pro-tech, pro-VC, pro-startup, pro-Big-Tech, etc. They don't grill you like an old-school journalist would about whatever the latest political controversy is, they ring a giant gong when their guest brings up a cool traction or fundraising number.

I would never use it as my only source of news for what's going on in tech, but with a lot of other tech journalism covering the downsides or problems with the industry, there is definitely a niche for them.

Just based on the number of very prominent guests they get to do interviews, they clearly have a lot of viewers in influential tech/vc circles, even if their total audience size isn’t huge.

An AI company owning a major tech podcast?

Wow, what’s next?

Ecommerce giants owning major newspapers? An aerospace company owning a microblogging platform? Startup accelerators owning tech news aggregators?