HN user

SeanAnderson

4,680 karma

Head of AI & Insights @ https://www.schoolstatus.com/

Previously - staff eng. / team lead @ https://collage.com/ (9 figure exit, shutdown post acquisition)

Built https://www.howbazaar.gg/ a popular gaming website for https://playthebazaar.com/

Earlier - built popular YouTube music extension called Streamus (https://en.wikipedia.org/wiki/Streamus)

email: Meo.DDR at gmail dot com

Posts39
Comments913
View on HN
downdetectorsdowndetectorsdowndetector.com 8mo ago

A down detector for down detector's down detector

SeanAnderson
203pts62
www.1x.tech 8mo ago

1x – safe humanoid robots that do your chores and offer personalized assistance

SeanAnderson
3pts3
fyrox.rs 9mo ago

Fyrox Game Engine 1.0 Release Candidate

SeanAnderson
2pts0
news.ycombinator.com 9mo ago

Ask HN: Most effective way to reduce excessive digital media consumption?

SeanAnderson
20pts40
nyx.run 9mo ago

Nyx – An Experiment in Artificial Survival

SeanAnderson
4pts2
news.ycombinator.com 11mo ago

Tell HN: ChatGPT 4o has been re-enabled

SeanAnderson
1pts0
franciscosan.org 1y ago

Francisco San Grants

SeanAnderson
2pts0
news.ycombinator.com 2y ago

Ask HN: Advice for deploying Arduino hardware into remote environments?

SeanAnderson
1pts0
www.forbes.com 2y ago

Nvidia is now more valuable than Amazon and Google

SeanAnderson
338pts369
news.ycombinator.com 2y ago

Ask HN: Is it possible to effectively discuss code architecture with AI yet?

SeanAnderson
2pts1
en.wikipedia.org 2y ago

Boquila

SeanAnderson
2pts2
en.wikipedia.org 2y ago

Flyby Anomaly

SeanAnderson
2pts0
news.ycombinator.com 2y ago

Ask HN: Do you think GPT-5 will release before Gemini Ultra? Will it be better?

SeanAnderson
16pts16
www.yowza.social 2y ago

Yowza – The new social network by Cards Against Humanity

SeanAnderson
2pts2
ant.care 2y ago

Show HN: Symbiants – digital ant farm / mental health companion

SeanAnderson
23pts12
news.ycombinator.com 2y ago

Tell HN: I took autonomous taxi rides to/from/around SF downtown yesterday

SeanAnderson
14pts20
www.stilldrinking.org 2y ago

Programming Sucks

SeanAnderson
3pts1
bevyengine.org 2y ago

Bevy Engine: A data-driven (ECS) game engine built in Rust

SeanAnderson
1pts3
en.wikipedia.org 2y ago

The Immovable Ladder

SeanAnderson
1pts0
news.ycombinator.com 3y ago

Ask HN: Is Paul Graham's article, “Change Your Name”, still relevant today?

SeanAnderson
2pts1
www.reuters.com 3y ago

Regional bank stocks plummet as First Republic's demise weighs

SeanAnderson
1pts0
www.wsj.com 3y ago

First Republic Bank Shares Sink 40% After Earnings Report

SeanAnderson
5pts6
news.ycombinator.com 3y ago

Ask HN: Do you feel OpenAI is in a league of their own?

SeanAnderson
6pts5
www.bloomberg.com 3y ago

Credit Suisse Finds ‘Material Weakness’ in Financial Reporting

SeanAnderson
3pts0
fred.stlouisfed.org 3y ago

Liabilities: Earnings Remittances Due to the U.S. Treasury

SeanAnderson
1pts1
tynansylvester.com 3y ago

The Simulation Dream (2013)

SeanAnderson
1pts1
news.ycombinator.com 3y ago

Ask HN: Will you be paying for ChatGPT+?

SeanAnderson
17pts53
news.ycombinator.com 3y ago

Ask HN: What simple facts have you learned surprisingly late in life?

SeanAnderson
41pts82
www.reuters.com 3y ago

India's Adani slammed by $45B stock rout

SeanAnderson
53pts18
news.ycombinator.com 3y ago

Ask HN: Why not default to non-paywalled links if they're allowed and desired?

SeanAnderson
13pts5
Claude Sonnet 5 22 days ago

Crazy. I just changed the default for our entire org to Opus because people were continually unimpressed with Sonnet's abilities. It's fascinating to think how varied people's experiences are when interacting with LLMs and how much the outcomes depend on how people approach interacting with the models.

Everyone says SpaceX IPO price is too high, but it's the most interesting IPO in a long time, it's critical to the US government, and America is, frankly, addicted to gambling. I'm not convinced it's going anywhere but up for a long while.

I disagree. Shipping functionality that works and users consume is all that matters. Everything else is noise. Technical debt can be viewed through this lens - it reduces the rate in which functioning code is shipped. That's bad, but it's only one of many dials.

The author states they feel that using LLMs allowed them to ship years faster. That's years of time in which they can collect feedback and iterate. They might even choose to scrap the entire project and rewrite it based on their learnings. The practicality of this is directly enabled by agentic coding.

Well, I'm commenting from a place of bias, as I'm Head of AI at our company and am in charge of rolling out agentic coding throughout the engineering org. So, bear with me a bit.

We're B2B SaaS in the Ed Tech space. It's very sales-driven. There's only so many players in the space, customers come with a laundry list of things they've seen others do and expect you to have those features, too. There are basic expectations that need to be met, some of those are compliance, but, sadly, a lot of what actually drives sales is just... flashy shit that looks good to those signing the checks not those using the underlying software. We lost a sale recently because someone was upset we didn't have the ability to give digital stickers to children - seriously.

You're more than welcome to tell the customer they're wrong and not give them their stickers. Or you can ask Claude to build stickers for you in two days and keep up with the Joneses.

Don't get me wrong. Customers aren't retained long-term with flashy shit. People churn out because of poor UX, security fears, pricing hikes, etc. Those frustrations tend to build over years and pain has to get pretty high because it's effortful to shift software providers. But, for getting new customers, sales is driven by flashy features and, at least in our experience, we need to be able to build those as quickly as our competitors or we lose out.

You get a discount for paying for a full year on Teams and Enterprise can involve contractual obligations. It's a lot of effort to get buy-in to change providers and to shift an entire organization. The winds change frequently in this space and the pain needs to get to a certain level before it's worth rolling the dice.

Your competition's behavior necessarily affects you unless your company has an unassailable moat.

If other companies are able to tolerate larger amounts of tech debt while shipping new features faster then you'll be out of a job at some point when your company loses market share.

It's fine if you disagree with the idea that AI lets established companies ship faster. I'm not here to argue that. But I think it's pretty easy to empathize with "why might one need to change their behavior due to this new technology?"

We encourage candidates to use AI on the homework and to be comfortable sharing the prompts they used and the workflow they engaged the AI with to get to their end result. We've experienced a wide range of proficiencies in using AI to solve the technicals. Anything from lazy one shots with 1k loc changed and 0 awareness of trade-offs to very surgical, 200 loc changed where the candidates broke down the problem and guided the AI step by step.

Whether to lean into or push back against using AI in the technical was a major point of discussion for us when building the hiring pipeline. Ultimately we decided it would be fighting against the current to try to prevent candidates from using AI and so we decided to assume they would and build questions in to evaluate their efficacy.

I'm also not sure it's fair to say we invest no time just because we use AI. We hop on a call with each candidate after they submit the technical and ask questions about their process, how they decided scope, and try to figure out how much awareness they have of what they coded.

My current job has me overseeing a few teams of engineers working on ~10+ y/o legacy software systems that have not been especially well maintained. As an example, one team had a completely broken CI pipeline due to numerous flaky tests. They had configured the CI pipeline to rerun tests multiple times and still the master branch had like.. a 40% pass rate. Super ugly, but the suite took ~40 minutes to run and they were demoralized enough to not want to investigate it anymore.

I came in, set Claude up, gave it read access to CI artifacts, had it build out some tooling to monitor the rolling pass/fail rate over the last 30 days, and let it loose. It identifies the worst offending flaky tests, forms hypotheses on whether it's a testing issue or a production issue, then tries to divide-and-conquer until it gets minimal reproduction steps. If it's not able to create deterministic reproduction then it'll make a best guess at fixing the issue and grind away at test re-runs all night until it can try to figure out if it fixed the issue with statistical confidence instead.

It's not perfect. I have to throw away some of the bad solutions, but shaved 20 minutes off their pipeline and improved pass rate by 35% in a handful of weeks. Very minimal oversight on my part - just letting it run while I'm asleep and reviewing PR proposals during the day between meetings.

We have an initiative to make an entire web application significantly more accessible in response to some government mandates. Tight deadline, tons of grunt work, repetitive patterns, some small nuances on edge-cases. The team was able to create a set of skills for doing the conversion logic, slowly build up and address all the edge cases, and are now able to work several magnitudes more quickly in modernizing the app.

A team had punted repeatedly on updating Jest to the latest version because it inherently came with a breaking change to JSDOM which made some properties unable to be spied upon. Took like 20 minutes to have Claude one-shot the entire conversion when they'd ignored it for months because it just felt too finicky prior to agents. In general, everything to do with testing infrastructure is easy to push forward with confidence.

Uhm, we have an active interview pipeline where we give a take-home technical assessment. After we got a few submissions, and manually evaluated them, I fed our analyses in and our grading rubric and had it generate assessments for incoming candidates following the rubric. After checking a few pretty carefully it became clear that it was good enough to trust - the take home wasn't groundbreaking and the problem space was understood enough to be able to identify obvious issues if there were any.

I was given a small team of semi-technical people who were being used to fetch numbers from DBs for product/marketing/sales and perform light data analysis on them. A lot of their day to day was just paper pushing SQL queries into Excel spreadsheets and then transforming them into PowerPoints with key takeaways. They didn't have any experience writing code. I had Claude build a gameified playground for them where I gave them a VSCode dev container, a SQLite DB full of synthetic data emulating what they'd encounter IRL, and a Jupyter notebook filled with questions they'd need to answer by writing code to interrogate the database and form insights. In a couple of weeks I was able to get them to the point where they were comfortable writing basic Python scripts with the help of Claude and they're now off automating all their paper-pushing workflows with deterministic scripts. When they're done we're going to move them to higher value work by having them do sleuthing against our data and surfacing proactive insights to propose to Product rather than just reactively fetching data and building reports.

I was asked to quickly build a prototype for some basic AI functionality we thought we might want to add to one of the products. I was able to go from "I have no idea what I should build" to "here's a prototype we can put in front of clients and see if this idea has any merit" in about 14 hours. Just riffing with Claude from product idea to functional/technical specs, implementation plan, then full working prototype was one shot, and then a tight iteration loop for a couple of hours with me guiding it on personal aesthetic choices to give it enough final polish. Obviously I wouldn't ship this code into production, but it's really nice not having any sunken cost biases when demoing a prototype. If customers don't like it? Great, I lost one day and half the time I was multi-tasking while Claude implemented specs. Even better - I had Claude write a script to extract all the conversations I had with it and include those in the prototype repo. Then I filmed a quick demo video of my process, shared that with the engineers, and they're able to review my Claude conversations to get inspiration for how to modify their own agentic coding strategies.

I don't think anyone is questioning all the benefits of using local LLMs. Those are readily apparent.

I just don't believe for an instant that they're anywhere in the same ballpark of capabilities as running Opus or similar. My time is the most valuable resource. Opus would need to be SIGNIFICANTLY more costly and unstable for me to start entertaining local models for day-to-day development.

Perhaps whatever work you're doing makes this trade-off more sensible, but I struggle to see how that could be true. I'm averse to running Sonnet on a large amount of software engineering problems - let alone Qwen.

I moved over to Linux a few months ago. Absolutely zero issues. My only thought was, "Wow. Why didn't I do this sooner?" There's nothing Windows can do to bring me back at this point.

Just a small project to assist with some stuff at work, but trying my hand at vibe-coding a "data science playground" to try and level-up a couple of people into feeling comfortable using Claude to write data analysis tools. I generated a bunch of synthetic data, that looks like stuff we might encounter on the job, and embedded trends into the data that can be revealed through statistical analysis. I encrypted the answers and put a lil LLM in front of the answer file. You submit answers to the LLM and it tells you warm/cold by looking at the answer file. Hoping to basically gamify the learning process to make it easier/faster to get data-driven results.

I think you're just trying to see ambiguity where it doesn't exist because the looser interpretation is beneficial to you. It totally makes sense why you'd want that outcome and I'm not faulting you for it. It's just that, from a POV of someone without stake in the game, the answer seems quite clear.

I think big players also have significant risk exposure during black swan events and the timeline of their operations makes those incidents not entirely unlikely to occur. It's sort of like insurance - most of the time they just get to extort rent, but sometimes they get crushed, too.

https://finance.yahoo.com/quote/HPP/

Check the 5-year on Hudson Pacific. They're down 96% and dropping. They own a significant number of downtown commercial properties in SF and LA. They're completely underwater, their spaces are barely half full, and they can't lower rents without violating their bank loan covenants.

Of course, if the commercial landscape hadn't shifted in a way nobody could predict then, yes, they'd likely have continued to print money for the foreseeable future. Instead, they're left holding a very heavy bag and will take it to the grave.

I don't even know what areas of the United States I would consider "walkable". I live in San Francisco, don't own a car, we have "pretty good" public transit, and it's still absolutely miserable getting around. It takes me 40 minutes to go from Outer Sunset to downtown by muni. There are many locations in this city that I can physically jog to faster than public transit.

I can appreciate this technology might negatively impact other countries more heavily, but, for me, it's easily the most exciting tech I interact with and I'm rooting for it whole-heartedly. I'm at around 1000 miles logged on Waymo and am part of their beta tester program for freeway usage.

I also think that post-Covid remote work has probably damaged incentives for increasing the density of cities more so than anything autonomous vehicles will do. San Francisco is actively cutting bus routes, bus density, and threatening to significantly cut BART stops due to budget constraints and reduction in ridership.

It's odd because I do get where you're coming from, and I feel like I should be your target audience, but, for me, the ship sailed so long ago that I struggle to relate to your position.