HN user

lanthissa

398 karma
Posts0
Comments111
View on HN
No posts found.

they're subsidized if you max them out, i'd imagine most users are paying $20 for maybe $2-5 of tokens.

anthropic probably has more customers that use more of their sub, but for open ai where a lot of their subs are consumers through chatgpt.com, they have a lot of free money to work with there

at least openai had the guts to call code red and improve.

releasing a model worse than luna is pretty bad. Its clear that internally they did not decide coding was a thing until relatively recently.

pi is the neovim of agentic harnesses, its barebones and extremely configurable. if you're the sort of person who likes that sort of things its a forever product, nothing is going to displace it because you have full control.

opencode builds a lot more in, which is better if you dont want to fiddle with config.

it really depends on the framing, some work, especially fun work that develops skills is more valuable than people realize.

From an org perspective the goal is to create the highest curve of performance over the lifetime engagement of the employee or from the employee perspective their career.

And a lot of that depends on teh relationship of the people involved. From my perspective its a net negative when if my movers worked out the day before, their muscles will be sore and they'll do a worse or slower job. From the moving companies perspective its good, they'll be stronger for more jobs. Unless they quit or are fired that day, in which case we're back to bad.

The real evaluation isn't the macro vs the sublime edit. its does the thought process of making them macro improve them in other things, and what were they doing before that. In my experience no one is going use the time they spent writing a macro or a learning vim to do real meaningful work, they're doing that because they're bored or burned out and want to think about something else they find fun at the time.

your problem isn't your employees choose to write random scripts, its that they dont have a sense of urgency or care about their current task.

i mean its a value exchange, the last mile matters a ton to the consumer, the value prop the average person gets from amazon vs shopping in 2000 is insane and scales up the more valuable your time is.

Not only are prices good, but if i lose my remote or need a shovel for the winter or whatever in 2000 im going to a store for that, that 15m of my time each way+parking+less choice.

Lets say i make $50 an hour, and lets say i value my free time at my working rate (i'd argue most people by definition value it more or they'd be working more hours).

Saving me 10m in the store 15m of driving both ways and 2-3m of transit is worth more than most items i purchase.

Amazons solved the last mile problem by having one vehicle bring each item to each home so its marginal cost of delivery is the distance between each home instead of the round trip between home and return that a customer has.

The more items you buy at one store the less valuable this is, which is part of why costco is well served by having such large product sizes.

the 3rd party isn't an intermediary, youtube content is fantastic because it found a way to pay people to produce something for every niche, and reward the ones doing well with more viewers and more money.

ofc you would prefer high quality content delivered directly to you for free, because you're ignore the producers perspective.

long term he doesn't want his own distribution though. youtube offers you new viewers constantly. peertube only makes sense if it has a viewerbase and an ad network, and at hte point we're back to google but they're going to pay more since they have a larger user base and more ads which means their monetization rate is going to be much higher.

disrupting youtube is super hard, you basically need to bring your own audience from something else to do it, or have a platform already existing that you expand into the us.

this is exactly right, people dont realize that they're getting a great deal on youtube. if you want to disrupt youtube you need to do it by winning over the creators which is a difficult thing to do since no new platform can subsidize to the scale of what youtube offers already due to its size, the only ones who really could are meta since they have the ad network and users already to funnel to it, or another company willing to eat a loss for a long time. The issue with eating hte loss, is video is a pretty painful loss to eat compared to text, so why not go into every text market first for places like meta.

i can tell claude to call haiku or sonnet.

in copilot i can could pass all my automated browser testing to 3.5 flash, could use dirt cheap models for simple tool that were better than claude.

i'd just use open router but @work we only have copilot and claude

i ran out of claude credits for the first time at work in months and had to fallback to copilot.

pleasantly surprised, claude's way ahead in tooling but the ability to designate what model your subagents use and having access to all models is a better feature than all of what claude offers combine atm.

The only limit on the amount of ai can consume in a month a work is dollars, so anything that helps with cost is the best model/harness for me.

It also did a better job at smart designating subagents itself where as claude often used higher cost models.

Claude Sonnet 5 22 days ago

so it doesn't get blocked. last time they said a model was great at cyber it didnt turn out well

the only people its relevant for is the people in first. We wont know what any other state would do until someone passes the us, if that happens.

it sucks that we're in a place where the us has an dishonest leadership, because the current situation would be pretty reasonable if any other admin was in charge.

let models go free, until one proves dangerous in the real world then require gov approval after that.

I don't think anyone rational would have the position everyone should have insta access at the same time to the highest model once it crosses the point of enabling actual dangerous things.

why would it? if you're the us gov and sam&greg your good boy giving you 25m

and dario's you naughty boy who you dont agree with politically.

Let 5.6 free, keep fable chained and anthropic instantly sees rev loss and has to cave.

AI is the first technology that doesn't incentivize offshoring, and incentivizes co-location of talent.

A NYC dev and a dev in india have the same ai costs, based the ratio tokens/salary it becomes less of comparative disadvantage to be in NYC.

Now combine that with the fact that AI makes the act of generating code less a % time of the job, and the ability to get/refine requirements more of the job and you have a decent shift.

Waymo Premier 1 month ago

in nyc any car based transportation is slower than subways often, but everyones so narrow minded they just think about their own life. if you're old in nyc cabs/ubers/waymo are a big deal, without them you're stuck walking to a bus stop or subway and that gets hard in your 70s and 80s.

Claude Fable 5 1 month ago

did they not pay them enough to get good ratings on the other 3 models?

whats the logic in claiming its a borked metric when everything listed is an anthropic model.

yes? the future for any verifiable task is the model attempts to verify initial state and a goal then decomposes its tasks in to every smaller verifiable subtasks, with /memory being the persistence between runs and then /dreaming on the results of those memory files + run data to introduce new ideas.

i think thats the path to async agi these labs are imagining. The only limit is that sensor data you have on the world or your system, how long your willing to wait, and how much you're willing to spend to parallelize it.

maybe once you start building out these verified workflows you can feed that back into training and hte model starts to get a feel for the world to the point that it can intuit things since it has these sub paths built.

my personal agi test is can a model, trained on video of someone knocking on a door and then open it encounter a microwave for the first time and open it when the foods done without knocking.

MAI-Code-1-Flash 2 months ago

i used to use opus for everything, thats not an option once you move to a multi agent system unless you're working on like high end research. I could easily spend 3k a day if i was using opus as just a normal dev.

As we build a better and better harness and better feedback/verifiers we're switching more to 3.5 flash. I think chinese models would work too, but we cant use those atm.

Generally theres a coordinator running opus and an ever growing set of skills and subagents that take actions using weaker models and output feedback to the coordinator opus.

I'm pretty convinced at this point we're past the level of intelligence needed for most tasks most devs do and that will trend down as we better build harnesses for our own codebases.

this is the finance team doing a fantastic job. keep in mind they're raising this cash right before 3 major ipos in their sector which people will need to raise money for and will fight against htem in the narrative.

If i was a google cfo and was trading at a premium to my peers before that, i'd want to raise the cash now. Look at MSFT, they're trading at 25 forward p/e and were buying back shares at 40. If they have to issue equity over the next few years the spread between teh performance of the 2 cfos could be 40-50b on that alone.

Just as a google shareholder, this company bought back shares hand over fist at a low p/e for a few years, issues 100 year debt at low rates, and is selling equity when its at a premium to its peers right before 2-3 major ipos of competitors put selling pressure on the stock for a while.

I don't know who's going to win the llm battle, but googles finance team has been doing their job fantastically.

flash 3.5 is the best price/performance model for what i'm doing. I had been using opus for everything but as we started running many agents at once, and then eventually agent managing sub agents frontier is not an option.

we started model testing the cost/performance of our skills and agents and flash 3.5 wins in most things.

As people develop harnesses for their codebase i think the intelligence required comes down a lot.