HN user

NichoPaolucci

137 karma
Posts0
Comments50
View on HN
No posts found.
OpenAI Presence 6 hours ago

“The challenge for enterprises is no longer proving that AI agents can work, it’s making them reliable enough to do high-value work in production.”

This reads poorly to me. So they’ve proven the models CAN work, but they also say in the next line that they CANNOT do high value work in production.

No other comments from me but that first two sentence opener should have been massaged a bit. I’m sure they have a PR / messaging person (agent?) though, so maybe I’m reading too far into it.

I wonder if this points at a “shared” future (or at least things will eventually converge there whether companies like it or not). Ultimately, if you’re going to release these models that are fundamentally built on shared data - it’s pretty wishful to assume you’ll be able to harbor that model and the data, forever, and profit from it.

It also leads me to think about things like the original release of Fable 5, people were complaining that it was safeguarded too much - if you lock the models down too much they cease to be useful. So it’s going to be increasingly difficult to protect a model from competition while ALSO keeping it useful.

Yes - precisely what I was getting at with the childish comment.

It felt a bit surreal to see a machines rendition of something that very closely maps to a younger human as they explore + understand more about the world. Scary, even (to me). This was the first time I’ve seen LLM output and thought “wow, maybe it is learning.”

As I looked through the images I was unimpressed entirely, at first. But, then I started thinking, these look a little... "childish" to me.

Childish as in... A newish artist who is drawing a concept rather than light / forms (Which is something artists typically do as they understand drawing more and more).

The rose in the vase specifically - some models understood that there was supposed to be shading, reflections, the concept of refraction - others just drew "blue = glass" and "green = stem" and "red = rose".

Really odd to look at, considering if I saw any of these drawings from a human kid, I would say "good job buddy" and put it on the fridge. I'm expecting these to get better as models improve, and perhaps the artistic progression will be there along with it...

5-10 feels like a lot. I can reliably maintain hands on 2-4 separate work streams. I still need to maintain a pretty close eye on changes to our system, as I’m one of 2 devs. I’ve tried doing 5+ things at once and I just cannot get it to work. I think I move faster when I’m more involved anyhow, since I have adequate “mini-context” of what each operation is.

I’m surprised people reach for plex, not sure what I’m missing - but a Jellyfin container set up has been fantastic for strictly at home usage (running on an old gaming PC).

It’s fantastic software - the cross platform applications are pretty useable. Don’t think we’ve even had any major issues over the last 4 years using it.

Many thanks to the builders and maintainers, hopefully Andrew can take some rest and know they help a lot of people!

Definitely cheating. I’ve taught my mostly-non-technical gf how to to load up sonarr and add something to the library.

Anything we can do to avoid the hodgepodge of 8 different streaming services and rotating availability…

Man we were scrambling this morning - intermittent issues, reports of all sorts of wonkiness. Still can't verify it was related to this but it was one of the last things we looked at (because they didn't update their status page early enough)

I drive a "Victory Red" 2005 Chevy Silverado. I always thought it was a "safer" color for a vehicle.

I have always assumed that, being in a larger vehicle that is bright red, people would be more likely to spot the vehicle from further away, notice it out of the corner of their eye, or that I would generally be MORE visible to other drivers.

I'm sure the correlation insurance companies are looking at is that the driver's of red vehicles are the cause of the higher accident rate.

This is why I feel prompt injection is going to continue to be an issue. Fantastic that “Hi we are Cloudflare, give us your personal data” works.

Either we stunt the models to the point where they are not useful, or we allow things like this to seep in and create one of the most insecure concepts the internet (and maybe tech as a whole) has ever seen: a robot that can be tricked.

Well - I think that’s a natural assumption but I’d say it’s probably not true. I’ve talked to who are fully enveloped in the AI buzz. Ask it everything. Claude can plan your day, strategize your roadmap, make your pitch deck, draft your email, read the reply, and it can tell you what to make for dinner when you get home!

I wish I was kidding but I work with people doing this. I think it’s going to be a real mess in the future, we’re going to completely disintegrate the critical thinking portions of our brains.

Fortunately, I’m seeing more articles about this, some of us are noticing and raising the flag asking “is this a good idea?”

I had a VP of engineering that loved to use “abstracty” engineering terms like Claude uses. Perhaps he was operating one level above what everyone else was doing.

Loved to use fancy words, speak at a “conceptual” level. Unfortunately it was mostly just tech mumbo-jumbo and he couldn’t actually back it up with real work - but I wonder if that’s why Claude does it. Makes it seem like a higher power, hand wavey abstractions that “seem” correct but don’t actually need to be rooted in truth or detailed.

“That’s exactly the type of seam we need to prepare for in a prod-like environment, if this change lands in the data plane, we’ve effectively shut down the load bearing critical path that was needed. It’s not over-engineered; it’s the right thing to do.”

Thanks Claude, whatever that means.

Recently read some LLM generated output that mentioned the “center of gravity” within a codebase.

Also have read the term “seam” dozens of times by now, when previously I saw it maybe once or twice over years. Very abstract term.

Your most expensive users consuming $1,000 dollars a month doesn't matter in the budget? That seems like FOMO activity or something, I feel like even big teams require proof of ROI for an investment like that (1M). BTW, the ROI of toilet paper + soap is pretty easy to prove (you GET to have employees if you provide those two things).

FWIW we are a smaller company and we had a user run through 500$ in a DAY. Had to put a stop to that. I'm hopeful that our company gets better at asking what the ROI is, what is being built, how much time is it taking, etc... It's no big deal when it's 20 / 100$ a month - but if the prices end up higher we will need to start seeing some returns other than "I feel faster".

FWIW I purchased one at the beginning of this year, secondhand, for about $250. It's been fantastic for me. I used to keep physical notebooks - but I was just taking small notes or keeping a running list (I still do this, just better now).

The editing is what really makes it useful - instead of wasting paper after writing some scratch math, I can just select the area and clear it - giving me a fresh page with all the info I already had on there. It's been pretty great for my use case, maybe a little expensive - but I've used it every day so far.

It is 1 level up from traditional pen + paper IMO, editing / moving things around / bulk erasing is a major upgrade that I didn't know I wanted.

These are probably mostly the enterprise customers - they may use the same amount of tokens as you do, but they have to pay the API price. From my experience the API is significantly more costly. We had one user ask for and receive usage credits on Claude, the bill the next day was to the tune of $400.

I agree with this sentiment.

Many folks have touted the "calculator" similarities as an argument, saying it's more of an efficiency gain / productivity enhancer. To me, LLMs are far more involved than this. Now, unknowingly (or knowingly), people are offloading the problem solving portion of small tasks.

- Creative Writing (Claude, make this email sound more professional)

- Coding (Handle this small logic bug for me)

- Note Taking (Generate a summary of this meeting recording)

- Strategy (Set up a roadmap for X project) and many other areas

- Design (Give me a powerpoint for a stakeholder meeting)

- Personal Life (Find a restaurant I can take my wife to for our anniversary)

Many people underestimate how many "simple" tasks required creative problem solving abilities, and we're actively handing more and more of that over to the thinking machine.

Perhaps it's human nature to give this up, and maybe it's in our best interests - but this is the first time I've ever seen people stop thinking for themselves en masse. Interesting times ahead, IMO.

I understand they put "normal" person in quotes, probably to reference the HN crowd or other computer enthusiasts - but it is very funny to think of my Uncle Rick (No smartphone, no computer knowledge), a "normal" person, coming to the family outing with microcontrollers rigged up to temperature sensors all over his body.

I saw a team build some payment software a while ago that had a similar, if not exact, vulnerability. If someone had enough time / effort to figure out a not-super-unique ID (Something like 2456733), they could acccess the payment portal for an order.

I notified them and they said that this was noted, skipped, and they didn't believe it was an issue. Worst case scenario an attacker could... Pay for someone elses order, if this happened the attacker would be found by their payment details. Likewise on the payment screen they only see the order's total, nothing about the customer, nothing else about the order, just the total. So - I'm not sure. Maybe they're right?

I just shrugged. I would've patched it, feels like poor design and is easy enough to fix - but I couldn't really argue other than to say it felt sloppy.

I agree with a lot of this. Sometimes it is easier to scan a code, zoom until the text is at the desired level, and scroll around the menu rather than opening a large leather book with tiny writing in a dimly lit restaurant.

Going to PAY with the QR code feels a little worse. I went out with some friends and I planned to pay cash, the waitress came over to tell us all that we should pay using the QR code on our receipts - missed me, and I had to wait 5 minutes to let them know I was using cash. My friends didn't have much trouble, but we're all younger folks so we are used to pulling up Apple Pay or whatever. I imagine some people do not enjoy that experience.

I will say that there is this inherent disdain towards automated systems, and I feel it's warranted - to a degree. Some experiences are improved with automation, others are stifled. Sometimes we just want to talk to other people who understand our problems.

I think a lot of developers probably FEEL like they are in super mode, but in reality they're just letting Claude drive the boat and they get to wear the captain's hat.

Maybe I'm wrong. Maybe AI Natives will be faster in the end and can build / do more, or building software really is a dead field - but I noticed that I was losing my brain and had to get back into the seat.

There are definitely great use cases for agents - but I think a lot of us aren't flexing our brains anymore and, even worse, some devs believe they are. I urge every developer to put Claude down for a day/week... see how well you can do in the "old" ways. It'll still be here when you get back, but my guess is it'll be a rude awakening.

I sometimes end up at or near this conclusion when considering the future.

Consider some software that was written by AI purely using markdown files. The spec was sufficient enough to classify all of the business logic, conditionals, etc... You might even end up creating "loops" that tell AI to do something over and over again. Some markdown files become standard, repeatable "functions" that are to be followed EXACTLY (determinism). Some markdown files become assertion tests. Heck, the markdown file might eventually invent some kind of "typing system" so that you know when you're working with the person "class", it's always going to have the same facets.

I love the concept of it going in this direction - we already had plenty of languages to tell a computer what to do, we've just generated MORE text at a higher level and made it less deterministic.

Just to clarify, I don't think it's likely that this is the end result, but it sure is funny to think about.

I bought an M4 Air about a year ago for under 1000$, it beat out my 2019 Intel MBP by quite a lot.

I fully expect the air to last me at least another 6 years or so for my use case. The thing is a beast.

Compare this to a Dell laptop I bought when I started college, that thing was 850 dollars and died on me within 3 years. For Apple, I could justify spending more (maybe even 20% more) considering both Apple computers I’ve had feel extremely fast. The only reason I dropped the 2019 MBP was battery fatigue (and I probably could have repaired it for 100$ and gotten another 3-4 years out of it. But the new air was just too attractive).

I think I’m on this side. I find it exceedingly unlikely that we just start producing “perfect” software all the time for everything, and at the same time start generating an order of magnitude MORE software.

When people finally offload 100% of their brain and forget how to use their creative reasoning abilities my guess is we’ll just all use Tailwind defaults across the board. No need to try new things, nobody will experiment because it’s so easy not to!

(Joking, mostly) but we did see this with Wordpress, Bootstrap, etc. the masses converge on simple web experiences because it’s pretty easy to get something that “just works”.

The Coming Loop 29 days ago

Man - what a ride this last year and a half has been. I feel for the juniors or newer developers, who really haven't had time to get into the seat. I don't see a great place to really "settle" into right now, as the field is unfolding rapidly. I wonder what things will look like in 10 years.

If I give an agent a sufficient spec, and it can one shot it, I imagine we won't need to loop, especially if we assume the tech is going to meaningfully improve in the coming years. In 5 years, "make no mistakes" and "add tests + review this code" will be baked into the agent or completely unnecessary, right?

Maybe I'm out of the loop.

Both companies offer "MAX" or "PRO" plans - and the best models were available to those customers. This new wave of "It's too dangerous for the public" is a new initiative from both companies.

I agree with your overall sentiment. Paying for "Claude Mini" doesn't get you "Claude Maximos".

However, the overall precedent that the companies have set is that if you pay for the top tier subscription, you get the top tier model. That's not true any more.

To me, they're selling the "power" of their product by mentioning the danger. "It's TOO powerful to even release yet!"

Whether or not that level of power exists, that is definitely how they're pitching it.