I had exactly this use case - in a grocery store in the Alps, no internet, fired up a local LLM on my phone to figure out what to cook and what to buy
HN user
timfsu
https://timsu.org hello@timsu.com
Wow, this is pretty scary. LLMs have made phishing attempts look so much more legit, and the damage they can do so much greater.
This is big, but until we have policy clarity we can’t trust it. I’ve always migrated all of our agents to Pi SDK, we aren’t going back
In contrast, I’m on the $200 max plan for Codex and I hit the 5 hour limit near daily at work. I typically am having it work on about 5 tasks an hour. I’ve never hit a 5 hour limit on Claude on the $200 plan but I have hit my weekly limit.
Yeah it’s hard to call that cheating from a model. Maybe “disqualifying” is more accurate
We saw this too with Gemini specifically. My favorite example - we built a hallucination detector (given the input, does the output make any false claims) in Gemini, and after the Seahawks won the Superbowl in February, it would consistently flag that as "not possible".
Imagine you try two products you’ve never heard of. You prefer one over the other. Was it marketing? That’s what’s happening here. Marketing can get you to try something you wouldn’t have otherwise, and it may suggest benefits you’d get if you tried it, but your preference of using one thing or the other is a subjective experience of your own.
This is dope! We basically built something very similar internally for our team and it's been a very natural and intuitive way to manage agents (as opposed to having a bunch of terminals to track). Not every task/conversation can be done in the background, so it's been helpful for us internally to be able to seamlessly transition between "interactive conversation" and "background job done by agents" even within a single card.
Yes, but they expire in ways that unions don’t
PSA - you can run something like `npm install -g npm@11.10.0; npm config set min-release-age=3` to update to a version of npm that supports the min-release-age configuration
Pnpm - installs are faster to boot. We haven’t missed anything
Did not know this was a thing, kudos to her for speaking out!
This is a really neat bridge between “looks cool” and “feels like you’re there”. Inferring real life properties like lighting is a cool trick and just the beginning I’m sure. I’m excited to explore new and dynamic worlds and bring the AAA experience closer to something you can build yourself.
I might call it a few different things, but spyware seems disingenuous until we learn that it’s actually spying…
The question is - if the SOTA model disappear - do these follow-on models have the ability to improve themselves without distillation?
I for one enjoyed this very long essay. It should've been a lot shorter, but you also didn't have to read it, it says right there in the title :)
Love this idea. Working with AI assistants, I find it easier to push to GitHub to look at the changes, rather than use my IDE. I wish that wasn’t the case, so this makes a ton of sense.
Fascinating article. Also extremely confusing (though probably not unexpected) that an important health researcher is named Nestle.
I for one would love this - if it’s done well - except that it would presumably be locked in to OpenAI agents
These narratives are so strange to me. It's not at all obvious why the arrival of AGI leads to human extinction or increasing our lifespan by thousands of years. Still, I like this line of thinking from this paper better than the doomer take.
I get the appeal, but it seems too early for one AI tool to be able to do "everything". I'm guessing there's some company out there trying to automate each of these tasks and dealing with the attendant complexity that comes with it - it's not clear to me that a single "do everything" AI would be able to do this pre-AGI
In my opinion, what you need is a person (or three), not a book :)
Someone who relies on you, whatever the context, is some of the greatest motivation out there.
Possibly best thing ever on Hacker News. There is something quite appealing about the simplicity of Boléro
It's not clear how much ChatGPT is investing in the discovery part of the app store experience, so this seems like mostly a way for users to install apps they're already familiar with and use them from inside a chat. For now, it seems like you have to explicitly @-mention an app to use it.
Unfortunately, I tried to use zed as my daily driver, but the typescript experience was subpar. While the editor itself was snappy, LSP actions like "jump to declaration" were incredibly slow on our codebase compared to VS Code / Cursor.
Really interesting article, unfortunately it seems quite lucrative to scam seniors, and very hard to prevent, especially when they don’t think they’re being scammed.
I understood it to be the reverse - they advertise on LinkedIn, and the trackers determine whether the users convert once they click through. Not great, but at least not as ill intentioned
Happy mrge user here - congrats on the launch! It’s encouraged our team to do more stacked PRs and made every review a bit nicer
Using a time-based expiration rather than a usage-based expiration should help
I appreciate the author’s openness in sharing their experience - it’s really worthwhile to share experiences where money isn’t everything, and that it can be a poor generator of meaning.
Speaking of meaning, I think it’s a task for everyone to find their life’s calling - something you’re uniquely suited to do. Sometimes that pays the bills and sometimes it doesn’t. I’m a bit surprised that “start another saas company” wasn’t really on the list for Vinay, but there’s probably a good reason for that. For me, I found that starting a family completely changed my life, as well as helping me appreciate more the family I already had - but I suppose that’s not something that lasts forever either. Good luck to the author.