An agent is an llm running on a loop until it decides it should stop.
/goal is a gimmick where you run a "parent" agent on top that runs the agent on a loop until the it decides to stop, just prompting it "nope, not done yet, continue".
HN user
An agent is an llm running on a loop until it decides it should stop.
/goal is a gimmick where you run a "parent" agent on top that runs the agent on a loop until the it decides to stop, just prompting it "nope, not done yet, continue".
both amazon and uber used that spending to deliver a network effect moat/almost monopoly.
But openai's chance of a moat on model quality is dropping as we go, not increasing
lol do you not have actual stuff to prioritize for your team rather than re inventing forges? there are quite a few open source alternatives if you want something quick. Fork them. No need to re invent the wheel
compute??
Github is struggling because of compute, which comes from everyone vibecoding and triggering actions 10x more.
I can vibecode an alternative, but once i have users, who is going to secure this amount of compute? Compute+talent to manage it(devops isnt vibecoded *yet) is a moat
Its quite funny that you are building an "infinite place profile", you both worked on products used by 100s of millions of people, and yet your website is down from 45 minutes of HN traffic!
Joking, but its a very good idea. Synchronization between the physical world information and digital has been a very hard problem for decades and im sure an agentic approach can 10x the value.
Just wanted to say i think most interpretability research it's just a smoke show nowadays but this is actually the first one that i think has a very serious potential. I love that the SAE is actually constrained and not just slapped unsupervised posthoc.
How granular can you get the source data attribution? Down to individual let's say Wikipedia topics? Probably not urls?
Would be interested to see this scale to 30/70b
You can edit it by describing changes
Even this is hard. Most people don't know what they want, and/or they don't know how to describe it/imagine it. They don't even know what a trend graph is.
They just want someone else to do the mental effort of creating a nice product. Hence iOS > android for most people. They don't want to customise basically anything other than colours.
That's why i predict Lovable/replit etc will not go mainstream. And why chatgpt will just offer you their UIs mainly. Artifacts weren't a big hit
The funny thing is that Anthropic is the only lab without an open source model
Mintlify is the best example of a product that is just nice. They don't claim to have a moat, or weird agi vibes, or whatever. It just works and it's pretty. 10m arr right there
Yes 250 gallons can save a lot of people from dying of thirst but have you considered that with that same 5kwh we can also produce 1 tiktok ai slop video of the queen boxing with mike tyson?
Never ask a woman her age, a man his salary, and an agent framework developer his long term plans
I wonder if using a local llm to override the ads would work. A finetuned one for removing ads will probably appear soon
Codex is more hands off, I personally prefer that over claude's more hands-on approach
Agree, and it's a nice reflection of the individual companie's goals. OpenAI is about AGI, and they have insane pressure from investors to show that that is still the goal, hence codex when works they could say look it worked for 5 hours! Discarding that 90% of the time it's just pure trash.
While Anthropic/Boris is more about value now, more grounded/realistic, providing more consistent hence trustable/intuitive experience that you can steer. (Even if Dario says the opposite). The ceiling/best case scenario of a claude code session is a bit lower than Codex maybe, but less variance.
which does pose interesting questions over nvidia's throne...
Zebra-Llama is a family of hybrid large language models (LLMs) proposed by AMD that...
Hmmm
It's definitely cool and engineering wise close to SOTA given lovable and all of the app generators.
But, assuming you are trying to be in between lovable and google, how are you not going to be steamrolled by google or perplexity etc the moment you get solid traction? Like, if your insight for v3 was that the model should make its own tools, so even less hardcoded, then i just dont see a moat or any vertical direction. What really is the difference?
All VC's have preferred shares, meaning in case of liquation like now, they get their investment back, and then the remainder gets shared.
Additionally, depending on round, they also have multiples, like 2x meaning they get at least 2x their investment before anyone else gets anything
because the secret is that the web runs on advertising/targeted recommendations. Brezos(tm) wants you to actively browse Ramazon so he can harvest your data, search patterns etc. Amazon and most sites like that are very not crawl friendly for this reason. Why would Brezos let Saltman get all the juicy preference data?
On the foundational level, test time compute(reasoning), heavy RL post training, 1M+ plus context length etc.
On the application layer, connecting with sandboxes/VM's is one of the biggest shifts. (Cloudfares codemode etc). Giving an llm a sandbox unlocks on the fly computation, calculations, RPA, anything really.
MCP's, or rather standardized function calling is another one.
Also, local llm's are becoming almost viable because of better and better distillation, relying on quick web search for facts etc.
A very obvious AI review with 80 points(?) plus a couple of more comments. Discussion also here https://old.reddit.com/r/MachineLearning/comments/1oyce03/d_...
Their page itself looks classic v0/ai generated, that yellow/orange warning box, plus the general shadows/borders screams LLM slop etc. Is it too hard these days to spend 30 minutes to think about UI/user experience?
I actually like the idea, not sure about monetization.
It also requires access to all the data?? And it's not even open source.
They introduced pay as you go recently. The limits on that is similar to the plans, 1 million tokens per minute, so if you stack a few keys and do a simple load balancing with redis, can cover a decent amount of traffic with no upfront cost. Eventually we would have to go enterprise though yes!
Great feedback thanks! We have added a synthetic e-commerce dataset as an example when you sign up so you can test it without your data first. Will also add a demo video ASAP.
I'm working on Flavia, an ultra-low latency voice AI data analyst that can join your meetings. You can throw in data(csv's, postgres db's, bigquery, posthog analytics for now) and you just talk and ask questions. Using cerebras(2000 tokens per second) and very low latency sandboxes on the fly, you can get back charts/tables/analysis in under 1 second. (excluding time of the actual SQL query if you are doing bigquery).
She can also join your google meet or teams meetings, share her screen and then everyone in the meeting can ask questions and see live results. Currently being used by product managers and executives for mainly analytics and data science use cases.
We plan to open-source it soon if there is demand. Very fast voice+actions is the future imo
Breaking news: For profit company chases profit, briefly pretends it's not while it is
This has nothing to do with what i said. I said they are addicted. The free limits are designed this way. If openai suddenly removed the free plan, i guarantee you a lot of people would buy. They dont have an alternative they cannot think independently anymore
Everyone. People are becoming dependent on chatgpt. They literally cannot function professionally or even socially without it. They will pay their last 20-30 dollars if needed. It's literally like a drug especially when it's asking you if you want to followup/continue.
If my mom gives me 1000 dollars for 1% of my lemonade stand, that doesn't mean my stand is worth 100k. Tether is in talks with investors to mayb raise 20b at a 500b valuation. Keep in mind also that crypto investors overvalue companies to create the hype and then lobby for better regulations etc. It doesn't mean at all that someone would be interested to buy 100% of tether for 500b. Now, if they were public is a different story, like Tesla etc
The whole point of embeddings and tokens are that they are a compressed version of text, a lower dimensionality. now, how low depends on performance, lower amount of vectors=more lossy (usually). https://huggingface.co/spaces/mteb/leaderboard
You can train your own with very very compressed, i mean you could even go down to each token=just 2 float numbers. It will train, but it will be terrible, because it can essentially only capture distance.
Prompting a good LLM to summarize the context is probably funnily enough the best way of actually "compressing" context
Context editing is interesting because most agents work on the assumption that KV cache is the most important thing to optimise and are very hesitant to remove parts of the context during work. It also sometimes introduces hallucinations, because parts of the context are with the assumption that eg tool results are there, but theyre not. Example Manus [0]. Eg, read file A, make changes on A. Then prompt on some more changes. If you now remove the "read file A" tool results, not only you break the cache, but in my own agent implementations(on gpt 5 at least) can hallucinate now since my prompt etc all naturally point to the content of the tool still beeing there.
Plus, the model got trained and RLed with a continuous context, except if they now tune it with messing with the context as well.
https://manus.im/blog/Context-Engineering-for-AI-Agents-Less...
Eh i mean often innovation is made just by letting a lot of fragmented, small teams of cracked nerds trying out stuff. It's way too early in the game. I mean, qwens release statements have anime etc. IBM, Bell, Google, Dell, many did it similarly, letting small focused teams having many attempts at cracking the same problem. All modern quant firms are doing basically the same as well. Anthropic is actually an exception, more like Apple.