This is now 404ing. Someone at BFL reads HN.
HN user
recsv-heredoc
A deep interest in philosophy, futurism and software.
https://mindfront.ai
Hopefully they've had enough time to improve things since then.
Open SoTA hasn't dramatically advanced at the top end (efficiency gains, sure, but nothing remarkable yet) since those releases.
Why do you still have a television in your home?
very load bearing suggestion.
Having to have Xcode installed is more than half the problem. It makes Visual Studio look lightweight.
It's unlikely that someone who understands this well enough to see its value wouldn't be completely comfortable with a terminal interface.
Great article. This is exactly what we're doing from a product perspective.
What if the frontier-minus-6-months assumption does not hold? The US has 5x the AI Capex of China, and 10x the EU. Assuming AI is compute limited (we certainly seem to be given the RAM crisis) - wouldn't it be reasonable to assume frontier models are likely to continue to pull ahead?
The harness is really important. It matters so much - possibly even more than the model. We had harness crashes after running many agents - granted we were doing quite a bit with it. Grok Build (as a product) review here:
UX friction behind is worse than Claude Code - but seems to be a strange positioning choice - they're more on the 'vibe' side than the 'agentic engineering' things.
Largest issue was actually reviewing output - but if you're going to largely make that opaque from the user, why choose a CLI-based interface that's so mouse-heavy?
There's also problems with the actual model. Thinking is visible, and every interaction goes like this:
"I would like you to investigate adding an API route to tackle x,y,z" *Grok, thinking: Okay - the user has asked me to add an API route to tackle x,y,z"
Also absolutely absurd other quirks - "I have no tools available in my context" being visible in the CoT.
The auto-approval (yellow, auto-mode) review of Claude Code via Opus is a killer feature - every build-it CLI should be offering this for long horizon tasks.
Messaged one of the engineers about our experience - no feedback.
You'd be better off with Claude Code 5x Max than the 300 USD/month subscription.
The cost of compute and inference efficiency keeps dropping. Deepseek has numerous advantages here that Anthropic has not likely implemented.
If we assume the current rate of 10x reduction every year or so, profitability is inevitable. It’s just a market-share cash-fight at the top right now.
It's very hard to compete with the massively-token-subsidized big players when our entire team is spending nearly 20x the claude code subscription cost in API token usage - it's impossible for anyone else to do it without eating huge losses.
Junie - their coding agent - was also a miss. I've had Rider for almost 10 years currently - but considering dropping it. Tradcoding is basically dead, and a lightweight text editor with tree-sitter has come a long way - and it's good enough to read/micro edit with anyway.
I feel bad for them as it's been such a stable product for decades of excellent development, but the world moves on.
Maybe that's not how it's being used though - Nobody needs photoshop to solve a specific focused problem.
Photoshop is a (formerly?) great toolbox. Toolboxes are good if you need to cater to a wide audience. An audience of one via bespoke software - the real revolution - doesn't need the full photoshop experience.
Countless examples previously requiring photoshop are now replaced with some ffmpeg and imagemagick pipelines written by AI daily.
CloudFlare offers excellent service for many of the open-weights models. It's fast, cheap and simple to set up. Can highly suggest as an LLM provider.
They serve gemma-4-26b-a4b-it.
There's so much tracking on this site it's even running WebGL to try and fingerprint the browser. Is that really necessary for a joke site, or does this ship by default with every CloudFlare site?
I can highly suggest visidata! https://www.visidata.org
It's an incredibly useful piece of software for data wrangling and exploration.
LLMs are rapidly becoming the first 'purely digital commodity'.
Being digital it's somewhat hard to apply any kind of trade protectionism or Chicken Tax onto them. Maybe there's a market for cruelty-free vegan non-GMO (low-water-use sustainable energy) LLM tokens as well as European ones?
I really like what Mistral did for open Models - but what is the plan to compete against the likes of Moonshot, DeepSeek in the global market? When you can get Kimi K2.6 served via cloudflare it raises tough questions on the economics of it all.
What exactly is Mistral's strategy is aside from niche regulatory requirements or a Eurocentric hedge for AI sovereignty? Do they even have ambitions to compete on the global stage?
I tried to find them on GitHub. Doesn't seem to be FOSS/OSS either.
Does address a real use-case - might be great as a library or a lightweight alternative to Mermaid.
As SaaS it's a very hard sell.
Yes - the interesting part is the decision that the “risk of losing internal comms to a ban is worth it” - even at that size.
According to one of the founders there’s no better way for them to reach a lot of low-skill part-time employees reliably.
It shows the need to bring AI to where people already are and onto the platforms they already use.
The world’s most successful one!
They do - but the utility is so high vs the risk (for a new number) that it’s worth doing anyway for many users and even organizations.
Just yesterday we spoke with a $50-100m ARR org org using baileys for internal messaging!
This is such a sorely needed point of integration. Cool to see Peter still shipping tools. It’s such a pity meta refuses to play ball like Telegram.
Either they’ll double-down and make this even harder -or- hopefully realise that WhatsApp is likely to be a really common control plane for AI systems in the next few years. Let’s hope the Llama energy strikes and it’s the latter.
How does WhatsMeow compare with Baileys?
This startup seems to have been at it a while.
From our look into it - amazing speed, but challenges remain around time-to-first-token user experience and overall answer quality.
Can absolutely see this working if we can get the speed and accuracy up to that “good enough” position for cheaper models - or non-user facing async work.
One other question I’ve had is wondering if it’s possible to actually set a huge amount of text to diffuse as the output - using a larger body to mechanically force greater levels of reasoning. I’m sure there’s some incredibly interesting research taking place in the big labs on this.
The market timing on this is perfect - it fills a major current gap I've seen emerging.
I've heard a few stories of QA departments being near-burnout due to the increased rate developers are shipping at these days. Even we're looking for any available QA resources we can pull in here.
No harm meant with the question - but what's the advantage over Claude Code + the GitHub integrations?
Setting your upstream to 1.1.1.1 or 8.8.8.8 should mitigate this, as these appear to respond with NXDOMAIN to invalid domains rather than fail silently which would close the window for the attack.
Of course - but all things being equal, your odds will certainly be better if you give it your best shot!
I'd intended my message to be one of actual hope and optimism - that those in the position to do so may affect a small positive change in the world at the individual level. This might be myopic by ivory-tower standards, but the intention was to empower and encourage people not to give up, and that their actions do have a measurable effect on their destiny.
We might not be able to change the circumstances easily (We certainly should be trying to do so), but parallel to this we can support and encourage those going through the tough times.
Many Gen-Zs are demonstrating remarkable resilience - this post was intended as a celebration of that, not to downplay the severity of the society-wide issues at play.
Thanks for the fair critique.
As in all things, there will be winners and those who fail to try with sufficient determination :)
For those with the means to do so, try hire the Gen-Zs out there who want to succeed despite the circumstances - especially the ones skipping college. They’re some of the most capable, self-motivated people you will ever have the chance to work with!
Congratulations to the founders on this batch!
Yeah for sure - that's exactly why we use that approach - it's unsurprising, simple and definitely works.
One difference is that you don't necessarily need structured data in, just output validation from the LLM. This is a big difference from ML because you're not having to worry too much about doing complex data engineering - or at least it solves the annoying ingestion problem in many cases.
Another observation is that most businesses don't have any ML engineering capabilities in-house - they're pretty much willing to pay a premium, because unlike the bespoke ML solutions, you can just do it with an off-the-shelf system (provided it's designed with the right validation loops).
The agent is in some ways an abstraction that just enables use and adoption - even if it would be orders of magnitude worse than normal ML solutions - it's competing against no solution, not ML-based ones.
Last thing is just around what level of autonomy people expect from these things. You can go pretty far, but like flipping N coins, the more you flip the greater the chance that something goes awry. Agents still need a lot of guidance and it's up to the system builders to bring that to them, either by connecting humans or very tightly integrated, well-designed tools.
From our observations on why - you need to have an extremely tight validation loop on everything you do for AI agents to be useful. They also need a ton of highly specific instructions and context. This requires a deep understanding of the platforms and tooling or a highly standard way of working (coding).
This is why tools like cursor work so great, they’re able to work in a super tight feedback loop with the compiler, linter and tests. They operate in a super well-known, documented environment.
If we can replicate the same thing on business systems… that’s when the magic happens - just very hard to do without deep knowledge of those platforms and agentic AI because everyone does stuff differently in each org. The overlap of people with skills in both AI and specific business ops areas is absolutely tiny.
An example of where we’re using this is in a fully AI native CRM (part of SynthGrid - see https://mindfront.ai). We don’t even have any way to interface with it outside of AI, but we’d also never want to do so again anyway because the efficiency gains are so huge for us.
The Pareto frontier will continue to inexorably advance forward, dragging even the complex or non-standardized domains in with it. For those tightly integrated business systems, we’ll probably see huge gain in utility, if not function, from the improved underlying models combined with the excellent tools. Be sure to try out Claude 4 Opus hooked into some systems if you haven’t already!
For the past almost 3 years - full-stack vertically integrated business AI systems. We got a nearly perfectly timed start on this.
We’re solving the problem of “How can agentic AI interface with legacy and existing business systems.” - if you’ve got a boring job and are tired of filling out forms in business software or swapping between 10 different systems, convince management to let us come and have AI do it for you.
Nice idea! Some possible really good reading here for you: https://substack.com/@jordanbryan - YC 2021 building out “git for lawyers”