Have you used Inkling enough to be able to tell? Or how can you be sure? Please add some substance before making such claims
HN user
juliangoetze
I thought HN was different. And yeah, wherever I go, my timeline is full of Opus is so bad today and I will switch from Fable to 5.6 Sol, it's 1.5x better and vice versa.
Non-public benchmarks (ideally suited to one's own use case) are probably the best way to judge, I agree.
This supposedly is better than KimiK2.7
How can you tell?
I just looked at the benchmarks and was kinda disappointed that it seems to be between KimiK2.6 and KimiK2.7 on most of the benchmarks.
Do you refer to what it feels like to use the model? Or are there other benchmarks I haven't seen?
I am very curious about how the "threat" of local inference and open-weights-models will change the trajectory of the model labs. The air will become pretty thin
I noticed that Claude's reasoning summaries show a phenomenon when I'm using the model in German (on claude.ai) - it mixes English grammar with German vocabulary!
"Evaluating Spülenposition gegen Wasseranschlussabstand" and "Analysierend die Platzierungskonflikte und Rohrleitungszwänge klären" are examples of generated summaries.
I wish someone could explain this - because I see no way that sentences like these would show up in training data. But maybe the RL for thinking has some quirks?
great point, we haven't really spent too much thoughts about that.. changing the logos to companies we know well makes more sense, we'll do that to save us trouble
tbh, we just looked at our users/ stargazers and picked the most interesting companies - and were cautious with the wording ("embraced by engineers")
so, i guess YOLOing it is the best way to put it
that gave the best results so far - tweaking the system prompt to make sure the agent respects the project's ui system, is aware of dark mode and responsiveness, etc.
We started with a setup that was quite GitHub-friendly (we set up https://github.com/changesets/changesets and added a small contribution guide).
People then found the courage to contribute themselves. We haven't really pushed for contribution so far!
Good catch - we could definitely put more effort into creating demos
It has VanillaJS out of the box! You can run `npx stagewise` and it will integrate into every web app you can imagine
Haha, u're right, I recently had a chat w a friend about how high our threshold for excitement has become. A few years ago, a new iPhone release would have meant the world
It's so funny how many people tell us that - it's one of the obvious ideas someone just needed to start with. What kept you from building the first version?
Have you tried the stagewise agent so far? If so, what are the core problems the agent still has?
Nice to hear that! What exactly do you imagine when talking about design controls?
Thanks! Would love to hear product feedback since there's still a ton we need to improve
Thanks! There are also some other tools out there that give 'vision' to your coding agent via MCP, but we figured that prompting the agent directly by clicking on elements is the most intuitive way to interact with it.
The whole 'vibe-coding' space seems a bit underserved for mobile or am I wrong?
It will work as long as the project is a web project and runs in the browser. The agent is a pure js snippet and can be injected into any web app
Yep, that's correct. stagewise only works with web applications.
Will do so!
Do you mean for UI-related changes that also include modifying the backend, or do you mean pure backend changes that don't involve UI at all?
Sounds like the experiences we've had in the industry - even though they're all 'logistics companies', their workflows and requirements differ greatly, even within a niche of a niche. This gave us a hard time building a scalable software that would be useful for as many companies in the industry as possible
Oh okay, seems like there was another issue then, I'll debug. Happy to see a run without issues later
What were you building? And did it evolve/ has been used by someone?
Haha, cool to hear!
We were tinkering with a transport management system because the industry was full of legacy software.
But before we were able to battle-test something, stagewise took off.
Still think that it's an industry worth going for (at least in Germany)
Hmm, the initial credits are usually enough to do 4-5 major tasks and a few more smaller ones. Was the log with 1.86/2 credits the latest?
Did the agent work well after the mail came through?
The Sign In option also serves as Sign Up.. I'm having a look at the email issue now!
I thought I was the only one.