A while back, I saw a similar feature land in Codex (I'm using the VS Code plugin) but it got removed quickly. What is the chance that an LLM recommended this same idea to the Claude PM or lead? I see LLMs across different providers converging on similar ideas or biases frequently.
HN user
m3h
Hello there
Chief Product Architect at APIMatic... I help API providers create the best Developer Experience for their APIs through awesome client libraries and documentation.
I spend a lot of time thinking about HTTP APIs and Code Generation.
Email at mehdi [at] apimatic [dot] io.
LinkedIn: https://www.linkedin.com/in/mehdi-jaffery/
- Mehdi Raza Jaffery
Kimi K3 is Kimi’s most capable model to date, with 2.8 trillion parameters.
This puts them on the top of the largest open models list:
Kimi K3 2.8T
DeepSeek-V4-Pro 1.6T (49B active)
Kimi K2.6 ~1T (32B active)
GLM-5.2 754B (40B active)
DeepSeek-V3.2 685B
Mistral Large 3 675B
That's one mighty large model! Moonshot is going to need the USD 500 million reportedly raised earlier this year to run this model.The author shared their experience building the first version in a month: https://themackabu.dev/blog/js-in-one-month
And then the follow up few months later: https://themackabu.dev/blog/ant-part-two
I'm not sure what the economics of building a new runtime and ecosystem from scratch are but it seems we're already in a phase where individual developers are creating software which previously took a whole team. And its only getting started...
We have an official pelican on a bicycle from the OpenAI livestream:
The speed up numbers based on their testing:
Codebase | TypeScript 6 | TypeScript 7 | Speedup
------------|--------------|--------------|--------
vscode | 125.7s | 10.6s | 11.9x
sentry | 139.8s | 15.7s | 8.9x
bluesky | 24.3s | 2.8s | 8.7x
playwright | 12.8s | 1.47s | 8.7x
tldraw | 11.2s | 1.46s | 7.7x
Congratulations to the team for pulling off this feat while doing a responsible migration (looking at you, Bun).Quick question: How does this affect downstream tools like tsdown and esbuild, which need to build the TypeScript codebase? Can I use TS 7 and current tsdown together?
When I reviewed the conversations affected by this issue, they did not always align with my feeling of "degraded output".
Some were definitely below par, and I recall having to iterate on the generated code more than I wanted to. However, it is only true for a very small number of conversations.
So we're looking at a small set of affected conversations, and even within that small set, only a few will have degraded output, likely because the model can compensate for the reasoning defect over the long conversation.
Indeed, it looks like my work has suffered from the clustering issue as well:
reasoning_output_tokens count percent
━━━━━━━━━━━━━━━━━━━━━━━━━ ━━━━━━━ ━━━━━━━━━
0 873 28.5948
───────────────────────── ─────── ─────────
8 64 2.0963
───────────────────────── ─────── ─────────
9 60 1.9653
───────────────────────── ─────── ─────────
11 54 1.7688
───────────────────────── ─────── ─────────
516 48 1.5722
───────────────────────── ─────── ─────────
12 45 1.4740
───────────────────────── ─────── ─────────
10 43 1.4085
───────────────────────── ─────── ─────────
17 40 1.3102
───────────────────────── ─────── ─────────
13 38 1.2447
───────────────────────── ─────── ─────────
14 36 1.1792
Created a script for this: https://github.com/thehappybug/codex-reasoning-token-checkAlso, kudos to the Z.ai team for adding Linux support from day one.
Z.ai documents integrations with nearly all the popular CLI-based agents: https://docs.z.ai/devpack/tool/others
If you're already used to your TUI coding agent, you don't need the desktop agent. Although it is nice that it is there for folks who prefer the Codex App/Claude App UI approach.
Correct. Albeit the nuance here is that a more capable model might solve problems more efficiently and faster, possibly saving you tokens.
As with any new model, you won't know the real impact until you start using it for your workload.
I didn't realize GPT 5.3 Codex was that good.
OpenAI claims to have made their new Terra model as good as GPT 5.5, but with half the cost per intelligence. Hopefully, this will bring it closer to the price you're expecting (or even better considering GPT models have good acceptance/success rates according to benchmarks).
I think you should try an OpenAI model like GPT 5.5. It is better at following instructions and boundaries set during prompt. It feels like a more capable "agent assistant" than Claude models but without loss of intelligence.
Most of my work involves "Agentic engineering" instead of fire-and-forget. I like to stay involved during the planning as well as review and ask a lot more questions from the agent than I've seen others doing. In a way, I'm using the agent in a sort of "hyper auto-complete" mode to fill in the blanks (rather big blanks) once I've set out the requirements, scope and design (sometimes specific module boundaries). This works best for me.
Important to note: "Sonnet 5 is an upgrade to Sonnet 4.6, but it uses an updated tokenizer that changes how the model processes text to improve performance (this is similar to the tokenizer change we introduced with Claude Opus 4.7). The tradeoff is that the same input can map to more tokens: roughly 1.0–1.35× depending on the content type. The introductory pricing is set so that the transition to Sonnet 5 is roughly cost-neutral."
Why is Claude Sonnet 5 allowed to be released but OpenAI Terra not? Are they not the same class of models?
If GPT-5.6 preview is not available outside US government approved "trusted partners", I don't see how the General Available can be trusted later.
Who knows what they will fix, block or change in the model between the preview and GA time. Open models can't arrive soon enough.
Or we could simply hallucinate that the packages are there at the three houses.
Hallucinations all the way down...
How are you measuring progress at your company?
Do you feel AI agents are helping you achieve company goals faster now?
Might I recommend trying APIMatic out: https://migrate-from-stainless.apimatic.io/
Helpful link: https://migrate-from-stainless.apimatic.io/
On a side note, for those impacted by the recent Stainless API wind down, APIMatic is offering 50% off for the first year: https://migrate-from-stainless.apimatic.io/
Congratulations to the Stainless team for their hardwork.
We are offering a 50% off for the first year subscription price at www.apimatic.io for companies impacted by this.
If you're looking for a solid long term SDK and docs partner, APIMatic is the OG CodeGen serving companies like PayPal, Maxio and PayQuicker for the past 10 years.
Reach out to mehdi@apimatic.io and I'll help you migrate.
PS: sorry for the shameless plug but sdks and APIs are my life and blood :-)
Why do major LLMs block china? Isn't that a potentially huge market for them?
Same in Pakistan: https://www.dawn.com/news/1924573
After COVID, grid electricity became hugely expensive, but the pushback was massive and unexpected, as people transitioned from a fixed supply to a hybrid online or offline (battery-powered) system.
Perplexity - for search (Google replacement), summarization and rewriting, basic research and making presentations (using Perplexity apps)
Granola - transcription and meeting notes, searching across notes, recalling action items
I've played around extensively with ChatGPT, but Perplexity now covers my use cases. I'm looking to test Claude, primarily because Perplexity does not currently support MCP servers, and I need an assistant who can answer questions across all my work files (Google Drives, Calendar, Slack messages, GitHub, etc.).
Then, assuming that the fake job posters are still there and that they now have AI help as well, the fall in the number of postings is more concerning.
Are they inviting you to apply?
And what happens if you do?
Yes, when taken in this context, the Monopoly argument makes sense:
The DOJ said Visa imposed “exclusionary” agreements on partners and smothered upstart firms.
How does an external authorisation service work without the knowledge contained in the application’s database? And vice versa, how does the application make the efficient correct queries from its database when the authorisation information has been externalised?
Are there any solutions for one-man SaaS to handle payments from enterprise customers? I'm assuming that the preferred mode of payment here is through wire transfers after some kind of PO process.
SST[1] looks pretty cool. Does it replicate the entire infrastructure Vercel attempts to provide you regarding hosting (such as CDN, caching, etc)?
[1]. https://sst.dev/