why can't this be a cli tool? then you can get an agent to write a script that programmatically calls the cli tool in addition to the agent calling it directly.
HN user
localhost
I work on making Python more awesome across the company, now including Excel!
My alias is jflam (at the company I work for).
there is an activation energy cost to so many activities - so those things just never got done. many times it is because the cost-benefit wasn't clear at the start (unknown unknowns) so it never got done. kudos to op for experimenting and showing us one way of making something like this happen.
then have ai write a deterministic transformation tool that turns it into the specific style and standard that is needed
i had gpt-5.4 write a 6502 font version of this font [1] in between working on other things. super fun and learned a lot by watching it do its thing.
tbf neither does anthropic
This is what ACP and https://github.com/zed-industries/claude-code-acp enables. ACP controls agents - there is native support in Copilot CLI and Gemini and adapters for claude code and codex.
Around 1980, while taking a "Saturday Morning Class" in Toronto - I discovered that there was a lab of ~24 Commodore PET 2001 (8K - blue phosphor, chiclet keyboards) at George Brown College. Spent as much time as I could there engaging with the early hacker community who all brought their shoeboxes of 5-1/4" floppies to trade programs. It was there that I had my first OMG moment when a much older kid showed me his floppy disk catalog program that could sort so much faster than mine did (he used quicksort).
First thank you for coming up with this idea. Love it. In the spirit of "tools for thought" I think this scratches an important itch.
Someone called the output "egregious, obviously referential LLM dog". And someone else called this trail an example of that: https://trails.pieterma.es/trail/tempo-gradient/
What struck me when reading it was that I parachuted into that trail with zero context. So I asked an LLM to generate some surrounding context for me: this is the verbatim prompt that I just gave gpt-5.2/medium running in my own agentic harness:
read this: https://trails.pieterma.es/trail/tempo-gradient/. then i want you to write a short essay that shows the relationship between the different quotes making sure to use each quote in the process. the narrative should help me understand the links between the different quotes in each book in context. feel free to use your own knowledge of the books and your general knowledge to make this an approachable thing. what another llm did here is cool - it found relationships between this set of quotes. now i want you to materialize them into a compelling essay that invites curiosity and teaches me something. write the essay to tempo-gradient.md and respond in chat with how you thought about the problem and whether you thought these relationships were meaningful.
This was my agent's response:
Wrote the essay to tempo-gradient.md.
How I thought about it: I treated “tempo gradient” as the shared mechanism behind all eight excerpts—mismatch in loop speed (notice → interpret → decide → act) caused by attention, infrastructure latency, organizational rigidity, system delays, stealthy manipulation, or runaway acceleration. Then I ordered the quotes from “inner time” (tennis/attention) → engineered milliseconds (HFT) → explicit doctrine (OODA) → industrial schedules (mobilization) → structural delays (Beer Game/economy) → slow covert tempo (cyber) → extreme fast tempo (AI takeoff), so each quote sets up the next.
Whether the relationships felt meaningful: yes, mostly. They’re not “the same idea” so much as the same shape appearing in different contexts—feedback delay and decision cadence determining who gets to set the terms of reality for everyone else. The only relationship that feels a bit more interpretive than direct is tennis ↔ geopolitics, but even there the bridge (“how many nows” as a controllable input to decision tempo) is surprisingly sturdy.
And this is the output: https://gist.github.com/jflam/49753b7da64a74f07e35f6e2465f1a...
Totally agree with this. I have seen many cases where a dumber model gets trapped in a local minima and burns a ton of tokens to escape from it (sometimes unsuccessfully). In a toy example (30 minute agentic coding session - create a markdown -> html compiler using a subset of commonmark test suite to hill climb on), dumber models would cost $18 (at retail token prices) to complete the task. Smarter models would see the trap and take only $3 to complete the task. YMMV.
Much better to look at cost per task - and good to see some benchmarks reporting this now.
I would say that "don't let perfect be the enemy of the good" here. Would universal be better? Sure. But what I saw is so much better than what we currently have here in the US.
The point is that OPEN (the name of the Delft library) is really a community center and not a library. Yes, it happens to have books. But it also has a stage for musical performances, art rooms, tables, wifi, washrooms, coffee. I would say that the only thing that is missing is a gym; there are small dance rooms in there but that's not quite the same.
But the essence here is walkable communities. Suburbs and exurbs are hostile to even small local stores because you have to drive everywhere to do anything. There is no community in visiting my Costco or even my QFC.
Take a look for yourself: https://www.opendelft.info
But why do social hubs need to be places of financial transactions?
I was in Delft recently and I really loved their library/community center. Full of music practice rooms, people playing board games on the ground floor, a coffee bar and it was full of people at 8pm. It is open from 9am - 11pm M-F.
You walk or cycle there (free indoor bicycle parking). There is a movie theater across the "street" (no cars).
one thing that i find works really well is to ask it to research things in the codebase and write a plan first. codex with gpt-5 is exceedingly good at doing this. then ask it to write a plan for what it would do with that information, i.e., i want you to research codebase for <goal>. then write a plan for how you would achieve <goal> given what you have learned.
even with t=0 they are stochastic. e.g., non associative nature of floating point operations
I used to own COM at Microsoft; I think that MCP is the current re-instantiation of the ideas from COM and that English is now the new scripting language.
The models are the same, but the actual prompts sent to the model are likely somewhat different because of the agentic loop - so I would imagine (without having done the experiments) there will be slight differences. Unclear whether they will be more or less than the variance in responses sent multiple times to the same experience (e.g., Claude.ai variance vs. Claude Code variance vs. variance between Claude.ai and Claude Code). Would be an interesting controlled experiment to try!
Right. But you can copy paste that into a separate doc and have Claude Code merge it in (and not a literal merge - a semantic merge "integrate relevant parts of this research into this doc"). This is super powerful - try it!
Did you use Claude Code to write the post? I'm finding that I'm using it for 100% of my own writing because agentic editing of markdown files is so good (and miles better than what you get with claude.ai artifacts or chatgpt.com canvas). This is how you can do things like merge deep research or other files into the doc that you are writing.
You can listen to him talk about it!
Python is the most popular language for data analysis with a rich ecosystem of existing libraries for that task.
Incidentally I've worked on many products in the past, and I've never seen anything that approaches the level of product-market-fit that this feature has.
Also, this is the work of many people at the company. To them go the real credit of shipping and getting it out the door to customers.
As a counterpoint to a lot of the speculation on this thread, if you're interested in learning more about how and why we designed Python in Excel, I wrote up a doc (that is quite old but captures the core design quite well) here [1]. Disclosure: I was a founding member of the design team for the feature.
[1] https://notes.iunknown.com/python-in-excel/Book+of+Python+in...
The second point makes sense. It gives Apple optionality to cut off the external LLMs at a later date if they want to. I wonder what % of requests will be handled by the private cloud models vs. local. I would imagine TTS and ASR is local for latency reasons. Natural language classifiers would certainly run on-device. I wonder if summarization and rewriting will though - those are more complex and definitely benefit from larger models.
It seems like this is an orchestration layer that runs on Apple Silicon, given that ChatGPT integration looks like an API call from that. It's not clear to me what is being computed on the "private cloud compute"?
I bought my shareware copy of DOOM from WBB, alongside many other books.
This is a giant dataset of 536GB of embeddings. I wonder how much compression is possible by training or fine-tuning a transformer model directly using these embeddings, i.e., no tokenization/decoding steps? Could a 7B or 14B model "memorize" Wikipedia?
How large is the set of binaries needed to do this training job? The current pytorch + CUDA ecosystem is so incredibly gigantic and manipulating those container images is painful because they are so large. I was hopeful that this would be the beginnings of a much smaller training/fine-tuning stack?
I wonder if enhanced operation of the lymphatic system might be causal to better mental health outcomes? The lymphatic system doesn't have a "pump", so it relies on muscle contraction to drive circulation. So more movement/exercise drives more lymphatic activity which may lead to better outcomes in people, especially if mental disorders are correlated with buildup of waste materials in the brain.
Then don't copy and paste your copy of Thinking Fast and Slow into your AI along with my prompt then?
This is exactly the kind of task that I want to deploy a long context window model on: "rewrite Thinking Fast and Slow taking into account the current state of research. Oh, and do it in the voice, style and structure of Tim Urban complete with crappy stick figure drawings."
That is, of course, on the roadmap. I use a Mac as my daily driver so this one's personal :)
For your use case, what if you could use Python in Excel with pandas, numpy and the Anaconda ecosystem readily available? How would that change things?
Disclosure: I was a founding member of the Python in Excel team and am looking for new problems that Python in Excel could solve.