Exactly! Number of turns, average tokens to achieve a task using your CLI, as well as average number of characters being returned per CLI command alongside other metrics: all important to both users and agents! I am working on allowing to accurately capture this at www.cliwatch.com! Feel free to request an example eval suite for a list of tasks you want to achieve with your CLI
HN user
climike
Working on testing and monitoring agent-readiness of CLIs with www.cliwatch.com, some interesting challenges in building test suites and analyzing data :)
Also worth pushing for a more standardized skills command for CLIs, similar to —help, but for (agent/human) workflows, https://cliwatch.com/blog/designing-a-cli-skills-protocol (if you ship these with your CLI, you also get versioning out of the box so to say)
Resource allocation based on your hackernews upvotes? Thanks in advance folks ;)
We are working on supporting agent harnesses @ www.cliwatch.com, so both 1. LLM model as well 2. LLM model + harness performance can be evaluated against your software/CLI. We also support building evals against your doc suite. End result is that you’ll feel more comfortable shipping CLIs that work for your agentic users!:)
Building www.cliwatch.com, so you can keep an eye on how agent-friendly your CLI is ;) feel free to request a benchmark against your CLI docs. Cheers
cliwatch.com, creating some benchmarks, reach out if you are interested :)
In a similar fashion it appears that article was automated - did the author read every word in their own article?
Not sure about the end of thinking, would say that this is the start of managing ever more stochastic systems