stripping to markdown with Jina Reader or Trafilatura before passing to the agent cuts that 68k down to ~3-5k for most Wikipedia pages, and handles the JS-rendered case too.
HN user
chonghaoju
Works well alongside my Robinhood MCP server! Do you store point in time vintages or only latest values? And what's pricing after the free tier?
If you want to live longer, take a nap but less than 20 min.
The moment that grabs me is that the verdict is internally settled several tokens before "not correct" gets verbalized.
What's the per-token cost of reading the lens at nine depths?
Interesting idea. Btw, I see 灵活就业规范 in en version.
Author here. The scraper is ~300 lines of Python stdlib. Happy to answer questions about the rate limits or the states I tried and couldn't get (Delaware was the most surprising — zero bulk access at any price)
Every agent run writes an audit record. Not for compliance theater — because when something breaks at 2am, you need to know exactly what happened and why.
curious how you handle sites that fingerprint headless browsers, most of the hard work in browser automation ends up being anti-bot evasion rather than the actual task execution
Representing variables as distributions rather than point values is underrated for telecom noise modeling — most engineers just slap a variance on top after the fact and call it done.
Prompt wording that implies the agent "should already know" something tends to suppress clarifying questions and leads to silent wrong assumptions