The token-economics for closed source models are different, they are optimizing for 200 USD tokens worth of software engineer monthly usage, they will increase per token price as models or harnesses are more optimized.
HN user
leroman
https://romansky.dev
Interesting, thanks!
What model are you working with where you still get good results at 25k?
To your q, I make huge effort in making my prompts as small as possible (to get the best quality output), I go as far as removing imports from source files, writing interfaces and types to use in context instead of fat impl code, write task specific project / feature documentation.. (I automate some of these with a library I use to generate prompts from code and other files - think templating language with extra flags). And still for some tasks my prompt size reaches 10k tokens, where I find the output quality not good enough
The biggest challenge an agent will face with tasks like these is the diminishing quality in relation to the size of the input, specifically I find input of above say 10k tokens dramatically reduced quality of generated output.
This specific case worked well, I suspect, since LLMs have a LOT of previous knowledge with HTML, and saw multiple impl and parsing of HTML in the training.
Thus I suspect that in real world attempts of similar projects and any non well domain will fail miserably.
Cool idea! but kind of wasteful.. I just feel wrong if I waste energy.. At least you could first turn it into markdown with a library that preserves semantic web structures (I authored this- https://github.com/romansky/dom-to-semantic-markdown) saving many tokens = much less energy used..
The title was so confusing to me, the reason I opened the link was to understand how you made the SSH tunnel manager learn the GO programming language
It's hilarious they put Claude 3.5 Sonnet in the far right corner while it scores the highest and beats most of Grok's numbers.
Thanks for sharing!!
Would be really helpful if you opened an issue in Github with a specific example, happy to look into that!
This is now supported, see here- https://github.com/romansky/dom-to-semantic-markdown?tab=rea...
Bumped this together with the side-by-side comparison task.. so will look into it :)
This is some great feedback, thanks!
1. there some crazy links with lots of arguments and tracking stuff in them, so it gets very long, the refification turns them into a numbered "ref[n]" scheme, where you also get a map of ref[n]->url to do reverse translation.. it really saves a lot, in my experience. It's also optional, so you can be mindful when you want to use this feature..
2. I tried to keep it domain specific (not to reinvent HTML...) so mostly Markdown components and some flexibility to add HTML elements (img, footer etc).
3. Not sure I'm sold with replacing the switch, it's very useful there because of the many fall through cases.. I find it maintainable but if you point me to some specific issue there it would help
4. There are some built in functions to traverse and modify the AST. It is just JSON in the end of the day so you could leverage the types and write your own logic to parse it, as long as it conforms to the format you can always serialize it, as you mentioned..
5. The AST is recursive so not flat.. sounds like you want to either write your own AST->Semantic-Markdown implementation or plug into the existing one so I'll this in mind in the future
6. Sounds cool but out of scope at the moment :)
7. This feature would serve to help with scraping and kind of point the LLM to some element? Then the part I'm missing is how you would code this in advance.. There could be some meta-data tag you could add and it would be taken through the pipeline and added on the other side to the generated elements in some way..
Ah, I suppose you mean a web page one could visit to see a demo :) Added to the backlog!
This totally makes sense, I will look into adding support for additional ways to detect the main content, super interesting!
By all means, you can be the first contributor :) You are welcome to either open an issue and brain storm together on possible approaches or send me a pull request with what you came up with and we start there
After removing the noise you can distill the semantic stuff where ever possible, like meta-deta from images, buttons, etc, and see some structures emerge like footers and nav and body.. And many times for the sake of SEO and accessibility, websites do adopt quite a bit of semantic HTML elements and annotations in respective tags..
Afraid to say that other than bumping into a talk about Deno, I haven’t played around with it yet.. So thanks for the heads up, will look into it.
Thanks for the bug report !
Please see here- https://github.com/romansky/dom-to-semantic-markdown/blob/ma...
Thank you! this is exactly why there's support for this specific use case- https://github.com/romansky/dom-to-semantic-markdown/blob/ma... (see `findContentByScoring`)
And if you pass an optional flag `extractMainContent` it will use some heuristics to find the main content container if there is no such tag..
You might find this useful- just added code & instructions on how to make it a global CLI utility- https://github.com/romansky/dom-to-semantic-markdown/blob/ma...
Thanks for sharing, will look into adding this as a flag in the options!
Markdown being a very minimal Markup language has no need for much of the structural and presentational stuff (CSS, structural HTML), HTML has many many artifacts which are a huge bloat and give no semantic value IMO.. It's the goal here to capture any markup with semantic value, if you have examples this library might miss, you are welcome to share and I will look into it!
Will add some side-by-side comparisons soon! the goal is not just to translate 1:1 HTML to markdown but to preserve any semantic information, this is generally not the goal for these tools. Some specific features and examples are in the README, like URL minification and optional main section detection and extraction (ignoring footer / header stuff).
Author here- it's a good point to have some benchmarks (which I don't have..) but I think it's well understood that minimizing noise by reducing tokens will improve the quality of the answer. And I think by now LLMs are well versed in Markdown, as it's the preferred markup language used when generating responses
Mostly look for an operator for that
By “eco system” I mean all the charts that get shared, when ever I try to look under the hood I instantly regret it..
Kustomize is basically a higher level file for K8s deployments, I have all the resources as declarative code that gets deployed when I apply the relevant directory. I have istio + ssl certs + services and any other resource, multiple projects with cross project communication and provisioning etc..
Somehow I am able to get bye with Kustimise. Cant stand the mess that is helm and its eco system.
No idea if its a valid approach but possibly train with a hidden layer containing a “role”?
Thanks to investing into k8s I was able to migrate a non trivial production with minimal downtime between cloud providers 3 times now, who offered us free credits, without too much friction.
GPT-4 is already hugely useful. If they are able to lower the cost of GPT-4 further, say 5x or 10x that in itself would be useful and huge.