It doesn't embed images, no. But that's a great idea for the roadmap!
HN user
tompec
thomas.io
Having looked at a lot of HTMLs, I noticed that sections are not really the default. I rely on headings (h1, h2, ...) to chunk each pages. Each chunk has its heading hierarchy attached to it. There are a lot of optimizations that could be done at that level.
I chunk pages and generate embeddings for each chunk. So there's no real size limit per page.
Gotta start somewhere :)
Sorry about that, a bit too much load at the moment
Thanks! I'm still figuring things out about pricing, but there will be small plans available.
It does respect robots.txt when crawling. I'll add more details about this in the docs.
Currently just a cloud-toy.
Thanks! The chat demo is actually just a small thing I put together as a preview of what can be done, but the main product is the API. But seeing that most users seem to like that, there's probably something there... If you want to email me at support at embedding.io with some requirements, I can see how to make that work for you.
You can group as many websites as you want into a collection. Then query that collection. Not sure what you mean by exporting; you would like to export the vectors themselves? Or just the chunks of text from the websites?
Apologies, please email me at support at embedding.io. If you have something you'd like embedded, please also mention it so I can set it up for you.
It currently will try to find a sitemap on its own. But I have on the roadmap to let users add their own.
It works with pretty much any website, and works well with docs hosted on GitBook yes, I have embedded a website that's hosted there.
Tech stack is a mix of serverless Laravel, with Cloudflare and AWS functions, and some Pinecone for vector storage. Still experimenting on a few things but don't want to over-engineer unless I know where I'm going.
Unless you own those sites, I'm afraid that's not going to be possible.
Give it URLs or domains, and it will crawl and extract their content, embed them in a vector database, and give you an endpoint that you can then query when doing RAG stuff or semantic search.
This is great! Congrats on shipping this. I made a basic version of this years ago (imageee.com) and would love to find a replacement. Something I could use in the same way. ``` <meta property="og:image" content="https://api.imgsrc.io/basic?title=This is a title&description=blahblahblah&logo=https://example.com/image.png"> ```
I made this service to solve my own customer support problem. GPT-4 generates draft replies using previous responses as a knowledge base. I've been using it for a few months and most replies don't need editing. Thoughts?
The content depends on where you stream from or the VPN you use. If you subscribed from India and then move to the USA, you will see content available in the USA for the price of Netflix India.
Not yet sorry :) Maybe one day!
Will do, thanks for the tip!
Glad you like the idea! Categories are on our list of features to implement. We're thinking about hashtags like Twitter does. Yeah, Signup with Twitter was mostly to accelerate the development time but an email signup will be introduced in the future as we're aware that some people prefer this way.
Hi there, My friend Nathan and I recently launched Predibly.com: a social platform to publicly share your predictions of the future and have interesting conversations about them. Nathan had the idea last week, I build the MVP and we'd love to know your thoughts!
This is awesome! Bookmarked for the next time I'll need a new place to eat :) A nice feature would be a map view of those places.
I like the concept! Small suggestion: you should hide the emails of the owners and use a contact form.
I had the domain since a while and was not using it. So instead of searching hours for a name, I just used what I had :D
Hi, please use the contact form on the website for support.
Thanks, appreciated!
It's fixed ;)
I'll add this feature really soon!