HN user

mfkhalil

146 karma

moe@webhound.ai

Posts18
Comments45
View on HN
moekhalil.substack.com 2d ago

The depth problem with agentic research

mfkhalil
2pts0
moekhalil.substack.com 5d ago

MCP onboarding is the most exciting thing happening in tech

mfkhalil
3pts0
www.skillhound.ai 1mo ago

Skillhound: Give your AI access to every public SKILL.md

mfkhalil
10pts2
www.skillhound.ai 1mo ago

Skillhound: A live index of every SKILL.md

mfkhalil
3pts1
www.webhound.ai 3mo ago

How we think about truth, verification, and "time to first trust" at Webhound

mfkhalil
1pts0
www.persuasion.community 6mo ago

What's causing populism around the world? It's the Internet Stupid (Fukuyama)

mfkhalil
10pts41
news.ycombinator.com 10mo ago

Launch HN: Webhound (YC S23) – Research agent that builds datasets from the web

mfkhalil
112pts80
news.ycombinator.com 1y ago

Seeking beta users for Webhound, an AI agent that builds datasets from the web

mfkhalil
1pts0
www.matterrank.ai 1y ago

Search could be so much better. And I don't mean chatbots with web access

mfkhalil
61pts61
www.matterrank.ai 1y ago

Show HN: MatterRank – Build search engines that rank results by your criteria

mfkhalil
3pts0
www.useremy.com 1y ago

Show HN: Remy – A Video Answer Engine

mfkhalil
2pts0
playechoes.vercel.app 1y ago

Show HN: Echoes – A simple web game about avoiding your shadow

mfkhalil
2pts5
nationalinterest.org 1y ago

Historians Predicted the Failure of Democracy (2019)

mfkhalil
6pts2
www.everydayacademic.com 1y ago

Show HN: Everyday's ArXiv Papers Explained at Your Level

mfkhalil
4pts0
thedailyroundup.vercel.app 2y ago

Show HN: Prank your friends with Opengraph previews

mfkhalil
3pts0
justtheanswer.vercel.app 2y ago

Show HN: Perplexity Without the Filler

mfkhalil
29pts6
www.playamnesia.com 2y ago

Show HN: Learning History via CYOA

mfkhalil
3pts0
babelfish.cc 2y ago

Show HN: Browse the Internet in Any Language

mfkhalil
1pts0

Built this internally for our coding agents but have been loving it so much that we decided to make it public.

The web UI is free to use and does not require sign up, and has every public SKILL.md on GitHub indexed, with the index refreshing every 48 hours.

Programmatic use is $20 a month for unlimited searches, and I can vouch for the fact that it's made my agents feel so much smarter and more knowledgeable. Only issue is right now you have to nudge it to use skillhound ("use skillhound first") otherwise it tends to try to figure out best practices on its own.

Hope it can be as useful to the community as it is to us, and would appreciate any feedback you have.

We're working on Webhound - budget controlled long-running deep research. You set a budget and Webhound will use that much in compute/LLM tokens to research your prompt, with built in verification cycles and optional added verification budget. Every claim is cited with evidence and a direct link to the tool calls that produced the claim

The goal is to build a deep research product for actual researchers, since we believe that it is an extremely powerful product that is still nascent but has enormous potential - which we've already seen with some early users.

https://webhound.ai

The least productive teams I've been a part of are the ones where everyone is waiting for their turn to say why an idea is bad. Sometimes being "too smart" can hold you back from building something genuinely new.

Hey, appreciate the feedback. Will address all your points.

Regarding Reddit, we have our own custom handler for Reddit URLs which uses the Reddit API, which we are billed for when we exceed free limits.

For Terms of Service, you're right, that is definitely an oversight on our part. We just published both our Terms of Service and Privacy Policy on the website.

When it comes to comparing with GPT-5 and Claude, we do believe that our prompting, agent orchestration, and other core parts of the product such as parallel search results analysis and parallel agents are improvements on just GPT-5 and Claude, while also allowing it to run at much cheaper costs on significantly smaller models. Our v1 which we built months ago was essentially the same as what GPT-5 thinking with web search currently does, and we've since made the explicit choice to focus on data quality, user controllability, and cost efficiency over latency. So while yes, it might give faster results and work better for smaller datasets, both we and our users have found Webhound to work better for siloed sources and larger datasets.

Regarding account deletion, that is also a fair point. So far we've had people email us when they want their account deleted, but we will add account deletion ASAP.

Criticism like this helps us continue to hold ourselves to a high standard, so thanks for taking the time to write it up.

Accuracy-wise we think it's almost there but probably still a few iterations away from being perfect. It's great at eliminating a lot of the collection time though.

Interestingly, we're working with B2B clients right now where we use Webhound to curate and then act as the "validation" layer ourselves. The agent lets us offer these datasets way cheaper with live updates, but still with human oversight.

Thanks for testing it! That's definitely a miss, sounds like it got confused about what you were looking for and went after board member pages instead of the actual meeting/document sites.

We're working on better query interpretation, but in the meantime you could try being more specific like "find BoardDocs or meeting document websites for each district" to guide it better. Also, you can usually figure out how it interpreted your request by looking at the entity criteria, those are all the criteria a piece of data needs to meet to make it in the set.

Good point. Our main differentiation is the shared workspace - users can step in and guide the agent mid-task, kind of like Cursor vs Claude (which can technically generate the same code that Cursor does). Firecrawl (or any crawler we may use) is only part of the process, we want to make the collaborative process for user <> agent as robust and user controllable as possible.

Thanks for the feedback! From what we've seen it's actually the other way around - once it gets a sense of where this information lives the latter stages of data collection go quicker, especially since it's able to deploy search agents in parallel to get information and doesn't need to do the manual work as much anymore. Having said that, it does sometimes forget to do that, and although we've added the critic agent to remind it to do that it can be inconsistent but usually if you step in and ask it to deploy agents in parallel that fixes it.

We use Gemini 2.5 Flash which is already pretty cheap, so inference costs are actually not as high as they would seem given the number of steps. Our architecture allows for small models like that to operate well enough, and we think those kinds of models will only get cheaper.

Having said all that, we are working on improving latency and allowing for more parallelization wherever possible and hope to include that in future versions, especially for enrichment. We do think that one of the weaknesses of the product is for mass collection - it's better at finding medium sized datasets from siloed sources and less good at getting large comprehensive datasets, but we're also considering approaches that incorporate more traditional scraping tactics for finding these large datasets.

Thanks a lot regarding UI and good point on the schema editing.

We've been having similar thoughts about pricing and offering unlimited, but since it is feasible for us in the short term due to credits we enjoy offering that option to early users, even if it may be a bit naive.

Having said that, we are currently working on a pilot with a company whom we are offering live updates, and they are paying per usage since they don't want to have to set it up themselves, so we can definitely see the demand there. We also offer an API for companies that want to reliably query the same thing at a preset cadence, which is also usage based.

For crawling we use Firecrawl. They handle most of the blocking issues and proxies.

We maintain a constant browser state that gets fed into the system prompt which shows the most recent results, current page, where you are in the content, what actions are available, etc. It's markdown by default but can switch to HTML if needed (for pagination or CSS selectors). The agent always has full context of its browsing session.

A few design decisions we made that turned out pretty interesting:

1. We gave it an analyze results function. When the agent is on a search results page, instead of visiting each page one by one, it can just ask "What are the pricing models?" and get answers from all search results in parallel.

2. Long web pages get broken into chunks with navigation hints so the agent always knows where it is and can jump around without overloading its context ("continue reading", "jump to middle", etc.).

3. For sites that are commonly visited but have messy layouts or spread out information, we built custom tool calls that let the agent request specific info that might be scattered on different pages and consolidates it all into one clean text response.

4. We're adding DOM interaction via text in the next couple of days, so the agent can click buttons, fill forms, enter keys, but everything still comes back as structured text instead of screenshots.

We currently use Firecrawl for our crawling infrastructure. Looking at their documentation, they claim to respect robots.txt, but based on user reports in their GitHub issues, the implementation seems inconsistent - particularly for one-off scrapes vs full crawls.

This is definitely something we need to address on our end. Site owners should have clear ways to opt out, and crawlers should be identifiable. We're looking into either working with Firecrawl to improve this or potentially switching to a solution that gives us more control over respecting these standards.

Appreciate you bringing this up.

Thanks! Unlike a lot of our competitors who use search-inspired UX, we went with an agentic approach inspired by tools like Cursor - basically iterative user control.

Instead of just search query → final result (though you can do that too), you can step in and guide it. Tell it exactly where to look, what sources to check, how to dig deeper, how to use its notepad.

We've found this gets you way better results that actually match what you're looking for, as well as being a more satisfying user experience for people who already know how they would do the job themselves. Plus it lets you tap into niche datasets that wouldn't show up with just generic search queries.

Thanks, glad to hear you had a good experience.

We were heavily inspired by tools like Cursor - basically tried to prioritize user control and visibility above everything else.

What we discovered during iteration was that our users are usually domain experts who know exactly what they want. The more we showed them what was happening under the hood and gave them control over the process, the better their results got.

Hey, we’re working on MatterRank which is pretty similar to this but currently works on web search. (e.g. I want to prioritize results that talk about X and have Y bias and I want to deprioritize those that are trying to sell me something). Feel free to try it out at https://matterrank.ai

Would also be interested in hearing more about what you’re envisioning for your use case. Are you thinking a browser extension that acts on sites you’re already on, or some sort of shopping aggregator that lets you do this, or something else entirely?

I think in a lot of cases that's because the meta with LLMs right now is to "have them do things for you", which generally means that obfuscating what's actually happening behind the scenes can make them seem "smarter" to the average user. Also, engineers are used to full control over deterministic input-output pipelines, which is a framework they try to force on LLM applications that fails miserably for the reasons you've listed.

In my opinion, the best applications of LLM UX will have full clarity for the end user (something we're trying to do with MatterRank). The non-determinism should be something the user can control to get better results, not something the engineer has prompted that takes control away from the user.

Now, if the use case you're looking for is "give me results with x text", then yes I agree with you that LLMs are just getting in the way. But that's not always the case.