If you want to try it without creating an account: https://platform.newscatcherapi.com/catchall/try?utm_source=...
HN user
artembugara
Building newscatcherapi.com
YC S22
artem [at] newscatcherapi.com
You can also check any signal from publicly available sources using tools like CatchAll.
For example, "CEO and CFO appointments at US public companies in the last two weeks" found 142 records [0]
You can also set up monitors to get updates.
[0] https://platform.newscatcherapi.com/catchall/example/gtm--ex...
By the law of large numbers, it's not a significant distance.
+1 for Mercury.
Another good thing about Mercury is that in case you’re stuck/not being treated fairly, you can just email/publicly mention Immad (CEO) and he’ll reply within minutes and will look into this
Chrome on iPhone
Do you mean it’s written by AI?
Or just my writing style?
It really makes sense, and the best part — customers love it. It’s the simple form of pricing, and it’s simple to understand.
In many cases though, you don’t know whether the outcome is correct or not but we just have evals for that.
Our product is a SOTA recall-first web search for complex queries. For example, let’s say your agent needs to find all instances of product launches in the past week.
“Classic” web search would return top results while ours return a full dataset where each row is a unique product (with citations to web pages)
We charge a flat fee per record. So, if we found 100 records, you pay us for 100. Of its 0 then it’s free.
We started doing quarterly RFC at Newscatcher, and it was a big game-changer. We're entirely remote.
I got this idea from Netflix's founder's book "No Rules Rules" (highly recommend it)
Overall, I think the main idea is: context is what matters, and RFC helps you get your (mine, I'm the founder) vision into people's heads a bit more. Therefore, people can be more autonomous and move on faster.
done
no no, they want to use it on external data, we do not do any internal data.
I'll give a few examples of how they use the tool.
Example 1 -- real estate PE that invests in multi-family residential buildings. Let's say they operate in Texas and want to get notifications about many different events. For example, they need to know about any new public transport infrastructure that will make specific area more accessible -> prices wil go up.
There are hundreds of valid records each month. However, to derive those records, we usually have to sift through tens of thousands of hyper-local news articles.
Example 2 -- Logistics & Supply Chain at F100 Tracking of all the 3rd party providers, any kind of instability in the main regions, disruptions at air and marine ports, political discussions around the regulation that might affect them, etc. There are like 20-50 events, and all of them are multi-lingual at global scale.
thousands of valid records each week, millions of web pages to derive those from.
Oh, I totally see your point.
We’re optimising for large enterprises and government customers that we serve, not consumers.
Even the most motivated people, such as OSINT or KYC analysts, can only skim through tens, maybe hundreds of web pages. Our tool goes through 10,000+ pages per minute.
An LLM that has to open each web page to process the context isn’t much better than a human.
A perfect web search experience for LLM would be to get just the answer, aka the valid tokens that can be fully loaded into context with citations.
Many enterprises should leverage AI workflows, not AI agents.
Nice to have // must have. Existing AI implementations are failing because it’s hard to rely on results; therefore, they’re used for nice-to-haves.
Most business departments know precisely what real-world events can impact their operations. Therefore, search is unnecessary; businesses would love to get notifications.
The best search is no search at all. We’re building monitors – a solution that transforms your catchALL query into a real-time updating feed.
Congrats on the HN Launch!
It's probably the best research agent that uses live search. Are you using Firecrawl, I assume?
We're soon launching a similar tool (CatchALL by NewsCatcher) that does the same thing but on a much larger scale because we already index and pre-process millions of pages daily (news, corporate, government files). We're seeing so much better results compared to parallel.ai for queries like "find all new funding announcements for any kind of public transit in California State, US that took place in the past two weeks"
However, our tool will not perform live searches, so I think we're complementary.
i'd love to chat.
oh, I totally understand that I'd need multiple GPUs. I'd just want to know what GPU specifically and how many
thanks, this part is clear to me.
but I need to understand 20 x 1k token throughput
I assume it just might be too early to know the answer
Disclamer: probably dumb questions
so, the 20b model.
Can someone explain to me what I would need to do in terms of resources (GPU, I assume) if I want to run 20 concurrent processes, assuming I need 1k tokens/second throughput (on each, so 20 x 1k)
Also, is this model better/comparable for information extraction compared to gpt-4.1-nano, and would it be cheaper to host myself 20b?
Oh it sounds exactly like what we’d need to embedding multi-page PDFs from government websites.
An absolute legend.
I missed my chance to listen to Black Sabbath in 2015 or 2016 during the Rock am Ring because the last day was cancelled.
I'm happy for what Ozzy did in his sixties and seventies, and what a way to go.
And let's not forget, the most likely reason he's been able to get this far with his lifestyle post-80s and 90s is Sharon
Congrats, @tndl
You guys rock! Big fan
What are some startups that help precisely with “feeding the LLM the right context” ?
Will, Jeff, I am a BIG Exa fan. Congrats on finally doing your HN Launch.
I think NewsCatcher (my YC startup) and Exa aren’t direct competitors but we definitely share the same insight — SERP is not the right way to let LLM interact with web. Because it’s literally optimized for humans who can open 10 pages at most.
What we found is that LLMs can sift through 10k+ web pages if you pre-extract all the signals out of it.
But we took a bit of a different angle. Even though we have over 1.5 billion of news stories only in our index we don’t have a solution to sift through as your Websets do (saw your impressive GPU cluster :))
So what we do instead is we do bespoke pipelines for our customers (who are mostly large enterprise/F1000). So we fine-tune LLMs on specific information extraction with very high accuracy.
Our insight: for many enterprises the solution should be either a perfect fit or nothing. And that’s where they’re ok to pay 10-100x for the last mile effort.
P.S. Will, loved your comment on a podcast where you said Exa can be used to find a dating partner.
Search the web is apparently using SERP.
It’s just breaks my head. We’ve build LLMs that can process millions of pages at a time. But what we give them is a search engine that is optimized for humans.
It’s like giving a humanoid robot access to a keyboard with a mouse to chat with another humanoid robot.
Disclaimer: I might be biased as we’re kind of building the fact search engine for LLMs.
Artem here, co-founder of NewsCatcher (YC S22), our data has been used for research.
Danny and team our old friends who are using our free/super-low pricing for academia and researchers.
AMA, or feel free to email artem@newscatcherapi.com
Co-founder of NewsCatcher (YC S22). There are some reasons for not having a dataset fully open sourced.
But we have free/very very low tiers for academia.
So in case you need access for your research, go to https://www.newscatcherapi.com/free-news-api
Or feel free to email me directly at artem@newscatcherapi.com
Yeah the link you provide is when it’s already official. And I’m more curious about cases when president just say “I’m gonna tax ‘em”
About 18 months after very modest Seed round. We were first bootstrapped and it took a year to go to 300r ARR after we started working full time
I’m a YC founder who did 0 to 2M ARR in founder led sales with absolutely 0 sales background. I’m basically a self-learned coder who had to take CEO role, therefore doing sales.
I find this video about enterprise sales from Pete Koomen (YC Partner) to be the best summary:
I open-sourced pyGoogleNews and wrote a quick blog about how you can reverse engineer google news RSS to turn it into an RSS feed of any website that is supported by Google News
Wow, that's one of the most orange tag-rich posts I've ever seen.
We're doing a lot of tests with GPT-4o at NewsCatcher. We have to crawl 100k+ news websites and then parse news content. Our rule-based model for extracting data from any article works pretty well, and we never could find a way to improve it with GPT.
"Crawling" is much more interesting. We need to know all the places where news articles can be published: sometimes 50+ sub-sections.
Interesting hack: I think many projects (including us) can get away with generating the code for extraction since the per-website structure rarely changes.
So, we're looking for LLM to generate a code to parse HTML.
Happy to chat/share our findings if anyone is interested: artem [at] newscatcherapi.com
this! I've been following Kadoa since its very first days. Great team.
Wow, Kyle, you should have mentioned it earlier!
We've been working on this for quite a while. I'll contact you to show how far we've gotten