HN user

artembugara

1,636 karma

Building newscatcherapi.com

YC S22

artem [at] newscatcherapi.com

Posts161
Comments323
View on HN
www.newscatcherapi.com 1mo ago

Web Search API Types: Three Architectures, One Confusing Name

artembugara
4pts0
platform.newscatcherapi.com 3mo ago

Show HN: CatchAll – slowest web search API that outperforms everything on recall

artembugara
7pts1
news.ycombinator.com 1y ago

Ask HN: A way to track all news regarding (potential) tariffs by US?

artembugara
2pts2
www.youtube.com 1y ago

Is Michael Scott in Founder Mode? [video]

artembugara
3pts0
aipredict.fun 1y ago

Show HN: Ask LLMs to predict anything based on news

artembugara
25pts9
www.pangram.com 1y ago

60k AI-generated news articles are published every day

artembugara
1pts0
news.ycombinator.com 2y ago

Tell HN: When to apply to Y Combinator (YC S22 alum advice)

artembugara
4pts0
open.spotify.com 2y ago

Parker Conrad, Founder of Rippling

artembugara
2pts1
mattbruenig.com 2y ago

When McDonalds Came to Denmark

artembugara
4pts2
weaviate.io 3y ago

Large Language Models and Search

artembugara
1pts0
techcrunch.com 3y ago

Why this startup chose to sell itself over raising a Series A

artembugara
1pts0
github.com 3y ago

Twitter search is only for logged in users now

artembugara
6pts0
www.elastic.co 3y ago

ChatGPT and Elasticsearch: OpenAI meets private data

artembugara
1pts0
www.sncf.com 3y ago

French train live perturbation updates by SNCF

artembugara
2pts0
www.politico.eu 3y ago

Britain’s plan for when Queen Elizabeth II dies

artembugara
4pts0
newscatcherapi.com 4y ago

Six Text Annotation Tools

artembugara
1pts0
newscatcherapi.com 4y ago

Streamlit: Crypto News Aggregator

artembugara
1pts0
www.youtube.com 4y ago

The Competitors Your Startup Should Worry About

artembugara
2pts0
news.ycombinator.com 4y ago

Ask HN: Have you applied to YC S22?

artembugara
2pts0
newscatcherapi.com 4y ago

Learning Natural Language Processing (NLP) Made Easy

artembugara
1pts0
newscatcherapi.com 4y ago

Custom Named Entity Recognition[NER] Model with SpaCy

artembugara
2pts0
newscatcherapi.com 4y ago

Annotate Company Name and Ticker with Spacy PhraseMacher

artembugara
2pts0
newscatcherapi.com 4y ago

Sentiment Analysis Using Python [Tutorial]

artembugara
3pts0
similarity-demo.newscatcherapi.com 4y ago

Show HN: Streamlit App to Compare Text Similarity Live

artembugara
5pts0
newscatcherapi.com 4y ago

Guide to Text Similarity with Python

artembugara
6pts0
newscatcherapi.com 4y ago

Mining Financial Stock News Using SpaCy Matcher

artembugara
2pts0
www.twitch.tv 4y ago

Twitch Is Also Down

artembugara
4pts0
www.youtube.com 4y ago

Cede and Co – The $54.2T Shadow Trust That Owns the World

artembugara
2pts0
time.com 4y ago

Elon Musk. Person of the Year 2021

artembugara
24pts4
newscatcherapi.com 4y ago

Named Entity Recognition with SpaCy [with code example]

artembugara
1pts0

+1 for Mercury.

Another good thing about Mercury is that in case you’re stuck/not being treated fairly, you can just email/publicly mention Immad (CEO) and he’ll reply within minutes and will look into this

It really makes sense, and the best part — customers love it. It’s the simple form of pricing, and it’s simple to understand.

In many cases though, you don’t know whether the outcome is correct or not but we just have evals for that.

Our product is a SOTA recall-first web search for complex queries. For example, let’s say your agent needs to find all instances of product launches in the past week.

“Classic” web search would return top results while ours return a full dataset where each row is a unique product (with citations to web pages)

We charge a flat fee per record. So, if we found 100 records, you pay us for 100. Of its 0 then it’s free.

We started doing quarterly RFC at Newscatcher, and it was a big game-changer. We're entirely remote.

I got this idea from Netflix's founder's book "No Rules Rules" (highly recommend it)

Overall, I think the main idea is: context is what matters, and RFC helps you get your (mine, I'm the founder) vision into people's heads a bit more. Therefore, people can be more autonomous and move on faster.

no no, they want to use it on external data, we do not do any internal data.

I'll give a few examples of how they use the tool.

Example 1 -- real estate PE that invests in multi-family residential buildings. Let's say they operate in Texas and want to get notifications about many different events. For example, they need to know about any new public transport infrastructure that will make specific area more accessible -> prices wil go up.

There are hundreds of valid records each month. However, to derive those records, we usually have to sift through tens of thousands of hyper-local news articles.

Example 2 -- Logistics & Supply Chain at F100 Tracking of all the 3rd party providers, any kind of instability in the main regions, disruptions at air and marine ports, political discussions around the regulation that might affect them, etc. There are like 20-50 events, and all of them are multi-lingual at global scale.

thousands of valid records each week, millions of web pages to derive those from.

Oh, I totally see your point.

We’re optimising for large enterprises and government customers that we serve, not consumers.

Even the most motivated people, such as OSINT or KYC analysts, can only skim through tens, maybe hundreds of web pages. Our tool goes through 10,000+ pages per minute.

An LLM that has to open each web page to process the context isn’t much better than a human.

A perfect web search experience for LLM would be to get just the answer, aka the valid tokens that can be fully loaded into context with citations.

Many enterprises should leverage AI workflows, not AI agents.

Nice to have // must have. Existing AI implementations are failing because it’s hard to rely on results; therefore, they’re used for nice-to-haves.

Most business departments know precisely what real-world events can impact their operations. Therefore, search is unnecessary; businesses would love to get notifications.

The best search is no search at all. We’re building monitors – a solution that transforms your catchALL query into a real-time updating feed.

Congrats on the HN Launch!

It's probably the best research agent that uses live search. Are you using Firecrawl, I assume?

We're soon launching a similar tool (CatchALL by NewsCatcher) that does the same thing but on a much larger scale because we already index and pre-process millions of pages daily (news, corporate, government files). We're seeing so much better results compared to parallel.ai for queries like "find all new funding announcements for any kind of public transit in California State, US that took place in the past two weeks"

However, our tool will not perform live searches, so I think we're complementary.

i'd love to chat.

Open models by OpenAI 12 months ago

oh, I totally understand that I'd need multiple GPUs. I'd just want to know what GPU specifically and how many

Open models by OpenAI 12 months ago

thanks, this part is clear to me.

but I need to understand 20 x 1k token throughput

I assume it just might be too early to know the answer

Open models by OpenAI 12 months ago

Disclamer: probably dumb questions

so, the 20b model.

Can someone explain to me what I would need to do in terms of resources (GPU, I assume) if I want to run 20 concurrent processes, assuming I need 1k tokens/second throughput (on each, so 20 x 1k)

Also, is this model better/comparable for information extraction compared to gpt-4.1-nano, and would it be cheaper to host myself 20b?

An absolute legend.

I missed my chance to listen to Black Sabbath in 2015 or 2016 during the Rock am Ring because the last day was cancelled.

I'm happy for what Ozzy did in his sixties and seventies, and what a way to go.

And let's not forget, the most likely reason he's been able to get this far with his lifestyle post-80s and 90s is Sharon

Will, Jeff, I am a BIG Exa fan. Congrats on finally doing your HN Launch.

I think NewsCatcher (my YC startup) and Exa aren’t direct competitors but we definitely share the same insight — SERP is not the right way to let LLM interact with web. Because it’s literally optimized for humans who can open 10 pages at most.

What we found is that LLMs can sift through 10k+ web pages if you pre-extract all the signals out of it.

But we took a bit of a different angle. Even though we have over 1.5 billion of news stories only in our index we don’t have a solution to sift through as your Websets do (saw your impressive GPU cluster :))

So what we do instead is we do bespoke pipelines for our customers (who are mostly large enterprise/F1000). So we fine-tune LLMs on specific information extraction with very high accuracy.

Our insight: for many enterprises the solution should be either a perfect fit or nothing. And that’s where they’re ok to pay 10-100x for the last mile effort.

P.S. Will, loved your comment on a podcast where you said Exa can be used to find a dating partner.

Search the web is apparently using SERP.

It’s just breaks my head. We’ve build LLMs that can process millions of pages at a time. But what we give them is a search engine that is optimized for humans.

It’s like giving a humanoid robot access to a keyboard with a mouse to chat with another humanoid robot.

Disclaimer: I might be biased as we’re kind of building the fact search engine for LLMs.

Wow, that's one of the most orange tag-rich posts I've ever seen.

We're doing a lot of tests with GPT-4o at NewsCatcher. We have to crawl 100k+ news websites and then parse news content. Our rule-based model for extracting data from any article works pretty well, and we never could find a way to improve it with GPT.

"Crawling" is much more interesting. We need to know all the places where news articles can be published: sometimes 50+ sub-sections.

Interesting hack: I think many projects (including us) can get away with generating the code for extraction since the per-website structure rarely changes.

So, we're looking for LLM to generate a code to parse HTML.

Happy to chat/share our findings if anyone is interested: artem [at] newscatcherapi.com