HN user

Major_Grooves

842 karma

Building: tilores.io

Posts69
Comments258
View on HN
medium.com 20d ago

I deduplicated 53,000 missing-persons reports from Venezuela's earthquake

Major_Grooves
3pts2
tilores.io 27d ago

Show HN: Entity Resolution on Your Desktop

Major_Grooves
2pts0
tilores.io 5mo ago

The Georgia voter data is clean

Major_Grooves
4pts0
medium.com 5mo ago

Virginia does not have a 33% duplicate voter registration rate

Major_Grooves
2pts0
www.opensanctions.org 1y ago

More Than a Technical Problem: Name Matching in Sanctions Screening

Major_Grooves
1pts0
www.opensanctions.org 1y ago

How Russian Data on Sanctioned Companies is disappearing and how we put it back

Major_Grooves
7pts0
github.com 1y ago

Show HN: IdentityRAG – customer identity resolution for LLM applications

Major_Grooves
6pts1
major-grooves.medium.com 1y ago

Don't Cheat on Product Hunt

Major_Grooves
4pts0
news.ycombinator.com 2y ago

Ask HN: Is there a need for entity-based RAG?

Major_Grooves
1pts2
tilores.io 2y ago

Deduplication of US voter data to identify voting fraud

Major_Grooves
9pts6
guitton.co 2y ago

How to Build and Interpret a Nomogram for Setting Better Running Goals

Major_Grooves
4pts0
webapps.stackexchange.com 2y ago

Meetup is unable to provide proper invoices

Major_Grooves
2pts1
blog.acolyer.org 2y ago

An overview of end-to-end entity resolution for big data

Major_Grooves
1pts0
major-grooves.medium.com 2y ago

I picked a Bangladeshi tech company over the Silicon Valley darling

Major_Grooves
6pts1
lantsandlaminins.com 3y ago

Is science becoming more novel?

Major_Grooves
1pts0
medium.com 3y ago

In Defense of Serverless

Major_Grooves
2pts0
github.com 3y ago

Show HN: CLI tool for sending batch requests to a GraphQL API (OSS)

Major_Grooves
4pts0
major-grooves.medium.com 3y ago

We turned the tables to catch our sister's Bumble stalker

Major_Grooves
2pts0
github.com 3y ago

Use ChatGPT Programmatically in Python

Major_Grooves
2pts0
tilores.io 3y ago

Compare Fuzzy Matching Algorithms

Major_Grooves
12pts3
major-grooves.medium.com 3y ago

Just how complicated could it be to register a German company?

Major_Grooves
228pts405
tilores.io 3y ago

Entity Resolution: Reflections on the most common data science challenge

Major_Grooves
60pts8
medium.com 4y ago

Is Elasticsearch suitable for entity resolution?

Major_Grooves
1pts0
major-grooves.medium.com 4y ago

I received emails when prospects tested our competitors

Major_Grooves
1pts0
www.youtube.com 4y ago

Using Dyson Sphere Program to tile the world for data geolocation

Major_Grooves
1pts0
tilodb.com 4y ago

A novel approach to entity resolution using serverless technology

Major_Grooves
63pts28
tilodb.com 5y ago

Show HN: TiloDB – serverless entity resolution technology

Major_Grooves
19pts10
www.indiehackers.com 5y ago

How to Build an Audience-Driven Business

Major_Grooves
23pts6
medium.com 6y ago

One half of Sleeping Giants leaving the movement

Major_Grooves
2pts0
lantsandlaminins.com 8y ago

Life in a lab (an under-appreciated comic)

Major_Grooves
5pts0

meanwhile you practically have to pay people to take away current used Ikea pieces.

The market for old DDR (East Germany) furniture here in Berlin is a bit crazy. Old cheap poor quality furniture (all MDF) selling for loads because it looks good on Instagram.

I know a guy that goes to house clearance sales in random places in East Germany, then brings pieces back to sell to hipsters in Berlin. Like me.

I'm the founder of Tilores, the entity-resolution tool used here - so full disclosure, this is my company's product. This wasn't a paid engagement or a case study. It started because my wife is from Venezuela, and she saw people on social media pointing out that the missing-persons lists had huge numbers of duplicates.

On the data: these are public citizen-lead efforts to crowd-source the names of the missing - hosted on websites and spreadsheets. There is no official verification process behind the individual entries, which is part of why the duplicate problem existed in the first place.

An issue we have now realised is "bad actors" trying to access the data...

Happy to answer anything - methodology, false positives, data handling, whatever.

Do you end up with lots of duplicates when you are scraping? If you also scrape IG, YouTube and LinkedIn, would you link them all to the same influencer?

That might be quite an interesting identity resolution challenge (disclosure: I build identity resolution tech).

I would not mind taking a look. Always interested to see how others are handling such data.

Hello HN! We (Steven, Hendrik and Stefan) built a real-time identity resolution system that can handle hundreds of millions of customer records, and recently launched a LangChain integration to use it as a RAG source for LLMs.

We built this while working at a European credit bureau, where we needed to deduplicate and match millions of monthly record updates from various sources. Traditional approaches using graph databases and Spark couldn't handle the scale, so we built our own solution using AWS Serverless.

Each identity is stored as an individual graph structure, using rules-based and ML matching. Performance: <300ms ingest (tested to 5,000/sec), <150ms search regardless of graph size. Several fintech companies use it for fraud detection, KYC, and customer 360.

Unlike vector databases which can blur similar entities together, IdentityRAG maintains distinct customer identities while pulling data from multiple systems - even when customer details differ across databases.

You can try it out with our sample chatbot in the Github repo (linked above). Free to sign up, we charge based on number of unified customer records (it is free for playing and testing). We would love to hear your comments and questions.

There is also a demo video in the repo and you can find more details about us here: https://tilores.io/

I experienced the same thing, and wrote nearly the same article two years ago: https://major-grooves.medium.com/just-how-complicated-could-...

Yes I know we can register a UG, but in the end you don't. And it is not just the share capital that is annoying it is everything else.

It literally costs 10x more to do the bookkeeping for a German company vs a UK one. Plus getting investors is much more difficult because of the notary requirements.

We used a SPV for our first round, but even that is annoying. I had some angel investors pull out purely because we are a German GmbH.

just to be clear - the 61 duplicate voting cases were only for Ohio and Pennsylvania - the 400k duplicate profiles were across all 7 states we looked at.

Indeed there is certainly not mass voter fraud. We were glad not to find that, but tbh surprised that we found any at all. Originally we were only going to look for duplicate profiles - it didn't even occur to us to look for actual fraud.

But why not make it a complete non-issue? It would be so easy to fix this data so there were no duplicates, then there would not even be any accusations like there were in 2020.

What I want to create is complete trust in the data to avoid the... bickering later.

/edit - as the poster below mentions, the 61 were just the ones that were manually confirmed. There were 1000 potential cases.

We come from Germany - where there is unlikely to be a big issue, as citizens have to be quite careful about registering where they live in one place only.

I suspect the data in the UK (where I originate) would be pretty messy. The voting lists there are a free for all, I reckon!

Indeed, unfortunately with the John Smiths of this world there will be false positive matches. What we could do is add that they need to be from the same town/postcode, but then that is quite an unreliable attribute too.

Similarly, there are a lot of false negatives where we know two records should match, but we could not because that would require a rule that would create more false positives.

In the end, it was the best we could do with the public view of the data. If we were working with the data Companies House actually holds itself, it would of course be much better.

I am so glad to hear about this. Readlang has always been one of my favourite language tools and I am sure you can grow it up to to be a pretty decent sized service, with decent MRR.

Let me know if you are ever in Berlin and we can meet up again!

exactly - under GDPR it would be fine to retain specific data for a limited period of time to prevent fraud (which is what this is, really).

Also, this type of user is not making deletion requests - they are deleting and then immediately reinstalling the app to appear "fresh".

Indeed the same though crossed our minds (as I write in the article) that it is unlikely to be his first offence.

We actually did not want to humiliate him, and believe me we had the opportunity, because we did not want to "push him over the edge". I kinda hope he has kept his job so he can try to get help and stop himself doing this again.

aye, exactly. I presume most of the people criticising our "citizen's arrest" idea are form the US, where I indeed would not have contemplated doing such a thing.

In Scotland I do not really have the same fear. Not sure I would do the same thing in Glasgow or London, but still, we are not idiots, we were not going to take any crazy risks but we figured the risk to us was worth it to stop the risk to our sister. We weighed the risks up.

"Complete idiots"might be a slight exaggeration, but hey no offence taken.

iirc they knew we had video, but the video of the car vandalism was never going to be enough to do anything. They had not seen that specific video. In the end it was only a small part of the evidence package.