HN user

dserban

573 karma

Well-Rounded Scala/Rust/Python Dev With Kafka, Flink, Spark Streaming And ETL / ELT Pipeline Experience.

E-mail: dserban01 => gmail

Posts14
Comments168
View on HN
  Location: Europe, Role Type: Fractional, Remote / Worldwide
  Remote: Yes
  Willing to relocate: Yes
  Technologies: Kafka, Flink, Spark, Cassandra, Scala, Rust / PyO3 / Tokio, Python, Zookeeper, Databricks, Delta Lake, BigQuery, Redshift, Hive, Kinesis, Airbyte, Airflow, Temporal, DBT, Aerospike, Snowflake, PrestoSQL, Trino, Clickhouse, SQL, Vector DBs, Golang, gRPC/protobuf, Terraform, CUDA
  Resume: https://drive.google.com/file/d/1lZuSpneLDVzzYoeaA8or-m5m2dfYAMB4/view
  Email: in the profile (please do mention "Via HN" in the subject for your email to land in the right place in my inbox)
[Remote Worldwide, Fractional] I'm pursuing a role in the high-scalability distributed systems space. I'm a well-rounded Scala/Rust/Python dev, well-versed in data engineering, with deep knowledge of the internals of distributed datastores. I have experience with data modeling for high throughput database activity and a strong understanding of which workloads and data access patterns are scalable and what datastore the data should reside in. I have a highly confident ability to lead a data / MLops project from start to finish. Hire me as fractional data engineer.

Core Focus Areas:

● Cassandra (Data Modeling, Troubleshooting Performance And Operational Issues)

● Apache Iceberg (Scaling, Tuning, Self-Hosted Setup)

● Stream Processing At Scale: Kafka, Flink, Spark Streaming, Storm

● Custom-Crafted Contextualized Embeddings, Vector-Based Semantic Search, Deep Intent Recognition In Search Engine Queries

● Languages: Scala, Rust, Python, SQL (proficient), Golang (ramping up)

Educational Background: Computer Science.

Solid experience working remotely and working with teams that are distributed geographically. I typically work Pacific Time hours.

Ask: $385K base, pro-rated.

Adding to your 6 and 7, Ed Zitron's Better Offline podcast has a good series on how the path was paved to the cost reckoning of the present day.

At some point, I needed to write a function which, given a collection of product titles, picks one that is neither the longest nor the shortest, it should pick the one which best captures the essence of the product while not being excessively verbose. For example, given the product titles below, it should pick "Portable Two-Way Translator, Handheld".

Portable Two-Way Translator, Handheld

Portable Handheld Translator

Handheld Two-Way Translator

Electronic Two-Way Portable Instant Voice Translator, 40 Languages, Handheld

Based on previous experience with centroid-based algos, the function I wrote does a first pass throwing all words from all product titles into one big bag, then computing a centroid (frequency histogram with low-frequency words removed). The second pass is to compute a cosine similarity score for each product title (its own frequency histogram against the centroid). Whichever product title is the most similar to the centroid wins.

That algo may have existed already in some academic paper somewhere, but I came up with it independently.

SEEKING WORK, Senior Data Engineer, Remote / Worldwide

Well-rounded Scala/Rust/Python dev, well-versed in data engineering, with deep knowledge of the internals of distributed datastores. I have experience with data modeling for high throughput database activity and a strong understanding of which workloads and data access patterns are scalable and what datastore the data should reside in. I have a highly confident ability to lead a data / MLops project from start to finish. Hire me.

Core Skills:

● Cassandra (Data Modeling, Troubleshooting Performance And Operational Issues)

● Apache Iceberg (Scaling, Tuning, Self-Hosted Setup)

● Stream Processing At Scale: Kafka, Flink, Spark Streaming, Storm

● Custom-Crafted TF-IDF / Embeddings, Vector-Based Semantic Search, Deep Intent Recognition In Search Engine Queries

● Languages: Scala, Rust, Python, SQL (proficient), Golang (ramping up)

Other Skills: Airflow, Aerospike, DBT, Snowflake, Qdrant, Databricks/DeltaLake, BigQuery, Redshift, Hive, Kinesis, PrestoSQL/Trino, ClickHouse, gRPC, Airbyte, Terraform, CUDA, Reverse-Engineering Search Engine Technologies.

Educational Background: Computer Science.

Solid experience working remotely and working with teams that are distributed geographically. I typically work Pacific Time hours.

E-mail address in the profile.

I'm not sure if I should be applying but before I do, I'm wondering what kind of novel or interesting data driven algorithms you refer to here. Are you referring to like adaptive query planners based on runtime statistics? Real-time detection of query latency spikes? Automated creation of secondary indexes based on usage patterns? Statistical smarts to decide what to cache and when? These would be sort of runofthemill for any modern database product, can you clarify if you are looking for any of these?

Don't rely on gut feeling or anecdotal evidence, there's good data you can look at.

Over a few years, I came to rely on proxy indicators that I found to be reliably in sync with the strength / weakness of the job market for software engineers.

One of these indicators is the business spending index (it has an official name, something about purchasing managers' sentiment index or similar). CNBC tends to show it a lot these days, you can't miss it.

Another one is "US Auto Loans Delinquent by 90 or More Days (I:USALD90)" which is a very good "finger to the wind" for how the overall economy is doing, since everyone needs a car.

Neither of these indicators are looking good these days.

ChatGPT Search 2 years ago

Looking at this, it's not long until OpenAI starts offering a version of this to ecommerce companies for their search backends, it's a space that is held back by corporate corruption, and it's ripe for disruption.

I'm still working through legal complications with my current employer, and in no mood to start my own company. The job market is also bad on the dataeng / datainfra / mlops side of things. My first guess is that interest rates still have to come down some. Giving up on software engineering and trying a different line of work is not an option for me, I was born with a passion for this.

https://getborderless.com/ would be one alternative I can think of that's in the same space as Wise.

But they too will behave in the same way at some point, due to compliance requirements. Banks and fintech companies need to know 1. who you are 2. that your income is from legit sources.

For #2, ask your customers (who paid you recently) to give you a copy of the banking app records of those payments, where the last 4 digits of the payer's bank account can be seen. Wise can then cross-reference those against what they see on their end, at which point they can declare your source of income verified. (I had to learn this the hard way myself.) Very importantly, those payment records have to be from the banking app that the accounts payable team is using, not Ripple or any internal payroll system.

The custom search engine space is absolutely ripe for disruption.

Whenever you need to add a search experience to your SaaS project, you don't really think of building it yourself, however the established companies in the search engine space are lazy, retrograde and corrupt, and they absolutely don't deserve your money. (I know because I work for one of them.)

OpenAI has recently announced it will slash the cost of calls to its embeddings API by a whopping 75%. This is a huge opportunity for a startup to disrupt the search engine space.

With access to OpenAI's now-affordable embeddings API and to a vector database, one mid-level engineer can build a highly scalable custom search engine for ecommerce and retail, in a few weeks.

And yet nobody is talking about this.

We are not at all in disagreement. In fact, you are making the same point I'm making, which is that there is nothing special about any of the commercially available search engines today. With access to OpenAI's embeddings API and to a vector database, a mid-level engineer can build a highly scalable search engine in a few weeks. As a startup, it makes sense today to build your own search engine rather than buy off the shelf.

The only thing companies like ElasticSearch and Algolia still have is their pre-existing customer bases, a few thin layers of marketing, and some network effects. Search engine companies are effectively marketing companies nowadays.

That explains why the corrupt management at my company thinks it's a better strategy to throw hissy fits on social media in the general direction of OpenAI, than to work on actually building useful tech.

Somewhat off-topic but congratulations, you have re-implemented (in the small) the search engine that powers HN. I say this as an employee of the company behind that engine, and I say this as someone who is intimately familiar with the codebase and with how the technology works.

With a little more effort towards scalability, and with some input from a devops co-founder, you'd be on your way to launching a state-of-the-art, market-competitive search engine.

My employer won't be too happy seeing me post this, but I don't care, this needs to be said.

I can relate to the 10-for-Google-90-for-ChatGPT ratio.

I use Google when I know the words that make up the name of a website, but not the URL itself. So Google is for unsophisticated searches. I also use Google instead of HN's own search function when looking for older submissions here.

What I like about about ChatGPT (and friends) is the ability to synthesize knowledge from unrelated bits and pieces in different sources.

In contrast, if it could verbalize what it's doing to you, a traditional search engine would say here's a link dump, have fun with it, and have a nice day.

If half the knowledge you're looking for is stored on web page A towards the bottom, and the other half on web page B in the first few paragraphs, a traditional search engine will dump on you links to web pages A and B, followed by some filler URLs pointing at documents that happen to vaguely match your search keywords. This is a very subpar experience by today's standards, but it's so ingrained into our habits that we're not complaining.

Haskell course at university. Had fun doing all the exercises in the curriculum, and couldn't stop. Explored on my own how to implement in Haskell the more advanced programming patterns that weren't covered in the course. Had fun with lazily evaluated infinite lists where you evaluate the second element of the lazy list that was passed as an argument, and at the end of the same lower order function you return the lazy list but with its head chopped off, to have the algorithm run in constant space.

I have come to regard lazily evaluated infinite lists as a must-have before you can call something a functional programming language.

In trying to build more complex systems, I noticed that you really have to know what you're doing with Haskell, otherwise you end up with leaks that are very difficult to trace back.

What really solidified FP for me was taking an elective category theory for computer scientists course.

Today I'm building stuff in Scala, which is the closest you can get to being able to pay the bills while doing principled FP.

I have a monster of a Linux cluster at home, which is busy running Cassandra and Spark for my other side projects. I'm piggybacking on that while this particular project is still in development.

At the rate LLMs are evolving, a reasonably priced cloud offering will probably exist for me to leverage for production, by the time I need it.

Although GPT is in the "peak of inflated expectations" phase of its hype cycle, there is immense opportunity now for startups to seriously disrupt even 10-year-old companies who falsely thought they had entrenched themselves and consequently became lazy and corrupt.

I predict that we're going to reach phase 5 some time soon ("the plateau of productivity"). And when the dust has settled, the damage to the entrenched-retrograde "IBMs" of our industry will be real. I know because I work for one of these "IBMs".

It's a two-step prompt-injection process. In the first step, I wrap the raw typed-in input inside of a dynamically-generated prompt template that takes some context into account (for example, the specifics of the product that the person is currently looking at, or the recent history thereof). The prompt generates a list of Meilisearch keyword search queries. I issue all queries and harvest the results. Scoring and ranking happens in the second (also prompt-injection powered) step.

So the answer to your question is pretty much yes.

It's Saturday and I have time to add some color to my comment above.

First things first: You might have seen all the news around the failure of Neeva, a search engine startup. There is a core structural reason why Neeva has failed, and I predict that all pure-play search engine companies not named Google will eventually suffer the fate of Neeva. Keep this in mind when you decide to chain the future of your technology stack to the fortunes of a search engine company. I can give this advice with confidence because I work in such a company, and our management is clueless.

Back to answering your question: any answer you get can't be complete, because the term you used, "search experience", is ill-defined.

Here are the 3 types of search use cases, in increasing order of complexity:

1. Traditional, barebones, keyword-matching search, with some level of tolerance for typos. Meilisearch is a rock-solid option if that is your use case.

2. Traditional search with deep intent recognition of search queries. This is where the snake oil search engine companies try to sneak their way in. I'll go into detail further down below.

3. Search-aided knowledge work, a problem for which LLM technologies (the ChatGPTs of the world) are the solution. Before November 2022, for lack of a better option, the space was being (very clumsily) served by traditional search technologies.

Details on number 2. My company spent a decade and burned several tens of millions of dollars on developing, refining and marketing a search engine codebase that was outmatched in an instant in November 2022, when ChatGPT became a thing. How will my company make all of that money back? Management doesn't know, but they hope against hope that they can charge startups like yours for access to the snake oil.

All of that engineering effort was expended on writing code by hand, without help from a copilot-type code generation tool. Based on what I have seen in the codebase, I estimate that close to 90% of that code is boilerplate that an AI could easily generate, so all of that effort can nowadays be easily replicated by one mid-level engineer working for a few weeks (search is a very well-understood problem, very competitive and very commoditized).

My point is, if you're going to buy instead of build, buy from a company that has benefitted from the AI revolution, as opposed to being disrupted by it.

To close out, I will leave you with this ranking page, those down-arrows should tell you everything you need to know about the search engine space:

https://db-engines.com/en/ranking/search+engine

Building it yourself is the better option nowadays.

Start with open-source Meilisearch, and don't look back.

The "buy" options people talk about are tired-old snake oil, and you'll end up with buyer's remorse.