HN user

prabhatjha

130 karma

Head of Engineering - Database and Vector Search at DataStax powered by Apache Cassandra. Previously I was CTO at Wootric, cofounded InstaOps - a mobile APM company which got acquired by Apigee. https://twitter.com/prabhatjha

Posts32
Comments38
View on HN
prabhatkjha.com 1mo ago

Execution Under Ambiguity and Uncertainty

prabhatjha
3pts0
thenewstack.io 2y ago

Impact of Vector Dimensions,PQ, BQ on ANN Relevance

prabhatjha
2pts0
www.datastax.com 2y ago

GPU-Powered KNN Ground Truth DataSet Generator for Vector Search

prabhatjha
3pts0
www.datastax.com 2y ago

One line of code change to move from Open AI to Astra Assistants API

prabhatjha
18pts8
cassio.org 3y ago

Integrating Apache Cassandra with Generative AI

prabhatjha
3pts0
segment.com 3y ago

Do's and Don'ts of software engineering internship

prabhatjha
4pts0
github.com 4y ago

A Lightweight CQL Proxy/Sidecar for Apache Cassandra Written in Golang

prabhatjha
1pts0
stargate.io 4y ago

Designing Data Gateway for Cassandra in Open Source

prabhatjha
18pts13
k8ssandra.io 4y ago

Helm charts, limitations and how K8s operator helps

prabhatjha
8pts0
medium.com 4y ago

Why and how of a multi-tenant Apache Pulsar as a service offering

prabhatjha
1pts0
www.datastax.com 5y ago

Multi-Cloud Streaming as a Service Built on Apache Pulsar

prabhatjha
1pts0
github.com 5y ago

Show HN: JMS with Apache Pulsar

prabhatjha
1pts0
www.datastax.com 5y ago

Serverless Cassandra from DataStax

prabhatjha
3pts0
www.datastax.com 5y ago

Why DataStax chose Pulsar over Kafka for their streaming platform

prabhatjha
19pts1
engineering.wootric.com 5y ago

Why we rewrote our JavaScript SDK in TypeScript

prabhatjha
9pts0
engineering.wootric.com 5y ago

Pragmatic code reuse in a Rails Monolith

prabhatjha
9pts0
engineering.wootric.com 6y ago

Taming Callback Pyramids in AnuglarJS App Using Async/Await

prabhatjha
14pts2
prabhatkjha.com 6y ago

Some recommendations for a fully virtual hiring process

prabhatjha
2pts0
engineering.wootric.com 6y ago

Wootric moved ML deployment pipeline from AWS to GCP

prabhatjha
28pts17
prabhatkjha.com 6y ago

Mobile app perf monitoring dream and Envoy Mobile

prabhatjha
2pts0
engineering.wootric.com 6y ago

Code Walkthrough of Bert with PyTorch for a Multilabel Classification in NLP

prabhatjha
13pts2
engineering.wootric.com 6y ago

Custom machine learning model and pipeline deployment on GCP AI Platform

prabhatjha
14pts0
engineering.wootric.com 7y ago

Understanding and explaining accuracy in machine learning systems

prabhatjha
12pts5
engineering.wootric.com 7y ago

Comparing AWS and GCP NLP API for Sentiment Analysis and a Case for Custom Model

prabhatjha
12pts4
www.wootric.com 8y ago

Our machine learning and NLP journey

prabhatjha
114pts27
www.wootric.com 8y ago

Announcing Command Line Surveys – Get NPS Feedback from Developers

prabhatjha
9pts0
www.wootric.com 8y ago

How We Migrated from Heroku Postgres to AWS RDS

prabhatjha
16pts0
www.wootric.com 8y ago

A Look Under the Hood: AI for Customer Experience at Wootric

prabhatjha
18pts3
cloud.google.com 9y ago

Wootric: Analyzing customer feedback using machine learning

prabhatjha
8pts0
www.wootric.com 9y ago

Engineering war stories and lessons learned in 2016

prabhatjha
43pts4

Initially I thought we need a dedicated vector database but as we tried to build even a simple gen-ai applicaitons we realized that we need another "regular" database to build a complete application.

Instead with DataStax's Vector Search, we designed a document style API and corresponding clients that give you a Vector Native experience to do CRUD of vector and meta-data as well CRUD of other data models. Here is one client ref for example https://docs.datastax.com/en/astra/astra-db-vector/clients/p...

Brad Cox has died 5 years ago

What a great tribute you have written. When I first found about swizzling through a seasoned iOS dev I was blown away. The swizzling capability in obj-c basically helped create my first startup, InstaOps, a long time back which allowed no code change to instrument an app to capture logs and network performance metrics.

This is a fantastic idea -- the kind you see and go why the heck this was not done before. Such a huge time saver.

Agree. You don't want to do Fashion Driven Development (FDD). You would end up chasing every shiny thing that comes out and not really master anything.

Pick any one stack, build some simple projects and then expand on it. You would also want to be able to read well written code from others on github and learn from those.

I am assuming you are talking about our deployment on GCP Cloud Run? We have thought about sending a heartbeat API call. It we notice any user experience friction because of this lag then we will definitely do that. As we said in the blog, this has not been a major pain point as of today.

We supported batch api calls in v0. But as those API calls increased a new instanced would get spun up but boot time was longer. To get around it, we would have to keep more instances running all the time which obviously costs more money.

Thanks for the tips. Having everything colocated in a k8s cluster will definitely help with latency and probably overall infra cost but it will be at the expense of engineering times spent on running a k8s cluster in prod.

Fingers crossed that we keep growing which would mean that we can justify working on v2 architecture.

Because we were loading all the models startup time was long which meant that server would return 5xx errors which created more instability. We would had to do some engineering around it with a mix of config and code changes.

The bigger issue was that he had to use bigger machine as we added more custom ML models for our customers. New architecture gives us huge $$ saving and more visibility into performance of each model.

I think that AWS are the "bad guy" here. I am happy to be proven wrong but they don’t have a track record of contributing code to the OSS projects they provide as a service. Whenever they are forced like it was the case with Mongo and now with Elastic they are using their power to have a competing OSS project. There is nothing wrong with this business wise but it's totally against spirit of OSS.

If you look at Red Hat on other hand when they decided that Kubernetes was the way to go for their OpenShift project they put lots of engineering resources for upstream k8s.

I totally agree. Deploying microservices and running k8s sound easy until you actually do it. For example, just see this section of k8s docs about exposing services: https://kubernetes.io/docs/concepts/services-networking/serv... . You need to understand many different concepts first to get this right. However, I think once you cross that hurdle, the traditionally harder stuff like auto scaling, rolling upgrade becomes relatively easier.

However, I would say that it's really early days for K8S and the ecosystem around it. As long as K8S does not try to solve every problem in the world and focus on the problems it's designed to solve, things will get easier and then may be a 60-min video can do some justice. ;-)

Working for a early stage startup depends on the stage of life, your financial freedom and your ability to hustle.

Early stage startup requires that you work a lot more than 40 hours of week with a big pay cut compared to what you would get at an established company. This usually is not a problem for individual or couple who don't have kids but very difficult for people with school kids. You have less time and less money -- both have direct impact on how your kids grow up. Is it worth taking the risk? This depends on your values and how you define success in life.

Your roles and responsibilities are not well defined. Even for a software engineer, you have to split your time helping sales, customer support and marketing. I personally find this aspect super exciting but I know a lot of people who don't like and wont thrive in this kind of environment.

Most startups fail. Founders and VCs can screw you -- intentionally or unintentionally. Odds are just stacked against you if you define success by financial gains.

I totally agree with this. We learned it pretty quickly that classification does not generalize across domains so we narrowed the problem space by focusing one domain at a time followed by predefined and fixed set of categories so that we can measure effectiveness of our solution as we experimented with different algorithms and deployment pipeline.

Ah -- I thought you were talking about classification of these tweets so that politicians know what their followers are talking about. Sentiment analysis is a very small part of what we do and as you said there are tons of examples on web that use Twitter's data in their model.

Forget about other languages, it's sometimes tricky for English as well. Someone said language is the world's oldest API but it's also the most complicated. ;-)

For couple of our Norwegian and Spanish customers we hit Google translate to translate feedback into English and then feed it through our ML engine to classify. Accuracy obviously is not as good as it should but it gives them good insight.