HN user

ctc24

585 karma
Posts31
Comments41
View on HN
news.ycombinator.com 7mo ago

Ask HN: What would you call a package whose purpose is to import data?

ctc24
7pts12
www.prequel.co 8mo ago

How to get a character from a codepoint in Spark SQL

ctc24
2pts0
www.cnbc.com 11mo ago

Google, Kairos Power plan advanced nuclear plant for Tennessee grid by 2030

ctc24
4pts0
www.prequel.co 1y ago

Could Open Table Formats End the Reign of Snowflake and Databricks?

ctc24
1pts0
dataengineeringcentral.substack.com 1y ago

S3 Tables are not as open as they may appear

ctc24
1pts0
www.nytimes.com 1y ago

Google Introduces A.I. Agent That Aces 15-Day Weather Forecasts

ctc24
49pts7
www.prequel.co 1y ago

How fast is Iceberg on Snowflake?

ctc24
3pts0
cloud.google.com 1y ago

BigQuery Introduces Continuous Queries

ctc24
3pts0
www.nytimes.com 1y ago

Russia Releases Evan Gershkovich in Sweeping Prisoner Swap

ctc24
7pts5
github.com 1y ago

Snowflake Polaris Data Catalog Now on GitHub

ctc24
3pts0
www.wsj.com 2y ago

Musk Says He Will Move X and SpaceX Headquarters Out of California

ctc24
33pts78
www.reuters.com 2y ago

Amex to buy restaurant booking platform Tock for $400M

ctc24
4pts0
techcrunch.com 2y ago

Databricks launches LakeFlow to help its customers build their data pipelines

ctc24
9pts2
www.bloomberg.com 2y ago

Chamber of Commerce Sues to Block FTC's Non-Compete Ban

ctc24
2pts1
www.nytimes.com 2y ago

Dish Soap to Help Build Planes? Boeing Signs Off on Supplier's Method

ctc24
2pts0
blog.staysaasy.com 2y ago

Advice That I Can't Get Out of My Head

ctc24
2pts0
www.prequel.co 2y ago

Writing an Immutable Slice in Go

ctc24
8pts0
clearbit.com 2y ago

Clearbit enters agreement to be acquired by HubSpot

ctc24
11pts0
news.ycombinator.com 2y ago

Ask HN: What's the biggest red flag you've encountered during a hiring process?

ctc24
185pts262
remotion.com 3y ago

Pushing the limits of NSStatusItem beyond what Apple wants you to do

ctc24
9pts3
www.bloomberg.com 3y ago

The Parking Reform That Could Transform Manhattan

ctc24
3pts0
www.prequel.co 3y ago

SQL Maxis: Constructing JSON Objects in Athena Using SQL

ctc24
5pts0
www.reuters.com 3y ago

Reid Hoffman's new AI startup Inflection launches ChatGPT-like chatbot

ctc24
8pts3
treeandforest.substack.com 3y ago

Letter to a Burned Out Founder

ctc24
1pts0
www.databricks.com 3y ago

AI Functions: Integrating Large Language Models with Databricks SQL

ctc24
3pts0
www.prequel.co 3y ago

SQL Maxis: Why We Ditched RabbitMQ and Replaced It with a Postgres Queue

ctc24
628pts361
seattledataguy.substack.com 3y ago

What I Don't Want to See in the Data World in 5 Years

ctc24
1pts0
news.ycombinator.com 3y ago

Show HN: Ingest data from your customers (Prequel YC W21)

ctc24
40pts13
news.ycombinator.com 3y ago

Show HN: Syncing data to your customer’s Google Sheets

ctc24
103pts33
www.prequel.co 3y ago

Database drivers: Naughty or nice?

ctc24
94pts44

Prequel | Software Engineers (backend or full-stack) | NYC (ONSITE) | $170k-$210k | prequel.co

Prequel enables software companies to sync data to their customers' data environments, at massive scale. With the rise of agents, syncing data to customers' data environments is becoming table-stakes for a lot of software companies. We make that incredibly easy for them.

We're a team of four engineers based in NYC. We're cash-flow positive and growing fast. We're solving a number of hard technical problems that come with syncing hundreds of billions of rows of data every day with perfect data integrity: building reliable & scalable infrastructure, making data pipelines manageable without domain expertise, and creating a UX that abstracts out the underlying complexity to let the user share or receive data. We're powering this feature at companies like Stripe (Metronome), Gong, Iterable, and more.

Our stack is primarily Golang/K8s/Postgres/DuckDB/React/Typescript and we support deployments in both our public cloud as well as our customers' clouds. Due to the nature of the product, we work with nearly every data warehouse product and most of the popular RDBMSs.

Apply here: https://www.ycombinator.com/companies/prequel/jobs/VNoKffl-s... or email jobs (at) prequel.co and reference this post.

Prequel | Staff Frontend / Full Stack Engineer / SWE Intern | ONSITE in New York City | Full Time | https://prequel.co

- Prequel is the customer data access platform. We enable SaaS companies like LaunchDarkly, Gong, and LogRocket to make data accessible to their customers.

- Transferring trillions of rows between data stores every month.

- Revenue is up 50% since Jan 1.

- We're launching a new product which is quite frontend heavy, and want to bring a pro onboard who can act as tech lead for it. This is not your standard CRUD app -- a big part of it is building complex SDKs and APIs used by engineers at top-tier companies.

- Frontend in Typescript/React, backend in Go.

- Team is stacked, with alums from Stripe, GIPHY, Google, Flatiron Health, and more.

- We're based in NYC with our HQ in Chelsea.

Apply here: https://www.ycombinator.com/companies/prequel/jobs/wdjx5KE-f... or email careers (at).

Disagree on the data silo issue. There's a growing trend of SaaS providers making data available to their customers by feeding it back into their DWs. It started with the likes of Segment and Heap, and has now grown to include companies like Stripe, Salesforce, and Zuora to name a few. I'd wager that making data accessible is only going to become more table-stakes over time.

That's a bit of a strawman argument. Per the post, you can't leverage this on a read replica, it has to be run on primary. So you're going to stand up and manage a full new Postgres instance for this?

I'm sure there are many cases when that makes sense, but there are many cases when that's also overkill. An in-memory cache inside your server will give you better performance, and a lot of less infrastructure maintenance complexity.

Why wouldn't you simply use SQLite (or some other in-memory flavor of SQL) instead of hacking the main Postgres db and adding load to the primary instance?

The author makes a valid point that there's something nice about using familiar tooling (including the SQL interface) for a cache, but it feels like there are better solutions.

Prequel | https://prequel.co | Senior/Staff Software Engineer | Full Time | GoLang, Postgres, Typescript, React, K8s | $150k-$200k + equity | ONSITE in NYC

Prequel is an API that makes it easy for B2B companies to sync data directly to their customer's data warehouse, on an ongoing basis.

We're solving a number of hard technical problems that come with syncing tens of billions of rows of data every day with perfect data integrity: building reliable & scalable infrastructure, making data pipelines manageable without domain expertise, and creating a UX that abstracts out the underlying complexity to let the user share or receive data. We're powering this feature at companies like LogRocket, Modern Treasury, Postscript, and Metronome.

// Full job posting here -- https://prequelco.notion.site/Senior-Software-Engineer-Prequ...

// To apply -- email jobs@prequel.co and include [HN] in the subject line

I don't thing there's necessarily anything there. Microsoft might be burning money because they've decided that browser adoption and usage is worth it to them. It doesn't have to involve OpenAI in any way.

Prequel | https://prequel.co | Senior Software Engineer | Full Time | GoLang, Postgres, Typescript, React, K8s | $150k-$180k + equity | ONSITE in NYC

Prequel is an API that makes it easy for B2B companies to sync data directly to their customer's data warehouse, on an ongoing basis.

We're a tiny team of four engineers based in NYC. We're solving a number of hard technical problems that come with syncing tens of billions of rows of data every day with perfect data integrity: building reliable & scalable infrastructure, making data pipelines manageable without domain expertise, and creating a UX that abstracts out the underlying complexity to let the user share or receive data. We're powering this feature at companies like LogRocket, Modern Treasury, Postscript, and Metronome.

Our stack is primarily K8s/Postgres/DuckDB/Golang/React/Typsecript and we support deployments in both our public cloud as well as our customers' clouds. Due to the nature of the product, we work with nearly every data warehouse product and most of the popular RDBMSs.

We're looking for a full stack engineer who can run the gambit from CI to UI. If you are interested in scaling infrastructure, distributed systems, developer tools, or relational databases, we have a lot of greenfield projects in these domains. We want someone who can humbly, but effectively, help us keep pushing our level of engineering excellence. We're open to those who don't already know our stack, but have the talent and drive to learn.

// Full job posting here -- https://prequelco.notion.site/Senior-Software-Engineer-Prequ...

// To apply -- email jobs@prequel.co and include [HN] in the subject line

The "how do we make money" section on their website is interesting.

We have not made any decisions about how we may charge for the product in the future. That said, we believe your personal AI should always be directly aligned to your interests. We therefore think it's crucial that you are the only person who pays for it, so that will likely be our primary default business model. However, it’s still early days for this new technology. We also recognize that some people would rather access a free service and would prefer to see adverts in return.

I'm sympathetic to the idea that startups need to iterate on their business model to be successful. At the same time, this sounds a whole lot like "we promise that our business model doesn't rely on selling your data, unless we decide otherwise."

We can detect deleted rows for incremental transfers (and propagate those) if they're soft-deleted in the source, whether through a deleted_at column or a is_deleted column.

For now, we only support maintaining current state in the target.

Yup! We support all common cloud file storage as destinations (S3, R2, GCS, Azure Blob Storage) as well as vanilla SFTP servers.

That can be part of the value-add, though for on-prem deployments, we never touch the credentials ourselves.

Not to sound like a consultant, but there's three value-adds I'd call out:

1. Handling the dialect, types, and connection modalities of many different databases. This takes a lot of time to build and there's a lot of nuance that's non-trivial to work through.

2. Replicating data and guaranteeing data integrity + reliability. There's again a lot of nuance here, especially once you start considering that data is eventually consistent in most sources, that you want to transfer it as efficiently as possible, etc.

3. Providing a clean UX that end-customers can use out of the box, such that the end-customer experience is clean and intuitive. We spend a lot of time thinking about how it makes sense for people to connect their data, so that our customers don't have to.

edit: fmt

It depends -- mostly on whether the vendor (the company receiving the data) is comfortable requiring the source to map some fields.

For low volume cases, we can operate with zero mapping of fields. In those cases, we run every transfer as a full refresh.

If the volumes are higher, then we'll typically ask the source to expose a primary key and last_updated_at timestamp field. In those cases, we run incremental transfers. We use the last_updated_at to figure out what data to transfer, and the primary key to merge it into the destination table without creating dupes.

Not sure if I'm understanding the analogy. The way I usually describe it is that it's like Census / Hightouch, but it's offered by the vendor as a first-party feature.

Let's take Salesforce as an example. Let's say they want to pull in data from their customer's database -- maybe so that sales reps can keep track of how much volume the customer did in the last month -- instead of requiring the customer to instrument their code with Salesforce API calls. Salesforce could use this tool to connect directly to all of their customer's databases / data warehouses, regardless of whether they're Postgres, Snowflake, Clickhouse, etc.

As far as why it's non-trivial: you have to support a lot of different databases / data warehouses, which all have slightly different query languages, type systems, and optimizations. Then you've got to move the data reliably, dealing with things like eventual consistency etc. We feel like that's the reason this hasn't been built yet.

Yup, that would be another valid approach. We felt that the product experience was a lot cleaner / more in line with our general philosophy if we could write directly to the user's sheet, rather than ask them to import data from somewhere else, which is why we went this route.

Not particularly. A large portion of our customers who sync data to Google Sheets use a daily frequency, so the theoretical upper limit is close to a half million sheets being written to (per GCP project). We have other projects available that we can start using once this gets close to becoming an issue.

Very cool to see a walkthrough with actual benchmarks. Not entirely surprised that Parquet shines here. Another big advantage of Parquet over CSV is that you don't have to worry about data integrity. Perhaps less relevant for GIS data, but not having to think about things like string escaping is rather nice.

"It would be great to see data vendors deliver data straight into the Cloud Databases of their customers. It would save a lot of client time that's spent converting and uploading files."

Hear hear! Shameless plug: this is exactly what we enable at prequel.co. If there are any data vendors reading this, or anyone who wants easier access to data from their vendor, we're here to help.

edit: quote fmt