HN user

Merick

124 karma
Posts52
Comments19
View on HN
medium.com 1y ago

Live Shard Data Archive: Export and Ingestion to StarRocks for Validation

Merick
2pts0
engineering.grab.com 1y ago

Building a Spark observability product with StarRocks

Merick
1pts1
starrocks.medium.com 1y ago

Demandbase Ditches Denormalization by Switching Off ClickHouse

Merick
1pts0
medium.com 1y ago

Why Did Databricks Open-Source Unity Catalog?

Merick
19pts2
medium.com 1y ago

Delivering Faster Analytics at Pinterest

Merick
1pts0
starrocks.medium.com 2y ago

Rockset Is Acquired by OpenAI. What Does It Mean for Its Users?

Merick
6pts0
medium.com 2y ago

Testing the High-Performance StarRocks Database

Merick
1pts0
starrocks.medium.com 2y ago

Why Starburst's Icehouse Is a Bad Bet

Merick
3pts0
celerdata.com 3y ago

What You Should Know Before Using ClickBench

Merick
3pts0
celerdata.com 3y ago

ChatGPT is now finding bugs in databases

Merick
212pts189
celerdata.com 3y ago

What You Should Know Before Using ClickBench

Merick
1pts0
www.starrocks.io 3y ago

StarRocks vs. ClickHouse: The Quest for Analytical Database Performance

Merick
3pts0
celerdata.com 3y ago

What You Should Know Before Using ClickBench

Merick
2pts0
starrocks.com 4y ago

Real Time Analytics: What, So What, and Now What?

Merick
17pts2
kyligence.io 6y ago

OLAP Analytics Is Dead. Really?

Merick
1pts0
kyligence.io 7y ago

Use Python for Data Science with Apache Kylin

Merick
1pts0
kyligence.io 7y ago

Augmented Analytics: The Future of OLAP

Merick
2pts0
kyligence.io 7y ago

After Salesforce Acquires Tableau, What’s Next?

Merick
3pts0
azure.microsoft.com 7y ago

Announcing General Availability of Apache Hadoop 3.0 on Azure HDInsight

Merick
1pts0
kyligence.io 7y ago

Kyligence Raises $25M Led by Coatue

Merick
1pts0
kyligence.io 7y ago

Apache Kylin – Yet Another Hadoop Query Engine?

Merick
1pts0
gojurnal.com 10y ago

Digital Over-Packaging

Merick
1pts0
gojurnal.com 10y ago

The Four Reasons Why To-Do Lists Fail

Merick
1pts0
gojurnal.com 10y ago

Episode 10: A Tale of Two Searches

Merick
1pts0
gojurnal.com 10y ago

Ad-Supported Chaos

Merick
1pts0
gojurnal.com 10y ago

Episode 9: Using Recommendations to Hide Instead of Help

Merick
1pts0
gojurnal.com 10y ago

Why We’ve Gone Back to Cave Painting

Merick
1pts0
gojurnal.com 10y ago

How Do You Decide?

Merick
1pts0
gojurnal.com 10y ago

Episode 8: Who’s Ad-Vising You?

Merick
1pts0
gojurnal.com 10y ago

A Tale of Two Searches

Merick
1pts0

YES!! Normalization was the BEST PRACTICE since DBMS was invented in like the 70s, and somehow, we just totally forgot about it the past 10 years ago for OLAP. 10x more expensive storage, impossible schema evolution because of data backfilling, the extra pipelines you have to build and maintain, and the cost scales with the data size and your business growth. I was just talking with a buddy of mine about this and how to run JOINs on the fly without denormalization pipelines using StarRocks in this case.

I'll always remember Rockset for their ridiculous comparison page: https://rockset.com/real-time-analytics-comparison/

Maybe they should rename it to their migration options page. Or maybe I'll just ask ChatGPT what the best alternative is...

Still, pretty useful stuff, but it also feels like Rockset had been moving a little too slowly in recent years, but congrats to them on finding a new home.

Apache Kylin was actually China’s very first top-level Apache project. It did come out of eBay, but the work all originated in China. It’s a really cool solution to query acceleration.

You can learn about it on the community page here: http://kylin.apache.org/

It’s pretty popular across China, and I’ve seen it come up a bunch in Europe/South America, but in the U.S. it’s pretty new to a lot of folks.

Very cool to see this! Kylin has been a super fun project to b a part of. A unique approach to OLAP/query acceleration that's pretty much the best way to deal with huge volumes of data.

It's great to see them collaborating with Women Who Code to get the word out and grow our open source family. If anyone is interested in learning more about the project, check us out here: http://kylin.apache.org/

Thanks for sharing this! Apache Kylin is an awesome project. It has been around for a few years and has been getting a lot of attention in Europe and Asia, but still seems to be relatively unknown by a lot of folks I talk to in the US.

If people are interested in getting involved with the community, or just want to check Kylin out, you can find the page for the project here: http://kylin.apache.org/

Very cool, thank you for sharing. I've been a part of the Kylin community for over a year now and it has been a great experience. The project is really pretty cool. I know some people think of it as just another OLAP engine - but it is way more than that.

If you're part of a team working with huge datasets, take a look at it, and if you're looking for a great open source project to get involved with, join us!

Well, for one I’d really love to see even more improvements in recognizing supply usage. As near as I can tell there is nothing out there that’s doing the job. Most inventory management still requires someone in the lab to track and report. There’s a lot of barcode-based solutions where folks scan stuff out, but people are so busy that this either never gets pick-up in the lab or not enough people commit to it and you might as well not even bother with the system.

Something that could track the removal/use of products without needing so much handholding or constant verification would save so much time and avoid so many urgent orders.

Also, something that ties together ELN work and related supplies alongside inventory, and that adjusts inventory levels and helps with reorder alerts at particular thresholds. That would smooth a lot of the pain around planning. Like if something could alert me that a reagent/consumable/chemical was out or low before I was about to commit to a project it’d sure avoid a lot of urgent orders which can be a huge cost driver due to shipping costs or having to pay more than average because only one vendor doesn’t have something on backorder or they’re the only ones who can ship in time to keep things moving.

HappiLabs is right that pricing is all sorts of messed up across the sciences. That anecdote about labs in the same building having different pricing is all an all too common story I hear from plenty of folks in the lab.

I see time getting wasted every week, and not just from folks managing the lab. Often, as soon as the weekly planning meeting ends even scientists will get out their laptops and start bouncing between VWR, Fisher, Sigma, etc's. websites to figure out where the best prices are, how long it'll take to get what they want, and shipping costs. The fun part is when the stuff arrives and they realize their inventory count was off and they're actually out of something they thought they had in stock...and then it's back to the same old websites.

It's good to see YC investing more towards solving this problem. I know Quartzy (YC S11 - https://www.quartzy.com/) has been working on this problem too, a bunch of labs I work with are using them for this same issue. They have a ton of partnerships with suppliers which has allowed them to consolidate all those vendors into one place. This has solved a lot of that price hunting, but I think there's plenty of room to expand with more automated inventory management since that's really at the heart of a ton of supply issues.

Honestly, thinking about it, there's a lot of stuff in the lab that automation would help with beyond the experimentation part which gets the majority of attention at the moment.

Working with labs, I've found it extremely eye-opening how archaic the supplier relationship is. A big challenge appears to be with the sales model and its reliance on contracts, opaque pricing, and exclusive relationships.

This is, of course, made even more complicated because not all supplies are created equal. Even simple plastics can have slight variations that are invisible without further testing which can impact experiments. You have to be careful, and many of the scientists I work with are superstitious when it comes to buying products, not wanting to risk their experiments.

An alternative I see a ton of labs using these days is Quartzy (YC S11) (https://www.quartzy.com/), which also provides great price alternatives, and also evaluates against other parameters people care about when it comes to their experiments. They have an entire team dedicated to vetting these products so scientists can order with peace of mind. Quartzy also provides some very useful supply ordering and inventory tools that are a huge boon for lab ops.

There's a lot of opportunities for disruption in the lab supply space. It's definitely needed.