HN user

gdcohen

135 karma
Posts72
Comments47
View on HN
clockwork.io 4mo ago

Decoding GPU Efficiency: The FLOPs Fallacy

gdcohen
3pts0
community.aws 2y ago

LLMs fought 314 Street Fighter matches. Here's who won

gdcohen
1pts0
www.nytimes.com 2y ago

One Arm or Two? How You Get Vaccinated May Make a Difference

gdcohen
1pts0
fablestudio.github.io 3y ago

A full episode of South Park generated by AI

gdcohen
144pts75
www.nytimes.com 3y ago

Airbnb Sues New York City over Limits on Short-Term Rentals

gdcohen
3pts2
news.airbnb.com 3y ago

Airbnb and 3 local Hosts filed separate lawsuits against the City of NY

gdcohen
2pts6
www.reuters.com 3y ago

Tesla recalls nearly 1.1M U.S. vehicles to update window reversing sw

gdcohen
2pts0
www.spiceworks.com 3y ago

ML Is a Game Changer for the Incident Management Lifecycle

gdcohen
1pts0
www.atlasobscura.com 4y ago

10% of the internet is encrypted by lava lamps

gdcohen
3pts1
www.economist.com 4y ago

The world’s most liveable cities

gdcohen
2pts0
venturebeat.com 4y ago

ML can solve root cause application failures for engineering and support

gdcohen
1pts0
sdtimes.com 4y ago

Using GPT-3 for root cause incident summarization of incidents

gdcohen
3pts0
www.lifewire.com 5y ago

Why the Internet Is Vulnerable

gdcohen
1pts0
www.zebrium.com 5y ago

A new machine learning approach for your Elastic Stack

gdcohen
2pts0
www.zebrium.com 5y ago

Zebrium vs. Elastic Machine Learning

gdcohen
1pts0
www.zebrium.com 5y ago

Try ML-driven RCA using a cloud-native microservices demo app

gdcohen
1pts0
www.linkedin.com 5y ago

Visualized Rise of Tesla

gdcohen
1pts0
www.zebrium.com 5y ago

Using GPT-3 for plain language incident root cause from logs

gdcohen
4pts0
www.cncf.io 5y ago

Machine learning for K8s logs and metrics

gdcohen
2pts0
www.zebrium.com 5y ago

Testing ML Incident Detection using a cloud native microservices demo app

gdcohen
1pts0
appleinsider.com 5y ago

Apple versus Epic Games 'Fortnite' App Store saga – the story so far

gdcohen
1pts0
www.zebrium.com 5y ago

Virtual tracing: Simpler alternative to distributed tracing for troubleshooting

gdcohen
1pts0
www.zebrium.com 5y ago

Single Sign-On with OAuth

gdcohen
1pts0
news.ycombinator.com 6y ago

Ask HN: What is the coolest prize to give to a Developer / Technical Geek?

gdcohen
1pts0
www.youtube.com 6y ago

Stephen Colbert: The Newest Zealander Visits PM Jacinda Ardern

gdcohen
1pts0
www.zebrium.com 6y ago

Zebrium and Grafana =

gdcohen
1pts0
www.forbes.com 6y ago

Forbes AI 50: America’s Most Promising Artificial Intelligence Companies

gdcohen
1pts0
www.zebrium.com 6y ago

You've Nailed Incident detection, what about Incident Resolution?

gdcohen
1pts0
www.zebrium.com 6y ago

The Problems with Log Management

gdcohen
1pts0
www.zebrium.com 6y ago

Busting the Browser's Cache

gdcohen
1pts0

I agree. If you can reliably recover from failures by migrating instead of restoring from checkpoints, then the need to checkpoint frequently becomes far less important and reduces the overhead of taking checkpoints. Similarly, a lot of users are experimenting with asynchronous checkpoints to reduce the blocking penalty, but there are big tradeoffs there with DRAM usage. So being able to take checkpoints far less frequently gives you more flexibility on the type of checkpointing you use as well.

Gavin from Zebrium here. We've found that if only you somehow knew what you were monitoring for in logs, they can be a great source of detecting (and then describing) the long tail of unknown/unknowns (failure modes with unknown symptoms and causes). Our approach is to be able to find these patterns in near real-time using ML. This blog by our CTO explains the tech with some good examples: https://www.zebrium.com/blog/is-autonomous-monitoring-the-an....

Gavin from Zebrium here. Completely concur with #1. We are big advocates of writing good logs and not having to worry about structured vs unstructured (and even if you structure your logs, you'll still probably have to deal with unstructured logs in third party components).

Our approach to deal with logs is to use ML to structure them after the fact (and we can deal with changing log structures). You can read about it in a couple of our blogs like: https://www.zebrium.com/blog/using-ml-to-auto-learn-changing... and https://www.zebrium.com/blog/please-dont-make-me-structure-l....

Gavin from Zebrium here. Again, we understand this concern. All data is encrypted in-flight and at rest and we have a lot of security controls in place (see our website). We also have the option of a dedicated VPC assigned to a single customer. Beyond this, we have a unique capability: since our machine learning structures and types everything, we have a feature that lets us hash (or delete) any sensitive information in logs (and we can record match to find other places this info occurs that you might not even be aware of). Not dodging the fact that we're SaaS, just pointing out what we do.

Also that the article talked about the colon as having a longer pause than a semicolon (different to how it is used today).

There is a trend towards not testing at all! Instead builds are deployed straight to production or canary (mini-subset of production) and then very carefully and closely monitored. If a problem is uncovered, a rollback is performed. If canary is done well, then the problem can be caught before it has widespread impact.

Sorry if this is pointing out the obvious, but the title is not a typo! We discuss an approach of using machine learning to automatically post-structure logs, rather than having to manually pre-structure them.

Implementing, and practicing, a framework like this is an incredible way to build a culture that makes every person feel like his/her opinion can count and will be listened to.