it wasn't as simple as it says
mind elaborating? we built loki for some pretty massive scale but I've always tried to make it work at super small scale to. what went wrong?
HN user
VP Technology, Grafana Labs; previously Founder of Kausal, Director Software Engineering at Weaveworks, SRE at Google, Founder of Acunu and engineer at XenSource.
it wasn't as simple as it says
mind elaborating? we built loki for some pretty massive scale but I've always tried to make it work at super small scale to. what went wrong?
Hi! Tom from Grafana Labs here, super excited to welcome the Pyroscope team and can’t wait to see what they achieve.
Continuous profiling is a next big thing IMO - easier to get started with than distributed tracing and delivers immediate value.
Its very much our aim to make this mix of self-hosted and cloud services as easy as going all-cloud; but I agree we're not quite there yet.
Do you mind if I ask what isn't super-easy about linking self-hosted loki search queries with SaaS-Prometheus? You should be e.g. able to add a Prometheus data source to your local Grafana (or securely expose your Loki to the internet and add a Loki data source to your Cloud Grafana)
I don't think so! I think thats being used in Tempo, but I'm not sure.
Sounds like Thanos is working well for you, so in your position I wouldn't change anything.
There are a bunch of other reasons why people might choose Mimir; perhaps they have out grown some of the scalability limits, or perhaps they want faster high cardinality queries, or a different take on multi-tenancy.
Do remember Cortex (on which Mimir is based) predates Thanos as a project; Thanos was started to pursue a different architecture and storage concept. Thanos storage was clearly the way forward, so we adopted it. The architectures are still different: Thanos is "edge"-style IMO, Mimir is more centralised. Some people have a preference for one over the other.
We tried to address this question on the Q&A blog post: https://grafana.com/blog/2022/03/30/qa-with-our-ceo-about-gr...
It doesn't have to mean the end for Cortex, but others will have to step up to lead the project. We've tried to put other maintainers in place to kick start this.
(Tom here; I started the Cortex project on which Mimir is based and lead the team behind Mimir)
Thanos is an awesome piece of software, and the Thanos team have done a great job building an vibrant community. I'm a big fan - so much so we used Thanos' storage in Cortex.
Mimir builds on this and makes it even more scalable and performance (with a sharded compactor and query engine). Mimir is multitenant from day 1, whereas this is a relatively new thing in Thanos I believe. Mimir has a slightly different deployment model to Thanos, but honestly even this is converging.
Generally: choosing Thanos is always going to be a good choice, but IMO choosing Mimir is an even better one :-p
I agree! Which is why I put one in the blog post ;-) https://grafana.com/blog/2022/03/30/announcing-grafana-mimir...
Yes! It’s something we’ve be mulling for a while, and I was just talking to one of the PMs about it this morning. This year for sure I hope.
For now, yes. Long term we're trying to offer everything we do both on premise and in the cloud. It's a bit tricky, so we can't say when....
Grafana Labs | Full-time | 100% remote (world) | Software Engineer
Grafana Labs is hiring! Come work on Grafana, Prometheus, Cortex, Loki, Tempo and more - lots of opensource source, with both a SaaS and an enterprise team. We use Golang, JS, Typescript, Kubernetes, Jsonnet, Tanka, CUE and more.
We're growing fast, have lots of happy customers and many exciting projects in the pipeline. We need good engineers to help us take Grafana and Prometheus to the masses.
PRs welcome!
That’s our motivation and experience over the last 6 months running it internally at Grafana Labs. Much cheaper and easier to operate.
Not to bash Jaeger though, its more powerful than tempo in that it allows you to search for traces. Tempo is about integrating with Grafana, Loki and Prometheus for finding traces.
Writes are batches up and committed asynchronously to s3 - this should add much if any latency to your services.
Super excited to launch Tempo, really starting to round out a prometheus-first observability stack. Kudos to Joe, Annanay and rest of the team!
Tom from Grafana Labs here! Super proud of this release, especially the tracing support. Let me know if you have any questions?
There are a bunch of different solutions out there; Thanos, Influx, federated Prometheus etc.
The local Cortex storage works pretty well but we have a very high bar for production worthiness. Right now I'd recommend using Bigtable of DynamoDB, and if you're on premise Cassandra. In the future the block storage will allow you to run minio.
Thats the "microservices" mode - you can run it as a single process and the architecture becomes super boring.
Its like looking at the module interdependencies of reasonably large piece of software; of course its going to look complicated.
wrapping prometheus and giving you that production readyness that they're claiming the OSS project won't give you out of the box
No! Prometheus is and has been production ready for many years. Cortex is a clustered/horizontally scalable implemention of the Prometheus APIs, and Cortex has just gone production ready. Sorry for the confusion.
Yes, Prometheus is an implementation - the HN text has a limited number of words, so I thought "Prometheus implementation" conveyed the fact Cortex was trying to be a 100% API compatible implementation of Prometheus, but with scalability, replication etc
I hope so! Goutham is apply for incubation as we speak..
Sure, check out this talk from PromCon I did with Bartek, the Thanos author: https://grafana.com/blog/2019/11/21/promcon-recap-two-househ...
Hi! Tom, one of the Cortex authors here. Super proud of the team and this release - let me know if you have any questions!
Tanka & jsonnet-bundler also work really well with Prometheus monitoring mixins, meaning we bundle up and share almost all the internal monitoring that we use at Grafana Labs to monitor our massive Cortex, Loki, Metrictank and Kubernetes deploys.
Cortex author here (Tom Wilkie). Great post that honestly highlights the differences between these systems - thank you!
The biggest take home here - and the first thing the post mentions - is the a single HA pair of Prometheus servers is enough for 80-90% of people. TLDR you probably don’t need Cortex (or Thanos, etc)...
...unless you run multiple, segregated networks (regions). Then something like Thanos (or Cortex) is useful - not for a the scale argument, but because you need a way to “federate” queries and get that global view. IMO!
Do the anonymous auth settings achieve this? https://grafana.com/docs/auth/overview/#anonymous-authentica...
We do that at Grafana Labs are part of Grafana Cloud...
Yes! Loads of idea floating around and I think Jon has been working on some improvements for 6.6. Its going to be a big project in 2020 for sure.
(Its something I'm really keen on myself - We use jsonnet internally to version control our dashboards and load them into config maps in our kubernetes clusters)
We want to use Jaeger as a datasource and show traces in Grafana - its all a bit early right now, but come see us at KubeCon and we'll have something to show!
Come see a demo at KubeCon! (right, David? ;-)