HN user

mateiz

216 karma
Posts9
Comments10
View on HN

Don't MLflow Projects exactly meet this use case? A project lives in a Git repo, which can include both code and data, and specifies its software environment (currently Conda but will eventually also support Docker): https://www.mlflow.org/docs/latest/projects.html. You can then run it wherever you want to run code: CI system, Kubernetes, cloud, etc. The reason MLflow doesn't force people to use Projects is because many users like to develop ML in notebooks, but we definitely expect engineering teams to use it with Projects.

While MLflow doesn't submit jobs to Kubernetes for you, it should be possible to integrate it with your favorite scheduler to do that. MLflow is designed to accept experiment results from wherever you are running your code, so you can just submit an "mlflow run ..." command to Kubernetes and have it report results to your tracking server.

Yup, this is a great explanation. Cuckoo hashes actually do much better in terms of load if you use buckets with multiple items, as we did here. The classical one with one item per location can get "stuck" with unresolvable cycles between elements at 50% load, but when you have these buckets, there are many more ways to resolve collisions, which lets you greatly increase the load. 99% is extreme but definitely doable. Writing a fuller post on this is definitely a good idea -- we may do it in the future.

Matei Zaharia (one of the PIs on DAWN) here. Snorkel, MacroBase and ASAP are already being used in production at several companies, and we intend to continue publishing everything as open source. We only started this lab a year ago, so a lot of the projects listed are still new.

Databricks -- San Francisco -- https://databricks.com

* Software Engineer (ONSITE)

* Software Engineer Intern (ONSITE)

* Product Manager (ONSITE)

Databricks was founded in 2013 by the team that started Apache Spark, meaning you might not only use Spark but also get to work on it :). We provide a cloud data processing platform based on Apache Spark used by customers including top 5 banks, healthcare and media companies. We've also done some really cool technical stuff, such as setting the 2014 GraySort record (http://www.wired.com/2014/10/startup-crunches-100-terabytes-...).

We are hiring engineers in the following areas:

* Backend (JVM, AWS, database engine)

* Frontend (React, D3)

* Machine learning

List of available positions: https://databricks.com/company/careers