HN user

ayw

438 karma

scale.com

Posts76
Comments65
View on HN
alexw.substack.com 2y ago

What I Learned in 2023

ayw
2pts0
scale.com 2y ago

How to Fine-Tune Llama2

ayw
2pts1
scale.com 3y ago

Why is ChatGPT so good? RLHF

ayw
1pts0
alexw.substack.com 3y ago

The AI War and How to Win It

ayw
11pts6
scale.com 3y ago

Show HN: Scale Forge – AI-generated product imagery in seconds

ayw
7pts1
www.notboring.co 5y ago

Scale: The Five-Year-Old, $7.3B Stripe for AI

ayw
7pts0
alexw.substack.com 5y ago

Information Compression – Why the wrong thing happens

ayw
1pts0
alexw.substack.com 5y ago

Information Compression

ayw
5pts0
techcrunch.com 5y ago

Scale AI raises at $3.5B valuation and breakeven

ayw
4pts1
alexw.substack.com 5y ago

Hire people who give a shit

ayw
5pts2
www.bloomberg.com 5y ago

Amazon Head of Robotics Joins Scale as CTO

ayw
17pts1
onezero.medium.com 5y ago

A Dataset Has Been Passing Its Bias to Algorithms for Almost Two Decades

ayw
1pts0
onezero.medium.com 5y ago

The Troubling Legacy of a Biased Dataset

ayw
19pts0
scale.com 6y ago

How to scale ML for annotation systems

ayw
1pts0
medium.com 6y ago

Active Learning with PyTorch Using Scale

ayw
2pts0
medium.com 6y ago

Labeling-in-the-Loop Deep Learning with PyTorch and Scale AI

ayw
6pts0
scale.com 6y ago

Interview with Christian Szegedy, discoverer of adversarial examples

ayw
1pts0
scale.com 6y ago

How to label 1M GPT-2 phrases per week

ayw
4pts0
scale.com 6y ago

How to label 1M data points per week

ayw
3pts0
scale.com 6y ago

Interview with Jeremy Howard (Fast.ai) on Future of AI

ayw
19pts1
scale.com 6y ago

Is Elon Musk Wrong about Lidar? A Quantitative Analysis

ayw
3pts0
scale.com 6y ago

Is Elon Musk Wrong about Lidar? A Quantitative Study

ayw
10pts3
scale.com 6y ago

Interview with Jeremy Howard (Fast.ai) and Alexandr Wang (Scale AI)

ayw
3pts0
scale.com 6y ago

Interview with Jeremy Howard, Founder at Fast.ai on Future of AI

ayw
2pts0
scale.com 6y ago

Scale AI (YC S16) raises $100M at $1B+ valuation to go beyond AI data labeling

ayw
4pts0
scale.ai 7y ago

Open Datasets for Autonomous Driving

ayw
74pts0
scale.ai 7y ago

Life of a Data Labeler

ayw
96pts28
scale.ai 7y ago

Interview with Christian Szegedy, Discoverer of AI Adversarial Examples [video]

ayw
28pts6
scale.ai 7y ago

Interview with Christian Szegedy, Discoverer of Adversarial Examples in AI

ayw
11pts0
scale.ai 7y ago

Interview with Christian Szegedy, Google AI, Inventor of Inception, BatchNorm

ayw
3pts0
AWS Data Exchange 7 years ago

At Scale (scale.com), we strongly believe that the “open-source” alternative to this is pretty critical.

We’ve built this index for autonomous driving datasets (https://scale.com/open-datasets) and are building that out for other domains right now.

Open source data has been a pillar to progress in ML (starting with ImageNet). It should continue to be the case that data that enables researches is sufficiently democratized.

Strongly agree with this.

One thing I’ll mention is that this is true both at the very early stages of a ML project, and even when an ML project is scaled up and in production. Oftentimes, the data pipeline is the true way in which a model will improve versus anything else, so it’s pretty critical that these data pipelines are setup to get an initial dataset but also to scale properly.

It’s one reason I started Scale (scale.com). It was viscerally clear that the real bottleneck to ML was getting the needed data, and in our case, annotating that data appropriately. It is very heartening to hear it echoed in this whole thread that data is very clearly what “matters” for ML.

1. While stereo depth estimation would work in theory, none of the self-driving cars actually have camera configurations that allow for stereo depth estimation (see here: https://electrek.co/wp-content/uploads/sites/3/2016/10/tesla...)

2. Stereo depth estimation is quite unreliable in practice because it requires you to match up pixels between the two images very precisely (1-2px difference can be a large disparity in distance), so it is not reliably used.

We have a rule when hiring people—we look for people with an internal locus of control. Roughly speaking, this means people who believe they have control over outcomes in their life, as opposed to external forces beyond their control.

It’s a small thing, but it’s surprising easy to spot once you look for it. And it really matters—startups are the business of building something from nothing. You need people who believe they can bend the earth.

The biggest change is your jobs goes from doing things (which makes sense) to building an incredible team that can do things (which is a more unintuitive job). In the limit, it’s always a people business.

Overcome many challenges, but per my last answer, building a team of the best people has been the most important and most challenging. That, and learning how to do sales ;)

Too many mentors. People in Silicon Valley are incredibly helpful. To name a few: Dan Levine, Mike Volpi, Nat Friedman, Adam D’Angelo, Ilya Sukhar, Jonathan Swanson, Albert Ni, Jeff Arnold, Charlie Cheever, and Drew Houston to name a few. I’m very very lucky.

Self-driving is one of many applications of AI/ML to the real world, each of which likely requires high-quality labeled data to truly be production-ready. This includes other robotics, self-checkout like Amazon Go, natural language understanding, and more.

Second, self-driving as a problem space will need labels for a very long time. In an application where (1) verifiable model performance is paramount, and (2) the models need to be extremely robust for cars to be safe, the need for labeled data is only magnified.

Re 1—It has been a bit of annoyance growing up (for example, Google autocorrects "Alexandr Wang" to "Alexander Wang"), but we run different circles ;)

Re 2—As with most companies working on ML these days, our stack is not fully proprietary. We don't take too strong an opinion on ML framework and use both Tensorflow and Pytorch currently. We generally use neural network architectures from the literature and then iterate on top of them to suit our unique problem requirements.

We do use AI and ML to help making the labeling process more efficient, but you are correct we do have scaled human insight that ensures very high quality.

One difference from "Not Hotdog" is that our data is used to power the algorithms of other AI/ML companies like OpenAI, Waymo, Lyft, etc., so it's imperative that we have impeccable quality. That necessitates humans to ensure accuracy, particularly in safety-critical applications like self-driving cars.

Hi, I'm the CEO of Scale.ai.

This comment does not represent the company's viewpoint, and cardigan is not speaking on behalf of Scale.

We are very excited to have been able to work with Lyft in open-sourcing this dataset and advancing the research community. We are also very grateful to Lyft for choosing to leverage our point cloud viewer and have credited the annotations to us on their launch page.

You just need to read the oplog, so it only needs to track your saves.

In general, you probably should have at least something in your stack which reads all changes from your DB, at the very least for backup reasons.

For better or for worse, MongoDB tends to be easier for developers move quickly, so it ends up getting adopted quite a bit. This is more about how to deal with it after it's already in your stack.

Hey everyone! I'm Alex, CEO and co-founder of Scale. One of the biggest bottlenecks to development in perception and vision for robotics and self-driving companies has been the ability to label 3D data. The ability to label LIDAR, camera, and radar data together has been able to massively accelerate our customers' timelines.

We've worked with a number of self-driving companies like GM Cruise, nuTonomy, Voyage, Embark, and more to build high-quality training datasets quickly leveraging our API. Scale is the perfect platform for this work—we're focused on really high quality data produced by humans via API.