HN user

ericjang

2,242 karma

evjang.com https://twitter.com/ericjang11

Posts61
Comments398
View on HN
docs.google.com 1y ago

How to Build a Vibrant Technology Industry

ericjang
2pts0
www.1x.tech 1y ago

1X World Model

ericjang
3pts0
www.youtube.com 3y ago

The Humanoid Robot Dream (Asianometry)

ericjang
2pts0
twitter.com 3y ago

“Think about this step by step; the person giving you the problem is Yann LeCun”

ericjang
10pts0
evjang.com 3y ago

How can we make robotics more like generative modeling?

ericjang
61pts3
www.natolambert.com 4y ago

Job Hunt as a PhD in AI / ML / RL: How It Happens

ericjang
18pts2
evjang.com 4y ago

Ranking YC W22 companies with a neural net

ericjang
16pts15
ystrickler.medium.com 4y ago

The Ownership Crisis

ericjang
23pts18
evjang.com 4y ago

Leaving Google Brain

ericjang
2pts0
ai.googleblog.com 4y ago

Can Robots Follow Instructions for New Tasks?

ericjang
2pts0
lukemetz.com 4y ago

On the Difficulty of Extrapolation with NN Scaling

ericjang
21pts2
evjang.com 4y ago

To Understand Language Is to Understand Generalization

ericjang
91pts38
www.independent.co.uk 4y ago

China ready for ‘friendly relations’ with the Taliban

ericjang
7pts1
blog.evjang.com 5y ago

Sovereign Arcade: Currency as High-Margin Infrastructure

ericjang
1pts0
blog.evjang.com 5y ago

Science and Engineering for Learning Robots

ericjang
2pts0
blog.evjang.com 5y ago

Don't Mess with Backprop: Doubts about Biologically Plausible Deep Learning

ericjang
91pts53
blog.evjang.com 5y ago

Software and Hardware for General Robots

ericjang
42pts32
togelius.blogspot.com 5y ago

How many AGIs can dance on the head of a pin?

ericjang
1pts0
blog.evjang.com 5y ago

Chaos and Randomness

ericjang
1pts0
youtu.be 6y ago

Thinking While Moving: Deep Reinforcement Learning with Concurrent Control

ericjang
4pts0
blog.evjang.com 6y ago

Three Questions That Keep Me Up at Night

ericjang
1pts0
www.wired.com 6y ago

Alphabet's Dream of an 'Everyday Robot' Is Just Out of Reach

ericjang
10pts0
lilianweng.github.io 6y ago

Self-Supervised Representation Learning

ericjang
2pts0
blog.evjang.com 6y ago

Robinhood, Leverage, and Lemonade

ericjang
3pts0
medium.com 6y ago

The Waymo Open Dataset

ericjang
7pts0
blog.evjang.com 7y ago

What I Cannot Control, I Do Not Understand

ericjang
4pts2
blog.evjang.com 7y ago

Uncertainty: a Tutorial

ericjang
5pts0
openreview.net 7y ago

Large Scale GAN Training for High Fidelity Natural Image Synthesis

ericjang
11pts2
blog.evjang.com 8y ago

Bots and Thoughts from ICRA2018

ericjang
1pts0
arxiv.org 8y ago

Connections Between Automatic Differentiation and Delimited Continuations

ericjang
2pts0

Ensoul | Member of Technical Staff | San Francisco, CA | ONSITE | $180,000-$300,000 + Equity + Benefits https://ensoul.inc/t/099a40

Ensoul's mission is to accelerate robotics research. We're a team of frontier-lab scientists and roboticists working to accelerate robotics research and unlock The Great Robotics Buildout. Our customer is the robotics researcher.

The company page is light on details, so here's a bit more about me: https://evjang.com, I previously led AI at 1X Technologies and co-led some of the efforts at Google Robotics that led to SayCan, RT-1, etc.

intra-distribution generalization is also not well posed in practical real world settings. suppose you learn a mapping f : x -> y. casually, intra-distribution generalization implies that f generalizes for "points from the same data distribution p(x)". Two issues here:

1. In practical scenarios, how do you know if x' is really drawn from p(x)? Even if you could compute log p(x') under the true data distribution, you can only verify that the support for x' is non-zero. one sample is not enough to tell you if x' drawn from p(x).

2. In high dimensional settings, x' that is not exactly equal to an example within the training set can have arbitrarily high generalization error. here's a criminally under-cited paper discussing this: https://arxiv.org/abs/1801.02774

I am an AI researcher. Most actual AI researchers and engineers use very few of these tools - the only one being model providers like OpenAI API and public clouds (AWS, Azure, GCP). The rest of these are infra-centric tools that a16z is highly incentivized to over-inflate the importance of.

I entered the title question of the article into ChatGPT, as I was curious whether it could factually synthesize information on the Internet similar to how this article did. Notably, it missed the detail about keeping vegetation away

<chatgpt> Railway tracks have crushed stones, commonly known as ballast, alongside them for several important reasons. Here are the main reasons for using ballast on railway tracks:

Stability and Load Distribution: Ballast provides stability to the railway tracks and distributes the load of the trains evenly. The crushed stones interlock with each other and form a stable base for the tracks. This helps prevent the tracks from shifting or sinking under the weight of the trains.

Drainage: Ballast allows for effective drainage of water from the tracks. The open spaces between the crushed stones allow water to flow away from the tracks, preventing the accumulation of water. Proper drainage is essential to maintain the stability of the tracks and prevent damage caused by water-related issues like erosion or track bed saturation.

Track Alignment and Adjustment: Ballast allows for easy alignment and adjustment of the tracks. The crushed stones can be added or removed as needed to maintain the proper track alignment and ensure a smooth ride for trains. This flexibility is especially important during track maintenance or when new tracks are being laid.

Noise and Vibration Damping: Ballast helps in reducing noise and vibrations generated by passing trains. The layer of crushed stones absorbs and dampens the vibrations, minimizing the impact on the surrounding environment and nearby structures.

Track Maintenance: Ballast provides a protective layer for the tracks. It acts as a barrier between the track infrastructure and the underlying soil, preventing soil movement or erosion that could damage the tracks. Additionally, ballast makes it easier to detect and repair any issues with the track, such as loose fastenings or track bed settlement.

Overall, the use of ballast alongside railway tracks is crucial for maintaining the stability, alignment, drainage, and longevity of the tracks, ensuring safe and efficient train operations. </chatgpt>

1X (Formerly Halodi Robotics) | https://1x.tech/ | Multiple Roles | Onsite in Oslo and Bay Area

1X is an engineering and robotics company producing androids capable of human-like movements and behaviors. The company was founded in 2014 and is headquartered in Norway, with over 50 employees globally. 1X's mission is to create robots with practical, real-world applications to augment human labor globally. We recently announced a $23.5M Series A2 funding led by OpenAI (https://1xtech.medium.com/1x-raises-23-5m-in-series-a2-fundi...)

Open Positions:

- Senior DevOps Engineer | Full-Time | Onsite | Oslo, Norway https://1x.tech/#job-1112510

- Electronics Hardware Engineer | Full-Time | Onsite | Oslo, Norway https://1x.tech/#job-1079027

- Full-Stack AI Resident | Intern | Onsite | Bay Area, CA, USA https://1x.tech/#job-1079027

Google DeepMind 3 years ago

Jeff was the first author on the DistBelief paper - he's always been big on model-parallelism + distributing neural network knowledge on many computers https://research.google/pubs/pub40565/ . I really have to emphasize that model-parallelism of a big network sounds obvious today, but it was totally non-obvious in 2011 when they were building it out.

DistBelief was tricky to program because it was written all in C++ and Protobufs IIRC. The development of TFv1 preceded my time at Google, so I can't comment on who contributed what.

Google DeepMind 3 years ago

Jeff was very early on in the "just scale up the big brain" idea, perhaps as early as 2012 (Andrew Ng training networks on 1000s of CPUs). This vision is sort of summarized in https://blog.google/technology/ai/introducing-pathways-next-... and fleshed out more in https://arxiv.org/abs/2203.12533, but he had been internally promoting this idea since before 2016.

When I joined Brain in 2016, I had thought the idea of training billion/trillion-parameter sparsely gated mixtures of experts was a huge waste of resources, and that the idea was incredibly naive. But it turns out he was right, and it would take ~6 more years before that was abundantly obvious to the rest of the research community.

Here's his scholar page (H index of 94) https://scholar.google.com/citations?hl=en&user=NMS69lQAAAAJ...

As a leader, he also managed the development of TensorFlow and TPU. Consider the context / time frame - the year is 2014/2015 and a lot of academics still don't believe deep learning works. Jeff pivots a >100-person org to go all-in on deep learning, invest in an upgraded version of Theano (TF) and then give it away to the community for free, and develop Google's own training chip to compete with Nvidia. These are highly non-obvious ideas that show much more spine & vision than most tech leaders. Not to mention he designed & coded large parts of TF himself!

And before that, he was doing systems engineering on non-ML stuff. It's rare to pivot as a very senior-level engineer to a completely new field and then do what he did.

Jeff certainly has made mistakes as a leader (failing to translate Google Brain's numerous fundamental breakthroughs to more ambitious AI products, and consolidating the redundant big model efforts in google research) but I would consider his high level directional bets to be incredibly prescient.

Good idea, but actually you'll probably want to look at the option markets as well, to take account the exact strike price at which Musk wishes to take Twitter private. If you only look at spot price (i.e. TWTR stock), then you need some way to factor out all the other beliefs market participants have about Twitter (e.g. stat-arb correlations with NASDAQ index).

As of time of writing, the delta on a $55 TWTR call option expiring 2 months from now is 0.297, representing a ~30% probability it will be in-the-money. But you still need to subtract the probability that the share price gets there without Mr. Musk's help.

You can also google "merger arbitrage" on google scholar to find some more maths on the subject.

Based on that, this is essentially what you could call a "DudeRank Classifier" because as The Dude in the Big Lebowski says, "Yeah, well, that's just like, your opinion, man" :)

Yes, but isn't human VC investing already just a big DudeRank classifier?

I have immense respect for Moxie, who has spent time building experiments and tinkering with a new technology, and as a result has a take on it that highlights very different issues than what most of the predictable web3 flamewar centers around. It makes you really think about who is really qualified to discuss said technology.

I agree with your point that the operative word here is "earn", and not "play".

To play the devil's advocate: many crypto developers continue to build in the space (or have done it in the past) despite large drawdowns in token prices (denominated in USD). Would that constitute a sufficient signal that "there is something real there", distinct from the question of "is the valuation too high"?

When I ask pro-CCP people the question "what might change your mind", I often hear a variant of what you just said - "consider the motives (of the American hegemony), it explains our observations".

But that doesn't answer my question (2). Would you please answer it directly? Either there is some hypothetical evidence that would change your mind, or there is no hypothetical that could change your mind. Providing the affirmative to one or the other would help us understand whether you are open to changing your perspective at all.

For what it's worth, I'm happy to volunteer a hypothetical that would convince me to believe more in the narratives that you provided.

Thanks for taking the time to write this out. Two questions, followed by a response to one of your comments:

1) On the Uyghur issue. If what you say is true, why is it difficult to get CSPAN footage in Xinjiang?

2) Is there a hypothetical future event or evidence that would shift your viewpoint to believe the "western narrative"? If so, what is it?

Women being lifted out of poverty became "genocide through birthrate suppression".

It can simultaneously be true that the heavy-handed actions (as you put it) successfully lifted women out of poverty, and also violated their individual human rights. The words we use to describe these actions aren't mutually exclusive truths.

you seem to be implying that

1) a transaction between the US and Lithuania occurred in which Lithuania agrees to provoke PRC.

2) Lithuania provokes PRC (presumably the act involves the embassy opening).

However, the article you linked publicizes the U.S. trade support which seems to have occurred after (2). The article you linked does not permit one to draw the causal inference that (1)->(2)

I agree with your comment "In many languages, if you choose any of those supposed rules you can probably construct an algorithm to generate odd, but understandable words that defy that rule." - it comes many forms, from Goodhart's Law to the "hot dog vs. sandwich" debate.

I do mention this in my blog post - although I think Generalization is Language, I don't think it's possible to create a formal framework of language, for precisely because of "adversarial examples" that can be supplied for any formal definition.

Natural language itself, ignorant of formality, is able to account for these exceptions insofar as language is sufficient for people to convey a bare minimum of meaning. I am proposing to define language and generalization via the implicit understanding of large language models, in the same way you might use an image classifier to define "cat images" or "hot dogs"

My hope is that sufficiently rich language models obviate the need for a lot of robot-language grounding data.

LfP (https://learning-from-play.github.io/) was a work that inspired me a lot. They relabel a few hours of open-ended demonstrations (humans instructed to play with anything in the environment) with a lot of hindsight language descriptions, and show some degree of general capability acquired through this richer language. You can describe the same action with a lot of different descriptions, e.g. "pick up the leftmost object unless it is a cup" could also be relabeled as "pick up an apple".

That being said, the LfP paper stops short of testing whether we can improve robotics solely by only scaling language - a confounding factor and central to their narrative was the role of "open-ended play data". We do need some paired data to ground (language, robot-specific sensor/actuator modalities), but perhaps we can scale everything else with language only data.

Thanks to the pointer on the Andreas paper! This is indeed quite relevant to the spirit of what I'm arguing for, though I prefer the implementation realized by the Lu et al '21 paper.

Fair enough, I agree that if we really examine the comment "word as a discrete unit of meaning", the edge cases start to accumulate and the semantics rapidly break down. But barring things like prefixes/suffixes/modifiers/composite word characters in traditional Chinese, words are fairly discrete and generally regarded as the primary layer for expressing singular units of "meaning"

Thanks!

Words are considered a "discrete unit of meaning", i.e. 3/4 of a word doesn't really mean much. So words like "red" and "grass" are "standalone" in the sense that the mean something by themselves. I agree that words are very much related to each other, in the sense that you can combine them.

I was trying to draw a connection that the "disentangled representations" ML folks often talk about are but a special few-word case of grammars for combining distinct concept.

There's a reason much (maybe even most) anger on the left revolves around a perceived denial of the existence of privilege, and perhaps most of the anger on the right ultimately boiling down to a perceived denial of the existence of responsibility.

Interesting. How do you use this "unified" framework then, to come up with your own political opinions? (as opposed to, say, post-modernism variants like critical race theory)