HN user

randomwalker

11,862 karma

Princeton prof: https://twitter.com/random_walker

Research: https://www.cs.princeton.edu/~arvindn/

Posts231
Comments427
View on HN
www.normaltech.ai 9d ago

What will be left for us to work on?

randomwalker
180pts265
www.normaltech.ai 13d ago

Up the Stack: How AI's Escape from the Commodity Trap Risks Enterprise Lock-In

randomwalker
3pts0
www.normaltech.ai 2mo ago

Did Google's AI agents build an operating system for $916?

randomwalker
4pts0
cruxevals.com 3mo ago

Open-world evaluations for measuring frontier AI capabilities [pdf]

randomwalker
2pts0
www.normaltech.ai 4mo ago

Towards a science of AI agent reliability

randomwalker
1pts0
cset.georgetown.edu 5mo ago

When AI Builds AI – Findings from a Workshop on Automation of AI R&D [pdf]

randomwalker
1pts0
static1.squarespace.com 8mo ago

The Longitudinal Expert AI Panel: Understanding Expert Views on AI [pdf]

randomwalker
1pts0
www.arxiv.org 9mo ago

Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation

randomwalker
1pts0
www.whitehouse.gov 12mo ago

America's AI Action Plan [pdf]

randomwalker
11pts0
www.aisnakeoil.com 1y ago

Could AI slow science? Confronting the production-progress paradox

randomwalker
2pts0
knightcolumbia.org 1y ago

AI as Normal Technology

randomwalker
239pts92
www.nature.com 1y ago

Why an overreliance on AI-driven modelling is bad for science

randomwalker
1pts0
www.aisnakeoil.com 1y ago

Is AI progress slowing down?

randomwalker
5pts1
knightcolumbia.org 1y ago

We Looked at 78 Election Deepfakes. Political Misinformation Isn't an AI Problem

randomwalker
5pts0
arxiv.org 1y ago

Inference Scaling FLaws: The Limits of LLM Resampling with Imperfect Verifiers

randomwalker
3pts0
www.aisnakeoil.com 1y ago

Is the UK's liver transplant matching algorithm biased against younger patients?

randomwalker
93pts62
arxiv.org 1y ago

Core-Bench: Computational Reproducibility Agent Benchmark

randomwalker
1pts0
www.aisnakeoil.com 1y ago

AI companies are pivoting from creating gods to building products

randomwalker
133pts195
www.aisnakeoil.com 2y ago

AI Agents That Matter

randomwalker
35pts10
arxiv.org 2y ago

AI Agents That Matter

randomwalker
4pts0
www.aisnakeoil.com 2y ago

Scientists should use AI as a tool, not an oracle

randomwalker
124pts106
www.aisnakeoil.com 2y ago

AI safety is not a model property

randomwalker
2pts0
www.aisnakeoil.com 2y ago

AI safety is not a model property

randomwalker
3pts0
crfm.stanford.edu 2y ago

On the Societal Impact of Open Foundation Models [pdf]

randomwalker
2pts0
www.aisnakeoil.com 2y ago

Will AI transform law? The hype is not supported by current evidence

randomwalker
2pts0
www.aisnakeoil.com 2y ago

Generative AI's end-run around copyright won't be resolved by the courts

randomwalker
4pts2
www.aisnakeoil.com 2y ago

Model alignment protects against accidental harms, not intentional ones

randomwalker
1pts0
www.aisnakeoil.com 2y ago

What the executive order means for openness in AI

randomwalker
2pts0
crfm.stanford.edu 2y ago

The Foundation Model Transparency Index

randomwalker
47pts16
www.cs.princeton.edu 2y ago

Evaluating LLMs Is a Minefield

randomwalker
3pts0

Yes, we're aware! Fortunately our book is not a broad indictment of AI :) And none of our claims are premised on tasks people can do remaining out of reach for AI. More here: https://www.normaltech.ai/p/faq-about-the-book-and-our-writi...

Our more recent essay (and ongoing book project) "AI as Normal Technology" is about our vision of AI impacts over a longer timescale than "AI Snake Oil" looks at https://www.normaltech.ai/p/ai-as-normal-technology

I would categorize our views as techno-optimist, but people understand that term in many different ways, so you be the judge.

Thanks! HN was part of the origin story of the book in question.

In 2018 or 2019 I saw a comment here that said that most people don't appreciate the distinction between domains with low irreducible error that benefit from fancy models with complex decision boundaries (like computer vision) and domains with high irreducible error where such models don't add much value over something simple like logistic regression.

It's an obvious-in-retrospect observation, but it made me realize that this is the source of a lot of confusion and hype about AI (such as the idea that we can use it to predict crime accurately). I gave a talk elaborating on this point, which went viral, and then led to the book with my coauthor Sayash Kapoor. More surprisingly, despite being seemingly obvious it led to a productive research agenda.

While writing the book I spent a lot of time searching for that comment so that I could credit/thank the author, but never found it.

Thanks for the comment! I agree — it's important to remain fluid. We've taken steps to make sure that predictively speaking, the normal technology worldview is empirically testable. Some of those empirical claims are in this paper and others in coming in follow-ups. We are committed to revising our thinking if it turns out that our framework doesn't generate good predictions and effective prescriptions.

We do try to admit it when we get things wrong. One example is our past view (that we have since repudiated) that worrying about superintelligence distracts from more immediate harms.

We do not assume a status quo or equilibrium, which will hopefully be clear upon reading the paper. That's not what normal technology means.

Part II of the paper describes one vision of what a world with advanced AI might look like, and it is quite different from the current world.

We also say in the introduction:

"The world we describe in Part II is one in which AI is far more advanced than it is today. We are not claiming that AI progress—or human progress—will stop at that point. What comes after it? We do not know. Consider this analogy: At the dawn of the first Industrial Revolution, it would have been useful to try to think about what an industrial world would look like and how to prepare for it, but it would have been futile to try to predict electricity or computers. Our exercise here is similar. Since we reject “fast takeoff” scenarios, we do not see it as necessary or useful to envision a world further ahead than we have attempted to. If and when the scenario we describe in Part II materializes, we will be able to better anticipate and prepare for whatever comes next."

I appreciate the concern, but we have a whole section on policy where we are very concrete about our recommendations, and we explicitly disavow any broadly anti-regulatory argument or agenda.

The "drastic" policy interventions that that sentence refers to are ideas like banning open-source or open-weight AI — those explicitly motivated by perceived superintelligence risks.

This paper is being misinterpreted. The degradations reported are somewhat peculiar to the authors' task selection and evaluation method and can easily result from fine tuning rather than intentionally degrading GPT-4's performance for cost saving reasons.

They report 2 degradations: code generation & math problems. In both cases, they report a behavior change (likely fine tuning) rather than a capability decrease (possibly intentional degradation). The paper confuses these a bit: they mostly say behavior, including in the title, but the intro says capability in a couple of places.

Code generation: the change they report is that the newer GPT-4 adds non-code text to its output. They don't evaluate the correctness of the code. They merely check if the code is directly executable. So the newer model's attempt to be more helpful counted against it.

Math problems (primality checking): to solve this the model needs to do chain of thought. For some weird reason, the newer model doesn't seem to do so when asked to think step by step (but the current ChatGPT-4 does, as you can easily check). The paper doesn't say that the accuracy is worse conditional on doing CoT.

The other two tasks are visual reasoning and answering sensitive questions. On the former, they report a slight improvement. On the latter, they report that the filters are much more effective — unsurprising since we know that OpenAI has been heavily tweaking these.

In short, everything in the paper is consistent with fine tuning. It is possible that OpenAI is gaslighting everyone by denying that they degraded performance for cost saving purposes — but if so, this paper doesn't provide evidence of it. Still, it's a fascinating study of the unintended consequences of model updates.

OP here. Unfortunately this thread is mostly misinformation. There were a bunch of viral threads from the growth hacker / influencer crowd, including this one, within hours of the code release with a very superficial understanding of the code (and how recsys work in general). That's partly what motivated me to write this article.

See here for a rebuttal of the main tweet in that thread (near the bottom of the article). https://solomonmg.github.io/post/twitter-the-algorithm/

Rebuttal: https://aisnakeoil.substack.com/p/a-misleading-open-letter-a...

Summary: misinfo, labor impact, and safety are real dangers of LLMs. But in each case the letter invokes speculative, futuristic risks, ignoring the version of each problem that’s already harming people. It distracts from the real issues and makes it harder to address them.

The containment mindset may have worked for nuclear risk and cloning but is not a good fit for generative AI. Further locking down models only benefits the companies that the letter seeks to regulate.

Besides, a big shift in the last 6 months is that model size is not the primary driver of abilities: it’s augmentation (LangChain etc.) And GPT3-class models can now run on iPhones. The letter ignores these developments. So a moratorium is ineffective at best and counterproductive at worst.

We don't expect it to be free -- please read the article. That's not the issue at all. It's like if you subscribe to a product that you need to do your job, and one day the company tells you that the product is going away in three days and that you need to switch to a different product (that isn't at all the same for your use case).

Sure, but the article is talking about a completely different meaning of reproducibility, where a researcher uses an LLM as a tool to study some research question, and someone else comes along and wants to check whether the claims hold up.

This doesn't in any way require the training run or the build to be reproducible. It just requires the model, once released through the API, to remain available for a reasonable length of time (and not have the rug pulled with 3 days' notice).

We're under no such misapprehension and we're keenly aware that this is an uphill battle. The issue is that LLMs have become part of the infrastructure of the Internet. Companies that build infrastructure have a responsibility to society, and we're documenting how OpenAI is reneging on that responsibility. Hindering research is especially problematic if you take them at their word that they're building AGI. If infrastructure companies don't do the right thing, they eventually get regulated (and if you think that will never happen, I have one word: AT&T).

Finally, even if you don't care about research at all, the article mentions OpenAI's policy that none of their models going forward will be stable for more than 3 months, and it's going to be interesting to use them in production if things are going to keep breaking regularly.

Addressed in the article:

"OpenAI responded to the criticism by saying they'll allow researchers access to Codex. But the application process is opaque: researchers need to fill out a form, and the company decides who gets approved. It is not clear who counts as a researcher, how long they need to wait, or how many people will be approved. Most importantly, Codex is only available through the researcher program “for a limited period of time” (exactly how long is unknown)."

OP here. Many people are reacting to the title of the paper. A few thoughts:

* The paper is 35 pages long and it's hard to convey its message in any single title. We make clear in the text that our point is not that predictive optimization should never be used.

* We do want the _default_ to change from predictive optimization being seen as the obvious way to solve certain social problems to being against it until the developer can address certain objections. This is also made clear in the paper.

* The title is a nod to a famous book in this area called "Against prediction". Most people in our primary target audience are familiar with that book, so the title conveys a lot of information to those readers. That's one reason we picked it.

* Despite its flaws, when might we want to use predictive optimization? Section 4 gets into this in detail.

Thanks for reading.

OP here. The full title of this article is "Students are acing their homework by turning in machine-generated essays. Good."

The last word was edited out by the mods, presumably under the belief that it's clickbait. Unfortunately, the headline now sounds like I'm complaining about this development, whereas my post is about how it will force much-needed improvements to education and free students from the drudgery of pointless essays that ask them to regurgitate content (as opposed to essays that teach writing skills or critical thinking, which remain valuable).

OP here. I totally agree that ideally authors should report most of this information in the paper itself. One advantage of a standalone document (we suggest putting it in an appendix) is that it's easy for reviewers to check that all of this information has been reported. Of course, authors could answer some of the questions by pointing to the sections of the paper in which they have been answered.

It's possible you may have misunderstood the title of the post. It isn't about the science of ML, or GPT-3, or brains. Rather, it's about using ML as a tool to do actual science, like medicine or political science or chemistry or whatnot. The first sentence of the post explains this.

Princeton University Center for Information Technology Policy | Princeton, NJ | Onsite | Full Time

Princeton CITP is a leading research center at the intersection of technology and public policy. We've conducted groundbreaking work on privacy, government surveillance, net neutrality, algorithmic fairness, dark patterns, and other high-profile topics. https://citp.princeton.edu/

We're hiring a data scientist who will collaborate with our world-class faculty, fellows, and students on interdisciplinary research projects and policy impact. If you live in New York City or New Jersey, are passionate about the societal impact of technology, and have an impressive resume in data science (broadly conceived), we want to hear from you.

Application: https://puwebp.princeton.edu/AcadHire/apply/application.xhtm...

FAQ: https://citp.princeton.edu/about/hiring/faculty-staff/faqs-s...

That's fair. We don't claim that this is a new problem; we are merely adding evidence and our perspective to a known problem. We do link to others who have reported similar problems when trying to disclose vulnerabilities. The sentence saying we "discovered two wider issues" was worded poorly; in the paper [1] we used the word "encountered", and I've now edited the post to use the same wording. Thanks!

Just as important, the post is a PSA that there are 9 websites whose users remain vulnerable, and people with accounts on these sites should check their 2FA and password recovery settings. The websites are: Amazon, AOL, Finnair, Gaijin, Mailchimp, PayPal, Venmo, Wordpress.com, and Yahoo.

[1] Link to paper: https://www.issms2fasecure.com/assets/sim_swaps-03-25-2020.p...

Princeton University Center for Information Technology Policy | Princeton, NJ | Onsite | Full Time

Princeton CITP is a leading research center at the intersection of technology and public policy. We've conducted groundbreaking work on privacy, government surveillance, net neutrality, algorithmic fairness, dark patterns, and other high-profile topics. https://citp.princeton.edu/

We're hiring a data scientist who will collaborate with our world-class faculty, fellows, and students on interdisciplinary research projects and policy impact. If you live in New York City or New Jersey, are passionate about the societal impact of technology, and have an impressive resume in data science (broadly conceived), we want to hear from you.

Application: https://puwebp.princeton.edu/AcadHire/apply/application.xhtm...

FAQ: https://citp.princeton.edu/about/hiring/faculty-staff/faqs-s...

Just to clarify, this is not a new finding, but an explainer of a study from 2016.

Note that this fingerprinting technique exploits differences in the behavior of the AudioContext API, but does not (and cannot) actually record audio.

Paper: https://webtransparency.cs.princeton.edu/webcensus/index.htm...

Demonstration (test your own audio fingerprint): https://audiofingerprint.openwpm.com

Discussion from 2016: https://news.ycombinator.com/item?id=11729438

Full list of websites where audio fingerprinting scripts were found (in March 2016): https://webtransparency.cs.princeton.edu/webcensus/audio_fp_...

Source: I'm an author of the research in question (but unaffiliated with this blog).

Note to mods: article title is "Audio Fingerprinting using the AudioContext API". Submitter title is "Sites are using audio (no permissions needed) to track users", which may violate the site guidelines.

Coauthor here. As it turns out, this is one of three papers released near-simultaneously that uncover the extent of tracking on TVs or IoT devices more generally. I've written up a survey of the three papers and what I thought were especially interesting findings, along with some thoughts on why targeted advertising as a business model for TV platforms is harmful to users: https://twitter.com/random_walker/status/1177570679232876544

Direct links to the other two papers:

https://moniotrlab.ccis.neu.edu/wp-content/uploads/2019/09/r...

https://arxiv.org/pdf/1909.09848.pdf

OP here. A small clarification: there are some legitimate criticisms of the Register piece that I cited in the first tweet, but I merely cited it as an example of why I think the hype is calming down. The arguments I make are independent of that piece; the limitations I point out have always been there, rather than something new that happened in 2018.

I also wanted to add a couple of points that I didn't get to in the Twitter thread.

Economics. Blockchain technologists seem to overestimate the extent to which new insights in economics are needed to understand cryptocurrencies and blockchains, as opposed to applying basic principles from economics and game theory. For example, a recent paper shows that thinking about miners and attackers in terms of stock and flow exposes important limitations of the security of Proof of Work. [1] I learnt of many other such examples at a recent conference on the economics of blockchains. [2] So I think a lot of the "cryptoeconomics" hype is misplaced.

Privacy. It's often taken for granted that decentralized architectures will improve privacy. This seems obvious given everything we've learnt about Facebook, but a better way to think about it is that decentralized systems exchange one set of privacy problems with another. I coauthored a paper a few years ago skeptical of the "decentralization ==> privacy" story in the context of social networks [3], but I think many of the arguments in that paper apply to blockchain/dApps that are being built today.

[1] http://faculty.chicagobooth.edu/eric.budish/research/Economi...

[2] https://bfi.uchicago.edu/events/cryptocurrencies-and-blockch...

[3] http://randomwalker.info/publications/critical-look-at-decen...

Thanks!

The main reason I'm using Twitter for this is that it's a bit too preliminary for a blog post. I don't yet have as good an understanding of the history as I would like. This way, when I discover new stuff, I can simply add at a tweet to the thread.

But TBH I think Twitter is underrated as a publishing medium. For example, I've had probably 10x the number of responses from other people as I would have gotten in the form of blog comments.

In any case, I'm definitely planning to make this more organized once I'm happy with my level of understanding. At least a series of blog posts; probably a paper and/or online lecture.

That's possible, but an alternative explanation for the cost overruns that I've read is that Babbage had terrible project management skills.

Wikipedia has this to say:

In 1991, the London Science Museum built a complete and working specimen of Babbage's Difference Engine No. 2, a design that incorporated refinements Babbage discovered during the development of the Analytical Engine. This machine was built using materials and engineering tolerances that would have been available to Babbage, quelling the suggestion that Babbage's designs could not have been produced using the manufacturing technology of his time.

https://en.wikipedia.org/wiki/Analytical_Engine

OP here. The number and variety of special-purpose computing devices that existed before general purpose computers is astounding. The surprising (to me) conclusion is that the main impediment to the development of computers wasn't technology. After all, Babbage's machine could have been built in his time if funding hadn't run out.

Rather, the limitation was that people didn't have the abstractions, vocabulary, and mental tools to properly conceive of general purpose computers as a concept and to understand their usefulness. They couldn't see that devices as seemingly disparate as tide prediction machines[1], census tabulation machines, and loom controllers were all instances of a single, terrifyingly general idea.

From what I can tell, Babbage mostly understood this, but it was Ada Lovelace who grasped it fully. But her writings weren't understood in her time and had to be "rediscovered" a century later. For example, she wrote [2]:

Supposing, for instance, that the fundamental relations of pitched sounds in the science of harmony and of musical composition were susceptible of such expression and adaptations, the engine might compose elaborate and scientific pieces of music of any degree of complexity or extent.

This leads me to wonder: what abstractions are we missing today that will be obvious to future generations?

BTW I have a follow-up thread on the optical telegraph, a form of networking that long predates the Internet. [3] My long-term goal is to teach a course on computing/networking/information processing before computers, with a view to extracting lessons that are still applicable today.

[1] https://en.wikipedia.org/wiki/Tide-predicting_machine

[2] https://googleblog.blogspot.com/2012/12/honouring-computings...

[3] https://twitter.com/random_walker/status/1037031465735860224

BlockSci is an academic research project at Princeton, but we're committed to maintaining it as open-source software, and we hope it's more broadly useful. If you're interested in using it or contributing to it, here's a list of ideas that we'd love to see implemented.

1. Create a Block Explorer. BlockSci would make a good backend for a block explorer website, because it would benefit from the built-in analysis library, with features like address clustering and parsing multisignature scripts.

2. Support more blockchains. BlockSci supports several blockchains, but there are limitations detailed in the paper [1]. For example, currently we don’t support any script operations not found in Bitcoin. Supporting more altcoins/blockchains would make BlockSci more useful.

3. Identify cold wallets and associated usage patterns. Cold wallet addresses could be identified by various patterns on the blockchain such as infrequent large withdrawals. After identifying these addresses, there are many interesting questions to ask such as studying the rate of deposits vs withdrawals.

4. Improve clustering heuristics. BlockSci’s address linking is based on the two heuristics from the Fistful of Bitcoins paper [2]. These heuristics have known limitations, leading to false positives and negatives; there’s a lot of room for improvement here.

5. Extract hidden messages. There are many messages encoded into the Bitcoin blockchain ranging from Wikileaks cables to Rickrolls [3]. We can find them if we can guess how they are encoded. But can we automatically extract and decode these hidden messages, say, by looking for address strings that look non-random?

[1] https://arxiv.org/pdf/1709.02489.pdf

[2] https://cseweb.ucsd.edu/~smeiklejohn/files/imc13.pdf

[3] http://www.righto.com/2014/02/ascii-bernanke-wikileaks-photo...