HN user

davmre

3,155 karma

cs.berkeley.edu/~dmoore

Posts30
Comments269
View on HN
sashachapin.substack.com 3y ago

Hey, why aren't we doing more research about spiritual awakening

davmre
8pts8
medium.com 8y ago

Theano, TensorFlow and the Future of PyMC

davmre
4pts0
fortune.com 9y ago

Why Ruth Porat Should Be Uber’s Next CEO

davmre
3pts0
medium.com 9y ago

Opening a new chapter of my work in AI

davmre
247pts74
acko.net 10y ago

Graphical Algebra and Fourier Analysis

davmre
2pts0
www.vox.com 10y ago

What Tim Cook doesn't want to admit about iPhones and encryption

davmre
2pts0
medium.com 10y ago

So you want to reform democracy

davmre
1pts0
multiverseaccordingtoben.blogspot.com 11y ago

Life Is Complexicated

davmre
1pts0
futureoflife.org 11y ago

Autonomous Weapons: An Open Letter from AI and Robotics Researchers

davmre
2pts0
www.whetlab.com 11y ago

Twitter acquires Whetlab (Bayesian hyperparameter optimization)

davmre
1pts0
amplab.cs.berkeley.edu 11y ago

David Patterson: Moore's Law Is Dead

davmre
2pts1
www.vox.com 11y ago

If the EU wins against Google, it could change the search engine forever

davmre
1pts1
www.vox.com 11y ago

Obama calls for municipal broadband

davmre
520pts275
lemire.me 11y ago

MOOCs are closed platforms and probably doomed

davmre
56pts27
bits.blogs.nytimes.com 11y ago

Warm West Coast Reception for China’s Web Czar

davmre
1pts0
www.washingtonpost.com 11y ago

Wikipedia’s ‘complicated’ relationship with net neutrality

davmre
108pts85
www.cs.berkeley.edu 11y ago

Rationality and Intelligence: A Brief Update [pdf]

davmre
47pts1
bits.blogs.nytimes.com 11y ago

Gift from Ballmer Will Expand Computer Science Faculty at Harvard

davmre
2pts0
amplab.cs.berkeley.edu 11y ago

Big Data, Hype, the Media and Other Provocative Words to Put in a Title

davmre
15pts1
www.vox.com 11y ago

HBO will let you watch their shows online – without having to buy cable

davmre
1pts0
www.zdnet.com 11y ago

Microsoft to Close Microsoft Research Lab in Silicon Valley

davmre
133pts84
bits.blogs.nytimes.com 12y ago

Large Round of Layoffs Expected at Microsoft

davmre
304pts245
www.huffingtonpost.com 12y ago

Transcendence: An AI Researcher Enjoys Watching His Own Execution

davmre
1pts0
talkingpointsmemo.com 12y ago

A Few Thoughts on Brendan Eich

davmre
6pts2
www.nytimes.com 12y ago

T-Mobile Turns An Industry On Its Ear

davmre
115pts80
www.arjunnarayan.com 12y ago

Valuing Spotify

davmre
1pts0
blog.mikiobraun.de 12y ago

How Python became the language of choice for data science

davmre
175pts79
www.talyarkoni.org 12y ago

Python: The homogenization of scientific computing

davmre
39pts10
www.slate.com 12y ago

Stop "Disrupting" Everything

davmre
48pts46
priceonomics.com 12y ago

The Accidental Breast Milk Entrepreneur

davmre
2pts0
Claude Fable 5 1 month ago

This sounds more or less unavoidable? Decompilers are inherently security-sensitive. If you take avoiding cyberattack uplift seriously as a goal, I don't see how you get around essentially refusing to work on them.

Obviously there are plenty of innocuous applications too, but it's not like the people building decompilers for nefarious reasons will be explicit about it. The LLM abstraction just inherently doesn't have enough context to distinguish your intentions or your broader use cases. This is why both Anthropic and OpenAI have had to create side channel mechanisms for security researchers to establish a trusted use context. It sounds like this makes this not a viable product for you, unfortunately, and it makes sense that that's frustrating. But I also don't see what different behavior one could reasonably expect given the constraints.

If it's any consolation, these restrictions only make sense for models that are ahead of the open-weights frontier, so open-source hackers will presumably get Mythos-level capabilities in the relatively near future anyway.

It's true that these are very different activities, but I think most ML researchers would agree that it's actually the creation of ImageNet that sparked the deep learning revolution. CNNs were not a novel method in 2012; the novelty was having a dataset big and sophisticated enough that it was actually possible to learn a good vision model from without needing to hand-engineer all the parts. Fei-fei saw this years in advance and invested a lot of time and career capital setting up the conditions for the bitter lesson to kick in. Building the dataset was 'easy' in a technical sense, but knowing that a big dataset was what the field needed, and staking her career on it when no one else was doing or valuing this kind of work, was her unique contribution, and took quite a bit of both insight and courage.

Find SF parking cops 10 months ago

Even if you charge $10/hr, or whatever the market rate would be for street parking spots, you still need an enforcement mechanism to prevent people overstaying.

In general, the idea of a "market rate" for any given property depends fundamentally on a system of property rights actually being enforced.

Senators, especially committee chairmen, have quite a bit of implicit leverage, beyond the direct leverage of subpoenas or directly cutting funding to an offending agency.

Any given Senator is to some extent constantly in a favor-trading game with executive branch officials. People from the President on down need congressional cooperation to get their pet provisions into bills, programs funded, nominees approved, etc. A Senator can tell a White House official "I'd love to help you with that, however I have this issue with this agency not responding to my requests". Assuming it's a reasonable thing, whoever at the agency is in charge of this then gets an irate call from their boss's boss's boss ordering them to cooperate.

Of course this mostly doesn't actually get played out, because everyone understands the dynamic that defying senatorial requests will ultimately cost the President in terms of cooperation on other issues. So the norm is mostly to comply with reasonable requests, unless you're quite sure that it's a top-level priority where the White House really wants to take a stand.

The public interest is in judging the trial process, not in judging the defendant.

Suppose the government charges you with murder, searches your house, and finds your sex toy collection. At trial they present some elaborate thesis about how you used a sex toy to kill someone, but do not convince the jury, so you're found not guilty. The public has a legitimate interest in judging that the trial was handled with integrity and that the correct verdict was reached. They do not have a legitimate interest in judging you based on whatever private information presented at trial might in some way embarrass you (eg, photos of your sex toy collection). On balance, it could be that the public-record interest does in fact justify making public the evidence of the sex toys, but you have to justify it on those terms. The transparency is not itself intended to be punitive.

Regardless of the effects, I don't think the case against MS was brought with the intent to "punish" MS through the trial process. The government brought the case because it thought it could win, it did win, and a judicial remedy was imposed. Trials are inherently unpleasant, but a just system tries to minimize this, not exploit it.

Any unjust policy (including just dispensing with trials altogether and allowing the executive to arbitrarily break up companies) will get to the 'desirable' outcome in some cases. That doesn't make it a just policy.

The specific allegation in the post is that the Trump administration will not appeal the verdict because Sundar gave $1M to Trump's inauguration. As far as I know, the government has not yet indicated whether it will appeal, so the claim that "Trump just paid him back, 40,000 times over" is in fact not true. (whether it becomes true at some point in the future, it was a falsehood at the time the author wrote it). It's also quite plausible that a Republican administration wouldn't appeal the verdict just due to being more pro-business in general, even without explicit corruption. But it's precisely because we have such a corrupt executive that it becomes all the more important to stick up for the rule of law. The correct response to authoritarianism is not to advocate for more authoritarianism!

Sure, there's a strong public interest in having proceedings on record. US civil cases are supposed to have a presumption of openness, which the judge weighs against other interests, like protecting trade secrets, confidential business information, privacy of third parties, etc.

The public record argument is fine; it's just a different argument than the extrajudicial punishment advocated by the original post.

If there's a general standard of transparency applied to all companies, fine. There are costs to increasing transparency, but certainly you could argue for that policy.

The argument that we should cheer on the use of government power to target a specific company, to selectively expose their dirty laundry as punishment for a crime they have not been convicted of, is what I found noxious in the original post.

The government doesn't have to win an antitrust trial in order to create competition. As the saying goes, "the process is the punishment."

Regardless of what you think of Google or this case specifically, this is an argument for authoritarianism: that it is legitimate for the government to "punish" any company at will, based only on them falling into political disfavor.

... the only punishment Google would have to bear from this trial would come after the government won its case, when the judge decided on a punishment (the term of art is "remedy") for Google.

Yes, this is called the rule of law. Punishment comes through the courts, after a guilty verdict. The government has to actually win the argument as to what remedies would be proportionate under the law. In this case the judge didn't buy it. It's fine to disagree with his reasoning (or with the law), but the fantasizing about extrajudicial punishment here is frankly un-American.

They're not proposing to apply tensor decomposition to an existing collection of weights. It's an architecture in which the K, V, and Q tensors are constructed as a product of factors. The model works with the factors directly and you just need to compute their product on the forward pass (and adjoints on the backwards pass), so there's no decomposition.

DeepSeek-R1 1 year ago

You're totally right there must be supervision; it's just a matter of how the term is used.

"Supervised learning" for LLMs generally means the system sees a full response (eg from a human expert) as supervision.

Reinforcement learning is a much weaker signal: the system has the freedom to construct its own response / reasoning, and only gets feedback at the end whether it was correct. This is a much harder task, especially if you start with a weak model. RL training can potentially struggle in the dark for an exponentially long period before stumbling on any reward at all, which is why you'd often start with a supervised learning phase to at least get the model in the right neighborhood.

Backprop itself doesn't invert the computation, but it does give you the direction for an incremental move towards the inverse (a 'nudge' as the article puts it). That is, given a sufficiently nice function f and an appropriate loss ||f(x) - y*||^2, gradient descent wrt x will indeed recover the inverse x* = f^{-1}(y*) since that is what minimizes the loss. I assume this what the article is pointing at.

If you want to be picky, it's true that the direct analogue of continuous optimization would be discrete optimization (integer programming, TSP, etc) rather than decision problems like SAT. But there are straightforward reductions between the two so it's common to speak of optimization problems as being in P or NP even though that's not entirely accurate.

Nitpicking, but for a technical audience it's worth noting that ibogaine is not at all a 'potent' psychedelic in the pharmacological sense of the term. A typical therapeutic dose is on the order of 500mg, which makes ibogaine something like 20 times less potent than psilocybin (typical dose ~25mg), which itself is 100 times less potent than LSD (typical doses less than 250ug).

Of course, this isn't really relevant to the subjective experience of taking ibogaine at its typical dose, which by all accounts is strange in ways that go beyond the classical psychedelics.

Altman was initially going to cooperate and even offered to help, until Brian Chesky & Ron Conway riled him up

I don't think the article supports this. All we know is that sama appeared cooperative when the board fired him. This was probably a reasonable posture for him to adopt regardless of his actual intentions at the time.

Counterpoint: SSRIs were transformative for my depression. Side effects were minor and manageable (eg, Wellbutrin worked well to prevent any sexual dysfunction). I was on them for several years and had no problem tapering off. My understanding is that this is a pretty typical experience. Rhetoric like this was actively harmful in dissuading me for years from trying what ended up being by far the most effective treatment for me.

(yes, I've tried psychedelics; they're fascinating and super promising, but at least for me, not transformative in the way that fluoxetine was)

No individual depression treatment works for everyone. SSRIs are not a magic bullet. Neither are psychedelics. But if you're depressed and haven't tried SSRIs, you owe it to yourself and everyone in your life to at least test the hypothesis that they might help.

Scott Alexander's page on SSRIs is a great, relatively objective resource, from a psychiatrist who regularly prescribes them: https://lorienpsych.com/2020/10/25/ssris/

Prophet has gotten a lot of attention since being released in 2017, I think because the idea of a fully automatic solution is very appealing to people. One of the original developers, Sean Taylor, recently posted a nice retrospective on the project's successes and failures: https://medium.com/@seanjtaylor/a-personal-retrospective-on-... He quotes one of his earlier tweets:

  If I could build it again, I’d start with automating the evaluation of forecasts. It’s silly to build models if you’re not willing to commit to an evaluation procedure. I’d also probably remove most of the automation of the modeling. People should explicitly make these choices.
Having worked on similar Bayesian time-series forecasting tools at Google, this matches my experience (though I've never used Prophet seriously, so please don't take this as any direct judgement of it as a software package). There is a lot of value in a framework that lets you easily experiment with different model structures (our version of this was the structural time series tools in TensorFlow Probability, see, e.g., https://blog.tensorflow.org/2019/03/structural-time-series-m...). But if you're forecasting something you actually care about, it's usually worth the time to try to understand yourself what structure makes sense for your problem, and do a careful evaluation on held-out data with respect to whatever metric you're really trying to optimize. A fully automated search over model structures is cute, but even when it works, it mostly just ends up rediscovering properties of the data you could or should have already known (e.g., of course traffic to your work-related website will have a day-of-week effect), so the cases where it really adds practical value are harder to find than you might like.

Even in the age of deep learning, I do think these relatively classical Bayesian models have a lot of value for many applications. Time-series forecasting tends to be a case where:

- you don't have a ton of iid data points (often, only a single time series),

- you'd like forecasts with principled uncertainty estimates, e.g., credible intervals, giving you a range of scenarios to plan for,

- you often do have a pretty good idea of what features are relevant to the process you're predicting, and

- you want to understand in detail what features the forecast is accounting for (and what it might be missing),

all of which play to the strengths of more classical, structured statistical models, compared to more data-hungry black-box deep learning models. So the basic ideas in Prophet and similar tools do still have a lot of relevance going forward, IMHO.

Before trying MDMA I had mostly written it off as a feel-good drug: an artificial high like cocaine or meth, useful only for escapism. Why bother chasing that sort of experience? But now having tried most of the commonly used psychedelics, I've come to believe that MDMA is the most profound of them all.

MDMA does feel good, of course, but it's not escapist. It’s a deep, wholesome, fundamentally healing sort of goodness. It is unconditional, redeeming love and forgiveness — the core of Christian spirituality. It is the revelation that you really are lovable, even your darkest, hidden parts, and that you are capable of love. Debatably, there is no more profound lesson to be learned about the human condition. It really is magical.

Even so, the experience is surprisingly subtle. It doesn't particularly force positive feelings ("ecstasy" is a total misnomer, IMHO). At first you don't necessarily even notice any effect at all, maybe just a mildly better-than-average mood. But gradually it becomes clear that this subtle sense of well-being is infinitely deep: nothing you might experience can possibly disturb it. All sense of shame and self-judgement, fear of rejection, hang-ups that get in the way of connecting with people --- all dissolve immediately on contact. And from that sense of absolute safety, the capacity to love emerges naturally. The drug doesn't generate it. It just helps you get out of your own way.

I've taken MDMA a few times now just on my own at home (lacking a rave community, although I'm sure that's a fantastic experience also), where it's relatively easy to implement harm reduction measures: stay hydrated, take protective supplements, get a full night's sleep before and after, and wait multiple months between doses (doing all these, I've never experienced a ‘hangover’, just a positive afterglow). I've found the most rewarding results from trying to keep my attention grounded in bodily sensation, gently returning to the body whenever I notice I've become lost in thought. Often, difficult memories or associations will surface of their own accord, sensing that it's safe to do so, and seeing them from a loving perspective can be immensely healing.

I really hope we can eventually find our way to making this experience legally and safely available to everyone who wants it. Yes, MDMA has sharp edges; it's not as physiologically benign as the classic psychedelics, but it's not addictive and it can be used safely. Not everyone has good experiences every time, but compared to the classical psychedelics, it's much more reliably positive. It apparently has some effectiveness as a medicine for specific illnesses like PTSD, but IMHO the real condition it treats is much broader: the universal human condition of feeling more walled off than we'd like to be.

One last galaxy-brain thought: if we ever figure out a way to replicate MDMA's pro-social effects that people could safely use on a day-to-day basis, it might be the most valuable thing we ever invent. One could even see it as the metaphorical second coming of Jesus, his kingdom on earth achieved through purely secular means. How's that for a career goal? :-)

I've been on the PhD admissions committee for the Berkeley CS PhD program. Students in many areas are, for all practical purposes, admitted by an advisor. It's not a formal arrangement in the same sense as European programs, but no one gets admitted unless at least one professor will commit to being willing to advise and fund you. For superstar students this is rarely an issue and you have your pick of advisors, but for many admits there is exactly one professor who's identified your background and interests as a good fit, so you're essentially admitted with the expectation you'll work with that professor.

Nothing you say is wrong; you're correct that everyone has to do course requirements. But culturally the course requirements are treated as a distraction from research, and students are expected to start working primarily on research from their first semester. I think this is more or less the case in all top American CS PhD programs, and there are reasons to set things up this way, but it's a very different environment from PhDs in most other fields.

What do you think about government-financed prizes for private-sector drug development, e.g., the government pledges $X billion if you develop an Alzheimer's cure and place it in the public domain? Bernie Sanders has a proposal like this: https://www.vox.com/2015/9/25/9397069/bernie-sanders-drug-pr....

It seems like that retains most of the advantages of the current system, incentivizes private-sector development, but also eliminates the problem of high prices at the point of delivery. (i.e., it correctly prices the marginal cost of treating someone at the marginal cost of producing the drug, which is generally low).

'Textbook' communism involves the government directly managing the means of production. Owning a minority stake in many corporations that are run for profit by incentivized managers in a market economy has, perhaps, some similarities but it's certainly not the same thing.

I'm not sure you can generalize Venezuela's collapse beyond a cautionary tale of corruption and populist looting; i.e., their economic problems are fundamentally political problems. A better comparison for a developed Western democracy with competent political system is Norway, which has an amazingly successful ($1 trillion) sovereign wealth fund https://en.wikipedia.org/wiki/Government_Pension_Fund_of_Nor... comprised of oil revenue used to fund social programs.

One possible mechanism is a social wealth fund that would gradually come to own (and distribute the proceeds from) a nontrivial portion of the national capital. Matt Bruenig wrote a NYTimes oped about this idea: https://www.nytimes.com/2017/11/30/opinion/inequality-social... with some suggestions for populating the fund:

Wouldn’t the enormous wealth that our increasingly productive society is generating, which now flows into just a few pockets, be a fair source? Some of the concrete ways this could happen are through the transfer of existing federal assets like land, buildings and portions of the wireless spectrum into the new fund. Other measures could include increases in taxes on capital that affect mostly the wealthy such as estate, dividend and financial transaction taxes and the creation of a new type of corporate tax that requires companies to directly issue new shares to the social wealth fund on an annual basis and during certain corporate moves such as initial public offerings, mergers and acquisitions.

Another way to bring assets into the fund would be to modify the way the Federal Reserve pumps money into the economy. Currently, the central bank does that by buying up Treasury bonds. If instead we used newly created money to buy up stocks that are then deposited into the social wealth fund, it would gradually socialize wealth ownership without the need to raise taxes on anyone. As Roger Farmer and Miles Kimball have argued, these kinds of asset purchases could also be ramped up during recessions, allowing the federal government to acquire significant portions of the national wealth relatively cheaply while also stabilizing financial markets and stimulating the economy."

Noah Smith makes the additional economic case that a social wealth fund is a buffer against decreasing labor share of income driven by technological change: https://www.bloomberg.com/view/articles/2017-12-05/robot-tak...

I haven't seen anyone here argue that it's difficult to invest foreign money in the US. Apple's only complaint is that they couldn't funnel the profits from such investment back to their US shareholders, without paying taxes.

Tensorflow's advantage is that once you build your model, you can run on everything from a massive cluster to a mobile GPU without significant modification. Because you're just writing a description of a computation graph, it's easy for backend systems to process that description and optimize the execution of your model.

PyTorch's imperative semantics (where the computation graph is implicitly defined at runtime by the execution of your Python code) definitely make it cleaner to do research prototyping. But AFAIK most PyTorch models need be reimplemented in lower-level code, or maybe something like Caffe2, before they can be used in production. That's a fairly significant tradeoff, which makes it hard to see PyTorch totally replacing Tensorflow anytime in the near future. That said PyTorch is obviously a great tool and it's exciting to see how it will develop and be used.

(disclaimer: I work for Google, opinions are my own)