HN user

dspoka

282 karma

ML Postdoc @ UCSD CMU PhD '24

Posts20
Comments29
View on HN
www.bloomberg.com 1y ago

OpenAI in Talks to Raise Up to $40B in New Funding

dspoka
2pts1
www.youtube.com 1y ago

Sequence to sequence learning with neural networks: what a decade

dspoka
89pts31
www.insidehighered.com 2y ago

Matching Toxic Posts to IP Addresses on Economic's Forum [pdf]

dspoka
11pts5
twitter.com 3y ago

Bing Token-Smuggling Jailbreak

dspoka
1pts0
www.latimes.com 3y ago

Strike by 48,000 University of California Academic Workers

dspoka
5pts0
research.contrary.com 3y ago

The Evolution of DevOps

dspoka
2pts0
octoml.ai 4y ago

OctoML CLI: A New DevOps Focused ML Deployment Tool

dspoka
19pts0
kunle.app 4y ago

Time to Money as a Competitive Advantage

dspoka
1pts0
kunle.app 4y ago

Children of Durbin

dspoka
1pts0
twitter.com 4y ago

Former OpenAI Team raise $580M Series B to build interpretable AI

dspoka
3pts0
www.detroitnews.com 4y ago

Rocket Mortgage to trim 8% of workforce as home-loan market shrinks

dspoka
282pts493
news.ycombinator.com 4y ago

Ask HN: Advice on Phone/Messaging/etc.

dspoka
4pts0
news.ycombinator.com 6y ago

Ask HN: What are your thoughts on unpaid internships in tech?

dspoka
2pts0
medium.com 6y ago

Carta’s CEO Covid-19 Layoff Message

dspoka
4pts2
techcrunch.com 6y ago

SoFi acquires banking and payments platform Galileo for $1.2B

dspoka
2pts0
www.lawfareblog.com 6y ago

Social Media Disinformation Run Out of Cyprus by Russians

dspoka
3pts0
news.ycombinator.com 7y ago

Ask HN: Where to learn about industries and markets?

dspoka
2pts0
www.argmin.net 8y ago

Reinforcement Learning, Policy Gradient Is Nothing More Than Random Search

dspoka
2pts0
www.theguardian.com 8y ago

Scientists capture exploding beetles' amazing escapes from toads' stomach

dspoka
2pts0
fairmlclass.github.io 8y ago

Fairness in Machine Learning UC Berkeley Class

dspoka
4pts1

Sensational title that misrepresents the message in paper.

However, when conducting more targeted automatic evaluations, we found that the imitation models close little to none of the large gap between LLaMA and ChatGPT. In particular, we demonstrate that imitation models improve on evaluation tasks that are heavily supported in the imitation training data. On the other hand, the models do not improve (or even decline in accuracy) on evaluation datasets for which there is little support. For example, training on 100k ChatGPT outputs from broad-coverage user inputs provides no benefits to Natural Questions accuracy (e.g., Figure 1, center), but training exclusively on ChatGPT responses for Natural-Questions-like queries drastically improves task accuracy.

Just because this might not be the way to replicate the performance of ChatGPT across all tasks, it seems to work quite well on whichever tasks are in the imitation learning. That is still a big win.

Later on this also works for factual correctness. (leaving aside the argument whether this is the right approach for factuality)

For example, training on 100k ChatGPT outputs from broad-coverage user inputs provides no benefits to Natural Questions accuracy (e.g., Figure 1, center), but training exclusively on ChatGPT responses for Natural-Questions-like queries drastically improves task accuracy.

My main qualm with the parent comment was this in particular, "Number one, is that forests work as long term carbon storage and sinks." Even with the surrounding context it sounded like this strategy will just "work".

The example you give with mangroves is a great one which does in fact work. Pragmatically and historically most of the attempts however, have not due to mismanagement and other unseen complications.

Seeing the further comments I see the point the parent was making is more around first principles of Forrests as carbon sinks not about its implementations.

I think the reaction from software devs on how Copilot's uses their code for ML is interesting in that all the ML companies have been doing this with all other forms of produced content: texts, posts, messages, photo captions, etc. And most likely even less care went into adhering to laws or ethics. Yes code has licenses and thus more distinct legal ramifications but on the other side are people who don't really understand that every time they interact with software or produce some content, everything is gathered and harnessed to power all these companies.

PyTorch 1.0 is out 8 years ago

Is there some sort of pytorch 1.0 migration guide or does anyone know if there is any breaking from .41 to 1.0 ?

So this idea seems intuitive at first but turns out to be one of the worst things to treat unfairness.

There are several reasons for this from both technical and legal perspective.

It is incredibly easy to find statistically significant correlations given just a few (more than 7) different views of the data. In general these ml models are not working with less than hundreds or thousands.

If the model learned this suppose racial bias, once, you deleting this column is not going to stop it from learning it again, and I believe some research showed that it actually can make the unfairness more severe.

from a legal standpoint a company that may or may not be infringing on rights could just say, oh we can't be because we don't have these fields in our data: which makes it harder to monitor and audit wrong doing.

most of the methods that I am familiar try to ease the effects of the learned biases as a post-processing step for the model.

The most promising area here to me seem like automl. The promise of the new machine learning was that we get to move away from tedious feature engineering and everything will work and be simple. It may have become simpler but training/debugging new DL models is still painful causing the focus to move to extensive hyperparameter search. automl may become the next step in abstraction, where we design single models/algorithms that are able to build viable networks for many tasks/purposes.

This seems like it's going to be a great class and topic taught my Moritz Hardt, a researcher at Google who has done research on several topics of fairness such as analysis of demographic parity.

I hope this becomes one of the most important and necessary topics for ML researchers as well as ML practitioners to consider when building models. There has been some serious discussion between researchers about adding some sort of licenses to help mitigate/ add accountability for researchers knowingly building biased ai.

For me getting alerted when there are new papers that cite papers that are relevant towards my current research topic would be ideal. Google scholars has alerts on authors and search queries but for me they don't have enough recall.

Its much easier to tell when a paper is relevant for me if it happens to cite 3 of the commonly used datasets for my particular task.

btw I use arxiv-sanity, its pretty great, thanks a lot!

Some people have already mentioned these but so far I'm using:

Karpathy's http://www.arxiv-sanity.com/library subscribe to archive email lists

Semantic Scholar (no notifications) is good for manually finding things

Google Scholar notifies you when your papers get citations... Unfortunately they don't have a way for you to get notified if the paper is not yours.. so I made a few fake accounts that add papers to the library as if they are the author and then I set up a forwarding to my email. (really wish they would just expand the notified of citations feature to your library and not just your papers but whatever)

I happen to live in Isla Vista, the college town of UCSB. Although not explicitly in writing, the rule of this town is that bikes have the right of way. We have somewhere around 14,000 cyclists that ride to campus and back everyday. From first hand experience this is a much safer system. The drivers are much more attentive at the wheel and the heavy bike traffic enforces that cars really don't go faster than the 25 mph speed limit

YC Fellowship 11 years ago

How important is it for this Fellowship to have a flushed out Business Model?

TrackR 12 years ago

but this is an actual product and tile is in production =)