HN user

natfriedman

4,377 karma

CEO of GitHub

Posts8
Comments146
View on HN

We think there's enough data.

There is a multi-century history of chemical methods destroying scrolls by the hundreds without reading them, so we're not eager to fund that work.

We're thinking about this too. The overall prize pool was about 6x smaller a week ago so we are still digesting this rapid influx of sponsorship.

The grand prize goes to the _first_ team to read 4 passages from the scrolls. But we could, for example, award something to the second team to do so. Or, we could award something to the team that reads the _most_ passages by the end of the year.

We deliberately did not allocate all of the recent sponsorships to the grand prize so we can solve for this exact challenge. So, we have about $500k in unallocated prize money, and might use a good chunk of it towards something like this. We're open to ideas, and consulting with experts from Xprize etc.

Hi folks, I'm the co-creator of the Vesuvius Challenge. We now have nearly $1.5M in prizes, thanks to a lot of amazing sponsors. Happy to answer any questions anyone has.

Vesuvius Challenge 3 years ago

It seems quite possible that the solution isn't fully automated. N is in the hundreds. And modern AI does, in fact, involve quite a lot of hand crafted data...

Vesuvius Challenge 3 years ago

Credit for those goes to Jonny Hyman, who also does animations for Veritasium, Dejan Gotić, who did the fancy 3d animations, and JP Posma, who directed the entire project!

GitHub Copilot 5 years ago

In general: (1) training ML systems on public data is fair use (2) the output belongs to the operator, just like with a compiler.

On the training question specifically, you can find OpenAI's position, as submitted to the USPTO here: https://www.uspto.gov/sites/default/files/documents/OpenAI_R...

We expect that IP and AI will be an interesting policy discussion around the world in the coming years, and we're eager to participate!

GitHub Copilot 5 years ago

We think that software development is entering its third wave of productivity change. The first was the creation of tools like compilers, debuggers, garbage collectors, and languages that made developers more productive. The second was open source where a global community of developers came together to build on each other's work. The third revolution will be the use of AI in coding.

The problems we spend our days solving may change. But there will always be problems for humans to solve.

GitHub Copilot 5 years ago

It shouldn't do that, and we are taking steps to avoid reciting training data in the output: https://copilot.github.com/#faq-does-github-copilot-recite-c... https://docs.github.com/en/early-access/github/copilot/resea...

In terms of the permissibility of training on public code, the jurisprudence here – broadly relied upon by the machine learning community – is that training ML models is fair use. We are certain this will be an area of discussion in the US and around the world and we're eager to participate.

GitHub Copilot 5 years ago

Hi HN, we've been building GitHub Copilot together with the incredibly talented team at OpenAI for the last year, and we're so excited to be able to show it off today.

Hundreds of developers are using it every day internally, and the most common reaction has been the head exploding emoji. If the technical preview goes well, we'll plan to scale this up as a paid product at some point in the future.

Any forks that make the same very minor changes that ytdl made -- not to provide specific instructions for copyright infringement -- will be reinstated. Or they can complete the DMCA counter-notice process, and if that is unchallenged or successful, be reinstated that way.

(Also, the person who wrote this article said that they contacted me, but I never received anything from them.)

The mitigations you suggest are all logical. However, there are legitimate reasons to run CI and tests for outside contributions without taxing maintainers with the cognitive load of having to evaluate whether each contribution is CI-worthy.

The attack vector in the article is not the main way miners try to steal CPU from the GitHub community. It's just an interesting one that the journalist chose to write about.

This is a cat and mouse game. We add code to detect and disable abuse – sometimes in very clever ways – and then the abusers come up with a new way of circumventing that detection. In order to prevent miners from creating long queues for legitimate free users of GitHub Actions, we have to stay on top of this all the time. So the miners are not just stealing CPU time, they are also stealing engineer time. Because without mitigations the miners will consume all available CPU, and because devising abuse countermeasures is, for whatever reason, a very powerful nerd snipe (including for me!). The sad thing is that it's displacing time that would be spent improving Actions in other ways.

(GitHub CEO)

It looked at first like this thread was complaining that GitHub doesn't work in something called Qutebrowser. So I downloaded Qutebrowser to try it out, and GitHub seems to work just fine.

We care quite a lot about web standards and broad browser support at GitHub, so if anyone is able to tell what the issue is, please let me know: nat@github.com. Thank you!

We have a really excellent policy regarding government takedowns: https://docs.github.com/en/free-pro-team@latest/github/site-...

Among other things, the policy requires that governments that want content removed from GitHub issue a lawful request to us, which we then push to a public repo: https://github.com/github/gov-takedowns

So you are able to see all of the government takedown requests that we have processed, there. You'll notice that there are only 3 directories in that repo: Russia, China, and Spain. When we do (reluctantly) take down content at the request of a government, we try to limit the takedown only to viewers in the country that made the request, rather than doing a global takedown.

It wouldn't help with sanctions. As I said in the blog:

The US has long imposed broad sanctions on multiple countries, including Iran. These sanctions prohibit any US company from doing business with anyone in a sanctioned country. (These sanctions can also apply to non-US companies whose activities directly or indirectly involve the US, including merely having payments that flow through US banks or payment mechanisms like Visa.)

The license is specific to GitHub.

However, we don't want this to be a competitive advantage for GitHub; developers should choose GitHub because it is better, not because it has a license from OFAC. So we have taken it upon ourselves to advocate for OFAC to allow developers in Iran and other sanctioned countries greater access to all platforms, and we will continue to do so.

This kind of change would likely require an update to OFAC’s regulations, the issuance of an updated general license, or the issuance of formal guidance from the agency. We hope that OFAC’s issuance of a license to GitHub will help pave the way for broader access to similar platforms.