HN user

themadprogramer

167 karma
Posts45
Comments17
View on HN
outsiderdata.netlify.app 3y ago

The Kakhovka Dam Disaster in Data

themadprogramer
2pts1
datahorde.org 4y ago

The Internet Archive and Alexa Internet

themadprogramer
4pts0
maribelhearn.com 4y ago

Archive of Royalflare, a 15-Year Touhou Scoreboard, by Maribel Hearn

themadprogramer
1pts0
datahorde.org 4y ago

YouTube to Permanently Replace Discussions with Community Posts in October

themadprogramer
3pts0
support.google.com 4y ago

YouTube to remove Discussion Tab on Oct 12

themadprogramer
2pts2
blog.youtube 4y ago

YouTube got unlisted videos thanks to a high school teacher (2010)

themadprogramer
3pts1
data-horde-blog.tumblr.com 4y ago

YouTube begins privating pre-2017 unlisted videos through staged roll-out

themadprogramer
2pts0
www.nicovideo.jp 5y ago

Bad Apple Carved into Apples [リンゴの魔術師/Ringonomajyutsu]

themadprogramer
2pts1
users.erols.com 5y ago

In 1982, 20% of the world lived under a military junta

themadprogramer
2pts0
readcoop.eu 5y ago

Transkribus Is a Handwritten Text Recognition Project

themadprogramer
2pts0
hai.stanford.edu 5y ago

Stanford Brain-Computer-Interface achieves 18 words per minute

themadprogramer
9pts0
datahorde.org 5y ago

YouTube Will Private Old Unlisted Videos Next Month

themadprogramer
3pts0
datahorde.org 5y ago

Rescuing a Forgotten Nick Gem: The Avatar Yield Project

themadprogramer
1pts1
datahorde.org 5y ago

Flash Player EOL, Adobe Announces 12-Day Grace Period

themadprogramer
67pts38
datahorde.org 5y ago

Over 1M Yahoo Groups now available on the Internet Archive

themadprogramer
1pts0
datahorde.org 5y ago

Back in a Flash: The Super Mario 63 Community

themadprogramer
1pts1
twitter.com 5y ago

Time to Celebrate Flashcember, Flash Player's Final December

themadprogramer
1pts0
twitter.com 5y ago

Help Archive Tank Online (Flash Version) by literally just playing the game

themadprogramer
2pts0
datahorde.org 5y ago

Some thoughts on Khan Academy retiring courses and why not to despair

themadprogramer
3pts0
datahorde.org 5y ago

A New Breed of Digital Archiving and Preservation

themadprogramer
1pts0
datahorde.org 5y ago

We Just Rescued Thousands of Unpublished YouTube Captions

themadprogramer
2pts0
datahorde.org 5y ago

Harmful language in the RIAA's DMCA against YouTube-dl

themadprogramer
33pts2
datahorde.org 5y ago

Yahoo Groups Archiving Status (October Update)

themadprogramer
1pts0
datahorde.org 5y ago

YouTube is now hiding Attributions to Captioners who WANTED to be credited

themadprogramer
3pts0
twitter.com 5y ago

YouTube now hiding translators, at the expense of diligent authors

themadprogramer
3pts1
datahorde.org 5y ago

YouTube Workaround: how to continue to use community translations

themadprogramer
3pts1
datahorde.org 5y ago

Shutdown September: Tales from Websites Shutting Down Across the Interwebs

themadprogramer
3pts0
github.com 5y ago

Archiving YouTube's “unpublished” Community Captions while we still can

themadprogramer
3pts1
datahorde.org 5y ago

A Tool for Saving Unpublished YouTube Community Contributions, Time Is Short

themadprogramer
1pts0
datahorde.org 5y ago

The Untold Story of Why YouTube Is Removing Community Contributions

themadprogramer
3pts0

If the IA didn't exist, we'd have to invent it.

You know, a part of the original company vision for YouTube prior to the Google acquisition was really something akin to the IA, in that they did pride themselves with hosting footage of the Indian Ocean Earthquake:

https://www.youtube.com/results?search_query=Indian+Ocean+ea...

Now, acting as diplomatically as I possibly can, I can say that your suggestions of the IA and YouTube interfacing together were at a previous point in time a continuous process. But a number of factors have made direct cooperation between the IA and Google (thereby YouTube) come to a screeching halt.

At this current point in time, we stand at a historical crossroad. And I'm only here to just act as a humble messenger ;)

You have to, however, also consider the multi-domainness of YouTube videos. Yeah sure, there's billions of hours of clips not one person can watch in a single life-time. But unlike your 2000-year old Roman shopping lists, we have footage of events that are anchored to a particular time-period. Or location.

One of the most impressive things you can do is, try searching up a landmark. My personal favorite is the [Jumping Stone](https://www.youtube.com/watch?v=u1TtMN8nXTM) on the Nias Island of Indonesia. What would have otherwise just remained a novel tourist attraction, forgotten by the modernity of the 21st century, is now essentially a "tag" which has hours of footage associated with it. Thousands of tourists travelling back and forth, locals growing old, new people being born, buildings being built and demolished around it. You can even just study how video quality improved in that particular region. That there IS something wholly unique to this era and definitely worth preserving. YouTube as a company has figured the logistics of storing it, but the question of how humans can hope to read such data remains yet unanswered.

Why hello there, it seems you've found your way into the So some more backstory on this blogpost: I'm trying to start a nostalgia wave, Flashcember, since this December is well Flash's last December. I'm sure a few the people on here are actually good artists, so would any of you be interested in drawing/sketching one of your favorite flash games? Or perhaps remixing/covering a song? I mean I'd genuinely appreciate it if you could do anything like this at all, but if you want to go the extra mile could you share it around with a #Flashcember hashtag or similar?

There is actually a very good reason for the "low usage" rates, YouTube made the publishing of closed captions a lot more difficult a year ago: https://twitter.com/TeamYouTube/status/1167565334917742593

Starting from August 2019, only uploaders can approve submissions, when previously other viewers or YouTube moderation could also publish. A lot of people seem to have forgotten this, and YouTube never really updated the UI/Help pages to reflect it.

I have to wonder if whichever analyst came up with those stats was aware of this but turning a blind eye, or simply never noticed.

YouTube has had a "community contribution" feature (akin to fan translations) since around 2014, so viewers can help caption or subtitle videos of channels they frequent, for deaf or international audiences.

2019: A controversy erupts due to a particular case of a troll basically adding spammy, graphic translations on big YouTubers' videos. Said Big YouTubers complain and YouTube begins restricting the feature to make the process of publishing submissions increasingly difficult. (Previously other viewers could give approval through consensus and YouTube changed this to require manual action from all uploaders, severely lowering the rate of published captions)

See https://twitter.com/TeamYouTube/status/1167565334917742593 for further context.

2020: YouTube decides to kill the feature for good (by September 28), in the process they will be deleting all unpublished drafts. Seeing as this decision was more so motivated on reputation than functionality, it didn't take long for projects such as http://youtubexternalcc.netlify.app/ to emerge, which will continue to offer the feature externally.

Our project is to try and grab these drafts in hopes of passing them along. There are people who rely on them as a function. No particular person would go looking for a specific translation of a certain video, but the ability to have human-written captions instead of an automatic translation is something people continue to look for. So here we are trying to rescue submissions YouTube deemed weren't worth assessing properly.

For a tl;dr:

While YouTube assures us that it's a very few number of channels which rely on community contributions, this paints a very blurry picture of how many users are reliant on these.

Not to mention use cases which aren't available to any of the alternative caption providers:

1) When a user uploads their captions or Google generates them, it’s often a one-time process. If there’s a mistake, no one is going to go back to fix it. But with community contributions, users can build off of previous work by correcting one another, not unlike how Wikipedia works.

2) Even though community contributions carry the risk of sabotage, they also provide viewers with the power to moderate. Someone snuck in a joke on a video you were watching? Just go into the caption editor and edit it out! Community Contributions spare channels of having to moderate these manually by relegating the task to channel viewers.

3) Where automated captions are stuck with plain text, and uploaders are limited by the time they’re willing to invest to stylize their captions, there are people out there waiting to tap into their potential. YouTube supports a plethora of caption/subtitle formats, which a seasoned captioner can use to add color, formatting and emphasis!

4) In translation, there are times when we don’t want it to be too precise. An example could be explaining the meaning of a word, the wordplay in a joke… And although there are techniques to recognize proper names, there are sentences that automated translation is not designed to handle. For an illustrated version of this explanation see The Impact of YouTube Removing Community-Contributed Closed Captions

OP here. Ok, thank you, this is the kind criticism I crave!

I suppose by "vague" you mean there weren't any hard numbers.

Well I kind of linked the archive team tracker but you had to do some clicking to find it, and unfortunately I didn't have a statistic on the SYG's progress at the time of posting. However since posting I do have a rough idea of things:

Archive Team's own Tracker reports 2.76 TB of data to have been saved. The SYG team hasn't fully tallied up their data yet, but have counted the number of groups they've retrieved and/or are retrieving from to be around 123K.

I have since inserted this data into the post and would like to thank you for your brutal honesty :)