Wouldn't it be easier if they published a list where they scraped their data from in the first place. Filling out forms, scanning id and sending it only to learn they didn't capture any of your data seems like such a waste of time.
On the other hand, they already know which sites they used to scrape data. So publish it, maybe with a handy lookup portal where you can enter urls to see if it got scraped.
I prefer an opt-in model, but that's not likely to happen any time soon, so this seems reasonable while this gets legally sorted out. Just because something is transmitted publicly doesn't mean it's without copyright. Otherwise any song broadcast on radio is up for grabs to be resold by anyone receiving it.