I'm not talking about parsing HTML. I'm talking about getting data out of HTML which is completely different.
One doesn't need to guess at all possible combinations of badness using a regex, you just need to find the data you need and adjust the regex to the badness of that particular page.
Do you know anyone that mines for data out of HTML "in the wild"? I didn't think so.