HN user

typeiierror

253 karma
Posts7
Comments13
View on HN

I know this is satire, but I have an adjacent problem I could use help with. In my company, we have some legacy apps that run, but we no longer have the source, any everyone that worked on them has probably left the planet.

We need to replatform them at some point, and ideally I'd like to let some agents "use" the apps as a means to copy them / rebuild. Most of these are desktop apps, but some have browser interfaces. Has anyone tried something like this or can recommend a service that's worked for them?

I've always wanted to do this - do you have any articles / videos that helped you get started? Is Fighter Factory Studio still the way to go?

As an aside, this feels like a fun application of Gen AI (generating sprites + movesets based on photos of friends or hand drawn characters, etc).

I've been looking for something like this to query / interface with the mountain of home appliance manuals I've hung onto as PDFs - use case being that instead of having to fish out and read a manual once something breaks, I can just chat with the corpus to quickly find what I need to fix something. Will give it a shot!

Try applying the problem to other issues to see the impact:

* An advertiser wants to place ads on sites / tv networks that have an audience that is more likely to buy their product upon seeing their ads. If they don't want to violate privacy, they run a survey. What if the response rate among a historical disenfranchised group (e.g. African Americans) is terrible? The modern "data driven" marketer would see little reason to advertise on Black media properties. This isn't a fictious example - it's a current problem in the media planning / agency industry.

* A local government has to decide between investing in more ESL resources in public education vs. other competing budget needs. They look at census / community survey data (which some Hispanic and immigrant populations are fearful of responding to d/t politicization) and decide to prioritize other asks due to undercounted demand. The data could also be skewed in other ways that warp their decision, like allocating budget to school zones that only represent specific immigrant communities that haven't historically been disenfranchised.

The big picture issue here is governments/businesses making decisions with bias information leading to incorrect conclusions, and the only know recourse currently is to scrap privacy.

The sample selection / non-response bias highlighted in this write-up is a _Big Idea_ problem I've been thinking about recently:

  Limitations...*Trust in surveys and political leanings:*
  About 95 percent of people contacted for the panel chose not to participate because of lack of trust in having a third-party application installed on their computer or other concerns for privacy.
Think about that - a reputable, privacy-first organization asked people to opt-in to fully consented, voluntary, compensated research and ~95% declined! I can't even imagine what hidden skews are present in the 5% that agreed. This issue is systemic in consumer research and impacts both public (e.g. election polling, U.S. census) and private (pharmaceutical trials, media/advertising research, voluntary AI/ML training daat) polling.

Governments and businesses make biased, potentially discriminatory decisions if a non-random segment of the population chooses to never be counted. The ad industry attempts to circumvent this through non-voluntary passive tracking, which trades off non-response bias with bulldozing user privacy. The headwinds are only growing too, as consumer awareness of privacy lapses and the politicization of polling continues to reduce who participates in opt-in research.

Finding a solution to this that doesn't resort to privacy-eroding tactics is a moonshot level problem in terms of the size-of-the-prize if solved.

Brian K. Vaughan's comic Private Eye [1] foreshadows what might happen if this dataset is breached. The premise is the digital cloud "bursts" - all private data is suddenly dumped and searchable - forcing people to completely abandon their identities and assume new ones - changing their name, appearance, re-starting their careers, etc.

When you consider this in the context of technologies like Voco [2] and Face2Face [3] that can fabricate a speech or make a fake "hot mic" video from a public figure, it makes you wonder if we'll ever be able to prove things are __true__ in the future, and what the value of our identity is if it can be shattered beyond repair due to negligence from a third party. What do we do then? How do you cryptographically sign yourself?

[1] http://panelsyndicate.com/comics/tpeye [2] http://www.bbc.com/news/technology-37899902 [3] http://www.graphics.stanford.edu/~niessner/thies2016face.htm...

Sometimes I wonder that when we shower criticism on Facebook about privacy concerns, we're missing the forest for the trees. The bigger issue I see is the sheer amount of eyeballs trained exclusively to Facebook's content.

What does it mean for society when Facebook can demote a challenging but important article (say, of war reporting) in your newsfeed so it can promote your friend's Wedding photos, because an algorithm says that challenging articles cause people to leave FB, reducing page views and ad revenue?

When you take into account the full range of tracking methods cited in the article, I think the impact of Tor would be limited. For example: you probably already have an app on your phone that has an SDK from Gimbal (beacon company in the article) or a similar firm that reads beacons and tracks your location passively. So while Tor would disguise the source of your web requests, the SDK still has free reign to send your location and your Ad ID back to the mothership.

A couple of reasons - data quality and usage restrictions are the top. Here's an example - say I'm Big Box Retailer & Co and I want to measure how many people who visit my store also visited my competitor, Small Box Retailer Inc. Where can I go to buy that data?

-Facebook or Google don't sell their raw user data. They can tell me how effective the ads I placed on their sites were at driving visits or sales by matching my customer's email & phone numbers to their users (see [1] and [2]), but they won't give me data on my competitor, and the data will always be aggregated. So on to the next idea..

-Mobile ad server data aggregators might sell me their raw data, but the quality isn't great. Sure, they track 100m+ devices...but how many times do they see each of those devices a day? For most devices it's 10 times or less, so you're going to have a big problem with false negatives (people who visited Big Box Retailer and Small Box Retailer, but due to the sparsity of the data I miss one of those visits). On to the next...

-Foursquare is newly in the data business, but despite having a big audience (50m MAU), only ~1.3m users have opted in for continuous measurement ([3]). On top of that, Foursquare's audience is pretty skewed - if Big Box Retailer's customers are older, I'm going to have trouble finding them in Foursquare's data set, which means the effective size of their 1.3m 'panel' is really something like 200k-500k. On top of that, I can't survey them to verify they actually visited Small Box Retailer instead of the McDonalds next store, since the GPS data Foursquare pulls is LastKnown instead of exact.

-...so if the only data I can buy on the open market is skewed, not really that big (not to mention collected under potentially dubious privacy policies), whats my alternative? Thats the need these panels are filling.

As an aside, if you're curious what apps on your phone are collecting your location data, you can use a self-hosted MITM server w/ SSL decryption to sniff your own traffic. Here's own for Android: https://code.google.com/archive/p/sandrop/

[1] http://www.wsj.com/articles/google-touts-mobile-ad-technolog... [2] http://www.adweek.com/news/technology/facebook-gives-retaile... [3] http://techcrunch.com/2016/02/22/attribution-by-foursquare/

I work peripherally in this industry, so I can provide more context to the data used here; based on the press release[i], Clear Channel is using three different types of data for this:

1) AT&T cell tower data. All major US carriers collect aggregate movement data, and some have productized it (check out Grandata and Streetlight Data if you're interested). They're likely providing something like a persons count by daypart to Clear Channel at some geography, likely census block group. AT&T likely provides course demographics as well (either by purchasing them from a data broker like Experian or Epsilon) or by looking up the aggregate demo characteristics reported by the US Census for the block group of the subscribers household. As an aside, current gen (4G) cell tower data isn't very precise - maybe 100m accuracy or worse.

2) Placed opt in GPS panel data. There are many market research companies that pay consumer run location tracking apps (mFour and Instantly are other examples). Placed is probably the biggest (~1mm panelists).

3) PlaceIQ mobile ad server GPS data. PlaceIQ, xAd, Factual, Verve, Ninth Decimal...all of these companies read the lat / long coordinates provided by mobile SSPs in mobile RTB bid stream to create location segment profiles associated with your phone's Ad ID. The data isn't very accurate (mobile ad fraud is a problem...an app change the GPS coord from rural Kansas to downtown Manhattan to juice their CPM in an auction; also, most of the GPS used for buying mobile inventory is via "LastKnownLocation" apis which are notoriously inaccurate). These guys generally use their data to group your Ad ID into a segment (if they see you at a Wendy's, they'll sell your Ad ID to mobile RTB bidders as a "Frequent Wendy's eater"). Clear Channel is probably using this to see if exposure to a billboard caused you to make a purchase that they can attribute to your Ad ID (say via the advertisers CRM database), or to augment the demo data from AT&T with demo segments they can buy from mobile data exchanges.

[i]: http://www.businesswire.com/news/home/20160229005959/en/Clea...