HN user

gwittel

429 karma

Computer security Anti-spam / malicious content filtering Backends Scaling

glwittel [at] gmail [dot] com

Posts0
Comments155
View on HN
No posts found.

Chronicle has a lot of great resources. I’ve been out of Java for a few years but when I had to write high perf code, the libraries and blog insights were invaluable.

A lot of the advice is good in general - keeping things simple, generating fewer objects, etc. Profiling with Yourkit or JMH to find and improve slow spots.

This let us build high performance software that had a vast difference in say p90 input size and p99. It involved rewriting third party libraries to let us 3x speed and substantially reduce GC. I think in the end we were around a p99 of 7ms and p90 well under that (the data spread was kbs to megabytes) across hundreds of millions of inputs per day.

Pretty much this. We see CF hosted/protected content all the time. CF does nothing for days or weeks. By then the campaign is gone and on a new account etc. Tycoon and Kratos are two phish kits that heavily use CF.

CF does have small anti-abuse teams, but it’s just not a business priority for the company to do better. We’ve tried many times to engage at a corporate level and the bottom line is they don’t care - they don’t want to police content as often stated by the executives.

"GitHub only gets better if people who give a shit stick around to make it better"

At a basic level I appreciate this sentiment. However, the common dysfunction I see in large corporation is its not the lack of people who give a shit. Its lacking a sufficient number of people in positions of power that give a shit -- such that they can actually make change happen.

All too often competing pressures (features, profit, delivery speed, politics) take precedence; not leaving time for things that would really move the needle. In essence, too many leaders are happy to ship garbage; they don't care (or don't know).

If Github were to put out a statement saying "service quality is our priority", it is fairly meaningless. If they added "here's how we'll get there", maybe it helps some. Moreso -- "from now on executive compensation is tied to these SLOs", then maybe something would actually happen.

Yes. In the past I helped sort out tooling like this for competitive analysts. There are a few ways this is done:

1) Check the businesses’ MX record. Often this points to a third party provider like Microsoft or Google. 2) Connect to the mail server identified in the MX record. Sometimes these have banners that identify the vendor (vs something generic like sendmail) 3) Email headers from messages sent to users in the company (or sometimes a bounce). Often these have headers from one or more providers. You’ll have to sort out the path to understand which bits were added by the sender/recipient path though.

These days often companies have multiple providers (security) so they might have one at the edge (mx) and more internal hops. You can usually see these in the headers.

Yes. I remember listening to it on the radio. The DJ used the handle Hard Hat Mack. It was pretty awesome to hear SID music over the radio.

I found this archive that has some of the shows recorded, and set playlists (I'm giving two links as the site is using frames so the top level page requires you navigate the menus to get to these):

  - Recordings: https://www.transbyte.org/SID/KDVS.html
  - Playlists: https://www.transbyte.org/SID/6581.html
The set playlists (using HVSC) works. For actual recordings they're 404 from this site -- arnold.c64.org is gone. But there are a few archives of the arnold.c64.org site! This should help re-construct the original links from the above page:
  - https://www.mmnt.net/db/0/0/arnold.c64.org/pub/sidmusic/lala/ra
  - https://archive.org/download/arnold.c64.org  (download the whole thing and dig into pub/sidmusic/lala/ra)
Due to the era, most of the files are in RealAudio format; with a few MP3s as well. Wonder if this could all be re-posted somewhere in modern formats to make it more accessible.

Its possible the authors are still around and have more copies; doubtful KDVS has archives, maybe tapes buried in the library.

Anyway, hope this helps! Its a cool piece of history and brings back a few memories.

I work a product that involves a security crawler (phish, malware detection, etc). It’s just a new arms race. Crawlers will adapt.

Cloudflare is already heavily abused by threat actors to host, and gate their malicious content. This means our crawler has to handle anti-bot and CAPTCHAs. It’s a pain. Cloudflare is no help.

They have a “verified bot” program but it’s a joke for security. You must register a unique, identifiable user agent, and come from a set of self declared IPs. Cloudflare users can check a box to filter these bots out. And now you're easily fingerprintable so the bad guys can just filter you even without Cloudflare’s help.

So now we have a choice. Operate above board and miss security threats. Or operate outside the rules (as opaquely defined by Cloudflare), and do right by our customers.

All of this on CFs side is to solve a real problem. Unfortunately by not working with the industry in a productive manner, Cloudflare is just creating new problems for everyone else.

Google search results are full of garbage pages populated with LLM generated content (the pages exist solely to serve ads and capture search results).

Search spam is not new, but the use of LLMs simplifies the ability to make pages that look like legit content (increasing the likelihood they’ll show up in search results).

They could police their content. Or if they don’t want to, they could meaningfully partner with the security industry - create a “security bots” program, respond to takedown requests in days not months, etc.

You can. Sort of. The good bots list is basically driven by a fixed user agent. And customers can set their preference to not allow “good bots”.

Not so good for security work.

It’s similar to their abuse reporting. They give your info to the site owner. Gee thanks, that’s just what I want to do.

I’m really mixed on this. Anti bot stuff is increasingly a pain point for security research. Working in this space, I have to work against these systems.

Threat actors use Cloudflare and other services to gate their payloads. That’s a problem for our customers who are trying to find/detect things like brand impersonation and credential phish. Cloudflare has been completely unhelpful. They just don’t care.

I was looking at [1] recently to understand omicron variant positivity length and they cite a few other papers. The article [1] is publicly available. I haven’t checked if all of the others are.

* Routsias JG , Mavrouli M , Tsoplou P , Dioikitopoulou K , Tsakris A . Diagnostic performance of rapid antigen tests (RATs) for SARS-CoV-2 and their efficacy in monitoring the infectiousness of COVID-19 patients. Sci Rep. 2021;11(1):22863. doi:10.1038/s41598-021-02197-z

* Currie DW , Shah MM , Salvatore PP , et al; CDC COVID-19 Response Epidemiology Field Studies Team. Relationship of SARS-CoV-2 antigen and reverse transcription PCR positivity for viral cultures. Emerg Infect Dis. 2022;28(3):717-720. doi:10.3201/eid2803.211747

* Korenkov M , Poopalasingam N , Madler M , et al. Evaluation of a rapid antigen test to detect SARS-CoV-2 infection and identify potentially infectious individuals. J Clin Microbiol. 2021;59(9):e0089621. doi:10.1128/JCM.00896-21

* Killingley B , Mann A , Kalinova M , et al. Safety, tolerability and viral kinetics during SARS-CoV-2 human challenge in young adults. Nat Med. 2022;28:1031-1041. doi:10.1038/s41591-022-01780-9

[1] COVID-19 Symptoms and Duration of Rapid Antigen Test Positivity at a Community Testing and Surveillance Site During Pre-Delta, Delta, and Omicron BA.1 Periods. https://jamanetwork.com/journals/jamanetworkopen/fullarticle...

The Inkplate 10 is great. I haven’t gotten a lot done other than toy stuff, but so far it’s been a mostly good experience.

Another nice entry point is micropython. Some of the getting started stuff has gaps but overall nice and simple if you’re more comfortable in Python. Major libraries have ports so it mostly is an easy dive in.

In a past job I’ve seen crappy crawlers from badly designed security applications do stuff like this. An an example one customer was using Trend CAS to scan all URLs in their inbound email. This causes big bursts of traffic on our systems.

The crawls came from Azure and AWS. Forged UAs, repeat hits in the same URL, etc.

Having worked in anti-abuse for nearly 20 years this is spot on. Even if it were possible, publishing “the algorithm” isn’t going to solve anything. It’s not like it can be published in secret or avoid being instantly obsolete.

All of this is an exercise balancing information asymmetry and cost asymmetry. We don’t want to add more friction than necessary to end users, but somehow must impose enough cost to abusers in order to keep abuse levels low.

Unfortunately for us, it generally costs far less for attackers to bypass systems than defenders to sustain a block.

As defenders we work to exploit things in our favor - signals and scale. Signals drive our systems be it ML, heuristics, signatures (or more likely a combination). Scale lets us spot larger patterns in space or time. At a cost. 99%+ effective systems are great, but at scale 99% is still not good enough. Errors in either direction will slip by in the noise; especially targeted attacks.

As a secondary step, some systems can provide recourse for errors. Examples might include temporary or shadow bans, rate limiting, error reporting, etc. Unfortunately, cost asymmetry comes into play again. It is far more costly to effectively remediate a mistake than it is to report one. We’re back to cost asymmetry.

All of this is suboptimal. If we had a better solution, it would be in place. Building and maintaining these systems is expensive and won’t go away unless something better comes along.

tl;dr version: assholes ruin it for everyone.

Definitely. Twitter seems to have not been doing a lot of standard best practices for a company of their size.

My intent was pointing out that engineers with high level access to their dev machines is pretty common in tech. Not that other controls like policy enforcement are also often absent in tech (esp in larger companies). Hard to know how common that is -- seems unusual at least in big tech.

In reading Mudges' complaint, it really paints the Twitter leadership (esp. Agrawal) as simply not caring about security enough to do anything about it. Instead you had an org with massive amounts of technical and operational debt, and leadership not willing to invest in it. There are always tradeoffs between fixing technical debt and building new features. Twitter leadership chose to ignore (and to some extent, hide) the problem rather than invest. They certainly aren't unique in having a security plan that is built around hope.

Engineers having full control over their dev machines up to and including preventing system updates is not ideal; but not out of the norm for tech. Poor data access controls, and out of date server fleets (where I'd expect updates to be pretty automated) are far more worrying to me.

For reference, here’s what we have: 36KBTU AIRHANDLER/HEATPUMP HI-STATIC M SERIES DUCTED SYSTEM INDOOR MOD# SVZ-KP36NA OUTDOOR MOD# SUZ-KA36NA2

It’s a slightly older model as we had height restrictions to work around. This prevented us from getting a newer or hyper heat model. IIRC ours had good efficiency into the 20F range which was plenty for us.

In comparison, the Trane dropped efficiency at 50F and needed heat strips at that temp (so pretty crap).

I used Mitsubishis site to find the local “diamond” contractors (those factory trained and do enough volume). Then cross referenced vs yelp.

https://www.mitsubishicomfort.com/find-a-contractor

I see a few in Portland and several in nearby zip codes. Hopefully one can work for you. There are different tiers of “diamond” so you can compare if the difference matters to you.

One thing that’s a bit different is the air handler and outside compressor run on one circuit (mine is a 3 ton unit). So there’s a power line between the 2 units. That threw off our city inspector. But it works out nice since I now have an extra 20A breaker free :)

Many of the Mitsubishi heat pumps work with central ducted systems just fine. I have one (replaced a central gas heat, electric AC system). It’s just a different air handler but the heat pump was the same as would have been used in a mini-split install.

When cross shopping the Mitsubishi vs Trane, the Mitsubishi was miles ahead. I didn’t even get the most cold weather efficient option (not needed for my climate).

The conversation we should be having is where are we now, and what is good enough?

Having worked in large scale anti-abuse detection for most of my career (~18 years), the points mentioned line up in the Twitter thread align with my experience. Scaling in this area is hard. 99% efficacy sounds great, until you say 99% out of millions/billions. The amount of FNs ('bad' or unwanted things) is still substantial enough for users to notice. Taking a 229m active user count [1], 99% fake account detection efficacy sets you at 2.2m fake accounts. Looking into tweets/day you've similarly large numbers if you want to look at content detection.

Twitter can most likely do better given the right resources, people, and leadership support (Facebook has similar problems aligning all 3 of those). Once they have those, the open question is how much better they can get. Each incremental increase in efficacy gets more expensive.

To top it off, as detection gets better, you think those abusing Twitter will sit still? Of course not, they'll change tactics (content, usage of hacked accounts, etc.).

[1] Twitter 10-Q 2022-Q1 - https://www.sec.gov/ix?doc=/Archives/edgar/data/0001418091/0...