HN user

justcool393

112 karma

[ my public key: https://keybase.io/justcool393; my proof: https://keybase.io/justcool393/sigs/tfoz2bW3SMnYfPPq5zD5tEljG6bLdkk_iwihcmZ0stA ]

Posts0
Comments60
View on HN
No posts found.

to be fair, i once accidentally ended up on shreddit or new reddit or whatever they call it nowadays and i think there's something for managing your posts on reddit and seeing analytics about that or whatever

uncovering modern bot operations

this significantly overestimates how sophisticated the spam waves are compared to like ability. the 80% of spam filtering basically never was really done as far as i can tell.

a thankless threadmill, and user engagement metrics from fake users are still user engagement metrics.

that's probably it tho

gallowboob in particular was an interesting case because he was very much a real person. and oh hell did a subreddit I mod know that way all too well

i think the guy had a like a keyword alert on his username because like one of my co-mods on a subreddit would talk about the guy and then we'd get reports for "It's targeted harassment against me" (which are reports that are sent to the admins) like a few hours later. much to the dismay of him, we had a chat with the admins later and it was like "as long as you're not saying to do vote manipulate or harass the guy it's fine."

i think a lot of it came from the fact that so like if you're modding a subreddit, a lot of people spend their time in the modqueue view rather than the comments so you see the targeted harassment reports on "xyz is a meanie head" and just click "remove" because it already is on the edge at best for most subreddits. this is how context gets lost. so people would see "unfavorable treatment" (not that it didn't happen, gallowboob's company's domain was soft-banned on reddit yet his subreddits had automod rules set to approve them) when if more people were as trigger happy on the report button a similar thing would happen

the admin problems with this are much worse because the comments tend to be looked at in isolation so saying "i'm gonna kill you", in isolation, looks without context pretty bad, but might be part of a joke chain or song meme that reddit likes to do every so often. take into account the fact that admins get whiny sometimes if your AEO removals are too high. then take into account the AEO guy's Tarot card reading and whether Mercury is in retrograde and you get a lot of mods who are a bit trigger happy, esp when people've gotten banned for approving stuff the AEO removed for dumb reasons

this somewhat led to a bit of an inflated ego with regards to reddit but eventually from what i see he left... at least under that username anyway.

"model" or "service" would be the term.

basically the point is that it's not protected because it doesn't fall under any classification of IP (it's not copyright or patent since it's mathematical outputs of a mechanical system, it's not trademark because obvious, and it's not a trade secret because it's not a secret)

the Unidan incident?

iirc it only got noticed at the time because of an argument between him and Ecka6 which led to the somewhat famous "here's the thing you said a jackdaw is a crow" copypasta

some more information perhaps?

banned_by true is more accurate to say "admin or automatic". in "admin mode," you can see these although not sure the UX for these nowadays now that it is spewing a gazillion lines of text into them).

Anti-Evil Operations removals (nee Trust & Safety) are generally human(-assisted) actions (although these actions can be applied en masse). there's some more information nowadays in the API which was really nice. it also helped because people stopped blaming "the mods" for removals when the spam filter slopped all over the place. this was also annoying because previously you had to previously guess from the API how it was removed even if you were a mod.

the 3 ways to remove a post/comment (i.e. in reply to: train_spam):

- remove not spam: removes it but doesn't train the spam filter, obvious

- spam: removes it and trains the spam filter, obvious

- confirm spam: only happens when you remove after removing for any reason, *does not* train the spam filter

- reinforce spam: trains the spam filter even if the spam filter already caught it. *does* train the spam filter. you can do this by doing `action: spam` in automod. not sure if there have been any more in the last few years

also you can tell the legacy of "removals", back in the day stories were "banned" instead of "removed" by moderators and administrators.

also also also... you can see a lot of the stuff from this article in the `approved_by` side of it as well. if you hover over a checkmark of someone who has been unshadowbanned, you'll see it says "approved by Reddit (shadowban removed)"

if an admin manually unspams someones stuff (say someone who got accidentally shadowbanned and got hit with an overzealous spam filter multiple times >.>), it'll say "approved by <username> (all)". there are some consequences to this. it approves stuff that has been "filtered" (as AutoMod filtering is a weird hack where it removes something but keeps in the modqueue).

spammit

i believe this is the thing that is "pretty similar to a naive Bayesian classifier"[1][2] that reddit used. /u/Deimorz iirc was a reddit dev at the time and it was somewhat public info. i say somewhat because you kinda had to be both interested in the this and probably be around the metasphere

iirc from some other comments i pieced together there are also per-subreddit spam filters. in the olden days sometimes they'd get way out of whack and you could ask an admin to reset it for you... or something idk

em

guessing em in this case btw refers to /u/hueypriest, who was reddit's GM at the time

would’ve been catastrophic for Reddit’s spam issues

the thing that surprised me at the time was just how bad reddit's spam filtering is. i did a small little thing at the time where i'd just look at stuff following some basic spam filtering rules (like stuff you'd probably get out of an artisinal spamassassin ruleset) and even that deluge was amazing to see.

like the ML stuff is cool and all but seriously 90% of this could probably still be solved with some basic rules. the profile hiding stuff didn't help either but that was way after my time.

[1] https://reddit.com/r/TheoryOfReddit/comments/10ko5h/comment/... (2012)

[2] https://www.reddit.com/r/modnews/comments/6bj5de/state_of_sp...

their IP

it's not IP, and it's certainly not their IP

the TOS

oh no, the terms of service how dare people break those. you don't get to claim fair use while CFAAing everyone's actual IP then whine about the tos, and then when called out on spying on users point to it as if it being in the tos somehow justifies it

a lot of other malware has a tos too but we still call it for what it is

in a lot of cases, the leaders of the communities are not following the rules. (see the ppl talking about ndas and such)

in any case, this isn't like "oh we don't want to build an apartment building because it might drop the value of a single family home halfway across town."

it isn't even like "we want to build this train line which will have some negative externalities but the positive effects (and externalities) are worth taking a hit in some areas"

the problems with the datacenters are that like (1) the service its providing (LLMs) has dubious societal value, (2) the direct negative effects such as noise pollution and such have been pretty well documented, (3) the indirect negative effects like massive strain on infra and (4) the people pushing them most heavily are effectively attempting to invade the communities, peddle conspiracy theories about "china" being behind the opposition, and demand to be specially treated because they were bankrolled by big tech, etc.

some people when this topic come up act like anyone opposed is some nimby who hates societal progress or smth and who is super concerned about that their home estimate might go down. but like communities do recognize the need for zoning and restricting certain things being built.

you need the thing being built to both (a) actually be a good that helps the community (or have a very very good reason why some damage to the community is justifiable (datacenter projects generally don't) and (b) need to contain negative externalities (which is why we don't put the chemical plant next to the elementary school even if it's the most economic option). people recognize these things on some level.

i.e. if the maintainer is serious enough to buy stars, is not in theory likely to spend time /money in maintaining /improving the project also ?.

i mean if maintainers clearly spend much more time and effort on fraud than actually improving the project, why should I at all believe they would, let alone trust their judgement with regards to other things such as technical choices for example

i mean ofc but like you can self-host pypi and the "Docker Hub" model isn't like VC-expected level returns especially as ECR and GHCR and the other repos exist

well no, (clean room )reimplementations of APIs have done since time immemorial. copyright applies to the work itself. if you implement the functionality of X, software copyright protects both!

patents protect ideas, copyright protects artistic expressions of ideas

it's even worse than that and i hope people recognize that it's not that he's a True Believer (though the TBs are often hilarious)

it's that he has no ethics to speak of at all. it's not that he's out of touch, it's that he simply does not care.

even 100 kB dynamically generated pages should be a piece of cake. if it's CRUD like (original op's site is), it should be downright trivial to transfer that much on like... shared hosting (although even a VPS would be much better).

(in original op's case, i clocked 197 requests using 20.60 MB while browsing their site for a little bit. most of it is static assets and i had caching disabled so each new pageload loaded stuff like the apple touch icons.)

honestly you could probably put it behind nginx for the statics and just use bog standard postgres or even prolly sqlite. nice bonus in that you don't have to worry about cold start times either!

why wouldn't you? these are easily compressible text files. storing even like 100x into a 400 day (at most, the default for GH is 90) box is downright cheap to do on even massive scales.

it's 2025, for log files and a spicy cron daemon (you pay for the artifact storage), it's practically free to do so. this isn't like the days of Western Union where paying $0.35 to send some data across the world is a good deal

yeah, i mean i guess what i'm trying to say is that the breaking point is very far up there as computers have gotten towards breakneck speeds, especially on the technology side, for the goal being achieved. it's downright difficult to hit the limits unless you're throwing effectively a DDoS at it.

i think the big thing though is that it's a community and so people are actually willing to support that even if it means the amount of 9s of availability is slightly fewer (although in practice, many providers bust right through their "9s" SLAs without a care in the world) and given a migration from a VM provider to the dedi occurred, migrations obviously can happen if failure presents itself.

i think that people vastly overstate the costs of this sort of thing and it's super bizarre. if you're treating this as a big official corporation™ and such and want to pay 500 devs like $200k/year or something to make work, then yeah you're gonna have problems.

but if you want to build a social network and aren't dreaming of being gazillionaires for it (which is quite reasonable), then you can get by very easily. how do I know this? because... well it's being done successfully. not was done successfully, is done successfully.

you can probably even get people to help out on it.

you can build a social network with a dedi running nginx hosting your Python application running on a Linux box backed against Postgres (and redis for session storage, although even that is a bit overkill) for like $80/month deployed with a "deploy.sh" script that you run to kick the damn thing into running (Docker is used in dev only, but could easily work here). should you probably add health checks or whatever? yeah. it still works really well.

this scales well past the 100k users mark.

what about video/images/etc? well, this nginx server happily sends out user uploaded video storing them as files on a bog standard ext4 filesystem. backups exist of the site.

the "stack" i mentioned here isn't fancy or particularly tightly optimized, it's in fact pessimized in a lot of ways. hell I know there were a gazillion ways we could improve performance of our application. show the backend app to a game dev and they'd probably want to start strangling people with how poorly optimized most of the actual app is.

and still, it scales well.

again, I stress that this isn't some theoretical idea, this is actively being executed. the entire venture makes money for the team from the users who willingly (and unforcibly in order to use the service, the actual site is free to use in its full form) give money. this isn't ZFS. this isn't Rust. this isn't using some blue-green deployment. this isn't spending hours toiling away at which sysctl to set to squeeze every last cycle out of each box. this isn't behind some massive CDN with "internet scale" boxen or even (for the video serving part) behind any anti-DDoS service.

it's just a matter of doing actual engineering and being willing to actually build the things you want to build.

The US definition of "derivative work" is quite broad, and seems to cover linking just fine.

the problem is the GPL view seems doubtful and has not only bad implications for software copyright but copyright of... well literally anything else. I mean, remember what linking actually is (especially dynamic linking), you're basically just making references to certain things.

the analogy that I can best describe is this: if you're writing a paper on something whatever, and you link to a page number of a book, that doesn't make your paper a derivative work of that thing per se.

if I say in the middle of my novel new text on foobars and fozzinators, hey "book A page 32" has instructions for how to confabulate your fozzinator or "book B page 42" has the values needed to valienate your foobaz, referring to those things in general makes no sense to consider this originally authored book a derivative of A, B, or A and B.

or for a more concrete example, saying Microsoft should be the final authority on who can interoperate with their products or saying that the people who publish research are automatically derivative works of other peoples research[1] papers or people who write articles can't even REFER to other articles in such a way.

[1]: research itself may come from derivative ideas of course, but I'm talking about the copyrightable elements here; i.e. not the facts necessarily presented within, but rather how such facts are presented and laid out. copyright does not cover facts (true or false[2]), but your presentation of such facts are.

[2] https://thowardlaw.com/2023/04/false-facts-denoted-as-actual...

it's also worth bringing up some arguments made by Theodore Tso over this very issue in 1998[1]:

Consider the following --- what defines "link"? Does an RPC call mean linking? What about shared libraries? What about making calls via the system call interface? What about running GPL'ed programs via the system() command from a commercial program? If you take things to extremes, a commercial program which uses the system() program will be interfacing with the GPL'ed /bin/bash on most systems --- is that considered "linking"?

And if not, what is the legal distinction between what /etc/ld.so does when it maps a GPL'ed library into memory and the thread of control is temporarily tranfered from propietary code to GPL'ed library code when a library function is called, and what happens when a propietary program calls system() and the kernel maps /bin/bash into system memory, and the thread of control transfers temporarily from the propietary program to /bin/bash? You can see how things can get quite ridiculous quite quickly.

[...]

The FSF assertion also a very dangerous legal argument to make. If this is true, does this mean that if you write code which happens to make use of interfaces developed by Microsoft and implemented by Microsoft DLL's, that Microsoft somehow has a claim over your code which it could enforce via copyright law? What about any i386 assembly code which makes use of the Intel machine language? Does Intel now have a copyright claim on all i386 object code, and can try to prevent people from executing i386 object code on non-Intel processors? (After all, when a Pentium interprets your object code, one could argue that it is "linking" your object code with the Pentium microcode, which is copyrighted by Intel....)

What the FSF is trying to advocate is one step down the slippery slope of interface copyrights, and we really, really don't want to go there.

i've seen arguments especially after the Google v. Oracle[2] decision and I think one in particular mentioned the sort of "reality distortion field"[3], which I found to be interesting, especially because a lot of open source projects that are GPL tend to rely on the good-naturedness of other users using their code in a way that's positively in spirit with the GPL. (which, to be fair, has probably helped open source immensely.)

but as Tso points out, therein lies a contradiction with GPL that at if its maximal interpretation to be correct, it's much much more dangerous, than if the GPL effectively is equivalent to the LGPL. but I don't think that (barring a world pre-this-case) this world is the one that is so. (no idea how the AGPL fits into this though, that sounds like a PITA.)

if I were ruler of the world, I'd say that symbol names are a matter of fact and thus should probably not be copyrightable by themselves. but then again I don't rule the world, and I'd probably have other things to change as well, even about copyright.

[1] https://yarchive.net/comp/gpl_linking.html

[2] https://www.supremecourt.gov/opinions/20pdf/18-956_d18f.pdf

[3] https://news.ycombinator.com/item?id=30404270

was it even reported? i heard a bunch of stuff that seemed to be hypothetical guessing like "satya must be furious" that seemed to morph into "it was reported satya is furious"

i've seen similar with the cloud credits thing, people just pontificating whether it's even a viable strategy.

it's hilarious how much people for no reason, want to defend the honor of Sam Altman and co. i mean ffs, the guy is not your friend and will definitely backstab you if he gets the opportunity.

i'm surprised anyone can take this "oh woe is me i totally was excited about the future of humanity" crap seriously. these are SV investors here, morally equivalent to the people on Wall Street that a lot here would probably hold in contempt, but because they wore cargo shorts or something, everyone thinks that Sam is their friend and that just if the poor naysayers would understand that Sam is totally cool and uses lowercase in his messages just like mee!!!!

they don't give a shit that your product was "made with <3" or whatever

they don't give a shit about you.

they don't give a shit about your startup's customers.

they only give a shit about how many dollars they make from your product.

boo hooing over Sam getting fired is really pathetic, and I'd expect better from the Hacker News crowd (and more generally the rationalist crowd, which a lot of AI people tend to overlap with).

we can only hope

i'm sick and tired of everyone sticking a chatbot on random crap that doesn't need it and has no reason to ever need it. it also made HN a lot less interesting to read

not the original poster but i want to chime in, the address is just poor. an IPv4 address is decent to remember and while DNS is useful, the issue sometimes is DNS, in which case you need the IP anyway.

12:141:::2315:14 or whatever is ugly and terrible syntax. firstly...

1. why colons? if you're typing in an IP address, you might also have to type in a TCP/IP port in colon format, why did IETF think overloading this was a good idea? it also makes the address scheme look ugly. the ways of getting around this by using brackets just look plain awful.

2. why do it in hexadecimal format? there are now 16 characters in each IP digit. i'd rather have (if we really need 128-bit addresses, which since we give out /64s, seems not to be the case) it be in dotted decimal. maybe even make the dotted decimals 16 bits if it's really an issue. do get rid of the myraid of other stupid ways to write IPs though, dotted decimal could easily be standardized.

3. DHCP is a good thing, why is it maligned? a central source of truth who is what is great, at the network level I can say "hey this person has the IP of X." it seems like a much better idea than SLAAC ever was.