HN user

mootothemax

9,482 karma

I'm a one-man web development machine by day, and work on various web projects in my spare time.

Terrible pianist.

Only properties include: https://fairinternetreport.com, https://crimerate.co.uk, and https://networkspy.co.uk

Twitter: @tbbuck

Posts180
Comments1,640
View on HN
fairinternetreport.com 2y ago

Visualizing 3.6 Million Internet Speed Tests: Animated Global Map

mootothemax
2pts0
fairinternetreport.com 3y ago

Mapping 24 hours of speed tests (~3.6 million) at 120 FPS

mootothemax
1pts0
fairinternetreport.com 3y ago

3.6M speed tests rendered at 120 FPS over a Mapbox layer

mootothemax
2pts0
fairinternetreport.com 3y ago

Over 3.6M speed tests happen every day. Here's what they look like.

mootothemax
3pts0
fairinternetreport.com 3y ago

The Heartbeat of the Internet: An Interactive Map

mootothemax
2pts1
fairinternetreport.com 5y ago

US is 23rd worldwide for internet speed, UK 31st, according to user speed tests

mootothemax
2pts2
broadbandbreakfast.com 5y ago

Trump Signs Executive Order on Artificial Intelligence

mootothemax
2pts0
fairinternetreport.com 5y ago

US internet speeds 91% faster in 2020 according to user speed tests

mootothemax
208pts261
news.ycombinator.com 10y ago

Ask HN: Anyone want campaignbar.com before it expires?

mootothemax
2pts2
twitter.com 10y ago

Clickbait Twitter Bot - Generates clickbait tweets based on current trends

mootothemax
2pts0
www.theguardian.com 10y ago

First almost fully-formed human brain grown in lab, researchers claim

mootothemax
3pts2
everynoise.com 11y ago

An algorithmically-generated scatter-plot of musical genres - with samples

mootothemax
7pts3
www.theguardian.com 11y ago

Astronomers solve mystery of the universe’s missing stars

mootothemax
2pts0
www.polygon.com 12y ago

Poland's Detroit: Life and (Computer) Games in Poland's Third-Largest City

mootothemax
1pts0
www.bbc.com 12y ago

Sony develops new 185 TB storage tape

mootothemax
65pts65
blog.twitter.com 12y ago

Twitter acquires Gnip

mootothemax
12pts3
www.coindesk.com 12y ago

Polish Bitcoin Exchange Bitcurex Targeted by Hacking Attack

mootothemax
1pts0
www.theguardian.com 12y ago

Man buys $27 of bitcoin, finds they're now worth $886k

mootothemax
50pts47
stripe.com 12y ago

Stripe in Ireland

mootothemax
92pts31
biasedphp.com 13y ago

PHP Commandments

mootothemax
57pts110
www.dadhacker.com 13y ago

Creating the Atari ST

mootothemax
6pts1
en.wikipedia.org 13y ago

The Scunthorpe problem

mootothemax
5pts0
en.wikipedia.org 13y ago

The Oh-My-God particle

mootothemax
9pts3
tbbuck.com 13y ago

Server backups: everything else you should take care of

mootothemax
1pts0
www.dadhacker.com 13y ago

How the Atari ST almost had Real Unix

mootothemax
2pts0
tbbuck.com 13y ago

Talking with a porn chat spammer, a lesson appears.

mootothemax
11pts3
news.ycombinator.com 13y ago

Ask HN: Who wants a couple of free domains?

mootothemax
9pts5
2012.jsconf.eu 13y ago

JavaScript is the new Punk Rock - Making Music with JS

mootothemax
3pts0
tbbuck.com 13y ago

Extraordinary support? You’d better blow my mind.

mootothemax
41pts25
press.freelancer.com.s3.amazonaws.com 13y ago

Freelancer.com acquires vWorker (formerly RentACoder.com)

mootothemax
2pts0

The anubis author has stated they recognize it's an arms race, but PoW scales.

The scraper wars are largely between script kiddies and people with both deep intimate networking and DOM knowledge. Yes greyhairs, I’m looking at you.

The problem is, you can’t PoW every page load and resource request because the user experience will suck and people will run away. And that window - the gap between what people will tolerate vs draconian enforcement - is exactly what the scrapers exploit.

And looking at the PoW options out there - I’ve seen at least one PoW WAF (honestly can’t remember if azure or amazon) have their PoW boil down to repeated trigonometric functions, ie very optimisable.

It’s a neat concept, but the answer and future to my eyes look bleak.

Can any LLM give you the rough pixel coordinates of an item it identifies in an image?

I found that while Claude, GPT etc could describe an image, there was no way to link the description back to specific pixels in the image itself. Not even to a bounding box or segment.

In your first comment, replace “until today” with “since then” and you’re good!

“Until today” is one of those English phrases that is particularly unfair on non-native speakers. You know “until” and you know “today” and so it’s completely natural to combine them in the way you did.

But as ever, English is dumb and annoying and hard work, all at the same time.

If you haven’t investigated storing in parquet format - and it doesn’t break other consumers that need your jsonl formatted files - it could be worth trialling for your use case. You’ll see vastly smaller file sizes (even more so if you use zstd compression), and querying time will shoot up.

Usual caveats apply, but as a general rule it’s held up well for me. Only downside is that inspecting the results moves from vi on the output file to duckdb and a select * from.

It’s great compression: Y sometimes a vowel, sometimes a consonant.

And while not encoded on a keyboard, it still blows my mind that English has a crazy number of past tenses - and a such a bad hack of a future tense that it’s hard to classify as such.

Linguistics is fun. The accents are alright.

Speaking from the scraper’s perspective, I like proof of work; a ten year old 96-core server will cost a couple of quid to run for a few hours and will grab an absurd number of pages thanks to the access granted by repeatedly solving proofs of work. Small slick codebases too!

Exactly. I’m constantly amazed at how little you actually need to bypass CF, Amazon, Azure WAFs and so on (Incapsula springs to mind too). When you look at the code you’ve come up with, it’s actually quite small and compact.

More to the point, these systems actually help scraping because proof of work unlocks essentially unlimited scraping, in my experience.

That said - from my experience on the other side, sure you can’t stop people like me or you, but you can stop 99% of the others. That’s more than worth it operationally.

I suspect that introducing the calibration concept might be a case of too much too soon for some people.

As far as I understand it, the various probability matrices boil down to: what token has the highest likelihood of coming next, given this set of input tokens. Which then all gets chucked away and rebuilt when the most likely token is appended to the input set.

Objective assessment of internal state - again, to my non-expert eye - doesn’t appear to have any way to surface to me.

Big-if my rough working understand is more or less correct - your calibration point makes a lot of sense to me. I’m not sure that it would make sense to someone who eg considers some form of active thinking process that is intellectualising about whether to output this or that token.

Inventing Cyrillic 3 months ago

The irony for me being that when I was first learning Polish and looking for any and all mnemonics - “ah, that word is the number nine, and that one is ten because it has an s in the middle and that’s next to t for ten in the alphabet”-levels of desperate - the false etymology helped me set word, słowo, in my head, and the rather delightful dosłownie, literally / to the word, has remained ever since.

(tho while on the subject, it’s hard to beat wieloryb as a wonder that I don’t want to know the true etymology of ever because if there’s even a chance that the word for whale derived from the words great as-in-size + fish, I want to hang on to it forever)

If you can find a way to combine this with local population to end up at pence per litre per thousand population, I bet you’d uncover some fun trends. Bet it’d also get interesting if combined with population within an X min drive too.

Tho really need some car population per road segment stats to drive the most out of it IMO.

With Maplibre or any modern map SDK this this is standard…

In practise, this doesn’t work out as visually pleasing as you’d like; labels repeat, or render partially or not at all, or become interfered with by other labels, or only work well at one given zoom. It’s easy to end up in a visually dissatisfying place that’s taking an unfathomable number of magic rules to get to.

The secret sauce to fixing this is creating separate label layers of perfect point locations or lines for labels to follow in advance. Added bonus is faster render and interaction times due to fewer rules.

Do you know game theory?

Never heard of it. The food there good?

The other devs with less moral can then outperform you.

I long for the days where it’s only my moral compass holding me back.

For what it’s worth, I cancelled my ChatGPT subscription, and every time I try debugging a Linux system issue, I feel sad that Claude is sooooooo confidently bad at it.

Claude is noticeably poor for my use case on this particular issue. That said, I imagine I’m not alone in refusing to continue paying OpenAI. We’re in for a wild ride.

Completely by accident, I have a setup that sends a pdf invoice to customers a couple of days after the sale. I’m pretty sure it’s a stripe option I must’ve misclicked.

Anyway- turns out that on the rare occasion someone’s had an issue, this gives them a really easy mechanism to write to me and tell me about it. They let off their steam in the email and then we make things good together. (Yet another reason why I always oppose noreply email addresses)

I still don’t know what or where the setting is, mind.

GPT-5.4 5 months ago

Huh, that’s interesting, I’ve been having very similar thoughts lately about what the near-ish term of this tech looks like.

My biggest worry is that the private jet class of people end up with absurdly powerful AI at their fingertips, while the rest of us are left with our BigMac McAIs.

Microgpt 5 months ago

For context, BERT is encoder-only, vs SLMs and LLMs which are decoder-only, and BERT is very much not about generating text, it’s a completely different tech and purpose behind it. I believe some multimodal variants nowadays may muddy the waters slightly, but fundamentally they’re very different things, let alone around been around for decades unless also including the history of computing in general.

While I could’ve written that better and with less attitude, gotta confess - and thx for pointing out my smugness - the AI stuff of the last few weeks really got under my skin, think I’m feeling all rather fatigued about it

Microgpt 5 months ago

> BERT isn’t a SLM Huh? BERT is literally a language model that's small and uses attention.

Astute readers will note what’s been missed here.

Fascinating, really. Your confidently-statement yet factually void comments I’d have previously put down to one of the classic programmer mindsets. Nowadays though - where do I see that kind of thing most often? Curious.

Microgpt 5 months ago

We had good small language models for decades. (E.g. BERT)

BERT isn’t a SLM, and the original was released in 2018.

The whole new era kicked off with Attention Is All You Need; we haven’t reached even a single decade of work on it.

Best Gas Masks 6 months ago

Oh that is beautiful :)

For myself, it’s the feeling of: thank fuck; the grownups have arrived. shoulders lower, everyone takes a deep breath

It’s a delight even to have a regulated source of all fuel station locations in the uk!

This might be a slight missing woods/trees moment but that aside - there is precious little open geospatial data in the uk that establishes see this dot here? That’s a fuel station, that is. That dot there? Oooooh no, that there’s a pub.

The uk govts of the time managed to hand both the address data and the this-is-what-it-is data off to separate commercial enterprises in the name of privatisation, and I genuinely believe it was by accident as it’s… err… quite a niche topic of knowledge.

So anything - anything! - that brings some of that back and truly open to the public is very much welcomed.