Slovakia mentioned, let's gooo. Ehm, exactly, we can achieve better smaller models for specialized tasks rather than using compute to improve a big model that does everything. There's a lingering philosophical question if better language processing capabilities translate to better image processing capabilities (i.e. having the vocabulary and experience to properly describe an image), but I still think that identifying tasks and splitting responsibilities saves a lot of effort.
HN user
ArcHound
Hello, I write my blog at blog.miloslavhomer.cz
I did a write-up at https://blog.miloslavhomer.cz/vibing-with-french-models-in-n....
It really can be a fancy auto complete, but more agentic usage moved out of the editors (and I think that's a good thing).
Brilliant. What I liked are the characters - it's hard to make every character motivation reasonable and so well communicated.
What I think is a bit of a missed opportunity is for the product to fail with "the pizza|cake|pastry is half-baked" and so customers still have to do the rest of the job anyway.
I remember when I thought security is a technical problem. I was shocked to see people neglecting "basic" principles. Don't get me wrong, doing secutis important. But we need to see the reason and work within our constraints.
So I've written up my thoughts on the topic. Since we're distributing limited resources to achieve the priorities decided, security is in fact a political problem.
We know that technical and political problems are different. I usually want to solve the technical problems while avoiding the political ones.
Sadly, security is a political problem, because in risk management you can always take the risk and so every technical solution might be abandoned the moment this decision is made.
Realizing and working with this fact helps me surviving another day in this line of work. Maybe it'll help you too.
I saw a game, where you played as a poor Soviet soldier that accidentally sent nukes to USA. To save the world, you had to navigate a phone call labyrinth to alert USA defense systems for missile interception. I haven't laughed that much in a hot minute.
ok, so I've parsed some logs. I do see the ALPNs pointing to http2, but I don't capture all of the headers. The only thing I capture is the user-agent, which is the major spoof anyway.
Now, to differentiate between spoofed and non-spoofed header, I need to check the "valid" JA4 signature for a given browser and then proclaim that the rest of them are wrong. The "valid" JA4 signature can be observed, but I've found that sometimes browsers tweak their handshake a bit, so it's not 100% consistent.
The JA4 DB was recently taken down, I've requested full access, but no response (as expected). There might be some issues in getting those valid headers for the browsers, the hardware and software varies a lot (PC, Mac, Android, Iphone of all kinds of versions and browsers).
I was hoping for a quick win to share, but it doesn't seem like so and I'll have to do it properly. That should be my next post on JA4.
As a quick note, approx 30% of traffic claims to use http2 and approx 60% of that traffic has a non-bot user-agent (you know, along the lines of "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/149.0.7827.102 Safari/537.36"). I suspect majority of those are spoofed as I know how many readers I have on my blog.
Back in the day I couldn't find a downloadable DB for offline checks, which is very much needed when looking at approx 10k different IPs. Even with an offline DB I might need to create this tree structure so that I can process the data fast.
I think we agree that JA4 is situational. It really saved me when investigating a credential stuffing attack - random logins with random chance of success spread into many ASNs, all had the same fingerprint.
From my experience, there are all kinds of levels of bots. Add them all together and they can produce a ridiculous load on a site (especially a fragile one that you have to secure anyway). So I look at the volume, trying to block anything stupid I can get away with.
It is a game of whack-a-mole. It also can cut down the overall traffic to a fraction of the original, which has tangible infra costs benefits.
And yes, captcha works better in a lot of cases. Fortunately I'm not selling JA4, I'm just curious.
And yes, IP rate limits and ASN checks work really well in plenty cases. Side note: I got a high-throughput free offline asn-checker too! https://blog.miloslavhomer.cz/asn-check/
This is the sad conclusion of the next part. JA4 is a great supplement, it can squeeze some additional info, but for a motivated attacker it can be avoided.
Now the question of how motivated are noisy AI scrapers is still open. Even a solution that cuts down 50 percent of the dumbest scraping attempts will still provide much needed relief to a struggling site.
I'll get back to you on this, I'll need to parse some logs. I should have at least ALPNs
Hello again! Yes it is. If you have an exotic client, I'm here for it :D
"Who solved the alignment problem for these superhumans?"
The gun, pointed to their head.
The article had a great opportunity to at least reference the newest encyclical from the Pope.
With a bit of click bait, this might be the first step of "Butlerian Jihad" -> "AI Crusade" and a founding text for the Orange CATHOLIC Bible.
I saw a take that you can now cite religious reasons to refuse working on and with AI - if more people try this, I wonder how it'd play out.
Interesting times.
Have you tried putting this behind a reverse proxy? This gives us a lot of features like rate-limiting and it should work well since it is https after all.
Thank you for the kind words.
DoH is a critical enabler of ECH, and getting it right isn't easy - especially dodging all of the free services provided by the giants.
In this article I take a look at the technical properties of Encrypted Client hello as well as some scenarios that are not really covered by the threat model proposed.
I argue that to get any tangible benefit you have to use the big providers, which places trust into entities that are behaving less trustworthy by the hour.
Simply put, you'll need algebra, linear algebra, number theory. So a lot of math with various degrees of depth.
Oh this brings me back to my uni days. I suppose that since this is the basis of post-quantum crypto it is a good time to learn this.
Seems to me that these lattices and error-correcting codes are very close to each other, but for some reason they are discussed separately.
I'd wager that there will be some reductions between those problems - maybe I could dig more around that.
Makes sense if you think about it: if all photons pass through you (invisible) then you can't capture them to get info (blind).
Pretty well if you consider the "bio" label, which is a set of practices not using all of the tech. They can ask for and usually get higher prices for the products.
Granted, it's more about chemicals than tractors, but still quite close to the spirit of the comments. Bio approach sacrifices some tech advances.
Ok, gotcha. So there's a demand for the additional features that are not bundled within git to be federated somehow.
I'd say we have emails, mailing lists and bug trackers. Or maybe: what is the missing killer feature that needs federation?
But, there are? I can host a repo on GitHub, Codeberg and self host it too. Then I need to watch over main to keep it consistent between those. After that's established, I can do updates from wherever. Link'em in the README.
Hello. For what it's worth, I'm in a similar role (Security Architecture) in a different sector. It's the same here.
The trick is to block your calendar with 30min slots to approx. 80% of capacity. If your calendar is private, these won't be challenged. This leaves you some space to do the work, some other space to get random meetings in and if you're lucky, everyone is happy.
I think this is the price to pay for a high-profile role.
You mean to tell me that companies that got rich by hoarding data are excited to hoard more data? Never would have guessed.
Also, why wouldn't anyone want to have data about everyone? Seems like a valuable asset.
Please tell me more. I'm looking at the VIVA and I really don't get why would anybody contribute to the "internal linkedin" and other features. Where did it come from? Where does it go?
I think you're right on the luxury brands being less durable.
To address the second airplane example, we really have to go through all that you're buying. Namely: more leg space, faster airport queue processing, more luggage, better in-flight service. Do I value these at 3x the cost? Maybe yes.
The core point is of course solid. By not updating on day 0, maybe somebody else spend the effort to discover this and you didn't. But there are plenty of other benefits for not rolling with the newest and greatest versions enabled.
I'd argue for intentional dependency updates. It just so happens that it's identified in one sprint and planned for the next one, giving the team a delay.
First of all, sometimes you can reject the dependency update. Maybe there is no benefit in updating. Maybe there are no important security fixes brought by an update. Maybe it breaks the app in one way or another (and yes, even minor versions do that).
After you know why you want to update the dependency, you can start testing. In an ideal world, somebody would look at the diff before applying this to production. I know how this works in the real world, don't worry. But you have the option of catching this. If you automatically update to newest you don't have this option.
And again, all these rituals give you time - maybe someone will identify attacks faster. If you perform these rituals, maybe that someone will be you. Of course, it is better for the business to skip this effort because it saves time and money.
I see your point, I do. It seems like all external software is going in the SaaS direction, where the vendor is keeping all of the data, so they are available over an API. So there are genuinely solid cases for Chromebooks.
The issue is how much power this gives to the vendors. I think we should be able to survive a vendor going poof, taking all our data with them. Having a general computing platform capable of mixing files and privileges seems to me like the only way of keeping this capability.
I guess I should set up a monitor alerting me if the two backup diffs are larger than 80% of the data size.
But yes, these are the practical problems we need to address.