HN user

nikcub

19,620 karma

Nik Cubrilovic - https://nikcub.me

squirrelscan - https://squirrelscan.com

open electricity - https://openelectricity.org.au

email nik at nikcub.me

@dir on twitter / x

Posts128
Comments2,538
View on HN
wordpress.org 5d ago

Two critical SQLi vulnerabilities in WordPress

nikcub
1pts0
blog.includesecurity.com 1mo ago

The Smart TV in Your LivingRoom Is a Node in the AIScraping Economy

nikcub
235pts105
www.bloomberg.com 3mo ago

Anthropic, OpenAI and Google sharing Intel to block Chinese distillation

nikcub
4pts0
www.bloomberg.com 5mo ago

New OpenAI funding to top $100B

nikcub
3pts1
www.cursor.com 1y ago

Cursor Release v1.0

nikcub
3pts0
www.nytimes.com 1y ago

They used Xenon to climb Everest in days – is it the future of mountaineering?

nikcub
97pts188
www.bloomberg.com 1y ago

Microsoft pulls back on datacenter ambitions

nikcub
8pts1
arxiv.org 1y ago

AI and the value of privacy-preserving tools to distinguish who is real online

nikcub
5pts1
www.documentcloud.org 8y ago

Ars and journalist sued by password manager vendor for reporting vulnerabilities

nikcub
1pts0
www.wired.co.uk 8y ago

A stolen painting appears for sale on the dark web

nikcub
114pts73
finance.yahoo.com 8y ago

“We've been breached” – Inside the Equifax hack

nikcub
3pts1
www.theguardian.com 8y ago

Mystery of sonic weapon attacks in Cuba deepens

nikcub
523pts277
www.washingtonpost.com 9y ago

Obama’s secret struggle to punish Russia for Putin’s election assault

nikcub
4pts1
www.nytimes.com 9y ago

Opioid Dealers Embrace the Dark Web to Send Deadly Drugs by Mail

nikcub
2pts0
medium.com 9y ago

Tactics, Techniques, and Procedures of the Yahoo Hack

nikcub
124pts10
www.nytimes.com 9y ago

India’s Call-Center Talents Put to a Criminal Use: Swindling Americans

nikcub
2pts0
brontecapital.blogspot.com 9y ago

Measuring how bad Twitter is

nikcub
129pts78
www.nytimes.com 9y ago

Ted Cruz to Fight Transfer of IANA to ICANN

nikcub
1pts0
www.patentbell.com 10y ago

Ghostery patent granted for displaying and blocking ad trackers in a browser

nikcub
2pts0
www.wsj.com 10y ago

Tech and Poverty Collide in San Francisco's Tenderloin

nikcub
2pts1
recode.net 10y ago

Billion-dollar exit of autonomous-tech darling Cruise brings legal headaches

nikcub
3pts0
www.nytimes.com 10y ago

Tech Companies Face Greater Scrutiny for Paying Workers with Stock

nikcub
58pts64
uncrunched.com 11y ago

The Venture Capitalists Behind Superfish

nikcub
5pts0
www.nikcub.com 11y ago

Large Number of Tor Sites Seized by the FBI were Clone or Scam Sites

nikcub
99pts27
www.documentcloud.org 11y ago

FTC civil complaint against Bitcoin mining manufacturer Butterfly Labs

nikcub
1pts0
www.nikcub.com 11y ago

Analyzing the FBI’s Explanation of How They Located Silk Road

nikcub
319pts121
www.wired.com 11y ago

The FBI Says How It ‘Legally’ Pinpointed Silk Road’s Server

nikcub
226pts157
www.nikcub.com 11y ago

Notes on the Celebrity Data Theft

nikcub
366pts274
www.texasmonthly.com 12y ago

The Murders at the Lake

nikcub
6pts0
thread.gmane.org 12y ago

Akamai release source to their custom secure_malloc for OpenSSL

nikcub
267pts80

I pretty sure OpenAI and Anthropic are doing the same or worse.

No they're not. It would end both companies if they were ever found to be doing that.

Their terms are clear - if you use the coding plans they can[0] train in return. Enterprise and API, absolutely not.

The argument here is that with the Chinese labs you have zero legal recourse.

[0] opt-in, thanks

This is a winner IMO. Lots of cost pressure on token spend atm within enterprises and tasks that don't require Opus / Codex class models.

These companies have hopefully captured all of their traces and now have enough to fine-tune an open model and host themselves.

Inkling feels like the right base - not obsessed with benchmaxxing on coding but rather being adaptable to the task required

For tasks like GTM, support, content writing etc. seeing 80%+ savings

The embedded networks collude with the builders and offer them the installation

I got into a dispute with my embedded provider because of a bad meter and came to discover through friends and family in the construction industry as well as speaking to a former sales person in the industry that there is a lot of additional corruption in the process with straight up payments being made to win installs with developers.

When it came time to switch providers in our building, strata was promised electric vehicle chargers as part of signing a new deal with a new provider. They never delivered because they found an escape clause because of fire safety approval.

We're now locked in for years (again) and they've already increased rates once in the first year.

Nobody in the entire chain works in the interests of residents or owners. It's a completely broken system and a thorn in the side of otherwise advanced and progressive Australian energy policy. It needs to be abolished ASAP.

I still pay more for my single apartment living alone in electricity than what family and friends do in full large homes with air conditioning, 4-6 residents, heated pools, etc. It's astonishing.

I'm glad there are more attempts at solving model routing, as costs (at API rates) has really become an issue. Some feedback:

1. Reiterate the cache issue from other comments already here. there is a lot of optimisation in harnesses around caching and a proxy model blows that up

2. Coding agents are model aware - they already route code discovery to mini / flash models, planning to heavy models, workflow design to ultra, implementation to mid / high etc. They know when they're exploring, planning, implementing, reviewing etc. and which model class to select and when it fails.

With a proxy you're breaking this control loop and feedback. It doesn't know, for ex. that it just attempted with deepseek v4 and it failed, lets try Opus?

3. How are you going to RL improvements and prevent the router becoming stale? You only have access to your own internal prompts and ~thousands of samples.

This is RL'd on one orgs codebase. There are going to be a lot of prompts you haven't seen before and have no insight to on how to route correctly, and you have no insight into users HF to improve your own model. Orgs aren't going to share their traces with you, so you need other sources to train on and improve

There are also new model releases every week that you need to keep up with - whats the story going to be here

4. Publish evals by running terminalbench / deepswe bench. Show us the performance / cost / time chart vs the other agent and model sets. If you can show gains there, you have a very simple value prop to sell where you can charge for a % of the saved costs

Om Malik has died 27 days ago

This is devastating. Om was the godfather of early tech blogging and lifted up so many people around him. He was kind, caring and compassionate.

When I first started blogging around 25 years ago, he would have been amongst the first 10 readers. He linked to me, emailed me privately with feedback, praised posts and would call bullshit when he saw it.

He was never competitive with other blogs or bloggers and was never tied up in drama. He was very often a mediator in behind the scenes conflicts and was obsessed with truth over getting the scoop.

He loved tech and startups and most of all loved seeing other succeed and didn't have a gram of resentment within himself.

Everybody from that post-dotcom crash era of tech owes Om a large debt of gratitude. He will be missed. RIP Om.

"a well written agents.md is very good for the agent"

while even a mildly bad agents.md can be _very_ bad for the agent. they rot very quickly which is why human curation is essential.

same with memory - a lot of the self-learning tools that are becoming popular now degrade agents over time - which is why you end up being able to run an eval with no context and it performs better

but that's why agentic coding can still be considered a "skill".

yes - far too many cases of throwing a kitchen sink of prompts, skills, tools etc. thinking the llm will sort it out. you need to constantly prune, eval, tweak, observe, update etc. in a loop

there are 150M+ of them and you'll be taking out a lot of human users with it

modern blocking is behaviour / heuristic based

Cloudflare are more likely to be undercounting bots - they don't really pick up many of the modern browser-driven bots and crawlers.

these old network security techniques don't really work anymore. the common bots are at known IP ranges, the problem bots are all on datacenter + residential proxies.

I believe the urgent deprecation timeline here may be related to ai labs using offline licensed Office in agents as part of workflows and Office integration. Microsoft wants _each_ agent instance to be a separate license[0]

There was always a probability that Microsoft were going to funnel offline users into O365 at some point - but I imagined that to take place over months / years not weeks and days.

Buying a single license for thousands of agents may have expedited that. It has resulted in non-Microsoft labs having better ai integration into their products than Microsoft.

edit: just read the detail of the note - so this is a cert expiry as part of Apple dist that is being warned about ~2 months before it happens. Standalone on Mac has a term limit.

[0] https://www.businessinsider.com/microsoft-executive-suggests...

Claude Opus 4.8 2 months ago

The syntax is the easier part - most programming tasks require the reasoning and understanding of a large world model to solve problems.

Fine tuning a 'lean and smart' model works really well for discrete, repeatable high volume tasks like support ticket triage, lead classification, content filtering, labelling, generating content with a voice, etc.

Inefficient token burn by throwing large models at everything is definitely a problem - it's like hiring Phd's to answer the phone or to wash dishes.

the CVE + patch data has been built into models for a few generations now. I actually thought the bug bounty companies were well positioned here, but they've been overtaken.

Mythos is a better hacker than we ever were

Finding the neeedle is easier when you remove the haystack

Or providing a map with a direction

There is a long history of high-value private vulns being rediscovered from scant details

There has been a lot of cynicism around mythos, that it's just the usual public models without guardrails, etc. etc. but this:

1,752 of those high- or critical-rated vulnerabilities have now been carefully assessed by one of six independent security research firms, or in a small number of cases by ourselves. Of these, 90.6% (1,587) have proved to be valid true positives, and 62.4% (1,094) were confirmed as either high- or critical-severity.

for anybody who has applied opus, codex or oss models for vuln scanning - the true positive rate and discovery volume are a clear step change[0]. The ~50 partners in Glasswing have largely all previously run harnesses with other models and many of them have come out and said - essentially - "ye, wow"

Question now is what a second and third phases of access looks like - deciding which class of systems to secure. Routers, firewalls, SaaS, ERP systems, factory controllers, SCADA systems, zero-trust VPN gateways, telecoms gear and networks, medical devices - there's just so much to do

This is why I believe mythos will remain private for the foreseeable future. There's such a large surface that needs to be secured and so much to triage, fix, deploy.

That may suit Anthropic as private models can't be distilled. There's also a runaway effect of model improvement from the discovery, triage and fix data. This is likely already the most potent corpus of curated offensive data ever assembled and will only get better.

I don't see how Chinese companies are given access soon, or ever. We're likely going to see a world soon of CISA mandated audits, and where to buy a mythos-proof VPN gateway or home router - you'll have to buy American[1].

[0] vs ~30% or so in regular audit tools

[1] or allied

Anthropic this quarter will have revenue of $10.9B, up from $4.8B last quarter[0]. They're paying SpaceX $1.25B per month for compute[1] - which is more than what SpaceX earn on space. SpaceX spent about $30-40B in capex on Colossus 1 & 2.

This is all real revenue, real spend, real usage.

Hetzner just aren't at this scale. Not even close. If they wanted to get into this business - first, they're late. Second, it's at a scale of ~10x of their total lifetime datacenter buildout. Third, they'd need to change their business to being one that is debt fronted.

xAI have proven out that being able to deploy compute is a very viable business (and difficult to pull off)

At some point AI cynicism clashes with reality, it must be exhausting maintaining it.

[0] https://www.wsj.com/tech/ai/mind-blowing-growth-is-about-to-...

[1] https://www.wired.com/story/spacex-ipo-anthropic-compute-fin...

This just incentivizes market for bio-mules, which already exists with world[0] - where prices stay low because it was rolled out to low-income countries.

Then there's the platform game theory. If you adopt you add friction which reduces signups, and there will always be a competitor who would risk the 10x fraud increase in order to capture 100x the market. Railway has seen hyper-growth because it's so easy to run from, and is recommended by, coding agents[1].

The solutions are here already just not well implemented or understood - probabilistic fraud detection, resource limits, service and automation limits, standard gov identity verification as a signal, enterprise sales channels with human relationships, etc.

There are tradeoffs with each platform choice that just aren't well understood. Most users shop on price and DX and don't see the abuse infra or problem until it hits them.

Google and GCP have a problem where they completely cook users who get flagged in their automated fraud net (this isn't news - or shouldn't be)

[0] https://www.coindesk.com/policy/2023/05/24/black-market-for-...

[1] and the problems that come with providing that simple interface, like sometimes dropping prod

This is the conflict at the center of running a hosting company - make it easy to signup and you get a lot of new users but also a lot of abuse.

Implement anti-abuse measures and you will hit some loud false positives (this may be the case with GCP here).

I don't envy anybody running a hosting co - the internet is a really ugly place under the surface.

edit: to add - AWS are really good here. Must be the ~30 years of retail fraud and abuse experience.