HN user

eis

4,400 karma
Posts30
Comments537
View on HN
www.reuters.com 15d ago

Beijing is looking at curbing overseas access to China's top AI models

eis
64pts11
www.reuters.com 5mo ago

US Supreme Court rejects Trump's global tariffs

eis
24pts1
connectrpc.com 2y ago

Connect RPC – A Better gRPC

eis
3pts0
www.reuters.com 3y ago

Trump hit with criminal charges in New York, a first for a US ex-president

eis
9pts1
www.bloomberg.com 3y ago

Apple Abruptly Shutters Store in North Carolina After Shootings

eis
3pts1
www.youtube.com 3y ago

Mobile CPU/GPU Power Efficiency Comparison [video]

eis
2pts1
www.reuters.com 3y ago

Twitter staff exodus accelerates amid Musk battle, whistleblower complaint

eis
7pts0
cryptobriefing.com 3y ago

Hodlnaut fires 80%, lost $190M due UST depeg despite claiming no exposure

eis
4pts1
danluu.com 3y ago

Files Are Hard

eis
5pts1
www.reuters.com 4y ago

Spanish prime minister's telephone infected by Pegasus spyware

eis
4pts0
www.reuters.com 4y ago

Elon Musk to buy Twitter for $44B

eis
9pts0
www.reuters.com 4y ago

Russian forces invade Ukraine after Putin orders attack

eis
2278pts1814
www.reuters.com 4y ago

Google faces a fine of up to 20% of Russian revenue this month

eis
2pts0
research.checkpoint.com 5y ago

Security probe of Qualcomm MSM data services

eis
1pts0
www.gizmochina.com 5y ago

Antutu bans the Realme GT after it finds evidence of benchmark cheating

eis
1pts0
www.anandtech.com 5y ago

Asus PN50 Mini PC with Ryzen 4000 CPU

eis
1pts0
www.reuters.com 6y ago

U.S. coronavirus supply spree sparks outrage among allies

eis
5pts1
www.postgresql.org 8y ago

Postgres dataloss due to Linux kernel forgetting writes on failed fsync

eis
7pts0
fortune.com 8y ago

China Enlists Its 'Great Firewall' to Block Bitcoin Websites

eis
1pts1
torrentfreak.com 8y ago

Cloudflare Bans Sites for Using Cryptocurrency Miners

eis
3pts1
twitter.com 9y ago

Golang: sub-millisecond GC pause on production 18gb heap

eis
119pts30
www.nytimes.com 10y ago

Supreme Court Says Police May Use Evidence Found After Illegal Stops

eis
4pts0
twitter.com 10y ago

Popular stickers among Palestinian techies and journalists in Gaza

eis
2pts0
torrentfreak.com 10y ago

Police Arrest Men for Spreading Popcorn Time Information

eis
27pts4
torrentfreak.com 11y ago

Under U.S. Pressure, PayPal Nukes Mega for Encrypting Files

eis
5pts0
www.reuters.com 12y ago

EU court rejects requirement to keep data of telecom users

eis
337pts77
www.reuters.com 12y ago

U.S. government shutdown begins after Congress fails to break impasse

eis
1pts0
klogk.com 12y ago

Use ASCII art to highlight code parts in Sublime Text

eis
6pts3
www.defenceindepth.net 14y ago

Cracking OSX Lion passwords

eis
281pts79
bitcoincharts.com 15y ago

500k Bitcoins traded in 1h, Mt.Gox market hacked + crash

eis
276pts245
Grok 4.5 14 days ago

Google wanted to release 3.5 Pro last month but because of the trouble Anthropic got with Fable they might have wanted to wait a bit for the dust to settle I could imagine. And now there is quite some competition. 3.5 Flash for me is a replacement to 3.1 Pro. It's more like a 3.2 Pro. It costs about the same (or more!) than 3.1 Pro, is a little bit smarter in many cases and a little bit faster. 3.5 Pro will be a lot more expensive and I expect it to juuuust be able to hang with Opus 4.8 and GPT-5.5.

I wish Google was able to actually push the industry further, either in terms of quality (intelligence) or quantity (price) but they've been playing catch up a lot.

They are playing the game a bit differently than all the others. The others have useable IDEs etc. while Google has a boatload of half-assed products.

Google better come out with a banger 3.5 Pro because who would have thought that Grok and GLM would be beating them?

I feel like it would be much better if the article focused on QuePaxa because IMHO it's an algorithm that finally brings some novel ideas to concensus (e.g. not relying on timeouts) by kinda coming at it from a gossip protocol angle and is not getting the attention it deserves. The post shouldn't have focused and introduced Meerkat which hasn't been fully developed and tried in production. If they clearly presented the pros and cons vs not just Raft (which is popular but doesn't even play in the same league because it is relies on a leader) but other leaderless or multi-leader concensus protocols that would have been of greater value. The Paxos family of algorithms are a much closer fit here and there's a reason why some serious large planet scale systems choose it over Raft.

E.g. 1. Intro about issues with concensus 2. Intro to QuePaxa 3. Comparison to other algos that are close to it 4. Mentioning active work on implementation via Meerkat and intent to bring to production with followup posts.

As always when it comes to concensus it's all about trade-offs. And with QuePaxa that might be the increase in messages (note: I don't mean message round-trips). We'll see how it goes but it will definitely be interesting.

But it's not a large-scale public deployment yet either. The article says towards the end that they just ran a proof of concept.

Maybe the blog post is just premature. It would be much more valuable if they posted it after actually having run it in production and validated the strengths and weaknesses with real world data.

I already gave up on Fable 5 because it sometimes was just not worth the editional price compared to Opus 4.8 and other times it flat out downgraded to Opus anyways for no good reason because it thought I'm looking for security vulnerability while working on the auth part of my app. In our company Fable 5 is not enabled because of the change in data retention being required.

And now this. How would they even enforce this restriction when they can't know what nationality the end user behind some API query belonging to a company account has? It seems like nobody is thinking things through anymore and the end result is total unreliability from every angle. What a huge mess all of it, sigh.

Here's a crucial mechanism that Paul Graham did not mention:

With a wealth tax using his calculation, the higher your returns, the lower the comparable income tax would be. If your returns are 10% you'll pay $1 on $10 capital gains which is 10% and you end up with $109. Conversely someone achieving a mere 1% cap gains would be essentially taxed for 100% of his return.

With income taxes it's usually the opposite: the more you earn, the higher the tax bracket you will be put into.

Somebody like Paul Graham surely has higher than 10% capital gains, otherwise he'd not be exactly a great investor.

Personally I'm against wealth taxes, I think capital gains taxes are a much more appropriate and fairer tool. I also think taxes in general are way too high, if you are part of the middle class and add up everything you pay in taxes, fees, insurance, duties and whatnot you can end up losing 70-90% of whatever you earn. It's extremely hard to actually accumulate wealth for the vast majority of people.

Gemini 3.5 Flash 2 months ago

3.5 Flash was more expensive than 3.1 Pro to run the Artifical Analysis test suite. $1551 for 3.5 Flash [0] vs $892 for 3.1 Pro [1]. That's 74% more cost while ranking lower. It's 2.5x as fast but I don't think the bang for the buck is there anymore like it was with 3.0 Flash. I'm a bit bummed out to be honest.

I did not expect such a huge (3x) price increase from 3.0 Flash and I bet many people will not just blindly upgrade as the value proposition is widely different.

One interesting point to note is that Google marked the model as Stable in contrast to nearly everything else being perpetually set as Preview.

[0] https://artificialanalysis.ai/models/gemini-3-5-flash [1] https://artificialanalysis.ai/models/gemini-3-1-pro-preview

Advantage for what exactly though? I'm not saying Elo Ranking doesn't give any information. It just doesn't give the information that the OP's project claims to be able to give: that models get nerfed over time. You could extract this kind of information from the raw results of each evaluation round between two models, ignoring any new model entries and compare these over time but not from the resulting Elo scores with an ever changing list of models.

New models are on average better than older models, the average skill of the population of models increases over time and so you are mathematically guaranteed that any existing model will over time degrade in Elo score even though it didn't change itself in any way.

It's like benchmarking a model against a list of challenges that over time are made more and more difficult and then claiming the model got nerfed because its score declined.

Elo is good at establishing an overall ranking order across models but that's not what this is about.

To detect nerfing of a model, projects like https://marginlab.ai/trackers/claude-code/ are much much better (I'm not affiliated in any way).

The Elo rating system measures relative performance to the other models. As the other models improve or rather newer better models enter the list, the Elo score of a given existing model will tend to decrease even though there might be no changes whatsoever to the model or its system prompt.

You can't use Elo scores to measure decay of a models performance in absolute terms. For that you need a fixed harness running over a fixed set of tests.

Interestingly NET is down 15%-ish in extended hours trading and was even down 20% at some point. Many times a stock will make a positive move when layoffs are announced.

Cloudflare is a growing company by most metrics so if efficiencies through AI were the reason for the layoffs they'd just take the boost and grow even faster.

It all doesn't check out and I think the real reason for the layoffs and the negative sentiment by the market on the news is that their revenue growth was not as fast as their expenses and they realized they overhired. Leadership doesn't want to dive too much into the red even if it would mean bigger growth down the line. They are now beholden to the near and mid term stock performance.

I've had the chance to talk to some SWEs working at Cloudflare off the record in recent months and the one concensus I heard was that there was many times some tension between the boots on the ground and the decisions from senior managment but of course nothing they could do and especially after this they'll make sure to be quiet should they remain. There seemed to be a lot of pressure to deliver features and new products but quality has been left behind which means the SWEs felt pressure to deliver while also having to deal with the ensuing issues to resolve.

Either way I wish everyone affected the best and a speedy job hunt - there'll be quite a few really good people on the market now for no fault of their own.

I don't think people are crediting Apple with inventing unified memory - I certainly did not. There have been similar systems for decades. What Apple did is popularize this with widely available hardware with GPUs that don't totally suck for inference in combination with RAM that has decent speed at an affordable price. You either had iGPUs which were slow (plus not exactly the fastest DDR memory) but at least sitting on the same die or you had fast dGPUs which had their own limited amount of VRAM. So the choice was between direct memory access but not powerfull or powerfull but strangled by having to go through the PCIE subsystem to access RAM.

The article is talking about one particular optimization that one can implement with Apple Silicon and I at least wasn't aware that it is now possible to do so from WebAssembly - so to completely dismiss it as if it had nothing to do with Apple Silicon is imho not fair.

Yes but that is just a tiny part of the whole CF worker ecosystem. The other services are not open source and so the lock-in is very very real. There are no API compatible alternatives that cover a good chunk of the services. If you build your application around workers and make use of the integrated services and APIs there is no way for you to switch to another provider because well, there is none.

How did you work around this problem? As in, how do you monitor for hung queries and cancel them?

You just wrap your DB queries in your own timeout logic. You can then continue your business logic but you can't truly cancel the query because well, the communication layer for it is stuck and you can't kill it via a new connection. Your only choice is to abandon that query. Sometimes we could retry and it would immediately succeed suggesting that the original query probably had something like packetloss that wasn't handled properly by CF. Easy when it's a read but when you have writes then it gets complicated fast and you have to ensure your writes are idempotent. And since they don't support transactions it's even more complex.

Aphyr would have a field day with D1 I'd imagine.

What about reads? We use D1 in prod & our traffic pattern may not be similar to yours (our workload is async queue-driven & so retries last in order of weeks), nor have we really observed D1 erroring out for extended periods or frequently.

We have reads and writes which most of the time are latency sensitive (direct user feedback). A user interaction can usually involve 3-5 queries and they might need to run in sequence. When queries take 500ms+ the system starts to feel sluggish. When they take 2-3s it's very frustrating. The high latencies happened for both reads and writes, you can do a simple "SELECT 123" and it would hang. You could even reproduce that from the Cloudflare dashboard when it's in this degradated state.

From the comments of others who had similar issues I think it heavily depends on the CF locations or D1 hosts. Most people probably are lucky and don't get one of the faulty D1 servers. But there are a few dozen people who were not so lucky, you can find them complaining on Github, on the CF forum etc. but simply not heard. And you can find these complaints going back years.

This long timeframe without fixes to their network stack (networking is CF's bread and butter!), the refusal to implement transactions, the silence in their forum to cries for help, the absurdly low 10GB limit for databases... it just all adds up. We made the decision to not implement any new product on D1 and just continue using proper databases. It's a shame because workers + a close-by read replica could be absolutely great for latency. Paradoxically it was the opposite outcome.

D1 reliability has been bad in our experience. We've had queries hanging on their internal network layer for several seconds, sometimes double digits over extended periods (on the order of weeks). Recently I've seen a few times plain network exceptions - again, these are internal between their worker and the D1 hosts. And many of the hung queries wouldn't even show up under traces in their observability dashboard so unless you have your own timeout detection you wouldn't even know things are not working. It was hard to get someone on their side to take a look and actually acknowledge and understand the problem.

But even without network issues that have plagued it I would hesitate to build anything for production on it because it can't even do transactions and the product manager for D1 openly stated they wont implement them [0]. Your only way to ensure data consistency is to use a Durable Object which comes with its own costs and tradeoffs.

https://github.com/cloudflare/workers-sdk/issues/2733#issuec...

The basic idea of D1 is great. I just don't trust the implementation.

For a hobby project it's a neat product for sure.

Quite strong results in the benchmarks but why Gemini 3 Pro instead of 3.1? Why only for a few of the benchmarks? Why is OpenAI not there in the coding benchmarks? Why Opus 4.5 and not 4.6? Just jumps out into my eye as a bit strange.

As always, we'll have to try and see how it performs in the real world but the open weight models of Qwen were pretty decent for some tasks so still excited to see what this brings.

These theorems apply to any system of axioms that are rich enough to state the liar's paradox.

Isn't that circular reasoning or tautological though? Rephrased: any system that can state something that these theorems apply to, can have the theorems applied to.

I think the word "rich" is too inaccurate in this context. It is not clear why there can't be a more "rich" system which does not suffer from this issue and can't state the liars paradox.

People are understandably a bit sensitized and sceptical after the last AI generated blog post (and code slop!) by Cloudflare blew up. Personally I'm fine with using AI to help write stuff as long as everything is proof-read and actually represents the authors thoughts. I would have opted to be a bit more careful and not use AI for a few blog posts after the last incident though if I was working at Cloudflare...

Deno Sandbox 6 months ago

What's with the pricing of these sandbox offerings recently? I assume just trying to milk the AI trend.

It's about 10x what a normal VM would cost at a more affordable hoster. So you better have it run only 10% of the time or you're just paying more for something more constrained.

A full month of runtime would be about $50 bucks for a 2vCPU 1GB RAM 10GB SSD mini-VM that you can get easily for $5 elsewhere.

We have also a service impacted by this. There is no status acknowledgement of the severe degradation. Most of the time queries just take 2-10s, some even 20s+ and a few times they even hang indefinitely. It's a CF-internal networking issue and not SQL execution issue because in the meta data you can see SQL processing taking like 1ms. Even a simple "SELECT 123" in their dashboard D1 console can take 10+ seconds.

Our company opened an issue but we have not heard back in 3 days. The issue has been going on for more than the 3 days, we've noticed it nearly a week ago.

You can also find people in the CF community forums or in other places complaining about D1 latency and networking issues going back as much as a year. Nobody from CF replies in their community forums.

I think this is 1. not an acceptable level of service 2. not acceptable level of support 3. very surprising for a company who is all about fast and reliable networking.

Unfortunately I experienced all kinds of problems with Cloudflare and their non-CDN services over the past couple of months and I've also never seen a major cloud service provider suffering from so many dashboard errors and bugs.

Cloudflare folks, if you read this: you need a Snow Leopard year. You've launched many services and products but they are riddled with bugs, performance issues and outright outages. The recent launch of your data warehouse offering which in our testing (and I think even in your presentation) showed data processing speeds measured in the kilobytes/s while not even supporting aggregations at launch would have been embarassing even 20 years ago.

Please please step up your quality game because you do have some interesting offerings with good potential. The two huge outages last year were just surfacing deeper issues that seem to permeate throughout the stack.

I think you can absolutely compare them and there is no added flexibility, in fact there is less flexibility. There is added convenience though.

For the huge factor in price difference you can keep spare spot VMs on GCP idle and warm all the time and still be an order of magnitude cheaper. You have more features and flexibility with these. You can also discard them at will, they are not charged per month. Pricing granularity in GCP is per second (with 1min minimum) and you can fire up firecracker VMs within milliseconds as another commenter pointed out.

Cloudflare Sandbox have less functionality at a significantly increased price. The tradeoff is simplicity because they are more focused for a specific use case for which they don't need additional configuration or tooling. The downside is that they can't do everything a proper VM can do.

It's a fair tradeoff but I argue the price difference is very much out of balance. But then again it seems to be a feature primarily going after AI companies and there is infinite VC money to burn at the moment.

Cloudflare Containers (and therefore Sandbox) pricing is way too expensive. The pricing is a bit cumbersome to understand by being inconsistent with pricing of other Cloudflare products in terms of units and split between memory, cpu and disk instead of combined per instance. The worst is that it is given in these tiny fractions per second.

Memory: $0.0000025 per additional GiB-second vCPU: $0.000020 per additional vCPU-second Disk: $0.00000007 per additional GB-second

The smaller instance types have super low processing power by getting a fraction of a vCPU. But if you calculate the monthly cost then it comes to:

Memory: $6.48 per GB vCPU: $51.84 per vCPU (!!!) Disk: $0.18 per GB

These prices are more expensive than the already expensive prices of the big cloud providers. For example a t2d-standard-2 on GCP with 2 vCPUs and 8GB with 16GB storage would cost $63.28 per month while the standard-3 instance on CF would cost a whopping $51.84 + $103.68 + $2.90 = $158.42, about 2.5x the price.

Cloudflare Containers also don't have peristent storage and are by design intended to shut down if not used but I could then also go for a spot vm on GCP which would bring the price down to $9.27 which is less than 6% of the CF container cost and I get persistent storage plus a ton of other features on top.

What am I missing?

Isn't that a bit of a holy grail though? If your software can fact check the output of LLMs and prevent hallucinations then why not use that as the AI to get the answers in the first place?