HN user

toraway

567 karma
Posts1
Comments244
View on HN

For me, once a couple words loaded the cadence/pace of the streaming words in the response was so recognizably “ChatGPT” it immediately lost the sorta eerie mysterious feel and almost veered into parody/comedy.

Like imagining the wizarding world full of Hogwarts students writing out prompts for “Write a 500 word history of the polyjuice potion, sound natural using my own voice, do not use em dashes, no mistakes.”

OP submitted a Substack newsletter as a source with recent posts including such topics as vaccine “deaths” during COVID, their link to autism, and Fauci bioweapon conspiracy theories implicating the entire scientific establishment.

With comments on the current post agreeing that the shadowy cabal will suppress yet another miracle treatment to keep for themselves, so the “goyim” can be sold poison.

Though it does seem safe to say your original claim of "way slower than the Neo" isn't correct. Considering it's losing in one benchmark and only ahead in Geekbench (that tends to show higher scores for Apple processors relative to other benchmarks anyway).

"Roughly equal" seems to be a more accurate description.

Nice comparison, I've been super impressed by both Deepseek V4 models, particularly Flash given the crazy value for price vs. performance.

It can definitely do "stupid" things and get off track at times but I've found it can easily handle routine web dev tasks like 9/10 times, and using Pro to handle any large refactors/tricky bugs/etc.

The only really negatives are both models (but particularly Flash V4) occasionally have a strange issue parsing instructions, almost like a "language barrier" where a clear instruction gets bizarrely misinterpreted in a subtle but very problematic way. It feels a bit like a SOTA model a year ago where they'd occasionally just miss the plot entirely while still being technically competent but misdirected.

Also not really a negative, but I can't handle watching the reasoning output on Pro anymore haha. It like actually started stressing me out and giving me heartburn watching it get something right on the first or second idea... and then spend like 5 minutes looping through a dozen extremely dumb guesses with "But wait.... Or... Unless..." lol.

Even if I knew it would (usually) end up where it should I just couldn't stand seeing it consider, like, deleting my prod DB and recreating tables manually/ripping out some critical dependency/etc without interupting it to say "Holy shit you had it right the first time, for the love of god just start doing the thing now and move on".

Having a static, immovable belief system about something like copyright that is unaffected by seismic shifts in the real world also doesn't seem very logical.

If like, Disney did a 180 overnight and bought rights from Google to scan every writer's saved work in Docs with some flimsy legal argument then a person saying "wait doesn't copyright actually protect that" would make sense. Even if you were previously upset about them suing schools for using 80 year art.

Fox to buy Roku 1 month ago

Yeah that's the Number 1 issue I have with Jellyfin.

It seems to be tolerating whatever semi-organized structure I give it until it just faceplants on some specific show and I have to tediously reorganize the directory structure/names and manual refresh until the metadata lines up correctly.

I like that I don't feel I'm about to be rugpulled on Jellyfin and the client is pretty solid for me but the library scanning is pretty aggravating at times.

Fox to buy Roku 1 month ago

Uh, it's a complete false dichotomy? There is literally no reason you need to participate in a botnet to stream content for free.

That's ... not a thing. Those sticks just glom on to free software maintained by other hardworking unpaid devs to steal residential IPs from unsuspecting buyers drawn to the "all-in-one" pitch for their sketchy VPNs and/or botnets. Then, eventually whatever API keys/endpoints they stole for streaming stop working and all you're left with is the botnet part of the deal.

This is like saying the included porn malware you got bundled with uTorrent from the first sponsored link on Google is a price worth paying to access The Pirate Bay and stick it to Netflix, lol.

Why earth would anyone voluntarily advocate for that/defend the malware authors instead of just downloading qBitorrent from Github like a normal person?!

Fox to buy Roku 1 month ago

Or you buy a non-scammy Onn stick for $20-$30 from Walmart instead, install a launcher like Projectivy/ATV Launcher Pro from Play Store (or Aurora/F-Droid), and either choose your streaming app subscriptions ... or remove them all and install Stremio/Kodi/Plex/Jellyfin etc for your own preferred "alternative" streaming sources via Usenet/Debrid/Torrents/etc.

Fox to buy Roku 1 month ago

I've used Projectivy for years on every Google TV stick I own (Google/Onn). Works perfectly with full customization/zero ads on the free version and has a workaround to take over the Home navigation to bypass the built-in launcher without ADB or rooting. I chip in for the premium version to help out the dev since I get so much value out of it but the freemium features are mostly just cosmetic and the free version has everything you need.

Fox to buy Roku 1 month ago

ADB is rarely actually a requirement unless you really want to do it the "right" way and fully remove the launcher.

I always use a custom launcher (Projectivy) on my Google TV devices, lately typically the $20 Onn stick and intercept the Home navigation to open the launcher either using the option built into Projectivy or with a free app from the Play Store/Fdroid.

Takes <5 minutes to setup everything once and then I basically forget the native Google TV launcher exists. Pretty much unbeatable value for a $20 ad-free Jellyfin/Plex/Kodi/Stremio setup. YMMV with different models but I also had no issues remapping the remote buttons from Netflix/etc to my own apps (including the "Free TV" button to launch Stremio which I always enjoy).

Also (somewhat ironically) the best smart TV OS to look for on cheap/subsidized TVs is built-in Google TV. Since they can easily be configured as 100% "dumb" on startup without any ads/nags/etc (it's the first question you're asked). The TV never hits Wifi to update and the remote/menus just do normal TV stuff without any "smart" features. Otherwise, it's luck of the draw how miserable/impossible the manufacturer makes avoiding Wifi/updates.

(Or you could do the same process installing the custom launcher on the TV's built-in Googe TV, but then you're at the mercy of the CPU/RAM the OEM included in BoM some # of years ago and lose the clean seperation between dumb TV/replaceable stick).

$20 Onn stick + $199 "smart" Google TV in dumb mode goes really far these days for a locally hosted setup without ads/annoyances.

GLM 5.2 Is Out 1 month ago

I still find it baffling how the idea that HN is "unashamedly anti-ai" gets repeated.

Every single model release gets submitted within minutes of an announcement and frequently break 1000+ points within an hour or two. Blog posts about vibe coding or the current flavor of harness/workflow/tool are constantly making the front page. Karpathy's latest writing/presentations or "Learn how LLMs work using X" are perennial front page content.

There were moments in 2023/2024 where all but a handful of posts on the front page were about AI (and not the Reddit r/popular "residents worried about infrasound and EM radiation near new datacenter" variety).

For example, the responses to this very recent post were overwhelmingly praising Gen AI's capabilities:

Ask HN: What was your "oh shit" moment with GenAI?

https://news.ycombinator.com/item?id=48406174

Or this post which rocketed to 2000+ points a year ago without bothering to steel man opposing arguments:

My AI skeptic friends are all nuts

https://news.ycombinator.com/item?id=44163063

There are counter examples of course but just because HN isn't exclusively AI hype at all times doesn't mean it's "unashamedly anti-AI".

I honestly can't think of any single topic other than the Snowden leaks in 2013/2014 that even comes close to dominating HN discussion like LLMs/GenAI from 2022 to present.

This just seems like a not great example to make that point though. Since whatever Claude tells the kid looking to build a reactor or even bomb is almost certainly going to be more grounded and professional than:

  Step 1. Obtain pliers 
  Step 2. Obtain 300 discarded smoke detectors 
  Step 3. Start yanking!
Instead it would send them on a wild goose chase for unobtainable isotopes, centrifuges, heavy water, etc where the biggest risk is probably getting reported to the police by some chemical or industrial equipment supplier. Which is a better outcome compared to contaminating their home with radiation and exposing anyone they interact with.

You'd maybe get a sketchy but near-viable plan that could be dangerous if asked for a dirty bomb, but there the danger would more be the conventional explosives and not where to source radioisotopes, as it was already common knowledge that most residential smoke detectors contained americium until recently.

  > However ~70% of the uninsured are eligible for Medicaid, subsidy, or employer insurance, so there's room to improve on getting those people signed up.
That number will decrease once Trump’s Medicaid work requirements take effect, and subsidies were also significantly reduced.

I’d love if our government saw “room to improve” there instead of doing the exact opposite and working overtime to reduce the number of fully insured people.

Notes on DeepSeek 1 month ago

Deepseek Flash V4 really was a "holy shit" moment and deserves the praise/hype it's been getting from users. I have a multi-tier subscription strategy I've maintained for the last year of: 1. $20-$30 plan from first Claude now Codex for "SOTA" 2. Gemini via the extra $10/mo or so from my Google One plan 3. a cheap fallback plan.

Together it gives me plenty of head room/model performance for $40ish/mo, plus letting me compare the various models over time.

Originally I'd been using the Z.AI plan (that I'm still grandfathered into for <1 yr) as my cheap plan but wasn't keeping up with the SOTA progress and is slow/limited now. So I subscribed to the Opencode Go plan and use Deepseek Flash V4 almost exclusively and it is insane how much usage I can get for $10/mo.

I did the math on my Flash usage vs. what I'm paying Opencode and I'm typically not even exceeding $10 in API costs! So it's actually sustainable not rugpull pricing at least for me. I can pound it with requests/agentic loops and have it running for 30 min doing whatever the fuck and check back and have spent literal pennies for what would have cost $30+ on my work's Github Copilot plan.

I know enterprise world works under different rules and isn't price sensitive in the same ways as an individual but I truly don't see how this is sustainable for the US AI giants in the long term to maintain like 25x+ markup for 1.25x performance benefit.

IMO it does help explain the recent emphasis on secret, scary "super models" like Mythos to muddy the waters for decision makers with hype and FOMO at at time when companies are beginning to seriously scrutinize their token spending for the first time.

Claude Fable 5 1 month ago

Changing a domain name doesn't actually amend federal law.

Just like how changing Kennedy Center letterhead to Trump Kennedy Center for a year didn't actually legally rename it.

Once a case with sufficient standing got in front of a judge it reverted to the actual legal name on the basis that only Congress can change the statutorily defined name.

Claude Fable 5 1 month ago

Tbf the first line of your first comment is:

  > Pelican for Fable 5 on default settings is a clear improvement on Opus 4.8
And doesn't contain any actual criticism within the comment (your blog post might, but just referring to what was posted on HN, which is a bit booster-y on its own).

Huh? That's just basic history of xAI. At no point was xAI being sold as a Coreweave-like middleman leasing out data centers to hyperscalars. That's a boring, regular business. The pitch was that xAI would develop groundbreaking AI models for Grok which would attract actual users and generate revenue.

That evidently did not work out, otherwise these deals wouldn't be happening. OpenAI and Anthropic aren't leasing out their datacenters, if they did it would be obvious something was grossly wrong with their projected growth.

Bad example but since it literally just happened a few hours ago:

Teams Copilot meeting assistant auto-renamed a meeting title/summary that’s now prominently placed at the top to “Month end close wrap up discussion“ because someone posted in chat “sorry can’t make the meeting, we’re wrapping up month end close”.

Really confused the next guy who joined the meeting and derailed things for a minute or two before we could get back on topic.

Yeah, I have a contract project for a webapp/integration to legacy Excel tool with an API endpoint for exchanging data with Excel. Over time, I notice issues or need to add functionality in the data processing and hadn't been closely watching the code changes Claude Code made to the API as long as it worked as expected/tests passed.

When I eventually read through the current state of the upload processing code it was like an absurd tree of checks on checks on fallbacks on triple checks added in response to whatever bug I reported in a bizarrely additive way and could be massively simplified (which would also make it less brittle to edge cases that then demanded more checks and workarounds).

The other issue is that for the upload API, there is documentation but not for every little bug or edge case so each time the model "wakes up" and loads everything into context it sees that crazy web of checks and edge cases as the only source of truth for the API so is hesitant to touch anything unless 100% necessary which then leads to more conservative behavior of additive code which makes the problem worse over time.

Codex seems a bit better but I still have to guide it towards proper abstractions/refactors to avoid that piling on cruft effect.

That account is setup pretty well compared to many I see on here, no giant red flags in any individual comment and they’ve prompted/filtered out the repetitive structure decently so they don’t all look exactly the same. But exhibits number of common attributes/failure modes you notice when you spend a while reading through the history of obvious bot comments.

Some of the signs in comment history:

1. Multiple uses of the notorious HN bot phrase “X is real” e.g “The tradeoff is real” Spam bot replies love that, at this point if I read that phrase on HN I immediately become suspicious. Also, ending a comment with an open ended question or random caveat/concern/qualifier too frequently.

2. Rapid fire commenting. Most new HN accounts run by actual people usually ease into commenting regularly after their first comment, spam bots go from 0-100 suddenly replying with detailed, uncannily on-topic comments 1-3x per day

3. Unnaturally on-topic but uncanny valley comments. Humans typically respond to an idiosyncratic idea of what the topic of a discussion is, but bots typically respond with a perfectly responsive summary and/or personal experience.

As for uncanny valley, this is the big one that caught my eye:

  > Love this approach. SQLite's WAL mode + litestream backup handles most durability needs beautifully. The simplicity of keeping state in a single file you can query with SQL is underrated. Going to try this for our next side project.
It’s just … bizarre haha. It’s replying with what is pretty much the most basic, fundamental description of what SQLite is, with zero actual opinion/experience/commentary.

It’s like responding to an article about PepsiCo Q2 Earnings with “Great read - my family of four enjoys drinking Pepsi products, the carbonation and various flavor choices are a big plus compared to other brands”. If it wasn’t a bot it would have to be someone ineptly karma farming, and either is trash deserving of the flags it got.

Plus, the initial comment in this thread while not enough evidence on its own is already suspicious because it’s confidently stating an anecdote about billing that just doesn’t really make sense as if it’s a super common way tokens are billed.

It could be true in their unique case but it’s just weird and presented like it’s common knowledge so comes across either as a human pretending to have experience or a bot slightly off target.

But it’s not enough without the full comment history, and even without a giant smoking gun having seen a lot of bot comment histories it fits the pattern closely enough in most ways to agree with the other commenter (and the flags/HN shadowbanning) that they are a LLM bot.

No, that's just OS war tribalism talking. I regularly use a M1 Macbook, Lenovo Ideapad 14 Pro with Windows 11, and an ARM Lenovo IP Slim 3 Chromebook. Each have their strengths and weakness at different price points.

Chromebooks (typical) strengths are 1. zero maintenance/instant updates 2. 10 years of OS support 3. battery life 4. touchpad 5. value/$ with options at very low price points.

I paid around $150 for my Lenovo ARM Chromebook (came out a few years ago) and got around 14-15 hours of battery life new with no noticeable fans/heat. Has virtually no self-discharge when in deep sleep and boots in <10 seconds after sitting around for 2+ weeks. Even with 4 GB RAM ChromeOS handles very well under memory pressure (with the memory saver tab mode turned on), that I can have multiple windows with dozens of tabs open before things start slowing down.

I use the Linux VM in ChromeOS for light dev work (and disabled Play Store/Android), the touchpad is absolutely fantastic (which isn't unusual on even cheap Chromebooks, Google actually prioritizes driver support for multitouch/palm rejection unlike the cheap Windows crap models), security is rock solid with essentially no risk of malware/viruses/etc and have literally no maintenance/stability issues that waste my time. Chromebooks are by far the best choice if the question is truly "how do I minimize Grandma needing help solving computer problems", even current locked down MacOS has so many more ways it can break/confuse compared to ChromeOS.

This is the fifth or so Chromebook I've owned over the years, having used both the ultra-premium end of the spectrum (original Pixelbook) and the very cheapo end, and this machine is one of my favorite tech purchases overall in the last few years. I'd definitely recommend 8GB of RAM if possible, but for the typical Chromebook casual web browsing use case 4GB is perfectly serviceable (especially on a newer ARM SoC).

A $150 Chromebook is not intended to replace a $3000 Macbook with 64GB of RAM to run a half dozen Docker images, etc so sure, they'll "suck" in that match up, but they are an extremely competitive option on most metrics for the "someone just needs to browse the web and I don't want to be pestered by IT issues" case.

You can still bypass the login requirement for Win 11 and that annoyance only happens once during install vs. every time you try to run a non-notarized app.

It’s easily in my top 3 most hated things about my MacBook. Plus, knowing Apple and the history of that “feature”, it will only ratchet towards becoming even more of a pain over time (it was actually tolerable back before they removed the hotkey to bypass).

For me, after running Win11debloat one time Win 11 disappears into the background 95% of the time, like an OS should. Unfortunately I don’t the luxury of doing something equivalent on MacOS without completely disabling SIP.

How is selling something and then removing the ability to use what was paid for not fraud? (setting aside the EULAs companies currently get away with using to sidestep the question)

There's even a specific term for scams where you pay money based on a specific description for an item being sold that is then changed after the time of sale known as a "bait and switch".

At my company the grassroots advocacy from devs has certainly been for Claude Code.

Unfortunately even though we have a degree or two of seperation from most federal contracts the punitive DoD blacklisting had enough of a chilling effect on our legal team to make them drag their feet on approving any contract involving Anthropic.

So I pitched OpenAI Business with Codex so we could drop our Github Copilot Business subscription before the billing change takes effect June 1st which was approved without pushback.

I felt some responsibility for finding an immediate solution to dump Copilot since I was the one who recommended adopting Copilot in the first place, ugh... Our prices would have quadrupled based on the single month Microsoft in their beneficence allowed previewing with their tool to simulate what the post-rug pull pricing would have looked like.

Codex becoming more or less a 1:1 replacement for CC made that a no brainer given our options and the exploitative value proposition of Copilot under the new pricing model (which Microsoft evidently hoped companies like us would just accept despite being a third tier option in the dev space these days).

  > Because how far does their stance against AI go? They won't accept music. What about if AI created a cure that could save their child?
The problem with this type of argument employing hyperbole ad absurdum to demonstrate irrationality is that it’s self negating.

If AI cured cancer then by definition it would no longer be the technology that’s primary use case is churning out various forms of derivative slop. And so the balance between its value vs the economic/social/environmental costs would immediately and fundamentally change.

Losing my job, spending 3x as much to replace my PC while my favorite websites devolve into a cesspool of spam might not feel worth it just because I can now vibe code a todo app in 2 minutes while listening to a 600 hour playlist of personalized elevator music.

But if it cured my dad’s cancer and my mom’s Parkinson’s? Well, that’s a different story…

But that isn't how it works, it's not a prompt like asking permission to use the camera allow/deny. The user gets presented with list of compatible devices and they have to select one themselves.

An attacker could try to convince users to select something specific but that depends on the actual devices that are present and the "default" option to a confused non-technical person is to just cancel out of the list.

The US Supreme Court effectively legalizing sports betting overnight provided an almost perfect real-world experiment to test that argument. And the result?

"Prediction market" ads running constantly on both sports channels/websites like ESPN. Shortly followed by mainstream cable news like CNN featuring Polymarket stats as a routine part of their horserace polling coverage. Gambling is now an omnipresent temptation for anyone even casually interested in following sports/political news.

And now millions of young men who previously would have had to seek out niche illegal venues to gamble have several dozen different apps on their phone offering to light their disposable income on fire by clicking a couple buttons every paycheck.

Wait, 11 vulnerabilities were discovered entirely in the timeframe after Mythos found 1? That seems like it would effectively debunk the theory that curl was so uniquely hardened that only 1 vulnerability even existed for Mythos to find, which I read numerous times back on the HN thread for the curl/Mythos blog post.