HN user

icyfox

1,564 karma

pierce [at] freeman.vc despite the email i mostly code

Posts37
Comments117
View on HN
github.com 2mo ago

Show HN: Rotunda - A browser built for agents with simulated typing

icyfox
13pts5
pierce.dev 6mo ago

A deep dive on agent sandboxes

icyfox
68pts20
adrift.today 6mo ago

Messages in bottles across the digital sea

icyfox
4pts1
pierce.dev 8mo ago

Automating our home video imports

icyfox
86pts41
nyx.run 10mo ago

Nyx – An Experiment in Artificial Survival

icyfox
3pts0
github.com 11mo ago

Show HN: Determystic, get agents to follow your coding conventions

icyfox
1pts0
pierce.dev 1y ago

Building a (kind of) invisible Mac app

icyfox
1pts0
pierce.dev 1y ago

Making an ASCII Animation

icyfox
6pts0
pierce.dev 1y ago

Under the Hood of Claude Code

icyfox
2pts0
pierce.dev 1y ago

Speeding up sideeffects with JIT in mountaineer

icyfox
2pts0
pierce.dev 1y ago

How Text Diffusion Works

icyfox
5pts0
github.com 1y ago

Show HN: Iceaxe – A modern, fast ORM for Python and Postgres

icyfox
2pts2
freeman.vc 2y ago

Generating database migrations with acyclic graphs

icyfox
2pts0
github.com 2y ago

Show HN: Mountaineer – Webapps in Python and React

icyfox
145pts57
freeman.vc 2y ago

Should you fine tune for JSON output?

icyfox
2pts0
github.com 2y ago

Lorax: Serve 100s of Fine-Tuned LLMs in Production for the Cost of 1

icyfox
5pts1
oxide.computer 2y ago

RFDs: A Tool for Discussion

icyfox
1pts0
rafalcieslak.wordpress.com 2y ago

Using LD_PRELOAD to cheat, inject features and investigate programs

icyfox
202pts108
freeman.vc 2y ago

Subclassing wheel builds for fun and profit

icyfox
1pts0
github.com 3y ago

Show HN: GPT-JSON – Structured and typehinted GPT responses in Python

icyfox
174pts72
freeman.vc 3y ago

Let's Talk about Siri in 2023

icyfox
3pts1
freeman.vc 3y ago

Greater sequence lengths will set us free

icyfox
2pts0
github.com 3y ago

Ulixee Hero - The web browser built for scraping

icyfox
3pts2
www.nytimes.com 3y ago

How California’s Bullet Train Went Off the Rails

icyfox
52pts10
troubles.md 3y ago

Acheiving warp speed with Rust (2017)

icyfox
2pts0
www.robinrendle.com 3y ago

The Futures of Typography (2017)

icyfox
1pts0
www.theatlantic.com 3y ago

Manchin’s New Bill Could Lead to One Big Climate Win

icyfox
1pts0
freeman.vc 3y ago

AWS vs. GCP reliability is wildly different

icyfox
545pts234
freeman.vc 3y ago

Headfull Browsers Beat Headless

icyfox
2pts0
www.newyorker.com 3y ago

The Pleasure of Mechanical-Keyboard Tinkering

icyfox
3pts0

1. Factory limits basically. There's a limit to the amount of fabrication lines that can create ram. Combined with the market incentives right now to make high bandwidth memory (HBM) over server memory (DRAM)... HBM starts as DRAM dies, so it competes with normal DRAM for wafer starts / cleanroom fab capacity.

2. Eventually more plants will come on line. Most of the main manufacturers have announced expansions but these can take O(years) to come online.

Not particularly. I'm not yet convinced people's mouse movements are unique enough to our identity that they're useful as a fingerprint, whereas it's very easy to classify whether something looks bezier or looks human.

Eventually I'm hoping to collect enough data here to train a biased decoding model, so you could input some randomized personality vector (which implicitly encodes slow movement, jerky motion, trackpad, mouse, etc) and have that impact the RNN generation. So in theory there would be infinite combinations from the larger subspace we're sampling from.

So much of what Apple has lost over the last 10 years is a lower bar for what counts as good enough.

You see this most obviously in software and marketing - the kinds of decisions where only a few people sign off at the end, and where "good enough" is whatever those few people decide it is. You see it less in hardware and procurement where there's a powerful review cycle and scrutiny at every level of the stack. Work there is more immediately measurable: benchmarks for performance, dollars for cost.

The "vibe" of software, or of a PDF [^1], is much harder to catch that way. There's no benchmark that flags it and most conventional executives aren't drilling down in that level of detail to see it either.

You want distributed decision-making, of course. But that only works well if it's distributed to people who've cultivated their own taste and who will make good calls under pressure. I'm not sure how much of that gets fixed by leadership change at the top. Taste isn't really something a CEO can decree into a 60,000 person org. But I've only heard good things about Ternus, so I'm optimistic. Fingers crossed for a bright new chapter.

[^1]: https://www.apple.com/promo/pdf/US_FY26_Earth_Day_Promo_Tand...

As far as I've seen, local OSS video understanding models just really aren't there yet. I briefly looked at facial recognition models but a good amount of signal was actually in the video's audio instead of the raw video frames. Depends on the accuracy you're looking for at the end of the day.

Digitizing my old tapes was one of the most rewarding side projects that I did over the last year. I managed to get in under the wire (pun intended) of Firewire compatibility on Sequoia and a long daisy-chain of adapters. But it was clear the days of this approach were numbered. I'm optimistic these 3rd party accessories will become more standardized into self-contained cheap boxes where people can easily transfer over their stuff before camcorders degrade.

My pipeline went camera -> dvrescue -> ffmpeg -> clip chunking -> gemini for auto tagging of family members and locations where things were shot.

We now have all our family's footage hosted on a NAS with Jellyfin serving over Tailscale to my parents Macbooks. I found the clip chunking in particular made the footage a lot more watchable than just importing the two-hour long tapes although ymmv.

I'm not commenting on the externalities. For that I'd also cite economic impact, job loss, occasional emergency services issues, etc. I'm saying the experience when you yourself are taking a ride. I haven't met a single person who's said "this sucked - I'm going back to Uber".

Waymo is such an interesting case study. For most other ~AI deployments you have strong public reaction to the proliferation of slop, non-human failure modes, cost cutting at the expense of quality, etc. But I haven't met a single person who doesn't like the experience of Waymo. They ended up cracking the code on what I suspect people really want:

- consistent car quality

- safety of the drive (conservative driving and potential fear of drivers)

- no randomly chatty driver

All of those feel like a breath of fresh air especially when stacked up against the current state of Uber & Lyft rides. People really just want consistency. I don't actually think you needed AI to get there (I've had occasional rides in black cars that provided the same experience). Waymo was just right time, right place, right price.

Fair point about the source, but the classification usually follows the mode of delivery, not the organism of origin.

Many plant-derived compounds function as venoms once introduced into the bloodstream (arrow coatings, darts, etc.), even if they’re also toxic when ingested. Curare is one example of a plant-based compound - lethal in blood, but largely harmless if eaten.

So while Boophone is absolutely a poison in the ecological sense, using it on arrows still fits the venom/toxin distinction better than a purely ingested poison. Otherwise why would people hunt with this if they got sick the second they ate the meat?

At the risk of being overly pedantic, topologists would typically classify this as venom.

Venom is inert if digested; it's only a problem if it gets in your blood stream. So arrows that were laced with venom and thereby contaminated meat were actually perfectly safe to eat.

Poison is different. If ingested, inhaled, or absorbed it will kill you.

Exactly half of these HN usernames actually exist. So either there are enough people on HN that follow common conventions for Gemini to guess from a more general distribution, or Gemini has memorized some of the more popular posters. The ones that are missing:

- aphyr_bot - bio_hacker - concerned_grandson - cyborg_sec - dang_fan - edge_compute - founder_jane - glasshole2 - monad_lover - muskwatch - net_hacker - oldtimer99 - persistence_is_key - physics_lover - policy_wonk - pure_coder - qemu_fan - retro_fix - skeptic_ai - stock_watcher

Huge opportunity for someone to become the actual dang fan.

We talked about this model in some depth on the last Pretrained episode: https://youtu.be/5weFerGhO84?si=Eh_92_9PPKyiTU_h&t=1743

Some interesting takeaways imo:

- Uses existing model backbones for text encoding & semantic tokens (why reinvent the wheel if you don't need to?)

- Trains on a whole lot of synthetic captions of different lengths, ostensibly generated using some existing vision LLM

- Solid text generation support is facilitated by training on all OCR'd text from the ground truth image. This seems to match how Nano Banana Pro got so good as well; I've seen its thinking tokens sketch out exactly what text to say in the image before it renders.

I used Serp via API many moons ago. The most interesting part of the company imo is their legal defense of different plans:

  Production - $150
  15,000 searches / month
  U.S. Legal Shield
ie. "Our U.S. Legal Shield protects your right to crawl and parse public search engine data under the First Amendment. We assume scraping and parsing liability for customers on most recurring plans unless your usage is illegal."

I imagine at least some portion of companies use them just for this liability shield.

I'm always a bit surprised how long it can take to triage and fix these pretty glaring security vulnerabilities. October 27, 2025 disclosure and November 4, 2025 email confirmation seems like a long time to have their entire client file system exposed. Sure the actual bug ended up being (what I imagine to be) a <1hr fix plus the time for QA testing to make sure it didn't break anything.

Is the issue that people aren't checking their security@ email addresses? People are on holiday? These emails get so much spam it's really hard to separate the noise from the legit signal? I'm genuinely curious.

Seems like the live demo is bear hugged - been waiting for ~5 minutes now. A bit ironic given their landing page: Don’t make your prospects wait–ever again

In its current iteration this demo might net discourage your future clients rather than encourage them.

I like the idea in general as an alternative to needing to book with a BDE. I'd always prefer to just self serve for a new product; anything that gates my time (sales calls, popover walkthroughs, etc) is something I'd prefer to skip. But I know non-engineering customers really love these calls to see the power of a new platform. I wonder if they'll be as engaged during an AI walkthrough versus when there's a person on the other end of the phone.

Gemini 3 8 months ago

I'm not sure how concerned people should be at the trend lines. If you're building a product that already works well, you shouldn't feel the need to upgrade to a larger parameter model. If your product doesn't work and the new architectures unlock performance that would let you have a feasible business, even a 2x on input tokens shouldn't be the dealbreaker.

If we're paying more for a more petaflop heavy model, it makes sense that costs would go up. What really would concern me is if companies start ratcheting prices up for models with the same level of performance. My hope is raw hardware costs and OSS releases keep a lid on the margin pressure.

Gemini 3 8 months ago

Pretty happy the under 200k token pricing is staying in the same ballpark as Gemini 2.5 Pro:

Input: $1.25 -> $2.00 (1M tokens)

Output: $10.00 -> $12.00

Squeezes a bit more margin out of app layer companies, certainly, but there's a good chance that for tasks that really require a sota model it can be more than justified.

AdaptiveDetector definitely did a better job, will append these new stats to the post:

precision 0.397, recall 0.727, F1 0.513, mean temporal error 0.307 s

Nice to see you on here! I used the ContentDetector with a threshold of 27.0 and otherwise default parameters. Realize I could have done a grid sweep to really hone in on a good param range, but because I had only one input video labeled I wanted something that would work well enough out of the box. I imagine this dataset is rather... heterogenous.

If you happen to know a better apriori threshold I would be happy to re-run the analysis and update the chart.

I bought a Sony TRV120 10+ years ago, back when I was doing this conversion project for the first time. It's built like a tank and still works today.

At the risk of being smited by professional archivists, I'm willing to wager that 99.9% of people can't tell the difference between a $5k archival rig and one of these higher quality camcorders. At this point it really feels like the biggest inhibiter to good quality digitization is the decaying of the tape versus the archival setup.

That said - for anyone with the time, patience, and soldering abilities I would love a more proper A/B test with RF signal capture software encoding. Something like this:

https://rastrillo.ca/digitizing-video8-tapes-with-vhs-decode...

Since the webapp is pretty opinionated to my setup (ie. linking against AVFoundation, using MPS for inference, always capturing an image after import) I didn't originally think it would be that useful to open source. Happy to do so - are you looking to get something specific out of it?

Most of my tapes did have pretty detailed narration and date overlays written directly to tape. But even without narration I still had luck doing basic event summarization and facial recognition of family members to build the tags.

Yes - garbage in / garbage out still holds true for most things when it comes to LLM training.

The two bits about this paper that I think are worth calling out specifically:

- A reasonable amount of post-training can't save you when your pretraining comes from a bad pipeline; ie. even if the syntactics of the input pretrained data are legitimate it has learned some bad implicit behavior (thought skipping)

- Trying to classify "bad data" is itself a nontrivial problem. Here the heuristic approach of engagement actually proved more reliable than an LLM classification of the content

It's a neat chip but I couldn't bring myself to spend in excess of the price of the UPS ($439.00 for RMCARD at time of writing). I ended up hooking my UPS via USB to my existing home server via NUT and it's been working well.

All of the products you mention already had research teams (in the case of ChatGPT and Claude that actually predated most of their engineers). So knowing how to build small language models was always in their wheel house. Scaling up to larger LLMs required a few algorithmic advancements but for the most part it was a question of sourcing more data and more compute. The remarkable part of transformers is their scaling laws, which let us achieve much better models without having to reinvent new architecture.

The folks at deno have really done a fantastic job at pushing a JS runtime forward in a way that's more easily plug and play for the community. I've used `denoland/deno_core` and `denoland/rusty_v8` quite a bit in embedded projects where I need full JS support but I can't assume users have node/bun/etc installed locally.

Not surprised to see yt-dlp make a similar choice.

From the jump, even without the emdashes, it was crystal clear that this post was written by Claude. I'm sure the OP put their own ideas into the prompt, provided some sources, etc. But reading some of these phrases provoked a pretty visceral sense that I'm just reading an LLM's output:

"The harsh reality" "he perfectly captured" "architectural decisions that make senior engineers weep" "fundamental issue"

It makes me wonder whether a whole class of writing is going to be deprecated because the cadence is just too similar to LLM outputs.

The LLM Lobotomy? 10 months ago

I'm responding to the parent comment who's suggesting we version control the "model" in Docker. There are infra reasons why companies don't do that. Numerical instability is one class of inference issues, but there can be other bugs in the stack separate from them intentionally changing the weights or switching to a quantized model.

As for the original forum post:

- Multiple numerical computation bugs can compound to make things worse (we saw this in the latest Anthropic post-mortum)

- OP didn't provide any details on eval methodology, so I don't think it's worth speculating on this anecdotal report until we see more data