HN user

seanw265

163 karma

building Easel @ https://tryeasel.dev

bsky.app/profile/seandubb.com

Posts0
Comments59
View on HN
No posts found.
DeepSeek v4 3 months ago

I take some issue with that testing methodology. It seems to me that you're conflating the model's performance with the reliability of whatever provider you're using to run the benchmark.

Many models, especially open weight ones, are served by a variety of providers in their lifetime. Each provider has their own reliability statistics which can vary throughout a model's lifetime, as well as day to day and hour to hour.

Not to mention that there are plenty of gateways that track provider uptime and can intelligently route to the one most likely to complete your request.

Kimi K2.6 also released today. I think it's fair to compare the two models.

Qwen appears to be much more expensive:

- Qwen: $1.3 in / $7.8 out

- Kimi: $0.95 in / $4 out

--

The announcement posts only share two overlapping benchmark results. Qwen appears to score slightly lower on SWE-Bench Pro and Terminal-Bench 2.0.

Qwen:

- Teminal-Bench 2.0: 65.4

- SWE-Bench Pro: 57.3

Kimi:

- Terminal-Bench 2.0: 66.8

- SWE-Bench Pro: 58.6

--

Different models have different strong suits, and benchmarks don't cover everything. But from a numbers perspective, Kimi looks much more appealing.

I'll piggyback on this to highlight Refactoring UI as well. It's an ebook by Adam and Steve, though I'm not sure if it's technically part of Tailwind Labs or not.

This book taught me so much about modern UI design. If you've ever tried building a component and thought to yourself, "hmm something about this looks off," you might benefit from this book.

These days some of the examples might be a little bit dated (fashions come and go), but the principles it teaches you are rock solid.

It's always up to the reader to determine which biases they themself care about.

If you're wondering at what point "we" as a collective will stop caring about a bias or set of biases, I don't think such a time exists.

You'll never get everyone to agree on anything.

I’d argue the goalposts have moved substantially over the past decade. The LLMs we casually use in ChatGPT today would have been described as AGI by many people 15, 10, maybe even 5 years ago.

This is not what the article says at all.

The article is about the constraints of computation, scaling of current inference architecture, and economics.

It is completely unrelated to your claim that cognition is entirely separate from computation.

I tend to agree. Cloudflare and Vercel were able to mitigate in the form of WAF rules, but it's not immediately clear what a user or vendor can do to implement mitigations themselves other than updating their dependencies (quickly!).

IMO the CVE announcement could have been better handled. This was a level 10. If other mitigations can are viable and you know about them, you have a responsibility to disclose them in order to best protect the safety of the billions of users of React applications.

I wonder how many applications are still vulnerable.

After reading the post I kept thinking about two other pieces, and only later realized it was Taylor who had submitted it. His most recent essay [0] actually led me to the Commoncog piece “Are You Playing to Play, or Playing to Win?” [1], and the idea of sub-games felt directly relevant here.

In this case, running a studio without using or promoting AI becomes a kind of sub-game that can be “won” on principle, even if it means losing the actual game that determines whether the business survives. The studio is turning down all AI-related work, and it’s not surprising that the business is now struggling.

I’m not saying the underlying principle is right or wrong, nor do I know the internal dynamics and opinions of their team. But in this case the cost of holding that stance doesn’t fall just on the owner, it also falls on the people who work there.

Links:

[0] https://taylor.town/iq-not-enough

[1] https://commoncog.com/playing-to-play-playing-to-win/

If they are serious they should realize that "80% accuracy" is almost meaningless for this kind of classifier. They should publish a confusion matrix if they haven't already.

I haven’t tried it myself, but if you’re asking specifically about the human models, the article says they’re not generating raw meshes from scratch. They extract the skeleton, shape, and pose from the input and feed that into their HMR system [0], which is a parametric human model with clean topology.

So the human results should have a clean mesh. But that’s separate from whatever pipeline they use for non-human objects.

[0] https://github.com/facebookresearch/MHR

Doable for http and https, but if you're running it in a browser environment, you'll eventually run into issues with CORS and other protocols. To get around this you need a proxy server running elsewhere that exposes the lower layers of the network stack.

Very cool! I'm curious as to how it compares with WASIX in terms of both compatibility and performance.

Also tangentially related: I'd love to see a performant build of Node.js compatible with this runtime (or really any flavor of WASM), but I think you'd run into the same issues that I have with WASIX. Namely build headaches, JIT, and wasm(-in-wasm) support. I'd explore it myself but I've already sunk way more time than is reasonable on that endeavor.

The designer obviously knows a thing or two. I enjoyed the fun presentation that others seem to dislike.

Where I ran into trouble was the readability of the annotations on the visuals. The tiny font combined with the low contrast was too much for me. I found myself squinting and trying to get close to my monitor. Eventually I had to move on, even though I was enjoying the content.

I’ve got a random subdomain hosting a little internal tool. About twice a year, Google Safe Browsing decides it’s phishing and flags it. Sometimes they flag the whole domain for good measure.

Search Console always points to my internal login page, which isn’t public and definitely isn’t phishing.

They clear it quickly when I appeal, and since it’s just for me, I’ve mostly stopped worrying about it.

Living life on the edge, huh?

Sometimes I notice myself go a bit too long without a commit and get nervous. Even if I'm in a deep flow state, I'd rather `commit -m "wip"` than have to rely on a system not built for version control.

Containers might be fine if you’re only sandboxing filesystem access, but once an agent is executing code, kernel-level escapes are a concern. You need at least a VM boundary (or something equivalent) in that case.

I'm still not sold on liquid glass as a whole. It can be quite beautiful, but in the demos provided (and even in Apple's promotional materials) I think readability of UI elements suffers tremendously.

That said, I've seen many attempt to recreate the effect on web but you've outdone them all. The variety and mathematical modeling of edge shapes elevates this implementation above the rest.

If you decide to continue with this, I would love to see:

1. chromatic aberration along displaced areas

2. higher resolution in the refraction

Many people discussing performance issues but this runs like butter on my M3 Pro.

Yes certainly! I've dealt with large datasets like this in the past and know firsthand how challenging it can be to wrangle them.

Something like this would be a great fit for my travel planner app if I knew I could trust that the results were high quality before prompting the user with them.

Btw I edited my earlier comment with a few more examples just before you replied.

Good luck!

I assume that payments from purchases come from you guys, rather than me needing to create and manage an affiliate account with each individual vendor?

You say that commissions average 5%, but what is the variability and where does it come from?

Last, a bit of feedback about the product.

I tried searching "nintendo switch 2" on your homepage and the results that came up kind of sketched me out. You mention that the products are US-only, but the first result clearly says "hong kong" in the title. And the store listed is "My Nintendo Store PT"; is that the official store? When I google that it takes me to the Portuguese version of the nintendo website, and that makes me even more confused.

The second result for the same search appears to be a dress, which is obviously completely unrelated to video games in general.

EDIT: I'm noticing irrelevant results for many queries. Searching "plain white pillowcase", the third result is a t-shirt, the seventh result is a dress, and the eleventh result is a light bulb.

Searching "men's wallet" the very first result is an outdoor picnic table.