HN user

daemonologist

2,445 karma

knçhwl7tg@mozmail.com (remove the cedilla from the 'c')

Posts2
Comments676
View on HN

My read is that they're announcing a plan to build non-folding wings to test the geometry they would eventually want with folding wings. Which makes sense but isn't very exciting.

It might be that a lot of the cost doesn't show up in the data because it comes out of your paycheck before it's "yours," in the form of your employer's contribution to health insurance premiums. (Anecdotally this is true for me - "my" portion of premiums would not be in the top three, but the total premiums are #2 after housing (for a while when I had roommates they were #1, which is ridiculous).)

(BLS gets this data from surveys: https://www.bls.gov/opub/hom/cex/home.htm )

GPT-5.6 13 days ago

There was a fad a while back of building insanely long prompts - tens of thousands of tokens - including having models write prompts for themselves. I always thought it was counterproductive, especially if you're going to use the prompt more than a couple of times. (That said, the e.g. Claude Code system prompt is insanely long, so if you genuinely have a lot of information to provide maybe it's beneficial. Like, shorter is better, but you don't want to be under-specified.)

FAANG Simulator 14 days ago

I think they're saying that you should build a small business which nets you 85k, not find an employer that underpays you.

(Whether this makes you more resistant to being "fired" is still up for debate of course.)

FAANG Simulator 14 days ago

It goes up - quite a lot - if you take promotions.

But, back in the day FIRE used to mean something beyond just being rich - there was an anti-consumerist bent (and expectation that you'd move away from your expensive city/former job) that usually went along with lower spending.

FAANG Simulator 14 days ago

Only if you seek them out (ski resorts and such). I've lived in both of the most expensive major inland cities (Chicago and Denver) and $85k is plenty in both, even for a small family.

A 7900 XTX is about $850, and the rest of the computer basically just needs to boot Linux. You could easily build such a machine for $1500.

Even that isn't strictly necessary - you can get perfectly acceptable performance by splitting a model between multiple older 12 or 16 GB cards.

I would definitely not recommend WordPress.

If you just want a website for cheap: Bearblog, carrd.co, etc.

if you want all the bells and whistles on a platter: Squarespace, Wix, etc.

if you want to supply all the HTML/CSS yourself: Github Pages or Cloudflare Pages.

(Later, if you want to host the above (except the "bells and whistles" tier) yourself: Hetzner, Digital Ocean, etc.)

Interesting that many of them lead with clams or oysters. (Perhaps this is still a thing at high-end restaurants, but to have them listed so frequently and prominently is completely foreign to me.)

Proof of what is possible with stripped down/optimized software. Imagine the battery life if they built one of these around something with better sleep power (e.g. an nrf52).

Ha, developers know perfectly well how to choose models - they know that the company is footing the bill and Opus will give them the best result and/or require the least clean-up work.

(We've received similar guidance. Even better is that our provider, GitHub Copilot, does not provide usage information to individual users if there is no per-user budget configured. So we just fire our requests into the black box and at the end of the month when IT gets the bill we maybe get a talking-to if it's excessive.)

Only Florida, Maryland, and New York require adults who are new drivers to take any kind of class/instruction. (Most states require it under 18, with a few under 21 or 25.) Everywhere else you can just walk in and take the test.

The perhaps greater problem though is that those tests are completely trivial.

At some point in late 2017 the paper was updated with this additional detail:

    Equal contribution. Listing order is random. Jakob proposed replacing RNNs with self-attention and started the effort to evaluate this idea. Ashish, with Illia, designed and implemented the first Transformer models and has been crucially involved in every aspect of this work. Noam proposed scaled dot-product attention, multi-head attention and the parameter-free position representation and became the other person involved in nearly every detail. Niki designed, implemented, tuned and evaluated countless model variants in our original codebase and tensor2tensor. Llion also experimented with novel model variants, was responsible for our initial codebase, and efficient inference and visualizations. Lukasz and Aidan spent countless long days designing various parts of and implementing tensor2tensor, replacing our earlier codebase, greatly improving results and massively accelerating our research.
In any case, if the authors considered their contributions equal, that's good enough for me.

The benefit of running the full precision version is negligible (probably not even measurable above the benchmark noise floor). Most common for cost-conscious users is to run something around 4-6 bits per weight, which would fit on a 24 or 32 GB card (as you mentioned).

There is no "coder" version of Qwen 3.6; I think they just mean it's a coding-focused model of similar size and performance (to Qwen 3.6 35B-A3B).

Regular Qwen 3.6 benchmarks slightly better and has much wider software support though, so this is probably of interest only to organizations which disallow models trained in China.

The allegation here is that it's not actually a fine-tune of Qwen, but instead an undisclosed mashup (merge) of someone else's fine-tune of Qwen and the original model. Rio subsequently said that the model was in fact a merge, that they did additional fine-tuning after the merge, and that they accidentally uploaded the base merge instead of the version with additional fine-tuning. But this seems like quite an oversight...

There are also significant economies of scale (namely: utilization and batching), which tend to make inference on a shared server more economical even after the operator takes a cut.

You can bounce the ball up slightly (presumably the spin from rolling is modeled or approximated, and gives lift when hitting a bumper), which might be enough to skip from the tee to near the end of the course. Not sure that should be considered for "par" though. Took me 14.

I admit I snorted when that was mentioned. It's frequently ranked as the most desirable place to live on earth.

Not to say the message of the article is completely without merit - there are things to see and do almost everywhere. But if I just get in the car and start driving I will 95% of the time find only strip malls and cornfields. Perhaps a suburban park with some trees.

Unfortunately Radxa and Milk-V are almost completely out of stock and not much cheaper. If you need more than a microcontroller there's no circumventing the memory shortage at this point.

Kicking myself for not buying the Q6A at the beginning of the year (I wanted three and arace would only sell one per customer, but one would've been better than none).

In the US, 99th percentile household wealth is ~$14M, which at historical rates of return is enough to live opulently indefinitely. (Of course although we're discussing a scenario where capital holds most of the cards, who knows if those returns would be dependable.)

    > can it be slower than without speculative decoding in worst case then?
Yes - running the draft model costs compute and memory bandwidth, and running the drafted futures through the main model costs compute. If the draft model were really inaccurate or you're already compute-limited (usually: running large batches) you would expect some slowdown.

In practice, for single-user (non-batched) inference with a working configuration, you pretty much always get some speedup. For non-coding tasks I've seen it be nearly a wash for some people, in which case you might want to avoid it due to the extra memory usage (you'd rather use that memory to run a bigger quant/model, even at a slightly lower speed).

The "library" UI has also gotten radically worse over time (in my family there is a 3G, an early Paperwhite, and a relatively recent base model, and each has a worse and sparser UI than the last). The pages turn faster though, due to improved display/display driver tech.

The tiles are not supposed to ablate - they're supposed to be ~fully reusable. That said I think it's plausible that the much higher iteration speed and lack of a need for human-rating (at least during reentry, for now) will allow for more success than the space shuttle saw with its similar approach.