HN user

mchinen

168 karma

Healthspan Capital (Longevity Biotech Fund) Manager Audio ML researcher/SWE ex-Audacity team member

audiomlpapers.com - latest and top audio ML paper feed machinelearningpapers.com - same concept for general ML longevitypapers.com - same concept for longevity biotech

blog at michaelchinen.com

Posts7
Comments58
View on HN

I enjoyed this.

I'm surprised there's no LoRa layer or auto RL or adversarial step to reduce the stock phrases as they pop up. Is it really so hard to push these out? Or is it just whack-a-mole no matter what you do?

GPT-5.6 13 days ago

I totally missed that, because in the charts they showcase for coding, the SWEBench score is not present, they only include it at the end of the post in tables. Hmm.

Great catch.

GPT-5.6 13 days ago

The frontier graph on all these benchmark are extremely in favor of 5.6 Sol over Fable, more than the best model comparisons in previous iterations.

I'd like to know how cherry-picked this is, and what tests it performed less overwhelmingly in, but I suppose that info is not going to be on this post.

If it pans out to be as good as it says, that's great. On the other hand, if this model is not overwhelmingly impressive over Fable, I will lose what remaining trust I had in these announcements.

I just meant a big innovation that reshapes everything. I should have used 'level' instead of 'type' here.

But there are a lot of analogies to computation in bio as a physical, atomic forces-driven, massively parallel computer, so it's possible there will be something related to electronics and computers that falls out. For example, there's also applications directly related to other fields including DNA storage of data and neuron-based computation.

Cool to see this from Brian Hie, who was doing interesting computational bio research at Meta's FAIR before they axed it. Interesting that this is work on the more physical/testing/manufacturing level than the computational, but it seems very useful.

It's hard to quantify the impact of new foundational tools like this at launch. Most of the time it falls flat, but even the successes are difficult. For example, CRISPR has led to interesting experiments and treatments on the way, but the effect does feel muted compared to the initial predictions. But there are many other related techniques that can be pulled out of this original research (e.g. dCas9 which lets you operate without cutting).

Similar story with cellular reprogramming.

Eventually one of these things will surface that will be GPU/transistor type innovations.

The audio side is even more interesting, as it seems they totally got rid of positional embedding are just doing a single linear transform to match the LLM input dimension and that's it.

Audio: We simplified audio processing even further. We removed the audio encoder entirely and projected the raw audio signal into the same dimensional space as text tokens.

This is interesting, because in the leaked code, it was found that they detected simple swearing keywords for analytics that get sent to Anthropic, but also had directions to keep the behavior the same for claude. I also have the feeling a 'wtf' does something, but it does feel good and might just be placebo, because 'that is still wrong' sometimes works the 4th time too. Or maybe they changed something.

Halt and Catch Fire 2 months ago

It's special for sure. For those on the fence, it has some writing and direction flaws, especially with minor characters, like the disgruntled neighbor and IP theft bit in the first season. But it grows as a show over time, and the 5 leads (including Toby Huss) smooth the problems out with their talent and chemistry.

They really captured the urge to build things in tech, and the problems that come with it. HACF, Silicon Valley, and The Soul of a New Machine are a trifecta.

Bowed instruments are very cool to model because of the nonlinear slip of the bow against the string. A bit curious why bowing was not discussed or used in the example of a violin, just plucking. Do luthiers test violins more by plucking than bowing?

Thank you for wording it better than I could. On the computational side, the integral as a function that measures how the another function accumulates intuitively lets you measure the area under the curve between two points only requires two variables.

Where it gets interesting is that if I were to naively construct this accumulator in a non-computational, infinitesimal manner, I would probably start using the idea that the infinitesimal accumulation is related to the instantaneous slope and using that somehow with the previous accumulation. But this is going the opposite direction of derivatives that the fundamental theorem uses.

I understand there are a few other types of integrals that slice things different ways that might be more intuitive. I vaguely recall some (non fundamental theorem) proof that made Riemann integrals click for a while.

I've studied the proofs before but there's still something mystical and unintuitive for me about the area under an entire curve being related to the derivative at only two points, especially for wobbly non monotonic functions.

I feel similar about the trace of a matrix being equal to the sum of eigenvalues.

Probably this means I should sit with it more until it is obvious, but I also kind of like this feeling.

Jeff is a cool guy to make a reasonable and level headed post like this. I would be struggling to not be more resentful.

He was also nice and approachable in person, which to be honest I didn't expect from the utilitarian no chit chat rules on SO.

Claude Opus 4.7 3 months ago

Does it run for you? I can select it this way but it says 'There's an issue with the selected model (claude-opus-4-7). It may not exist or you may not have access to it. Run /model to pick a different model.'

Claude Opus 4.7 3 months ago

These stuck out as promising things to try. It looks like xhigh on 4.7 scores significantly higher on the internal coding benchmark (71% vs 54%, though unclear what that is exactly)

More effort control: Opus 4.7 introduces a new xhigh (“extra high”) effort level between high and max, giving users finer control over the tradeoff between reasoning and latency on hard problems. In Claude Code, we’ve raised the default effort level to xhigh for all plans. When testing Opus 4.7 for coding and agentic use cases, we recommend starting with high or xhigh effort.

The new /ultrareview command looks like something I've been trying to invoke myself with looping, happy that it's free to test out.

The new /ultrareview slash command produces a dedicated review session that reads through changes and flags bugs and design issues that a careful reviewer would catch. We’re giving Pro and Max Claude Code users three free ultrareviews to try it out.

I've been feeling the squeeze too. I've tried switching between different models as a test, I can at least say it feels like the limits are about half of what they used to be a few months ago. I'd be totally willing to concede that this is just my perception if Anthropic would only release some tools for measuring your usage.

In theory the /stats command tells you how many tokens you've used, which you could use to compute how much you are getting for your subscription, but in practice it doesn't contain any useful info, it may be counting what is printed to the terminal or something - my stats suggest my claude code usage is a tiny amount of tokens, but they must be an extremely underestimated token count, or they are charging much more for the subscription than the API per token (which is not supposed to be the case).

Last week's free extra usage quota shed some light on this. It seems like the reported tokens are probably are between 1/30th to 1/100th of the actual tokens billed, from looking at how they billed (/stats went up 10k tokens and I was billed $7.10). With the API it should be $25 for a million tokens.

This is similar to my experience. I find many people on the forums that can't resolve the crashes in Red Dead Redemption 2, so I suspect a lot of it depends on the specific games you've picked.

It is clearly getting better globally, so I expect in 3 years or so things will be ready for me to try again.

This has clearly triggered a lot of people, with the full spectrum of arguments for and against having kids below. I'm bookmarking this as an example of driving engagement by taking a fairly benign topic ('helping others was rewarding!') with an extreme view on a topic that everyone has a stake in ('an entire generation lost the plot on life').

How amazing and ironic or a reminder it is that the comments below that seem the most reasonable and avoid generalizing an entire group of people or way of life are the ones that are the least likely to drive more comments because they are perfectly reasonable.

I wrote a bit about how the kind of tech culture in HACF feels more relevant in light of the LLMs even 2 years back before I heard of mourning a craft. Here's an excerpt:

One thing I liked about HACF is despite using a decent amount of technobable, it plausibly captures the approach and spirit of hacking and coding, like reverse engineering a memoryless chip by rigging up a hex LED system to read out the values for each of the 65536 inputs to a ROM.

The expertise of the coders are demonstrated mainly by others admiring the structural complexity of their code as objects of beauty. This is something that feels extra nostalgic now.

https://michaelchinen.com/2023/12/31/halt-and-catch-fire-pos...