HN user

deepsquirrelnet

1,797 karma
Posts5
Comments503
View on HN

I think this is only accurate when no external ideas are used, but I'd like to suggest that nearly all new discovery is built on a combination of old ideas and LLMs are really good at the latter.

If you bring something new to the table, then in my experience, AIs are really good at helping you ground it old ideas. If you want to set it and forget it, then you will get the mean. If you want to do something new, in my experience, they are enablers and not blockers.

Can anybody find trustworthy stats that these actually reduce crime? All I see are occasional anecdotes about how they were used to find one person one time.

Skeptical me seriously doubts this is an effective solution for crime. But maybe that's because this country has a history of being willing to do a million expensive and privacy violating things, and only if it's a punitive measure.

Because I love swiping, but all my problems with it come from the fact that the QWERTY layout is far from ideal for it. I am 100% willing to learn a new layout if anyone will develop an optimal one for English so that swiping has a 99.9% accuracy rate instead of what currently feels more like 90% or 95%.

90-95% is a very good estimate! That's about what we measure on our test set. I have good news for you, and we will have a blog post about it soon. Because of how our models are built, we are able to optimize for detection accuracy directly by constructing synthetic swipes on each layout for ~50k words, and then testing them through the model. We tested around 800,000 layouts this way.

The biggest issue with QWERTY is that there are far too many words that swipe colinear or obtuse angle letter trigrams. These are both hard to detect and frustrating for swipe users, because you can't clearly indicate the letters you're gesturing. Neural swipe models (at least ours) look for indicators in the gesture pattern that suggests a user was targeting a specific letter, rather than trying to match a gesture shape like algorithmic detection does.

The shape of the keyboard can significantly improve the way the gestures are formed so that there is better indication of letters. The model can still respond to dwell times because unlike shape matching it uses the temporal information. But dwell interrupts flow, and in my opinion should be minimized in swipe layouts.

I'd been working with language models for several years before LLMs were a solution to this kind of problem. These are some ideas "off the top of my head" about how you can do classification in various ways. There's really a lot of ways to tackle it now, and a lot of trade-offs you can learn by experimenting with them.

There's even more options still, especially if you go further back toward more traditional methods. Static word vectors like GloVe or fasttext (optionally more modern equivalents like WordLlama or Model2Vec). Then there's sklearn-style stuff too. Those can be really small/fast but have more accuracy-level tradeoffs.

If you want to go deeper on language models, try these project ideas:

- Zero-shot encoders like tasksource or GliNER

- Natural language inference: https://huggingface.co/blog/dleemiller/nli-xenc-ways-to-use

- GRPO training

- GEPA prompt tuning Qwen 0.6B (or GEPA, then GRPO)

- Use an embedding model and train a classifier (MLP, logistic, svm)

- Use a larger LLM to generate a synthetic dataset (beware of lack of diversity, mine "seed text" from real sources first)

- Synthetically generate "hard examples" where more than one category may be valid and DPO tune your preferred responses

Unions exist to benefit the median and bring up the floor, but it stifles competition among those who really do desire to be at the top. And in doing so while it brings up the floor, it also brings down the ceiling because people who would normally be motivated enough to move up would not have much incentive to do so anymore.

I think people tend to fixate on the worker-to-worker differences inside of unions. Yes, that is the most visible part of a union when in place, and at least in the US has valid arguments about meritocracy.

What is missed when limiting the scope to just that is the population-level abuses of workers that no amount of meritocracy will fix. When corporations engage in collusion against workers (now common and nearly unpunished in the US) the top-level wages are suppressed industry wide.

The whole pay band alignment that comes out of that undermines the meritocracy argument, and doesn't even begin to address the wage-fixing that has gone almost unchecked in tech for decades[1,2]. As a merited employee, you might have more options to where you can go, but it won't protect you from predatory hiring/layoff cycles and it certainly won't guarantee that you'll receive a truly competitive wage.

On paper, meritocracy sounds great. I have worked many places in tech and never once observed it, personally. Best case, if you have warmed a seat for enough years, then you advance that way. Worst, your employer knows they can just take advantage of you because you're willing to work without a dangling carrot.

As before, either the government frees itself from corruption and enacts justice or unions will come back. That is point we are at.

[1] https://www.npr.org/sections/alltechconsidered/2015/01/16/37...

[2] https://conversableeconomist.com/2025/10/31/the-silicon-vall...

Not that I’d want to work there given what they do, but every time I’ve been contacted by a recruiter there, it seems like it’s within a month of a mass layoff they’ve had… which is maybe just because they seem to have mass layoffs every quarter now.

They also seem to have adopted a no-remote hire policy and are in an extreme high CoL location. It’s a truly awful mix for trying to attract outside talent. I don’t know why they even bother.

Any discussion related to this topic always seems to assume everyone uses code the same way and for the same function, and then forces the rest of the world through that lens.

So here we walk around the circle one more time again, voicing our anxieties, talking past each other, waiting for the next opportunity for commentary to come in half an hour.

Mistral Medium 3.5 3 months ago

I would love to be able to run frontier locally, but I think the larger importance of open weight models is price accountability.

In the US with our broken system of capitalism, it’s the only way we can tether these companies to reality. Left to their own devices, I’m not convinced they would actually compete with each other on price.

Buy nobody like to talk about how “moat” building is fundamentally anti-competitive, even in name.

Funny that self proclaimed capitalists hate the system in practice. Commodity pricing is what truly terrifies them.

This is my own take, directly related to this that I posted a little while back. The one thing that I think the article missed is the geopolitical angle they’re also working:

* We need to completely deregulate these US companies so China doesn't win and take us over

* We need to heavily regulate anybody who is not following the rules that make us the de-facto winner

* This is so powerful it will take all the jobs (and therefore if you lead a company that isn't using AI, you will soon be obsolete)

* If you don't use AI, you will not be able to function in a future job

* We need to lineup an excuse to call our friends in government and turn off the open source spigot when the time is right

They have chosen fear as a motivator, and it is clearly working very well. It's easier to use fear now, while it's new and then flip the narrative once people are more familiar with it than to go the other direction. Companies are not just telling a story to hype their product, but why they alone are the ones that should be entrusted to build it.

In a provocative GitHub post, machine-learning engineer Han-Chung Lee argued that even rosy internal numbers that do show AI-assisted productivity gains are suspect, as they’re produced to hit adoption targets no one can effectively audit.

Isn't this fundamentally what MBAs do with their time? Keep going with this analysis, because it goes much deeper... In my experience, BI is often a house of cards. A lot of times it's just narrative crafting, just like we're all encouraged to do when we write our resumes.

Can you embellish a story? Can you invent a convincing political narrative? As far as I can tell, that's the fundamental unit of US corporation.

Traditional call noise canceling relies on those small onboard neural networks and can have difficulty isolating your voice in very noisy environments, which results in ambient noise leaking through or voices getting highly compressed, making it difficult to hear. Anker says the larger neural network available on the Thus chip, plus eight MEMS (micro-electromechanical systems) microphones and two bone conduction sensors to focus in on your voice, in its yet-to-be-announced earbuds will have significantly cleaner call audio, regardless of the environment.

Anyone who likes good noise cancellation, which is a lot of people.

Back in the day we just called it ML. But now you have to stop for a minute to read and determine what they’re talking about, because “AI” is primarily a marketing term.

I tried it on openrouter and set max tokens to 8192, and every response is truncated, even in non-thinking mode. Maybe there's an issue with the deployment, but in your link also shows it generates tons of output tokens.

Claude Opus 4.7 3 months ago

My tinfoil hat theory, which may not be that crazy, is that providers are sandbagging their models in the days leading up to a new release, so that the next model "feels" like a bigger improvement than it is.

An important aspect of AI is that it needs to be seen as moving forward all the time. Plateaus are the death of the hype cycle, and would tether people's expectations closer to reality.

Cuyahoga Valley: There is nothing wrong with Cuyahoga Valley. Statistically, you’re from Ohio, so why not?

In college, I took an interim elective course on geology of the national parks. On the first day of class, the professor asked an icebreaker for students to say which national park they lived closest to. I said Ohio - Cuyahoga Valley.

Well some snot nosed boy scout confidently piped up that there were mostly certainly no national parks in Ohio, and the professor agreed. This is a deep personal grudge that I still hold to this day.

CnakeCharmer - https://github.com/dleemiller/CnakeCharmer

https://huggingface.co/datasets/CnakeCharmer/CnakeCharmer

This project started from a belief that llms should be better at doing python to cython code translations than they are. So we started setting a large set of parallel implementations.

Then I realized that Claude code was much better at working on the data using tools (mcp) to check and iterate. The scope transformed into an platform for creating the SFT agentic trace dataset using sandboxed tools for compilation, testing, linting, address sanitizing and benchmarking.

We still need to bulk up the GRPO dataset with a large number of good unmatched python examples. But early results using SFT only on gpt-oss 20b are quite good.

There are so many reason if you look at how it's being sold.

* We need to completely deregulate these US companies so China doesn't win and take us over

* We need to heavily regulate anybody who is not following the rules that make us the de-facto winner

* This is so powerful it will take all the jobs (and therefore if you lead a company that isn't using AI, you will soon be obsolete)

* If you don't use AI, you will not be able to function in a future job

* We need to lineup an excuse to call our friends in government and turn off the open source spigot when the time is right

They have chosen fear as a motivator, and it is clearly working very well. It's easier to use fear now, while it's new and then flip the narrative once people are more familiar with it than to go the other direction. Companies are not just telling a story to hype their product, but why they alone are the ones that should be entrusted to build it.