We take a lot of shortcuts when speaking, it's actually much harder to transcribe phonemes than to transcribe words, even when aware of the language being spoken. Some models have been trained for the task (e.g. look at https://huggingface.co/spaces/KoelLabs/IPA-Transcription-EN ), but the error rate is really high.
HN user
shenberg
You always either go left or down so total 40 steps, choose 20 to be down (or 20 to be right)
Some example >1B companies off the top of my head: DataDog, Sentry, Snowflake, Okta, MongoDB
moondream is a beast
Seems 100% AI generated and automated, the judge also seems suspect - in the first one it's actually GPT-5.5 pro which has the correct email RE: the deepseek one will match a@b.com1 as "a@b.com" while 5.5 will correctly require a word boundary at the end of the email. I quit after this. No test-cases = useless judge.
11% MFU does not mean 89% of GPUs are idle, it means that they're using the GPUs ineffectively.
Using existing enterprise apps probably - this solution is scalable for the vendor and it's easier to sell using existing software as-is than to start out by writing new custom tools.
Mid-way I realized this was AI writing (took me a while), then I read a quote in the text about a comment that "The tragedy isn’t that they cheated; it’s that the system was designed to let them thrive for a decade before anyone bothered to look at the data." I didn't find this comment in EJMR, or anywhere on the internet except the OP post, for that matter.
Moshi was an amazing tech demo, building the entire stack from scratch in 6 months with a small team was an amazing show of skill: 7B text LLM data + training, emotive TTS for synth data generation (again model + data collection), synth data pipeline, novel speech codec, rust inference stack for low latency, audio LLM architecture incl. text "thoughts" stream which was novel.
But, this piece is a fluff piece: "underfunded" means a total of around $400 million ($330 million in the initial round, $70 million for Gradium). Compare to Elevenlabs who used a $2 million pre-seed for creating their initial product.
A bunch of other stuff there is disingenuous, like comparing their 7B model to Llama-3 405B (hint: the 7B model is a _lot_ dumber). There's also the outright lie: team of 4 made Moshi, which is corrected _in the same piece_ to 8 if you read enough.
Location: Paris, France (US citizen, EU resident)
Remote: Yes
Willing to relocate: Not until 2027
Technologies:
ML / DS: PyTorch, CUDA, distributed training & inference, performance profiling/optimization (audio & speech focus, some LLM inference acceleration)
Systems: C/C++/asm, low-level performance work, reliability/scale engineering
Backend / Infra: Python/Java/C# prod services, ETL / data pipelines, k8s (incl. operator work)
Roles: Tech Lead, Research Engineer (training and/or inference of large models)
Résumé/CV: https://www.linkedin.com/in/roeeshenberg/
Email: roee.shenberg@upai.dev
I’m a hands-on engineer who’s spent the last 6 years doing freelance ML + data science, primarily in audio/speech, and before that 10+ years in startups building and scaling production systems.
I’m looking for where research meets real systems: training and/or inference for large models, especially roles that value end-to-end ownership. Open to freelance engagements or full-time roles.
There are two ingredients that don't fit in the "attention-is-kernel-smoothing" as far as I can tell: positional encoding and causal masking (another way to say positional encoding, I guess)
Also, Simplical attention is pretty much what the OP was going for, but the hardware lottery is such that it's gonna be pretty difficult to get competitive in terms of engineering, not that people aren't trying (e.g. https://arxiv.org/pdf/2507.02754)
I don't understand how using group-theory language to describe number-theoretic properties provides extra insight in this case (e.g. conjecture: all perfect numbers are even is more concise than the group-theoretic description given in the page). Can you expand on why you believe the tools of group theory have something to say about this? (e.g. for polynomial roots, the connection with symmetry groups comes from symmetries of factorized polynomials, while there's no obvious-to-me connection here as there is no unique-up-to-symmetry integer factorization)
ssh exe.dev works
The short and unsatisfying answer is that an LLM generation is a markov chain, except that instead of counting n-grams in order to generate the posterior distribution, the training process compresses the statistics into the LLM's weights.
There was an interesting paper a while back which investigated using unbounded n-gram models as a complement to LLMs: https://arxiv.org/pdf/2401.17377 (I found the implementation to be clever and I'm somewhat surprised it received so little follow-up work)
When countries like North Korea, which depends on cybercrime to fund itself, are signatories, you have to wonder whether this agreement means what its title says.
The reality of meetings in most places I've seen is that key stakeholders have already formed an opinion beforehand, the meeting is a place to disseminate decisions that have already been made and align the organization.
When I read "51% fewer false positives" followed immediately by "Median comments per pull request cut by half" it makes me wonder how many true positives they find. That's maybe unfair as my reference is automated tooling in the security world, where the true-positive/false-positive ratio is so bad that a 50% reduction in false positives is a drop in the bucket
The DeepSeek v3 model had a net training cost of >$5m for the final training run, the paper lists over 100 authors[1], meaning highly-paid engineers. This is also one of a sequence of models (v1, v2, math, coder) trained in order to build the institutional knowledge necessary to get to the frontier , and this ends up still far above the $10m mark. It's hardly a "trio of super-smart engineers".
That's really not true, e.g. the wikipedia page on population transfer in the Ottoman empire[1]. This dates way back to the Assyrian and Persian empries explicitly moving conquered peoples around in their empires in order to safeguard their rule. This book on population transfer in the Ottoman empire[2] explicitly states, with references, that the Ottomans habits were inherited from the steppe Turks, the Byzantines (=the Romans) and the Arabs.
[1] https://en.wikipedia.org/w/index.php?title=Population_transf... [2] https://websites.umich.edu/~gocek/Work/ja/Gocek.Muge.ja.popu...
Anecdotally, a pro-audio software company I worked with had to fire 1/3 of the company when their copy-protection was cracked and sales tanked immediately afterwards, and recovered once a new copy-protection scheme was developed and applied. And just to be clear, software licenses in direct-to-user sales are not that company's only revenue stream (they sell hardware and software to OEMs).
This is to say, the evidence in this natural experiment points towards piracy reducing sales by a lot.
Under the leaderboard tab, if the "Solution" column has an icon, it's clickable. 2nd place solution is by Jeremy Howard (of fast.ai fame), which I'd summarize as TrueSkill Through Time (Microsoft Research paper) + some overfitting on the public leaderboard (1st place was #26 in the public leaderboard).
The CLIP plot (Fig. 2) is damning, however some of the generative models show flat responses in Fig. 3 (e.g. Adobe GigaGAN, DALL-E-mini). While those are on the one hand technically linear relationships, but are also exactly what we'd want: image generation aesthetic score that doesn't care about concept frequency. Maybe the issue is with the contrastive training target used in CLIP?
I would have expected a sham-treatment arm to the experiment, because how do you differentiate between "an intervention 30 minutes beforehand caused improved learning" and "our specific intervention 30 minutes beforehand caused improved learning."
I suspect that weight initializations are geared towards inputs being normal random variables with mean 0 and variance 1. Deviating from that makes the learning process unhappy.
"since 2020, the US has printed nearly 80% of ALL US Dollars in circulation" - I've seen this notion repeated and I assume it's a reference to M1 as published by FRED: https://fred.stlouisfed.org/series/M1SL
The actual story, as far as I can tell, is that money that had previously been considered as M2 (=less liquid) is also counted as M1 due to rule changes regarding savings accounts.
To see this is the case, you can plot both together. If in fact, new money was printed, you would expect M2 to have the same jump as M1, as M2 is M1 + more stuff. However, you see a much smaller jump:
https://fred.stlouisfed.org/graph/fredgraph.png?g=1dwhY
The rule-change coincided with COVID relief measures which did include money-printing, but at a much smaller scale than implied.
If you're seriously suggesting that attacking unarmed civilians intentionally, killing parents in front of their children and then kidnapping the children, slaughtering defenseless party-goers, etc. is what I any resistance movement would do, that's ridiculous. If Hamas would only have attacked military targets, there would be no legitimacy to Israel's actions. However, what actually happened was that they attacked plenty of civilian targets, in a premeditated fashion, in areas that are recognized internationally to be part of Israel.
There's a confusion this article isn't helpful with: there are physical electrons, the actual physicalparticles. They move in the metal very slowly. But, their motion propagates very quickly, and turns out that the change in motion acts almost exactly like an electron itself, up to having a different mass. This is the "electron" quasi-particle, which is the abstraction that's breaking down. this only shows up about a screen or two deep into the article.
Prefixes are modifiers to specific instructions executed by the processor, e.g. to control the size of the operands or enable locking for concurrency.
In terms of practical algorithms, the Strassen algorithm (O(n^2.8)) is the only one that has runtime advantages for matrix sizes that aren't enormous, and even then, it's not always used because it has two non-trivial costs: reduced numerical stability and more memory space requirements for intermediate results.
Really reminds me of this short story by Ted Chiang: https://waldyrious.neocities.org/ted_chiang/liking-what-you-...