HN user

xavriley

96 karma
Posts4
Comments38
View on HN
MAI-Thinking-1 2 months ago

“ We trained it from the ground up on enterprise grade, clean and commercially licensed data, without distillation from third-party models.”

I went down a similar rabbit hole at the start of my PhD and I wish I’d written more of it up. One of my theories is that they combined effects quite often. For example, “harder better faster stronger” seems more likely to be a talk box recorded for a single note, then looped, then run through an AutoTune rack unit with MIDI inputs to repitch it. I mention this a little bit in a talk I have at ADC 2022 https://youtu.be/uX-FVtQT0PQ?feature=shared

This is cool - there’s some similar work here https://arxiv.org/pdf/2402.01571 which uses spiking neural networks (essentially Dirac pulses). I think the next step for this would be to learn a tonal embedding of the source alongside the event embedding so that you don’t have to rely on physically modelled priors. There’s some interesting work on guitar amp tone modelling that’s doing this already https://zenodo.org/records/14877373

This is a hypothesis put forward by Gerald Langner in the last chapter of “The Neural Code of Pitch and Harmony” 2015. I personally think he was on to something but sadly he died in 2016 before he could promote the work

I’m the author of the high resolution guitar model posted in a comment above. I have a drum transcription model that I’m getting ready for release soon which should be state of the art for this. I’ll try to update this thread when I’m done

Hydrofoil from Sorrento to Capri in choppy seas, on our honeymoon. Was the stuff of nightmares. My wife said we’d have to live on Capri because she was never setting foot on a boat again

It sounds like you’ve found it already but th original pYin implementation is in the VAMP plugin. Simon Dixon is my PhD supervisor but he’s quite busy. Feel free to email me questions in my the meantime. j.x.riley@ the same university as Simon. There’s also a Python implementation in the librosa library which might have a better license for your purposes.

High latency - agreed but it depends on whether a GPU is available or not. If it is then theoretically CREPE could be real-time. The error rates for pitch recognition are still quite good though for the full CREPE model. I’m interested to see the data on this claim.

Simple techniques like autocorrelation can still recover a missing fundamental. To answer the GP post, using neural networks for this task is overkill for simple, clean signals but it can be desirable if you need a) extremely high accuracy or b) robust results when there are signal degradations like background noise

how does authorization between the host and the forked work?

On fly.io you get a private network between machines so comms are already secure. For machines outside of fly.io it’s technically possible to connect them using something like Tailscale, but that isn’t the happy path.

how do I make sure that the unit of work has the right IAM

As shown in the demo, you can customise what gets loaded on boot - I can imagine that you’d use specific creds for services as part of that boot process based on the node’s role.

It’s not been mentioned yet, but if you play music then going to jam sessions is a great way to meet people. You’re all on a journey together toward improving as musicians which helps things to gel. As a jazz musician I can find a jam session in pretty much any city I go to. If you don’t play you can always go just to listen, watch and be inspired

There’s a model for music transcription (audio to midi) called MT3 which takes an end-to-end transformer approach and claims SOTA on some datasets. However, from my own research and comparing with other models it seems that MT3 is very prone to overfitting and the real world results are not as impressive. A similar story seems to be playing out in the comments here

Someone in my PhD lab looked at this and commented that they weren’t that impressed. The authors didn’t account for the fact that ballads and uptempo numbers have vastly different swing ratios (in both cases practically straight) which skews the results. I think rhythmic phenomena and perception are worthy of study but this isn’t a great example imo

For anyone wondering, a lot of work on Sonic Pi recently has gone into integrating an Elixir backend to handle distributed jamming. It has Ableton Link support so it can easily be synced with a DAW and other apps. It can also control external devices via MIDI and OSC protocols more reliably as a result.

There are several methods for pitch tracking of audio inputs that reach up to 99.9% accuracy on certain datasets. CREPE and pYin are just about state of the art. Granted they don’t have a batteries included way of doing real-time voice-to-midi but it’s more an issue of packaging. Electroglottography is cool but isn’t necessary for this task.

Thank you for posting this. I feel like it gives a better context. My interpretation is that the focus is on monks (as opposed to lay people) in that they might be preaching about their virtues while continuing to have vices and earthly desires. This makes sense in a religion in that you’d want your monks to strive for something like “ideal” behaviour (even if it’s not reachable) otherwise what is the point of a monk?

Stating it as a list of facts (as in the OP) seems akin to literal interpretations of the bible where the world was created 4000 years ago etc. A more reasonable view would be to take the message in context. I don’t think Buddha had anything against laypeople playing chess

Perfect pitch for a jazz improviser is a bit like a super power. Audiation - the audio equivalent of visualisation - is a key part of jazz improv. Having absolute pitch means you don’t have work quite as hard to translate those ideas onto your instrument (it’s still work though).

I’ve studied Pat’s playing a lot and transcribed several of his solos. The “muscle memory” seems to be a huge part of his playing as he has so much great dorian material under his fingers and he knows how to connect all those ideas really well, and how to repurpose them over various other chord types. WRT perfect pitch I think that must have helped him to build all that vocabulary in the first place. I’m just wondering whether he retained the ability after the op and how he related to it in his playing.

Coda - “absolute pitch” is a biological phenomenon. People recognise a C as easily as we recognise the colour red. “True pitch” is a variation where people play an instrument for so long that they can remember the sound of a pitch on that instrument. It’s a slower, less reliable process. Many people get the two confused

Asked not to is different from forbidden though, no? I’m not a legal expert but I doubt this clause is one they wanted to put in and is probably driven by some legal counsel with an eye on international law.

Harmonic intervals are interesting but they are also misunderstood. Humans are able to distinguish out of tune notes down to a value of about 1%. To get more accurate tuning we tend to listen for beating (a kind of amplitude modulation) instead. This means we can tolerate tuning systems other than just intonation. Another thing to consider - the missing fundamental phenomenon suggests that the ear/brain is actually doing something like autocorrelation. This makes more sense than the idea that we have a template for the harmonic series wired in our brains. Finally auto correlation works for chords too, not just intervals. Every chord has a fundamental period of repetition - shorter periods are widely ranked as more consonant. There are lots of grand music theories that fixate on the harmonic series. The maths is fun, but it can get in the way of more effective alternatives for organising sounds.

I found my cooking improved massively when I focussed on more scientific methods. For me personally The Food Lab and Salt, Fat, Acid, Heat were both great. Reducing the variables to a couple of axes made it much easier to decide what to add or change to get a result. For example if a sauce tastes too greasy, add the appropriate acid to lighten things up

This is just my two cents, but if you break it down into two components - frequency response and latency - then it's easier to imagine. With the latest modelling techniques (neural nets etc) it's possible for solid state amps to get very, very close to the frequency response of pretty much any tube amplifier in my experience. Expensive modelling amps and plugins are really nice. The second issue though is latency - when players say it "feels" different playing through a tube amp, my hypothesis is that they are noticing the latency (or lack thereof). Any kind of digital processing is going to impose some kind of delay compared with a tube amp which is a direct electrical signal. I'm aware that the delay will be on the order of milliseconds (sound travels roughly 1 foot per ms through air) and the effect could be replicated by moving the amp further away etc. It would also help to explain the change in sympathetic vibrations within the instrument. Shorter delays and high SPLs could have a non-trivial effect on the resonance of the guitar body for example.