Warning: the photos are nightmare fuel and not safe for bedtime
HN user
bkitano19
notable omission of deepgram models in comparisons?
+1 to running. If you run consistently, you'll learn to believe in your body as something that naturally improves if you train it well, and that belief will cross over to your mind and heart.
You can use voice prompting; it's supported on ElevenLabs and Hume.
Awesome post!
you might like https://en.wikipedia.org/wiki/Noether%27s_theorem
Related work:
Interpreting Modular Addition in MLPs https://www.lesswrong.com/posts/cbDEjnRheYn38Dpc5/interpreti...
Paper Replication Walkthrough: Reverse-Engineering Modular Addition https://www.neelnanda.io/mechanistic-interpretability/modula...
hume.ai specializes in expressive prosody for TTS (disclaimer - I work here)
Time to first token is as important to know for many use cases, rarely are people reporting it
this is nuts
+1, had the fortune to work with him at a previous startup and meetup in person. Our convo very much broadened my perspective on engineering as a career and a craft, always excited to see what he's working on. Good luck Simon!
https://transformer-circuits.pub/2022/in-context-learning-an...
there is a lot of evidence to suggest that they are performing induction
Had this exact problem (Heroku Postgres to RDS) at my old co. Data migration went as bad as it possibly could (dropped indices, foreign keys, everything but the data itself). This would have saved us months of pain.
https://hootdoogs.com/ - "find food between you guys"
High integrity maintainer, big respect.
Karpathy covers this in Makemore, but the tl;dr is that if you don’t normalize the batch (essentially center and scale your activations down to be normally distributed), then at gradient/backprop time, you may get values that are significantly smaller or greater than 1. This is a problem, because as you stack layers in sequence (passing outputs to inputs), the gradient compounds (because of the Chain Rule), and so what may have been a well behaved gradient at the end layers has either vanished (the upstream gradients were 0<x<1 at each layer) or exploded (the gradients were x>>1 upstream). Batch normalization helps control the vanishing/exploding gradient problem in deep neural nets by normalizing the values passed between layers.
Wow, great catch. I will update this in the morning!
edit: bearblog getting ddos'd, here's the repo https://github.com/bkitano/llama-from-scratch
working on it hehe
Hey! This is what I've been working on, would love to chat, feel free to email
I write about tech, personal growth and some random stuff.
I really love bearblog, it was exactly what I was looking for to get started, shout-out Charlie Meyer for getting me started!
I'm in a ton of LLM Discords, and I follow the announcements and events for updates!
I think the point of differentiating between the blinded and unblinded groups was to show that placebo alone does not account for the effects of microdosing.