if the ai is the product, and the product isnt trustable, isnt that a product issue??
HN user
senseiV
Does a TPU have XLA-graph for GPUs Cuda-graphs? Not sure on TPU theory
Ive noticed the same on extremely small models aswell, magnitude is a positional encoding or a couple tokens, so its easy to grok?
well world model in the context of the tulip fields, so models could be finetuned+sheared to drop size and remain effective
claude.ai
Theres a startup doing that named galileo_ai
Part of an FRC team building a Vision system from scratch, quite fun and nearly complete, just need to recalibrate some angle formulas
I just saw a markdown mode show up today, but only partially, like bold and italics in markdown
not sure if this is just chatgpt, but analogous evolution is interesting to see
yes the size is different, but training a diffusion model and a language model are really different, like how RL models can be small but take a long time to train aswell
Looking into the nordic pile maybe? There are some datasets
NLP is not the industry, and a lot of research still goes into other things, like RL
I've worked with several transformers competitors, and it def wont stay centralized on them
GPT 2 and 3 used the p50K right? Then GPT-4 used cl100K
simulating entire AI-based societies.
Didnt they already have scaled down simulations of this?
replit/codesandbox maybe?
? its better than GPT 2 for sure...
V5 7b is out, close to hyena, gets 1400 t/s on a 3090, while an h100 llama 7b 8bit is 1200 t/s
They Do, the latest rwkv v5, matches mamba at 3b scale, and from the benchmarks I see, its similar to hyena
Just make a throwaway google?
The so called "AI People" built the entire architecture, something people didn't think was possible at the scale and quality a year ago, and the matter of "artists should get whatever they want" because it trained on their works isn't the point. Diffusion Models don't rip parts of pictures together, they happen to be trained to make art out of noise, finding patterns in art. same things happening with LLM's in court with the LLama model and book authors claiming it only makes books.
I still remember that piece of art that was submitted to an art contest and won, only to be announced an SD prompt later
Bruh its simple physics, does one end or the other get lighter, by all measures we care about, not really, the mass of a proton or electron is beyond any consumer hardware measurement. I doubt it would matter beyond extreme scenarios or controlled experiment
He's talking about llama 2 superhots, and mistral derivatives that can be uncensored
No, the orca 2 paper mentions more of a counter point towards NSFW and stuff, like if you gave it a NSFW prompt, it would retort back against it, which is arguably a good thing, but really lost in RLHF
Where can you find those? I'm in the same situation as him, I've never heard of a 3d dataset better than objaverse XL.
Got a public dataset?
Oh no do not use that. That was servo based, AI drones, which I think is the real "safety issue"
Nice, whats the text to video model? Also, you could try to go for a 1b llm for the browser, would fit.
ah yes RWKV, always great to mention, crazy about how no one talks about it, it literally the most powerful multilang model at 1b and 3b scales, probs going for 14b and 7b too
Nope Actually, networks like alpha zero learned with nothing. If only we could get that to training data
Do you inference NanoGPT with say a Flask app? I made a WebUI for tiny Llama, but it doesn't work for NanoChatGPt
Nice!
I remember seeing this for the first time on HN with urlpages, inspired me to build my own version of these