one of a kind single-transformer block layer, high throughput. The new generation of transformer-based lightweight models for common NLP tasks?
HN user
fblgit
16 karma
juanako.ai
Posts1
Comments4
Single-layer transformer model "HarEmb" showcasing PII SOTA performance 3 months ago
Mistral "Mixtral" 8x7B 32k model [magnet] 3 years ago
doesn't require much data, in a 7B can take a couple hours ~
Mistral "Mixtral" 8x7B 32k model [magnet] 3 years ago
Correct. UNA can align the MoE at multiple layers, experts, nearly any part of the neural network I would say. Xaberius 34B v1 "BETA".. is the king, and its just that.. the beta. I'll be focusing on the Mixtral, its a christmas gift.. modular in that way, thanks for the lab @mistral!
Mistral "Mixtral" 8x7B 32k model [magnet] 3 years ago
UNA: Uniform Neural Alignment. Haven't u noticed yet? Each model that I uniform, behaves like a pre-trained.. and you likely can fine-tune it again without damaging it.
If you chatted with them, you know .. that strange sensation, you know what is it.. Intelligence. Xaberius-34B is the highest performer of the board, and is NOT contaminated.