Show HN: Bertina – 3M Bert (ITA)
https://huggingface.co/mascIT/bertina-3MEnough with LLM news. it may sound boring but I've pretrained a tiny version of BERT, approx 3M params, from scratch on a bunch of italian data (wiki mostly).
According to benchmarks it performs just 5-8% worse than a fully fledged 100M vanilla ita BERT.
I've used a 15k vocab. size and tweaked the hidden size. Nothing too fancy, it's been a good excercise though.