I'm not sure what kind of point you're trying to make. There are projects to train competent modern LLMs in which the entire pipeline (data, training process, final weights) is all completely transparent, shared, and reproducible by anyone with the compute to try it out.
Or is your definition of "open source" mean that a small indie dev should be able to reproduce the entire pipeline? Because that would disqualify more than just LLMs, but also hardware platforms like Arduino where you need to pay for manufacturing to get the underlying stuff built... is Arduino "open source?"