Perhaps it's not a moat.
However, if the advantage is due to things like inference infrastructure to support a massive model, that isn't easy to duplicate.
I would also say that the quality of these smaller models are good, but we also may not be measuring them correctly. Recent papers suggest that these smaller LMs dont fully capture ChatGPT quality in ways that may not have appeared with crowd worker ratings [1]. It's easy to have your inputs be inside a happy distribution for a paper but fail in the real world in ways that GPT-4 doesnt.
Lmsys would love to compare with bigger models but have limited resources. Contributions are welcome [2]
[1] https://arxiv.org/abs/2305.15717
[2] https://lmsys.org/blog/2023-05-25-leaderboard/#next-steps