Our next-generation model: Gemini 1.5 2 years ago
Regarding how they’re getting to 10M context, I think it’s possible they are using the new SAMBA architecture.
Here’s the paper: https://arxiv.org/abs/2312.00752
And here’s a great podcast episode on it: https://www.cognitiverevolution.ai/emergency-pod-mamba-memor...