We are already near the limits of what we can do
Hard disagree. If I had a million Claudes worth of compute I'd be livestreaming my entire reality feed to a local server 24/7 and having it organize my observations and thoughts, synthesize new ideas, implement prototypes and discard infeasible ones while I sleep. If you're in the business of knowledge creation, a million Claudes isn't enough. Text is an easy modality, I want foundation models that operate on text, images, audio, video, streaming point clouds, ...