I suppose all copies of all pages are theft, if you choose to look at it that way. Me saving a .html file locally, Google caching their search results, archive.org saving copies of pages.
The internet is built around this, and we've all collectively decided that it's OK.
But to your point, their robots.txt does allow this kind of access, so this argument is moot - they explicitly allow bots like archive.org's to crawl & index these pages' content.