This got a downvote and I understand why: because I didn't describe the test, which is to ask it "Please recite Jabberwocky".
This is actually difficult because there are so many invented words in the poem which have extremely low frequencies in the training data. So a model that can do it properly is likely to be very good in other ways. Qwen-3.6-27B can do this until it gets overly quantized.