SmolDocling: An ultra-compact VLM for end-to-end multi-modal document conversion 1 year ago
OCR is not the task being solved here, though. This is supposed to help you when dealing with complex layouts where text is not just read left-to-right, top-to-bottom.
But I agree that accurate OCR is kind of a prerequisite for adaptation.