Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Datalab's Marker and OCR models for document parsing and text extraction are now available on Replicate. Marker converts PDF, DOCX, PPTX, and images to markdown or JSON, supports structured extraction via JSON Schema, and is benchmarked against other models on olmOCR-Bench. OCR detects text in 90 languages. Both models are priced per page.
From the source
Datalab’s state-of-the-art document parsing and text extraction models are now on Replicate.
replicate.com