feat(crawler): add settings to enable OCR for images and scanned PDFs - #3520
Merged
Merged
Conversation
Tika's Tesseract OCR parser was shipped but excluded in tika.xml, and scanned PDFs never reached it because PDFs are extracted with PDFBox. - tika.xml no longer excludes TesseractOCRParser. - New fess_config.properties keys: crawler.document.ocr.enabled (default false), crawler.document.ocr.language (default eng) and crawler.document.ocr.timeout (default 120 seconds). - FessTesseractOCRConfig, registered as the tesseractOCRConfig component, applies these settings as the default Tesseract config of every TikaExtractor. When OCR is disabled it sets skipOcr, so a host that has the tesseract command installed keeps the previous behaviour. When OCR is enabled it also preserves interword spacing, so that Tesseract does not put a space between CJK characters, which would keep Japanese words from matching. An invalid language disables OCR with a warning. - FessPdfExtractor sets tikaExtractor as the fallback for PDFs from which PDFBox extracts no text, when OCR is enabled. It first checks that Tika can really run Tesseract; if the tesseract command is missing or a customized tika.xml still excludes the parser, it logs a warning instead. Requires the fess-crawler change that adds the default Tesseract config and the PDF fallback extractor.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Makes Tesseract OCR easy to turn on for images and scanned PDFs.
Before this change, Tika's
TesseractOCRParserwas shipped but excluded intika.xml. Scanned PDFs never reached it anyway, because PDFs are extracted with PDFBox.tika.xml: no longer excludesTesseractOCRParser.New settings in
fess_config.properties:crawler.document.ocr.enabledfalsecrawler.document.ocr.languageengjpn+engcrawler.document.ocr.timeout120FessTesseractOCRConfig: registered infess.xmlas thetesseractOCRConfigcomponent next totikaConfig. With feat(extractor): add a default OCR config and a fallback for text-less PDFs fess-crawler#217, everyTikaExtractorin the crawler container uses it as its default, including subclasses from plugins.skipOcr, so a host that hastesseractinstalled behaves as before.preserve_interword_spaces. Without that, Tesseract puts a space between CJK characters ("請 求 書"), and Japanese words do not match.config.tika.tesseract.configparameter still takes precedence.FessPdfExtractor(viacrawler/extractor+pdfExtractor.xml): when OCR is enabled, it setstikaExtractoras the fallback for PDFs from which PDFBox extracts no text. Tika then renders the pages and OCRs them.tesseractcommand is missing, or a customizedtika.xmlstill excludes the parser, it logsOCR is enabled, but Tesseract OCR is not available. ...instead of silently re-parsing.OCR is enabled: language=..., timeout=...s.Depends on codelibs/fess-crawler#217, so CI needs that change in the fess-crawler snapshot first. Docs: codelibs/fess-docs#562. Docker recipe: codelibs/docker-fess#89.
Upgrade note: RPM/DEB installs keep a modified
/etc/fess/tika.xml, and a customizedapp/WEB-INF/conf/tika.xmlalso stays. Those files still exclude the parser and must be edited to use OCR. The new warning points at this.Tests
FessTesseractOCRConfigTest: disabled, enabled, blank timeout, invalid language, andinit()readingFessConfig.FessPdfExtractorTest: fallback set only when OCR is enabled and available, and the availability check withskipOcr.jpn+eng):tika.xmlthat still excludes the parser: the warning is logged once, there is no PDF fallback and no OCR, and indexing is otherwise unaffected.