docs(15.9): add an OCR configuration page - #562
Merged
Merged
Conversation
Add config/crawler-ocr.rst in every language and list it after the thumbnail page in the crawler section. It covers what OCR applies to, how scanned PDFs are handled, installing Tesseract, the crawler.document.ocr.* settings, the Docker recipe in docker-fess, the per-crawl config.tika.tesseract.config parameter, operational notes, and the tika.xml exclusion that a customized file from 15.8 or earlier still carries.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds
config/crawler-ocr.rstto the 15.9 tree in all seven languages and lists it after the thumbnail page in the crawler section ofconfig/index.rst. The page covers:tesseractcommand is installed and OCR is enabled. It applies to images, to images embedded in documents that Tika handles, and to PDFs that have no text layer.crawler.document.ocr.enabled,crawler.document.ocr.languageandcrawler.document.ocr.timeoutsettings, also settable as-Dfess.config.*options.compose/tesseractrecipe in docker-fess.config.tika.tesseract.configparameter, which takes a classpath resource name.tika.xmlfrom 15.8 or earlier still excludesTesseractOCRParser, so OCR stays off until that line is removed.OCR is enabledandTesseract OCR is not availablelines infess-crawler.log.The documented behaviour comes from the companion fess and fess-crawler changes, and was checked on a live Fess 15.9.0-SNAPSHOT with Tesseract.
config/properties.rstis not regenerated here. Regenerating it from the current fess tree also picks up unrelated keys added since its last run, so it is better done as a separate sweep after the fess change is merged.Checks
tools/check_headings.pyfinds no heading problems in the new pages.index.rstwere built with Sphinx in a scratch tree for every language, with no warnings from the new page.