Skip to content

docs(15.9): add an OCR configuration page - #562

Merged
marevol merged 1 commit into
mainfrom
feat/ocr-enable
Oct 1, 2026
Merged

marevol merged 1 commit into
mainfrom
feat/ocr-enable

Conversation

@marevol

@marevol marevol commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Summary

Adds config/crawler-ocr.rst to the 15.9 tree in all seven languages and lists it after the thumbnail page in the crawler section of config/index.rst. The page covers:

  • Scope: OCR runs when the tesseract command is installed and OCR is enabled. It applies to images, to images embedded in documents that Tika handles, and to PDFs that have no text layer.
  • Scanned PDFs: when PDFBox gets no text, Fess extracts the PDF again with Tika and OCRs it. PDFs that already have some text are not OCR'd.
  • Installation: installing Tesseract and its language data on Debian/Ubuntu and RHEL-family hosts.
  • Settings: the new crawler.document.ocr.enabled, crawler.document.ocr.language and crawler.document.ocr.timeout settings, also settable as -Dfess.config.* options.
  • Docker: the compose/tesseract recipe in docker-fess.
  • Per-crawl override: the config.tika.tesseract.config parameter, which takes a classpath resource name.
  • Operation: notes on CPU cost, the default excluded URLs of web configs, size limits and OCR accuracy.
  • Upgrade note: a customized tika.xml from 15.8 or earlier still excludes TesseractOCRParser, so OCR stays off until that line is removed.
  • Checking the setup: a short procedure, including the OCR is enabled and Tesseract OCR is not available lines in fess-crawler.log.

The documented behaviour comes from the companion fess and fess-crawler changes, and was checked on a live Fess 15.9.0-SNAPSHOT with Tesseract.

config/properties.rst is not regenerated here. Regenerating it from the current fess tree also picks up unrelated keys added since its last run, so it is better done as a separate sweep after the fess change is merged.

Checks

  • tools/check_headings.py finds no heading problems in the new pages.
  • The new pages and the updated index.rst were built with Sphinx in a scratch tree for every language, with no warnings from the new page.
  • The de, es, fr, ko and zh-cn pages are translations of the en and ja pages.

Add config/crawler-ocr.rst in every language and list it after the
thumbnail page in the crawler section. It covers what OCR applies to,
how scanned PDFs are handled, installing Tesseract, the
crawler.document.ocr.* settings, the Docker recipe in docker-fess, the
per-crawl config.tika.tesseract.config parameter, operational notes,
and the tika.xml exclusion that a customized file from 15.8 or earlier
still carries.
@marevol marevol self-assigned this Oct 1, 2026
@marevol
marevol merged commit 72aa2e8 into main Oct 1, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant