Skip to content

feat(crawler): add settings to enable OCR for images and scanned PDFs - #3520

Merged
marevol merged 1 commit into
mainfrom
feat/ocr-enable
Oct 1, 2026
Merged

marevol merged 1 commit into
mainfrom
feat/ocr-enable

Conversation

@marevol

@marevol marevol commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Summary

Makes Tesseract OCR easy to turn on for images and scanned PDFs.

Before this change, Tika's TesseractOCRParser was shipped but excluded in tika.xml. Scanned PDFs never reached it anyway, because PDFs are extracted with PDFBox.

  • tika.xml: no longer excludes TesseractOCRParser.

  • New settings in fess_config.properties:

    Key Default Meaning
    crawler.document.ocr.enabled false Run OCR on images and scanned PDFs
    crawler.document.ocr.language eng Tesseract languages, e.g. jpn+eng
    crawler.document.ocr.timeout 120 Seconds per Tesseract run
  • FessTesseractOCRConfig: registered in fess.xml as the tesseractOCRConfig component next to tikaConfig. With feat(extractor): add a default OCR config and a fallback for text-less PDFs fess-crawler#217, every TikaExtractor in the crawler container uses it as its default, including subclasses from plugins.

    • OCR disabled: it sets skipOcr, so a host that has tesseract installed behaves as before.
    • OCR enabled: it applies the language and timeout, and turns on preserve_interword_spaces. Without that, Tesseract puts a space between CJK characters ("請 求 書"), and Japanese words do not match.
    • Invalid language: OCR is disabled with a warning.
    • A per-crawl config.tika.tesseract.config parameter still takes precedence.
  • FessPdfExtractor (via crawler/extractor+pdfExtractor.xml): when OCR is enabled, it sets tikaExtractor as the fallback for PDFs from which PDFBox extracts no text. Tika then renders the pages and OCRs them.

    • It first checks that Tika can really run Tesseract with this config.
    • If the tesseract command is missing, or a customized tika.xml still excludes the parser, it logs OCR is enabled, but Tesseract OCR is not available. ... instead of silently re-parsing.
    • On success it logs OCR is enabled: language=..., timeout=...s.

Depends on codelibs/fess-crawler#217, so CI needs that change in the fess-crawler snapshot first. Docs: codelibs/fess-docs#562. Docker recipe: codelibs/docker-fess#89.

Upgrade note: RPM/DEB installs keep a modified /etc/fess/tika.xml, and a customized app/WEB-INF/conf/tika.xml also stays. Those files still exclude the parser and must be edited to use OCR. The new warning points at this.

Tests

  • Unit tests:
    • FessTesseractOCRConfigTest: disabled, enabled, blank timeout, invalid language, and init() reading FessConfig.
    • FessPdfExtractorTest: fallback set only when OCR is enabled and available, and the availability check with skipOcr.
    • The full unit test suite passes locally.
  • End to end: run on a live Fess 15.9.0-SNAPSHOT with OpenSearch 3.8.0 and Tesseract 5.5 (eng, jpn, osd), crawling a PNG, an image-only PDF, a text PDF and a text-less image.
    • OCR enabled (jpn+eng):
      • The PNG and the scanned PDF are OCR'd, and 見積書, 請求書, and English tokens from both files are found.
      • The text PDF is unchanged.
      • Tesseract runs once per image or page, including the text-less image.
    • OCR disabled, with Tesseract installed: the OCR words are not found, all files are still indexed, and Tesseract is never invoked.
    • OCR enabled, with a tika.xml that still excludes the parser: the warning is logged once, there is no PDF fallback and no OCR, and indexing is otherwise unaffected.
    • Webapp start-up: no errors from the new component.

Tika's Tesseract OCR parser was shipped but excluded in tika.xml, and
scanned PDFs never reached it because PDFs are extracted with PDFBox.

- tika.xml no longer excludes TesseractOCRParser.
- New fess_config.properties keys: crawler.document.ocr.enabled
  (default false), crawler.document.ocr.language (default eng) and
  crawler.document.ocr.timeout (default 120 seconds).
- FessTesseractOCRConfig, registered as the tesseractOCRConfig
  component, applies these settings as the default Tesseract config of
  every TikaExtractor. When OCR is disabled it sets skipOcr, so a host
  that has the tesseract command installed keeps the previous
  behaviour. When OCR is enabled it also preserves interword spacing,
  so that Tesseract does not put a space between CJK characters, which
  would keep Japanese words from matching. An invalid language
  disables OCR with a warning.
- FessPdfExtractor sets tikaExtractor as the fallback for PDFs from
  which PDFBox extracts no text, when OCR is enabled. It first checks
  that Tika can really run Tesseract; if the tesseract command is
  missing or a customized tika.xml still excludes the parser, it logs a
  warning instead.

Requires the fess-crawler change that adds the default Tesseract
config and the PDF fallback extractor.
@marevol marevol self-assigned this Oct 1, 2026
@marevol marevol added this to the 15.9.0 milestone Oct 1, 2026
@marevol
marevol merged commit 2f7cf43 into main Oct 1, 2026
2 of 3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant