Skip to content

AP-859 add basic worker task - #2

Merged
jason-raitz merged 7 commits into
mainfrom
AP-859_implement_worker
Sep 25, 2026
Merged

jason-raitz merged 7 commits into
mainfrom
AP-859_implement_worker

Conversation

@jason-raitz

@jason-raitz jason-raitz commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Very basic reimplementation of the perl tesseract caller for a given filelist, list of languages, and output path

  • (temporarily) mounts tests for quicker development
  • turns on autodiscovery of celery tasks
  • sketches out a run_tesseract_job with expected params
  • sample tests for the job
  • Note: meta params added to job's update_state. This is optional and can be omitted if we don't want it for debugging/tracking.
  • added tesseract to worker along with a few language training files

- (temporarily) mounts tests for quicker development
- turns on autodiscovery of celery tasks
- sketches out a run_tesseract_job with expected params
- sample tests for the job
- Note: meta params added to job's update_state. This is optional and
  can be omitted if we don't want it for debugging/tracking.
@jason-raitz jason-raitz self-assigned this Sep 23, 2026
@jason-raitz jason-raitz changed the title add basic worker task AP-859 add basic worker task Sep 23, 2026
@jason-raitz
jason-raitz marked this pull request as draft September 23, 2026 19:31

@anarchivist anarchivist left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i know you're still working on this, but it's looking good so far!

 - adds tesseract to worker
 - also adds a couple languages for local dev/testing
 - mounts a files folder, but this will probably change?
 - updates tasks.py to actually use tesseract on a filelist
 - adds a couple of fixture files for future tests
- also handles an output ending in `.pdf`
@jason-raitz
jason-raitz marked this pull request as ready for review September 24, 2026 21:21

@steve-sullivan steve-sullivan left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

r+ w/a few questions and a tweak!

Comment thread Dockerfile

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since fra is one of our default languages, I'm assuming we'd want to include tesseract-ocr-fra here as well?

Comment thread quiabo/tasks.py

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we verify that the output directory exists, or just trust that the caller has already validated it?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems to me like a good idea to ensure the output directory exists, since that's a clear error that could have a much simpler message than CalledProcessError with a bunch of Tesseract output.

Comment thread test/unit/test_tasks.py

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just to make sure I'm following this correctly — with subprocess.run monkeypatched here, you're not actually calling Tesseract to process the file, right? You've added an actual .png file, so is there an intent to have some integration(ish) testing later?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think we need to run integration tests, but I thought it would be helpful for local development to have a sample image with text available.

@awilfox awilfox left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

r+wc agree with Steve's suggestions, after that this looks good to me as a first pass.

Comment thread quiabo/tasks.py

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems to me like a good idea to ensure the output directory exists, since that's a clear error that could have a much simpler message than CalledProcessError with a bunch of Tesseract output.

@anarchivist anarchivist left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

rw+c (i think just Steve and Anna's remaining comment to verify that the output directory exists). plus a couple of comments - i don't think those are necessarily blocking, but i'd like us to make a conscious decision on them.

Comment thread quiabo/tasks.py
if not output_dir.exists():
raise FileNotFoundError(f"Output directory {output_dir} does not exist.")

# The state update does not need any meta, unless we want it for debugging or tracking.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this because the task already has the metadata on it?

Comment thread quiabo/tasks.py
Comment on lines +25 to +26
command = ["tesseract", "-l", "+".join(languages), filelist, output, "pdf"]
subprocess.run(command, check=True, capture_output=True, text=True)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is there a reason why we're invoking tesseract ourselves using subprocess.run() rather than doing it through pytesseract?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From what I understand, the way I do it here is basically equivalent to using pytesseract as pytesseract is essentially just a wrapper. In the interest of not blocking things, I'm going to merge this for now, and I'll take a look at using pytesseract in my follow-on work.

Comment thread Dockerfile
Comment on lines +56 to +61
tesseract-ocr-eng \
tesseract-ocr-spa \
tesseract-ocr-deu \
tesseract-ocr-chi-tra \
tesseract-ocr-ita \
tesseract-ocr-fra \

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are we sure that we want to install these directly into the image? i thought we'd decided that we wanted to mount a tessdata directory with the .traineddata files into the container to keep the image size small. that said, these specific language data files only add about 100 MB to the image, so maybe there's a tradeoff about which we include in the image vs. which we have to bring in separately.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we should keep them for now for development. I'm guessing it's the chinese language data that's bloated the size.

@jason-raitz
jason-raitz merged commit b4a4f8b into main Sep 25, 2026
5 checks passed
@jason-raitz
jason-raitz deleted the AP-859_implement_worker branch September 25, 2026 19:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants