Repository navigation
Raise on unparseable submissions so they surface as 422 - #13
Merged
Merged
Conversation
Unparseable responses previously returned a 200 with is_correct=False and parse_error feedback, so lf_toolkit's invalid-submission handling never fired. FeedbackException now subclasses ValueError and is no longer caught, so lf_toolkit reports it as an invalid submission (422). An unparseable answer is a fault in the task rather than the student's submission, so it raises RuntimeError instead and is reported as a 500. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Pins lf_toolkit to toolkit-python 9d89276 (feature/invalid-submission-error), which maps a ValueError raised by the evaluation function to a JSON-RPC 422 error, so unparseable submissions can be tested end to end on staging. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 5, 2026
toolkit-python v1.2.0 and shimmy main now include the invalid-submission (422) handling, so switch from the test builds to the released toolkit tag and the edge base image. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Responses that can't be parsed (e.g.
"A n ") used to come back as a 200 withis_correct: falseandparse_errorfeedback. Because the function returned normally, the new invalid-submission handling in lf_toolkit and shimmy never ran.FeedbackExceptionnow subclassesValueError, andevaluation_functionno longer catches it. lf_toolkit maps aValueErrorfrom the handler to an invalid-submission error, and shimmy returns that as 422VALIDATION_ERROR.RuntimeError. That's a problem with the task setup, not the student's input, so it comes back as a 500.preview_functionhasn't changed: it still catchesFeedbackExceptionand returns inline feedback.Behaviour change
A student who submits something unparseable now gets an error response instead of a graded result with
parse_errorfeedback.Dependencies
lf_toolkitis pinned to v1.2.0, which includes Report unprocessable submissions as invalid-submission errors (422) toolkit-python#14.python:edge-3.12, which is built from shimmymainand includes Return 422 when the evaluation function rejects a submission shimmy#38.Testing
test_returns_is_correct_false_not_parseablewith tests checking that an unparseable response raisesValueErrorand an unparseable answer raises a non-ValueError.pytest evaluation_function: 20 passed, plus 2 subtests.🤖 Generated with Claude Code