fix(fairness): count a masked 0/inf ratio as exceeded in fairness_check - #586
Open
shaurya416 wants to merge 1 commit into
Open
shaurya416 wants to merge 1 commit into
shaurya416 wants to merge 1 commit into
Conversation
calculate_ratio() masks a subgroup ratio that is exactly 0 or inf to NaN, so calculate_parity_loss()'s log stays finite. universal_fairness_check() reused that same masked frame for the pass/fail comparison, and NaN compares False both ways in numpy, so a masked ratio was never counted as exceeded -- even though 0 or inf is the most extreme disparity a metric can show. A model that never predicts the positive class for one subgroup could be reported as fair. universal_fairness_check() now recomputes the ratio from the raw scores already stored on GroupFairnessClassification (self.metric_scores, unmasked) to tell apart why a ratio is NaN: masked because it was really 0 or inf (now counts as exceeded), or genuinely 0/0 because both groups scored exactly 0 (stays excluded, since that case really is undecidable). GroupFairnessRegression has no metric_scores and no such masking, so it is unaffected (hasattr guard). calculate_ratio() and calculate_parity_loss() are unchanged. Adds test_fairness_check_masked_ratio: a subgroup that never predicts positive (TPR/FPR/STP ratios exactly 0, PPV genuinely 0/0). Asserts the printed verdict names TPR/FPR/STP as exceeded, not PPV, and does not say the model is fair. Full FairnessTest suite (23 tests) passes.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #585.
Problem
calculate_ratio()masks a subgroup ratio that comes out as exactly0orinftoNaN, socalculate_parity_loss()'slogstays finite.universal_fairness_check()reused that same masked frame for the pass/fail comparison, andNaN > x/x > NaNare bothFalsein numpy, so a masked ratio was never counted as exceeded — even though a ratio of0orinfis the most extreme disparity a metric can show. A model that never predicts the positive class for one subgroup can be reported as fair.Fix
universal_fairness_check()now recomputes the ratio from the raw scores already stored on the object (self.metric_scores, unmasked) to tell apart why a ratio isNaN:0orinf→ counts as exceeded (masked_disparity)0/0(both groups scored exactly0) → stays excluded, since that case really is undecidableGroupFairnessRegressionsharesuniversal_fairness_checkbut has nometric_scoresand no such masking (calculate_regression_measuresis a different computation entirely) — guarded with ahasattrcheck so it's unaffected.calculate_ratio()andcalculate_parity_loss()are unchanged.Testing
Added
test_fairness_check_masked_ratio: a subgroup that never predicts positive (TPR/FPR/STP ratios exactly0, PPV genuinely0/0). Asserts the printed verdict names TPR/FPR/STP as exceeded, but not PPV, and does not say the model is fair.Ran the full
FairnessTestsuite (23 tests, includingGroupFairnessRegression) — all pass. As a control, reverting only the fix hunk makes the new test fail exactly as expected (prints "No bias was detected!" on data with three 0-ratio metrics).