Skip to content

FIX: Include matcher configuration in decoding scorer identity - #3075

Merged
Roman Lutz (romanlutz) merged 2 commits into
microsoft:mainfrom
Ayoubhm07:fix/decoding-scorer-matcher-identity
Oct 11, 2026
Merged

Roman Lutz (romanlutz) merged 2 commits into
microsoft:mainfrom
Ayoubhm07:fix/decoding-scorer-matcher-identity

Conversation

@Ayoubhm07

Copy link
Copy Markdown
Contributor

Description

Closes #3074 .

DecodingScorer instances with different matcher settings could produce opposite verdicts but share both hash and eval_hash, allowing cached evaluation metrics to be reused across different behavior. This is the same defect #2977 described for SubStringScorer, in the sibling class that #2979 did not touch: both hold a TextMatching instance at the same point of the same method, but only SubStringScorer._build_identifier() passes its get_identifier_params() through.

The fix copies the two lines #2979 added to the sibling, so the built-in matchers' case sensitivity, whitespace handling, threshold, and n-gram size now reach the identity. Custom matchers that only implement is_match() remain supported through the same callable() guard. Built-in matcher hashes change intentionally, as they did for SubStringScorer; no shipped scorer metric registry references DecodingScorer, so no published eval_hash moves.

Scorer categories reach the emitted score_category but are absent from both scorers' identifiers. #2979 left them out for SubStringScorer too, so I kept this change symmetric with the sibling rather than widening it; happy to cover categories for both in a follow-up if you want them included.

Tests and Documentation

  • Added the sibling's two identity regressions, adapted to this scorer: three parameterised cases covering case sensitivity, threshold and n-gram size, each asserting that opposite verdicts no longer share an identity, plus equivalent-configuration stability for both built-in matchers.
  • Before the fix: 3 failed, 11 passed in tests/unit/score/test_decoding_scorer.py. After: 14 passed.
  • The full tests/unit/score/ suite reports the same 47 pre-existing failures before and after, with passes going from 3422 to 3427, which is exactly the five added cases. Those failures are confined to test_azure_content_filter.py and test_local_refusal_classifier_scorer.py and are unrelated to this change.
  • ruff check and ruff format --check at the pinned v0.16.10, git diff --check, check_async_suffix.py and check_no_rest_roles.py all pass. ty check reports the same single pre-existing diagnostic on line 50 before and after, which substring_scorer.py also reports. I ran pytest directly rather than make unit-test, and did not run the full repository suite.

DecodingScorer._build_identifier recorded only the text matcher's class
name, so the built-in matchers' case_sensitive, ignore_whitespace,
threshold and n settings were absent from both the component hash and
the evaluation hash. Two scorers returning opposite verdicts for the
same input therefore shared an evaluation identity, and
ScorerEvaluator._should_skip_evaluation could reuse metrics from a
different matcher configuration.

Pass the matcher's get_identifier_params() through, as microsoft#2979 did for the
sibling SubStringScorer. Custom matchers that only implement is_match()
remain supported through the same callable guard.

Add the sibling's identity regressions, adapted: three parameterised
cases over case sensitivity, threshold and n-gram size, each asserting
that opposite verdicts no longer share an identity, plus
equivalent-configuration stability for both built-in matchers.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@romanlutz
Roman Lutz (romanlutz) added this pull request to the merge queue Oct 11, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Oct 11, 2026
@romanlutz
Roman Lutz (romanlutz) added this pull request to the merge queue Oct 11, 2026
Merged via the queue into microsoft:main with commit 83ce11c Oct 11, 2026
55 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

DecodingScorer identifiers omit text matcher configuration

2 participants