Describe the bug
SubStringScorer._build_identifier() records only the text matcher's class name. The built-in matchers' case_sensitive, ignore_whitespace, threshold, and n settings are absent from both the component hash and evaluation hash.
Two scorers can therefore return opposite verdicts for the same input while sharing an evaluation identity. ScorerEvaluator._should_skip_evaluation() looks up existing metrics by this hash, so an evaluation can reuse metrics from a different matcher configuration when the remaining cache conditions match.
Steps/Code to Reproduce
import asyncio
from unittest.mock import MagicMock, patch
from pyrit.analytics import ExactTextMatching
from pyrit.memory import CentralMemory, MemoryInterface
from pyrit.score import SubStringScorer
async def reproduce_async():
with patch.object(CentralMemory, "get_memory_instance", return_value=MagicMock(spec=MemoryInterface)):
a = SubStringScorer(substring="Hello", text_matcher=ExactTextMatching(case_sensitive=True))
b = SubStringScorer(substring="Hello", text_matcher=ExactTextMatching(case_sensitive=False))
print((await a.score_text_async("hello"))[0].get_value())
print((await b.score_text_async("hello"))[0].get_value())
print(a.get_identifier().hash == b.get_identifier().hash)
print(a.get_identifier().eval_hash == b.get_identifier().eval_hash)
asyncio.run(reproduce_async())
Expected Results
Verdicts are False and True; both hash comparisons are False because the behavioral configurations differ.
Actual Results
Verdicts are False and True, but both hash comparisons are True. Reproduced for all five built-in matcher settings listed above (case sensitivity for each matcher).
Versions
PyRIT main at b6a18a20, Python 3.12.3, Linux. No external model or service is needed.
Describe the bug
SubStringScorer._build_identifier()records only the text matcher's class name. The built-in matchers'case_sensitive,ignore_whitespace,threshold, andnsettings are absent from both the component hash and evaluation hash.Two scorers can therefore return opposite verdicts for the same input while sharing an evaluation identity.
ScorerEvaluator._should_skip_evaluation()looks up existing metrics by this hash, so an evaluation can reuse metrics from a different matcher configuration when the remaining cache conditions match.Steps/Code to Reproduce
Expected Results
Verdicts are
FalseandTrue; both hash comparisons areFalsebecause the behavioral configurations differ.Actual Results
Verdicts are
FalseandTrue, but both hash comparisons areTrue. Reproduced for all five built-in matcher settings listed above (case sensitivity for each matcher).Versions
PyRIT main at
b6a18a20, Python 3.12.3, Linux. No external model or service is needed.