Faithfulness#
Metrics against an annotation check whether a highlight matches what a human marked.
They cannot check whether the highlight is what the predictor read, a model can match the annotation and rest its prediction on something else.
faithfulness adds two measures that test the claim by changing what the predictor sees:
Registry.from_key(TOY_TASK, faithfulness=True).run()
# results.json gains test_sufficiency and test_comprehensiveness
Writing x for the input, h for the highlight and x \ h for the input with the highlight removed:
sufficiency = p(y_hat | x) - p(y_hat | h)
comprehensiveness = p(y_hat | x) - p(y_hat | x \ h)
Sufficiency asks whether the highlight carries the signal alone, so lower is better. Comprehensiveness asks whether anything the class rests on was left outside it, so higher is better. Two extra predictor passes per test batch pay for both: the highlight pass is the model’s own output, already computed.
Off by default. They are two more columns rather than a correction, and a registered reproduction should report what its paper reports.
Every architecture is scored partly off its training distribution, and which part differs. Nothing is trained on an input with a hole in it, so comprehensiveness is off-distribution for everyone.
The other passes divide by family: a full-input classifier is trained for p(y_hat | x) and never sees h alone, while a select-then-predict predictor is trained for p(y_hat | h) and never sees x.
MCD is the exception, it trains both passes by design.
So this is a property of the measure, not a demerit of a model, and published numbers carry the same distortion with the terms exchanged.
Comprehensiveness can be uninformative rather than low. Both of its terms are the full-length pass.
For a select-then-predict model that is the pass it never trains on, so both are degraded and their difference is compressed toward zero whatever the highlight contains.
Take a document of thirty-five words with four kept.
Removing the four words leaves thirty-one, which is still a long off-training input, so the prediction barely moves and the column reads near zero.
That is a fact about the measure applied to this model class, not a finding about the highlight.
Do not read a low value here as a demerit unless p(y_hat | x) is a pass the model is trained for.
Two consequences to know before reading a column:
y_hatis the class predicted from the highlight, not from the full input. ERASER takes it from the full input because for a full-input classifier that is the model’s prediction; the intent is the class the model predicts, and here that comes fromh. Anchoring on the full-input pass would anchor on the one pass such a model was never trained for. A documented deviation, and the reason a column here is not interchangeable with a published one.Negative sufficiency is expected.
p(y_hat | h)is the trained pass andp(y_hat | x)is not, so a select-then-predict model can score better on its highlight than on the whole input. Not a defect.
The library says highlight where the literature says rationale, and writes h where it writes r.
The metric names stay as published, so a column still matches a paper’s.
API#
Faithfulness: whether the highlight is what the prediction rests on.
A model that reports a highlight makes a claim about itself, and metrics against an annotation do not check it: a highlight can match the annotation and still not be what the predictor read. Two measures test the claim by changing what the predictor sees, both from DeYoung, Jain, Rajani, Lehman, Xiong, Socher and Wallace, 2020, ERASER: A Benchmark to Evaluate Rationalized NLP Models.
Writing x for the input, h for the highlight, x \ h for the input
with the highlight removed, and y_hat for the class scored:
sufficiency = p(y_hat | x) - p(y_hat | h)
comprehensiveness = p(y_hat | x) - p(y_hat | x \ h)
Sufficiency asks whether the highlight carries the signal on its own, so lower is better. Comprehensiveness asks whether anything the class rests on was left outside the highlight, so higher is better. They are not two views of one number: a model can highlight three words that suffice while ten others would have sufficed too. It is sufficient, and not comprehensive.
The literature calls h the rationale and writes it r. This library
says highlight throughout, as its loaders, models and metrics do.
Every architecture is measured partly off its training distribution, and which part differs. Nothing is trained on an input with a hole in it, so comprehensiveness is off-distribution for everyone. The other two passes divide by family:
Trained on |
|
|
|
|---|---|---|---|
full input |
in |
off |
off |
select-then-predict |
off |
in |
off |
MCD (both, by design) |
in |
in |
off |
So this is a property of the measure rather than a demerit of a model, and the published numbers carry the same distortion with the terms exchanged.
Two consequences worth knowing before reading a column:
``y_hat`` is the class predicted from the highlight, not from the full input. ERASER takes it from the full input because for a full-input classifier that is the model’s prediction; the intent is the class the model predicts, and for a select-then-predict model that comes from
h. Anchoring on the full-input pass would anchor on the one pass such a model was never trained for. A documented deviation, and the reason a column here is not interchangeable with a published one.Negative sufficiency is expected here.
p(y_hat | h)is the trained pass andp(y_hat | x)is not, so a select-then-predict model can score better on its highlight than on the whole input. That is not a defect.
- pyhighlights.components.faithfulness.evaluate(model, loader)[source]#
Mean faithfulness terms over
loader, as the model currently stands.Runs outside the Lightning loop: the terms need extra forward passes with masks of their own rather than another binding over the fields a step already produced. The model is left in the mode it arrived in.
The caller owns where the model sits. Batches are moved to
model.device, and the model is read where it is: a stage outside every Lightning loop is one Lightning no longer places, soscore()moves the model to the run’s device before calling this.- Return type:
Dict[str,float]- Parameters:
model (Model)
loader (DataLoader)