Metrics#

What a run scores, and the one field a class-imbalanced corpus cannot be read without.

Metrics are registered in two layers, as the losses are: a torchmetric is the scoring object, and a metric binds it to the fields it reads. The binding is what a task names, since the same F1 scores classes or tokens depending on what it is handed.

Key

Reads

Reports

ACCURACY_METRIC

class_logits, y_true

Classification accuracy, two classes

F1_METRIC

class_logits, y_true

Macro F1, two classes

MULTICLASS_ACCURACY_METRIC

class_logits, y_true

Accuracy over three classes, for HateXplain

MULTICLASS_F1_METRIC

class_logits, y_true

Macro F1 over three classes

HIGHLIGHT_F1_METRIC

highlight_mask, highlight_true

Token F1 against the annotation

HIGHLIGHT_IOU_METRIC

highlight_mask, highlight_true

Token intersection over union

SELECTION_RATE_METRIC

highlight_mask, mask

Share of the document the selector kept

SELECTION_SIZE_METRIC

highlight_mask, mask

Words kept, averaged over samples

SELECTION_SPANS_METRIC

highlight_mask, mask

How many contiguous spans the selection falls into

CLASS_F1_METRIC

class_logits, y_true

F1 of one class rather than an average over all of them

HIGHLIGHT_PRECISION_METRIC, HIGHLIGHT_RECALL_METRIC

highlight_mask, highlight_true

The two halves of the highlight F1, which a study reports when a selector is precise and short or broad and complete

Both classification metrics use task="multiclass" even for a two-class corpus: a predictor emits one logit per class, and the "binary" task wants a single score per sample instead.

Highlight metrics ignore positions annotated with -1, which is what an unannotated split is padded with, those positions score nothing rather than counting as negatives. The selection metrics read mask instead, so they report what the selector kept whether or not the corpus is annotated at all.

BINARY_METRICS collects the set a two-class corpus wants, and Beer, Hotel, Movies and Toy use it as it stands.

CLASS_F1_METRIC is the one to reach for where a corpus is imbalanced enough that a macro average hides the answer: a corpus in which one class is 99.5% of the rows is scored at roughly 0.5 by a model that never predicts the other one. The selection metrics need no annotation at all, which is what makes them the first thing to read on a corpus that ships none, and Interlocking and what the bottleneck guarantees is what they are read for.

Class weights#

Where a corpus is class-imbalanced, the loss needs one number per class, and where that number came from decides whether the run can be repeated. ClassWeightsTask is a run whose whole result is those numbers:

from pyhighlights.configurations.keys import TOY_CLASS_WEIGHTS_TASK

Registry.from_key(TOY_CLASS_WEIGHTS_TASK, save_path="results").run()
{
  "split": "train",
  "weights": [1.0, 1.0],
  "counts": {"0": 32, "1": 32},
  "rows": {"train": 64, "val": 16, "test": 16},
  "labels": {"train": {"0": 32, "1": 32}, "val": {"0": 8, "1": 8},
             "test": {"0": 8, "1": 8}}
}

It trains nothing and takes no seeds. What it writes is the usual results.json and manifest.json, so the weights arrive with the key of the corpus and of the preprocessing that produced them, which is what makes them worth copying into a CrossEntropyConfig, where every training run’s manifest then records them.

The reading itself is ClassWeights, a preprocessor that changes no row. Being a step of the pipeline is the point: it runs over the split the study trains on, after whatever filtering and aggregation came before it, since that is what changes the frequencies.

Declaring the numbers rather than computing them at training time is deliberate. A fixed split has fixed frequencies, and a declared weight is in the manifest of every run that used it, where one computed inside training exists only for the length of the process.

API#

class pyhighlights.utility.metrics.BinaryHighlightF1Score(pos_label=1, ignore_index=-1, **kwargs)[source]#

nan where there is nothing to score, deliberately.

The denominator is zero when no annotated position was marked, either by the corpus or by the model. Two ways to reach it: the metric was never updated, or every position it saw was a true negative: a split with no annotations, or one whose rows annotate that nothing is a highlight and a model that selected nothing on them.

Both are the same statement, that no highlight was asked for and none was offered, and nan is what says it. Returning 0.0, which is what torchmetrics’ own zero_division default does, would be a score, and a score of zero says the model got everything wrong. Returning 1.0 would say it got everything right for selecting nothing. Neither happened.

HighlightMetric.update() masks predictions by valid, so a row the corpus does not annotate contributes to no counter at all.

Parameters:
  • pos_label (int)

  • ignore_index (int)

class pyhighlights.utility.metrics.BinaryHighlightIoU(pos_label=1, ignore_index=-1, **kwargs)[source]#

nan where there is nothing to score; see BinaryHighlightF1Score.

Parameters:
  • pos_label (int)

  • ignore_index (int)

class pyhighlights.utility.metrics.BinaryHighlightPrecision(pos_label=1, ignore_index=-1, **kwargs)[source]#

Of the positions the selection marked, the share the corpus annotates.

Reported beside recall because a sparsity target moves the two in opposite directions and their F1 hides it: a selector squeezed below the annotation’s own length buys precision with recall, and a column of F1 alone reads as a model that got slightly worse rather than one that changed what it does.

nan where there is nothing to score; see BinaryHighlightF1Score. Here the denominator is zero when the model marked no annotated position, which is a different statement from marking the wrong ones.

Parameters:
  • pos_label (int)

  • ignore_index (int)

class pyhighlights.utility.metrics.BinaryHighlightRecall(pos_label=1, ignore_index=-1, **kwargs)[source]#

Of the positions the corpus annotates, the share the selection marked.

nan where there is nothing to score; see BinaryHighlightF1Score. The denominator is zero on a split that annotates nothing.

Parameters:
  • pos_label (int)

  • ignore_index (int)

class pyhighlights.utility.metrics.BoundMetric(name, metric, inputs=('class_logits', 'y_true'))[source]#

Binds a torchmetrics metric to the fields feeding it.

Mirrors Loss: the metric stays a plain Metric, and the binding says which fields of the step namespace it scores.

Parameters:
  • name (str)

  • metric (RegistrationKey[Metric])

  • inputs (Sequence[str])

class pyhighlights.utility.metrics.ClassF1Score(pos_label=1, num_classes=2, **kwargs)[source]#

F1 of one class rather than an average over all of them.

Macro F1 over two classes is half the majority class, and on a corpus where one class is nearly every row that half carries the score: a model that answers “negative” to everything reports about 0.50 while finding nothing. Averaging is the wrong summary there. What the run is about is the rare class, so this reports it alone.

pos_label names that class. The predictor emits one logit per class, which is why this is a multiclass metric restricted to one class rather than task="binary": binary wants one score per sample.

Parameters:
  • pos_label (int)

  • num_classes (int)

class pyhighlights.utility.metrics.EmptySetAccuracy(threshold=0.5, ignore_index=-1, **kwargs)[source]#

Share of examples annotated with nothing that were named nothing.

Reported on its own rather than folded into an average, because the examples that instantiate nothing can be the overwhelming majority of a corpus, and would otherwise carry every number they entered. An example grounded in an entry it has no business in has invented a reason, which is a scoring error and not an ambiguity.

Parameters:
  • threshold (float)

  • ignore_index (int)

class pyhighlights.utility.metrics.ExactSetMatch(threshold=0.5, ignore_index=-1, **kwargs)[source]#

Share of examples whose named set is exactly the annotated one.

Strict, and the number a domain expert cares about: naming two of the three entries that explain an example explains it two thirds of the way, which is not what an explanation is for.

Parameters:
  • threshold (float)

  • ignore_index (int)

class pyhighlights.utility.metrics.HighlightMetric(pos_label=1, ignore_index=-1, **kwargs)[source]#

Token-level confusion counts over labelled highlight positions.

Parameters:
  • pos_label (int)

  • ignore_index (int)

class pyhighlights.utility.metrics.KnowledgeSetMetric(threshold=0.5, ignore_index=-1, **kwargs)[source]#

Per-example agreement between the entries named and the entries annotated.

Both scores here are about the set, not about the individual links. Per-link F1 rewards naming one correct entry out of several and stopping; what a reader wants to know is whether the model got the explanation right, and a partial set is a partial explanation.

An example the corpus does not annotate scores nothing: ignore_index marks it, and it is skipped rather than counted as an empty set. That distinction is the whole point of the knowledge axis carrying -1 and 0 as different values.

nan where no example was annotated, as for BinaryHighlightF1Score: zero would report every set as wrong.

Parameters:
  • threshold (float)

  • ignore_index (int)

counts(preds, target)[source]#

Predicted and annotated sets, over the annotated examples only.

Parameters:
  • preds (Tensor)

  • target (Tensor)

class pyhighlights.utility.metrics.SelectionMetric(**kwargs)[source]#

Per-sample selection statistic over the tokens a document actually has.

target is the padding mask, 1 for a real token and 0 for padding, and not an annotation. A selection statistic is about the document, so it is defined whether or not the corpus annotated anything, which is why the registered binding names mask rather than highlight_true.

The distinction is the whole metric. Counting padding makes a rate depend on the widest row in the batch rather than on the document: a 6-token selection out of a 34-token document is 18%, and reads as 6% once 67 columns of padding join the denominator. Sizes are unaffected, since padding adds zero to a sum, and rates are not.

nan where no document had a token, as for BinaryHighlightF1Score: zero would report a model that kept nothing.

reduce(selected, length)[source]#

One statistic per document, from what it kept and how long it is.

selected is [B, T] with every padded position already zeroed, and length is [B]; both cover only the documents that have a token. A subclass sums, divides, or reads the positions themselves. A count of spans is not recoverable from a total, which is why this takes the row rather than its sum.

Return type:

Tensor

Parameters:
  • selected (Tensor)

  • length (Tensor)

class pyhighlights.utility.metrics.SelectionRate(**kwargs)[source]#

Share of a document the selection kept, so bounded by zero and one.

class pyhighlights.utility.metrics.SelectionSize(**kwargs)[source]#

How many tokens the selection kept, which nothing bounds above.

class pyhighlights.utility.metrics.SelectionSpans(**kwargs)[source]#

How many contiguous runs the selection falls into.

Contiguity is a penalty in pyhighlights.utility.losses and was never a reported number, so a run said how much it kept and never whether the kept words sit together. Six words in one span and six scattered over a document are the same selection size and not the same highlight: the first can be read as a phrase, the second is what a model keying on punctuation produces.

A run of ones is counted at its first position, so a document that selects nothing scores zero and one that selects everything scores one.

The mask has to be hard. Every positive entry opens or continues a run, so a probability of 0.01 counts as kept. SelectionRate and SelectionSize sum their input and stay meaningful on a soft mask; this one does not, and reports the number of runs of non-zero entries rather than anything about the highlight a threshold would produce.

Metric registrations, and the bindings that feed them named fields.

Two layers, as with the losses: a torchmetric is the scoring object, and a metric binds it to the fields of the step namespace it reads. The binding is what a model or a task names, since the same F1 scores classes or tokens depending on what it is handed.

Classification metrics are registered per class count, because torchmetrics needs to know: Beer, Hotel, Movies and Toy have two classes, HateXplain has three. Both use task="multiclass": a predictor emits one logit per class, including when there are two, and the "binary" task wants a single score per sample instead. Highlight and selection metrics are class-agnostic, since a token is selected or it is not, so they are registered once and used by every corpus.

class pyhighlights.configurations.metrics.ClassF1Config(**data)[source]#

F1 of one class, for a corpus an average would flatter.

pos_label is the class the run is about: the rare one, on a corpus skewed enough for macro F1 to be mostly the majority class.

Parameters:
  • pos_label (int)

  • num_classes (int)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.metrics.ClassF1MetricConfig(**data)[source]#

Still reported as f1: which F1 it is, the manifest’s key says.

Parameters:
  • name (str)

  • metric (RegistrationKey[Metric])

  • inputs (Sequence[str])

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.metrics.EmptySetAccuracyConfig(**data)[source]#

Share of examples annotated with nothing that were named nothing.

Parameters:
  • threshold (float)

  • ignore_index (int)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.metrics.ExactSetMatchConfig(**data)[source]#

Share of examples whose named set is exactly the annotated one.

Parameters:
  • threshold (float)

  • ignore_index (int)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.metrics.HighlightF1MetricConfig(**data)[source]#

Scores the selected tokens against the annotated ones.

Parameters:
  • name (str)

  • metric (RegistrationKey[Metric])

  • inputs (Sequence[str])

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.metrics.HighlightMetricConfig(**data)[source]#

Token-level scores over the positions a corpus annotated.

ignore_index marks a position that carries no annotation, which is what an unannotated split is padded with; those positions score nothing rather than counting as negatives.

Parameters:
  • pos_label (int)

  • ignore_index (int)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.metrics.KnowledgeMetricConfig(**data)[source]#

A link metric reads the gate and the annotation behind it.

Parameters:
  • name (str)

  • metric (RegistrationKey[Metric])

  • inputs (Sequence[str])

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.metrics.LinkF1Config(**data)[source]#

Per-link F1, micro-averaged: every (example, entry) link counts once.

The comparable headline number. Recall is the one to watch inside it: an example can instantiate many entries, so a model that names one correct entry and stops looks precise and has missed the case.

Parameters:
  • num_labels (int)

  • average (str)

  • ignore_index (int)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.metrics.LinkMacroF1Config(**data)[source]#

Per-link F1 averaged over entries rather than over links.

Reported beside the micro average rather than instead of it, and it is the one that can see a rare entry. Micro weights every link equally, so the entries that fire often carry it. The entry that decides a case is frequently the one that fires on a handful of examples.

Parameters:
  • num_labels (int)

  • average (str)

  • ignore_index (int)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.metrics.MetricConfig(**data)[source]#

A scoring object bound to the namespace fields it reads.

Parameters:
  • name (str)

  • metric (RegistrationKey[Metric])

  • inputs (Sequence[str])

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.metrics.SelectionRateConfig(**data)[source]#

Share of a document the selector kept, averaged over samples.

No ignore_index: the metric is bound to mask, so what it has to exclude is padding rather than an unannotated position, and a mask says that with a zero.

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.metrics.SelectionRateMetricConfig(**data)[source]#

Reports what the selector kept, whether or not the corpus is annotated.

mask and not highlight_true, so the denominator is the document’s own length: the statistic is about the selection, and a corpus with no annotation still has one.

Parameters:
  • name (str)

  • metric (RegistrationKey[Metric])

  • inputs (Sequence[str])

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.metrics.SelectionSizeConfig(**data)[source]#

Tokens the selector kept, averaged over samples.

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.metrics.SelectionSpansConfig(**data)[source]#

Contiguous runs the selection falls into, averaged over samples.

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)