Configurations#

pyhighlights.configurations.keys holds every RegistrationKey the other configuration modules register against, so model keys such as GRU_GENSPP and external trainer keys such as GRU_GENSPP_TRAINER are imported from one place.

Every registration#

The cards below are built from the registry itself, once per documentation build, so a key here is a key that exists and its defaults are the ones a run would get. Each card carries the registration name, its tags as coloured chips, the component it builds, and every parameter with its type and default. The search box filters by any of those, so transformer mcd narrows to the MCD registrations on a Transformer backbone and loader narrows to the corpora.

123 registrations

analyzerhighlightpositionrunnable

pyhighlights.components.analyzers.HighlightPositionAnalyzer

ParameterTypeDefaultWhat it does
directorystr | None
patternstr'predictions-seed=*.pkl'
binsint10
absoluteboolFalse
latestboolTrue
analyzerlabel-studiorunnable

pyhighlights.components.analyzers.LabelStudioExporter

ParameterTypeDefaultWhat it does
directorystr | None
patternstr'predictions-seed=*.pkl'
splitstr'test'
latestboolTrue
model_versionstr'pyhighlights'
labelsSequence'highlight'
onlyint | None
columnstr'label'
stemstr'label-studio'
analyzermetricsrunnable

pyhighlights.components.analyzers.MetricsAnalyzer

ParameterTypeDefaultWhat it does
directorystr | None
metricsSequence
splitstr'test'
pairsboolFalse
latestboolTrue
analyzerpredictionrunnable

pyhighlights.components.analyzers.PredictionAnalyzer

ParameterTypeDefaultWhat it does
directorystr | None
patternstr'predictions-seed=*.pkl'
splitstr'test'
latestboolTrue
backbonefrozentransformer

pyhighlights.components.models.spp.implementations.TransformerBackbone

ParameterTypeDefaultWhat it does
pretrained_model_cardstr'distilbert-base-uncased'
num_featuresint | None
freeze_transformerboolTrue
backbonegensppgru

pyhighlights.components.models.spp.implementations.GRUBackbone

ParameterTypeDefaultWhat it does
vocab_sizeint10000
embedding_dimint128
hidden_sizeint16
freeze_embeddingsboolTrue
num_layersint1
bidirectionalboolFalse
dropout_ratefloat0.0
backbonegenspptransformer

pyhighlights.components.models.spp.implementations.TransformerBackbone

ParameterTypeDefaultWhat it does
pretrained_model_cardstr'distilbert-base-uncased'
num_featuresint | None
freeze_transformerboolTrue
backbonegru

pyhighlights.components.models.spp.implementations.GRUBackbone

ParameterTypeDefaultWhat it does
vocab_sizeint10000
embedding_dimint128
hidden_sizeint128
freeze_embeddingsbool | None
num_layersint1
bidirectionalboolTrue
dropout_ratefloat0.0
backbonegrutransformer

pyhighlights.components.models.spp.implementations.StackedBackbone

ParameterTypeDefaultWhat it does
pretrained_model_cardstr'distilbert-base-uncased'
hidden_sizeint128
num_featuresint | None
freeze_transformerboolTrue
num_layersint1
bidirectionalboolTrue
dropout_ratefloat0.0
backbonetransformer

pyhighlights.components.models.spp.implementations.TransformerBackbone

ParameterTypeDefaultWhat it does
pretrained_model_cardstr'distilbert-base-uncased'
num_featuresint | None
freeze_transformerboolFalse
benchmarktoyrunnable

pyhighlights.components.benchmarks.Benchmark

ParameterTypeDefaultWhat it does
tasksListtask[toy]
namestr'toy-benchmark'
save_pathstr | None
strictboolFalse
task_argsDict{}
callbackcheckpointloss

pyhighlights.components.callbacks.WarmupModelCheckpoint

ParameterTypeDefaultWhat it does
monitorstr'val_loss'
modestr'min'
callbackcheckpointscore

pyhighlights.components.callbacks.WarmupModelCheckpoint

ParameterTypeDefaultWhat it does
monitorstr'val_score'
modestr'max'
callbackearly_stoppingloss

pyhighlights.components.callbacks.WarmupEarlyStopping

ParameterTypeDefaultWhat it does
monitorstr'val_loss'
modestr'min'
patienceint5
callbackearly_stoppingscore

pyhighlights.components.callbacks.WarmupEarlyStopping

ParameterTypeDefaultWhat it does
monitorstr'val_score'
modestr'max'
patienceint5
callbackgeneralization_lossscore

pyhighlights.components.callbacks.GeneralizationLossScore

ParameterTypeDefaultWhat it does
qualitystr'val_f1'
lossstr'val_loss'
coefficientfloat2.0
namestr'val_score'
comparerentailment

pyhighlights.components.models.spp.grounded.EntailmentComparer

ParameterTypeDefaultWhat it does
hidden_sizesList128
criterioncontiguity

pyhighlights.utility.losses.ContiguityPenalty

No parameters.

criterioncross_entropy

pyhighlights.utility.losses.CrossEntropy

ParameterTypeDefaultWhat it does
weightOptional
criterionjs_div

pyhighlights.utility.losses.JSDiv

No parameters.

criterionkl_div

pyhighlights.utility.losses.KLDiv

No parameters.

criterionknowledgemasked_bce

pyhighlights.utility.losses.MaskedBinaryCrossEntropy

ParameterTypeDefaultWhat it does
ignore_indexint-1
pos_weightOptional
criterionmasked_bce

pyhighlights.utility.losses.MaskedBinaryCrossEntropy

ParameterTypeDefaultWhat it does
ignore_indexint-1
pos_weightOptional
criterionmasked_cross_entropy

pyhighlights.utility.losses.MaskedCrossEntropy

ParameterTypeDefaultWhat it does
weightOptional
criterionsparsity

pyhighlights.utility.losses.SparsityPenalty

ParameterTypeDefaultWhat it does
thresholdfloat0.15
datasetbeer

pyhighlights.components.loaders.BeerLoader

ParameterTypeDefaultWhat it does
directorystr | None
splitsOptional
urlstr'https://people.csail.mit.edu/yujia/files/r2a/data.zip'
sha256str | None'23fcb4cac883ec1de86d83a7747294d7fdae10061d3803fd4c34c930e66f25de'
taskstr'beer0'
datasetbeertask=beer1

pyhighlights.components.loaders.BeerLoader

ParameterTypeDefaultWhat it does
directorystr | None
splitsOptional
urlstr'https://people.csail.mit.edu/yujia/files/r2a/data.zip'
sha256str | None'23fcb4cac883ec1de86d83a7747294d7fdae10061d3803fd4c34c930e66f25de'
taskstr'beer0'
datasetbeertask=beer2

pyhighlights.components.loaders.BeerLoader

ParameterTypeDefaultWhat it does
directorystr | None
splitsOptional
urlstr'https://people.csail.mit.edu/yujia/files/r2a/data.zip'
sha256str | None'23fcb4cac883ec1de86d83a7747294d7fdae10061d3803fd4c34c930e66f25de'
taskstr'beer0'
datasethatexplain

pyhighlights.components.loaders.HateXplainLoader

ParameterTypeDefaultWhat it does
directorystr | None
urlstr'https://raw.githubusercontent.com/hate-alert/HateXplain/01d742279dac941981f53806154481c0e15ee686/Data/dataset.json'
divisions_urlstr'https://raw.githubusercontent.com/hate-alert/HateXplain/01d742279dac941981f53806154481c0e15ee686/Data/post_id_divisions.json'
sha256str | None'63bb3340fee0ec469b09690d04cb68f7c187787dd8b83807f071892c084967fb'
divisions_sha256str | None'c2fb0d89862e7897b11ea3e9380753f15a793482b4b70ad0532dfb1212212835'
datasethotel

pyhighlights.components.loaders.HotelLoader

ParameterTypeDefaultWhat it does
directorystr | None
splitsOptional
urlstr'https://people.csail.mit.edu/yujia/files/r2a/data.zip'
sha256str | None'23fcb4cac883ec1de86d83a7747294d7fdae10061d3803fd4c34c930e66f25de'
taskstr'hotel_Location'
datasethoteltask=hotel_Cleanliness

pyhighlights.components.loaders.HotelLoader

ParameterTypeDefaultWhat it does
directorystr | None
splitsOptional
urlstr'https://people.csail.mit.edu/yujia/files/r2a/data.zip'
sha256str | None'23fcb4cac883ec1de86d83a7747294d7fdae10061d3803fd4c34c930e66f25de'
taskstr'hotel_Location'
datasethoteltask=hotel_Service

pyhighlights.components.loaders.HotelLoader

ParameterTypeDefaultWhat it does
directorystr | None
splitsOptional
urlstr'https://people.csail.mit.edu/yujia/files/r2a/data.zip'
sha256str | None'23fcb4cac883ec1de86d83a7747294d7fdae10061d3803fd4c34c930e66f25de'
taskstr'hotel_Location'
datasetmovies

pyhighlights.components.loaders.MoviesLoader

ParameterTypeDefaultWhat it does
directorystr | None
taskstr'movies'
splitsOptional
urlstr'https://www.eraserbenchmark.com/zipped/{task}.tar.gz'
sha256str | None'66e18d4e6c9df9e9f5544572b0bfe92a39673f74ecbfc3859b46cedb2f5b2dee'
datasettoy

pyhighlights.components.loaders.ToyLoader

ParameterTypeDefaultWhat it does
directorystr | None
sizesOptional
triggersList'aa', 'bc'
lengthint20
contaminationsint0
min_chunkint2
vocabulary_sizeint20
seedint0
detectorleakage

pyhighlights.components.leakage.LeakageDetector

ParameterTypeDefaultWhat it does
keystr'text'
normalize_keysboolTrue
detectorshortcut

pyhighlights.components.shortcuts.ShortcutDetector

ParameterTypeDefaultWhat it does
max_lengthint4
separatorstr''
tokensstr'tokens'
labelstr'label'
seedint0
permutationsint30
guiderattentiongru

pyhighlights.components.models.spp.grat.AttentionGuider

ParameterTypeDefaultWhat it does
backboneRegistrationKeybackbone[gru]
predictorRegistrationKeypredictor[mlp]
noise_sigmafloat1.0
guiderattentiontransformer

pyhighlights.components.models.spp.grat.AttentionGuider

ParameterTypeDefaultWhat it does
backboneRegistrationKeybackbone[transformer]
predictorRegistrationKeypredictor[mlp]
noise_sigmafloat1.0
lossalignmentclassification

pyhighlights.utility.losses.Loss

ParameterTypeDefaultWhat it does
namestr'alignment_classification'
lossRegistrationKeycriterion[cross_entropy]
inputsList'aligner_class_logits', 'y_true'
coefficientfloat1.0
enabledboolTrue
lossclassification

pyhighlights.utility.losses.Loss

ParameterTypeDefaultWhat it does
namestr'classification'
lossRegistrationKeycriterion[cross_entropy]
inputsList'class_logits', 'y_true'
coefficientfloat1.0
enabledboolTrue
lossclassificationcomplement

pyhighlights.utility.losses.Loss

ParameterTypeDefaultWhat it does
namestr'complement_classification'
lossRegistrationKeycriterion[cross_entropy]
inputsList'complement_class_logits', 'y_true'
coefficientfloat1.0
enabledboolTrue
lossclassificationfull

pyhighlights.utility.losses.Loss

ParameterTypeDefaultWhat it does
namestr'full_classification'
lossRegistrationKeycriterion[cross_entropy]
inputsList'full_class_logits', 'y_true'
coefficientfloat1.0
enabledboolTrue
losscontiguity

pyhighlights.utility.losses.Loss

ParameterTypeDefaultWhat it does
namestr'contiguity'
lossRegistrationKeycriterion[contiguity]
inputsList'highlight_mask', 'mask'
coefficientfloat2.0
enabledboolTrue
lossdiscrepancy

pyhighlights.utility.losses.Loss

ParameterTypeDefaultWhat it does
namestr'discrepancy'
lossRegistrationKeycriterion[kl_div]
inputsList'class_logits', 'full_class_logits'
coefficientfloat1.0
enabledboolTrue
lossdiscrepancyremaining

pyhighlights.utility.losses.Loss

ParameterTypeDefaultWhat it does
namestr'remaining_discrepancy'
lossRegistrationKeycriterion[kl_div]
inputsList'complement_class_logits', 'full_class_logits'
coefficientfloat-1.0
enabledboolTrue
lossguide

pyhighlights.utility.losses.Loss

ParameterTypeDefaultWhat it does
namestr'guide'
lossRegistrationKeycriterion[masked_bce]
inputsList'selection_logits', 'guide_target', 'mask'
coefficientfloat1.0
enabledboolTrue
losshighlight

pyhighlights.utility.losses.Loss

ParameterTypeDefaultWhat it does
namestr'highlight'
lossRegistrationKeycriterion[masked_cross_entropy]
inputsList'highlight_logits', 'highlight_true', 'mask'
coefficientfloat1.0
enabledboolTrue
lossjsd

pyhighlights.utility.losses.Loss

ParameterTypeDefaultWhat it does
namestr'jsd'
lossRegistrationKeycriterion[js_div]
inputsList'class_logits', 'guider_class_logits'
coefficientfloat1.0
enabledboolTrue
lossknowledge

pyhighlights.utility.losses.Loss

ParameterTypeDefaultWhat it does
namestr'knowledge'
lossRegistrationKeycriterion[masked_cross_entropy]
inputsList'knowledge_logits', 'knowledge_true', 'knowledge_valid'
coefficientfloat1.0
enabledboolTrue
lossknowledgesparsity

pyhighlights.utility.losses.Loss

ParameterTypeDefaultWhat it does
namestr'knowledge_sparsity'
lossRegistrationKeycriterion[sparsity]
inputsList'knowledge_mask', 'knowledge_valid'
coefficientfloat1.0
enabledboolTrue
lossknowledgesupervised

pyhighlights.utility.losses.Loss

ParameterTypeDefaultWhat it does
namestr'knowledge'
lossRegistrationKeycriterion[knowledge, masked_bce]
inputsList'knowledge_score', 'knowledge_true', 'knowledge_valid'
coefficientfloat1.0
enabledboolTrue
losssparsity

pyhighlights.utility.losses.Loss

ParameterTypeDefaultWhat it does
namestr'sparsity'
lossRegistrationKeycriterion[sparsity]
inputsList'highlight_mask', 'mask'
coefficientfloat1.0
enabledboolTrue
metricaccuracy

pyhighlights.utility.metrics.BoundMetric

ParameterTypeDefaultWhat it does
namestr'accuracy'
metricRegistrationKeytorchmetric[accuracy]
inputsSequence'class_logits', 'y_true'
metricaccuracymulticlass

pyhighlights.utility.metrics.BoundMetric

ParameterTypeDefaultWhat it does
namestr'accuracy'
metricRegistrationKeytorchmetric[accuracy, multiclass]
inputsSequence'class_logits', 'y_true'
metricclassf1

pyhighlights.utility.metrics.BoundMetric

ParameterTypeDefaultWhat it does
namestr'f1'
metricRegistrationKeytorchmetric[class, f1]
inputsSequence'class_logits', 'y_true'
metricempty_set

pyhighlights.utility.metrics.BoundMetric

ParameterTypeDefaultWhat it does
namestr'empty_set_accuracy'
metricRegistrationKeytorchmetric[empty_set]
inputsSequence'knowledge_mask', 'knowledge_true'
metricexact_set

pyhighlights.utility.metrics.BoundMetric

ParameterTypeDefaultWhat it does
namestr'exact_set_match'
metricRegistrationKeytorchmetric[exact_set]
inputsSequence'knowledge_mask', 'knowledge_true'
metricf1

pyhighlights.utility.metrics.BoundMetric

ParameterTypeDefaultWhat it does
namestr'f1'
metricRegistrationKeytorchmetric[f1]
inputsSequence'class_logits', 'y_true'
metricf1highlight

pyhighlights.utility.metrics.BoundMetric

ParameterTypeDefaultWhat it does
namestr'highlight_f1'
metricRegistrationKeytorchmetric[f1, highlight]
inputsSequence'highlight_mask', 'highlight_true'
metricf1link

pyhighlights.utility.metrics.BoundMetric

ParameterTypeDefaultWhat it does
namestr'link_f1'
metricRegistrationKeytorchmetric[f1, link]
inputsSequence'knowledge_mask', 'knowledge_true'
metricf1linkmacro

pyhighlights.utility.metrics.BoundMetric

ParameterTypeDefaultWhat it does
namestr'link_macro_f1'
metricRegistrationKeytorchmetric[f1, link, macro]
inputsSequence'knowledge_mask', 'knowledge_true'
metricf1multiclass

pyhighlights.utility.metrics.BoundMetric

ParameterTypeDefaultWhat it does
namestr'f1'
metricRegistrationKeytorchmetric[f1, multiclass]
inputsSequence'class_logits', 'y_true'
metrichighlightiou

pyhighlights.utility.metrics.BoundMetric

ParameterTypeDefaultWhat it does
namestr'highlight_iou'
metricRegistrationKeytorchmetric[highlight, iou]
inputsSequence'highlight_mask', 'highlight_true'
metrichighlightprecision

pyhighlights.utility.metrics.BoundMetric

ParameterTypeDefaultWhat it does
namestr'highlight_precision'
metricRegistrationKeytorchmetric[highlight, precision]
inputsSequence'highlight_mask', 'highlight_true'
metrichighlightrecall

pyhighlights.utility.metrics.BoundMetric

ParameterTypeDefaultWhat it does
namestr'highlight_recall'
metricRegistrationKeytorchmetric[highlight, recall]
inputsSequence'highlight_mask', 'highlight_true'
metricselection_rate

pyhighlights.utility.metrics.BoundMetric

ParameterTypeDefaultWhat it does
namestr'selection_rate'
metricRegistrationKeytorchmetric[selection_rate]
inputsSequence'highlight_mask', 'mask'
metricselection_size

pyhighlights.utility.metrics.BoundMetric

ParameterTypeDefaultWhat it does
namestr'selection_size'
metricRegistrationKeytorchmetric[selection_size]
inputsSequence'highlight_mask', 'mask'
metricselection_spans

pyhighlights.utility.metrics.BoundMetric

ParameterTypeDefaultWhat it does
namestr'selection_spans'
metricRegistrationKeytorchmetric[selection_spans]
inputsSequence'highlight_mask', 'mask'
modeldargru

pyhighlights.components.models.spp.dar.DAR

ParameterTypeDefaultWhat it does
namestr'dar'
selector_backbonesRegistrationKeybackbone[gru]
selectorsRegistrationKeyselector[mlp]
predictorRegistrationKeypredictor[mlp]
predictor_backboneRegistrationKeybackbone[gru]
temperaturefloat1.0
compactboolFalse
select_overstr'word'
encoder_lrfloat | None
optimizerRegistrationKeyoptimizer[adam]
train_metricsOptional
val_metricsOptional
test_metricsOptional
lossesListloss[classification], loss[sparsity], loss[contiguity]
aligner_backboneRegistrationKeybackbone[gru]
aligner_lossRegistrationKeyloss[alignment, classification]
pretrain_epochsint20
modeldartransformer

pyhighlights.components.models.spp.dar.DAR

ParameterTypeDefaultWhat it does
namestr'dar'
selector_backbonesRegistrationKeybackbone[transformer]
selectorsRegistrationKeyselector[mlp]
predictorRegistrationKeypredictor[mlp]
predictor_backboneRegistrationKeybackbone[transformer]
temperaturefloat1.0
compactboolFalse
select_overstr'word'
encoder_lrfloat | None
optimizerRegistrationKeyoptimizer[adam]
train_metricsOptional
val_metricsOptional
test_metricsOptional
lossesListloss[classification], loss[sparsity], loss[contiguity]
aligner_backboneRegistrationKeybackbone[transformer]
aligner_lossRegistrationKeyloss[alignment, classification]
pretrain_epochsint20
modeldrgru

pyhighlights.components.models.spp.dr.DR

ParameterTypeDefaultWhat it does
namestr'dr'
selector_backbonesRegistrationKeybackbone[gru]
selectorsRegistrationKeyselector[mlp]
predictorRegistrationKeypredictor[mlp]
predictor_backboneRegistrationKeybackbone[gru]
temperaturefloat1.0
compactboolFalse
select_overstr'word'
encoder_lrfloat | None
optimizerRegistrationKeyoptimizer[adam]
train_metricsOptional
val_metricsOptional
test_metricsOptional
lossesListloss[classification], loss[sparsity], loss[contiguity]
scale_floorfloat0.05
modeldrtransformer

pyhighlights.components.models.spp.dr.DR

ParameterTypeDefaultWhat it does
namestr'dr'
selector_backbonesRegistrationKeybackbone[transformer]
selectorsRegistrationKeyselector[mlp]
predictorRegistrationKeypredictor[mlp]
predictor_backboneRegistrationKeybackbone[transformer]
temperaturefloat1.0
compactboolFalse
select_overstr'word'
encoder_lrfloat | None
optimizerRegistrationKeyoptimizer[adam]
train_metricsOptional
val_metricsOptional
test_metricsOptional
lossesListloss[classification], loss[sparsity], loss[contiguity]
scale_floorfloat0.05
modelfrgru

pyhighlights.components.models.spp.fr.FR

ParameterTypeDefaultWhat it does
namestr'fr'
selector_backbonesRegistrationKeybackbone[gru]
selectorsRegistrationKeyselector[mlp]
predictorRegistrationKeypredictor[mlp]
predictor_backboneOptional
temperaturefloat1.0
compactboolFalse
select_overstr'word'
encoder_lrfloat | None
optimizerRegistrationKeyoptimizer[adam]
train_metricsOptional
val_metricsOptional
test_metricsOptional
lossesListloss[classification], loss[sparsity], loss[contiguity]
modelfrtransformer

pyhighlights.components.models.spp.fr.FR

ParameterTypeDefaultWhat it does
namestr'fr'
selector_backbonesRegistrationKeybackbone[transformer]
selectorsRegistrationKeyselector[mlp]
predictorRegistrationKeypredictor[mlp]
predictor_backboneOptional
temperaturefloat1.0
compactboolFalse
select_overstr'word'
encoder_lrfloat | None
optimizerRegistrationKeyoptimizer[adam]
train_metricsOptional
val_metricsOptional
test_metricsOptional
lossesListloss[classification], loss[sparsity], loss[contiguity]
modelgensppgru

pyhighlights.components.models.spp.genspp.GenSPP

ParameterTypeDefaultWhat it does
namestr'genspp'
selector_backbonesRegistrationKeybackbone[genspp, gru]
selectorsRegistrationKeyselector[mlp]
predictorRegistrationKeypredictor[mlp]
predictor_backboneRegistrationKeybackbone[genspp, gru]
temperaturefloat1.0
compactboolFalse
select_overstr'word'
encoder_lrfloat | None
optimizerRegistrationKeyoptimizer[adam, genspp]
train_metricsOptional
val_metricsOptional
test_metricsOptional
lossesListloss[classification]
modelgenspptransformer

pyhighlights.components.models.spp.genspp.GenSPP

ParameterTypeDefaultWhat it does
namestr'genspp'
selector_backbonesRegistrationKeybackbone[genspp, transformer]
selectorsRegistrationKeyselector[mlp]
predictorRegistrationKeypredictor[mlp]
predictor_backboneRegistrationKeybackbone[genspp, transformer]
temperaturefloat1.0
compactboolFalse
select_overstr'word'
encoder_lrfloat | None
optimizerRegistrationKeyoptimizer[adam, genspp]
train_metricsOptional
val_metricsOptional
test_metricsOptional
lossesListloss[classification]
modelgratgru

pyhighlights.components.models.spp.grat.GRAT

ParameterTypeDefaultWhat it does
namestr'grat'
selector_backbonesRegistrationKeybackbone[gru]
selectorsRegistrationKeyselector[mlp]
predictorRegistrationKeypredictor[mlp]
predictor_backboneRegistrationKeybackbone[gru]
temperaturefloat1.0
compactboolFalse
select_overstr'word'
encoder_lrfloat | None
optimizerRegistrationKeyoptimizer[adam]
train_metricsOptional
val_metricsOptional
test_metricsOptional
lossesListloss[classification], loss[sparsity], loss[contiguity], loss[guide], loss[jsd]
guiderRegistrationKeyguider[attention, gru]
guider_lossesListloss[classification]
pretrain_epochsint10
guide_decayfloat0.0001
guide_lossstr'guide'
jsd_lossstr'jsd'
modelgrattransformer

pyhighlights.components.models.spp.grat.GRAT

ParameterTypeDefaultWhat it does
namestr'grat'
selector_backbonesRegistrationKeybackbone[transformer]
selectorsRegistrationKeyselector[mlp]
predictorRegistrationKeypredictor[mlp]
predictor_backboneRegistrationKeybackbone[transformer]
temperaturefloat1.0
compactboolFalse
select_overstr'word'
encoder_lrfloat | None
optimizerRegistrationKeyoptimizer[adam]
train_metricsOptional
val_metricsOptional
test_metricsOptional
lossesListloss[classification], loss[sparsity], loss[contiguity], loss[guide], loss[jsd]
guiderRegistrationKeyguider[attention, transformer]
guider_lossesListloss[classification]
pretrain_epochsint10
guide_decayfloat0.0001
guide_lossstr'guide'
jsd_lossstr'jsd'
modelgroundedgru

pyhighlights.components.models.spp.grounded.GroundedSPP

ParameterTypeDefaultWhat it does
namestr'grounded'
selector_backbonesRegistrationKeybackbone[gru]
selectorsRegistrationKeyselector[mlp]
predictorRegistrationKeypredictor[mlp]
predictor_backboneOptional
temperaturefloat1.0
compactboolFalse
select_overstr'word'
encoder_lrfloat | None
optimizerRegistrationKeyoptimizer[adam]
train_metricsOptional
val_metricsOptional
test_metricsOptional
lossesListloss[classification], loss[sparsity], loss[contiguity], loss[knowledge]
comparerRegistrationKeycomparer[entailment]
modelgroundedtransformer

pyhighlights.components.models.spp.grounded.GroundedSPP

ParameterTypeDefaultWhat it does
namestr'grounded'
selector_backbonesRegistrationKeybackbone[transformer]
selectorsRegistrationKeyselector[mlp]
predictorRegistrationKeypredictor[mlp]
predictor_backboneOptional
temperaturefloat1.0
compactboolFalse
select_overstr'word'
encoder_lrfloat | None
optimizerRegistrationKeyoptimizer[adam]
train_metricsOptional
val_metricsOptional
test_metricsOptional
lossesListloss[classification], loss[sparsity], loss[contiguity], loss[knowledge]
comparerRegistrationKeycomparer[entailment]
modelgrumcd

pyhighlights.components.models.spp.mcd.MCD

ParameterTypeDefaultWhat it does
namestr'mcd'
selector_backbonesRegistrationKeybackbone[gru]
selectorsRegistrationKeyselector[mlp]
predictorRegistrationKeypredictor[mlp]
predictor_backboneRegistrationKeybackbone[gru]
temperaturefloat1.0
compactboolFalse
select_overstr'word'
encoder_lrfloat | None
optimizerRegistrationKeyoptimizer[adam]
train_metricsOptional
val_metricsOptional
test_metricsOptional
shared_lossesListloss[sparsity], loss[contiguity]
predictor_lossesListloss[classification], loss[classification, full]
generator_lossesListloss[discrepancy]
modelgrumgr

pyhighlights.components.models.spp.mgr.MGR

ParameterTypeDefaultWhat it does
namestr'mgr'
selector_backbonesListbackbone[gru], backbone[gru], backbone[gru]
selectorsListselector[mlp], selector[mlp], selector[mlp]
predictorRegistrationKeypredictor[mlp]
predictor_backboneRegistrationKeybackbone[gru]
temperaturefloat1.0
compactboolFalse
select_overstr'word'
encoder_lrfloat | None
optimizerRegistrationKeyoptimizer[adam]
train_metricsOptional
val_metricsOptional
test_metricsOptional
lossesListloss[classification], loss[sparsity], loss[contiguity]
inference_headint0
loss_reductionLiteral'sum'
modelgrumrd

pyhighlights.components.models.spp.mrd.MRD

ParameterTypeDefaultWhat it does
namestr'mrd'
selector_backbonesRegistrationKeybackbone[gru]
selectorsRegistrationKeyselector[mlp]
predictorRegistrationKeypredictor[mlp]
predictor_backboneRegistrationKeybackbone[gru]
temperaturefloat1.0
compactboolFalse
select_overstr'word'
encoder_lrfloat | None
optimizerRegistrationKeyoptimizer[adam]
train_metricsOptional
val_metricsOptional
test_metricsOptional
shared_lossesListloss[sparsity], loss[contiguity]
predictor_lossesListloss[classification, complement], loss[classification, full]
generator_lossesListloss[discrepancy, remaining]
modelmcdtransformer

pyhighlights.components.models.spp.mcd.MCD

ParameterTypeDefaultWhat it does
namestr'mcd'
selector_backbonesRegistrationKeybackbone[transformer]
selectorsRegistrationKeyselector[mlp]
predictorRegistrationKeypredictor[mlp]
predictor_backboneRegistrationKeybackbone[transformer]
temperaturefloat1.0
compactboolFalse
select_overstr'word'
encoder_lrfloat | None
optimizerRegistrationKeyoptimizer[adam]
train_metricsOptional
val_metricsOptional
test_metricsOptional
shared_lossesListloss[sparsity], loss[contiguity]
predictor_lossesListloss[classification], loss[classification, full]
generator_lossesListloss[discrepancy]
modelmgrtransformer

pyhighlights.components.models.spp.mgr.MGR

ParameterTypeDefaultWhat it does
namestr'mgr'
selector_backbonesListbackbone[transformer], backbone[transformer], backbone[transformer]
selectorsListselector[mlp], selector[mlp], selector[mlp]
predictorRegistrationKeypredictor[mlp]
predictor_backboneRegistrationKeybackbone[transformer]
temperaturefloat1.0
compactboolFalse
select_overstr'word'
encoder_lrfloat | None
optimizerRegistrationKeyoptimizer[adam]
train_metricsOptional
val_metricsOptional
test_metricsOptional
lossesListloss[classification], loss[sparsity], loss[contiguity]
inference_headint0
loss_reductionLiteral'sum'
modelmrdtransformer

pyhighlights.components.models.spp.mrd.MRD

ParameterTypeDefaultWhat it does
namestr'mrd'
selector_backbonesRegistrationKeybackbone[transformer]
selectorsRegistrationKeyselector[mlp]
predictorRegistrationKeypredictor[mlp]
predictor_backboneRegistrationKeybackbone[transformer]
temperaturefloat1.0
compactboolFalse
select_overstr'word'
encoder_lrfloat | None
optimizerRegistrationKeyoptimizer[adam]
train_metricsOptional
val_metricsOptional
test_metricsOptional
shared_lossesListloss[sparsity], loss[contiguity]
predictor_lossesListloss[classification, complement], loss[classification, full]
generator_lossesListloss[discrepancy, remaining]
optimizeradam

torch.optim.Adam

ParameterTypeDefaultWhat it does
lrfloat0.001
weight_decayfloat0.0
optimizeradamgenspp

torch.optim.Adam

ParameterTypeDefaultWhat it does
lrfloat0.01
weight_decayfloat0.0
predictormlp

pyhighlights.components.models.spp.implementations.MLPPredictor

ParameterTypeDefaultWhat it does
hidden_sizesList
num_classesint2
preprocessoraggregatorhatexplain

pyhighlights.components.preprocessors.AnnotationAggregator

ParameterTypeDefaultWhat it does
labelsSequence'hatespeech', 'normal', 'offensive'
highlightsstr'majority'
tiesstr'drop'
preprocessoraggregatorhatexplainhighlights=intersection

pyhighlights.components.preprocessors.AnnotationAggregator

ParameterTypeDefaultWhat it does
labelsSequence'hatespeech', 'normal', 'offensive'
highlightsstr'majority'
tiesstr'drop'
preprocessoraggregatorhatexplainhighlights=intersectionties=keep

pyhighlights.components.preprocessors.AnnotationAggregator

ParameterTypeDefaultWhat it does
labelsSequence'hatespeech', 'normal', 'offensive'
highlightsstr'majority'
tiesstr'drop'
preprocessoraggregatorhatexplainhighlights=union

pyhighlights.components.preprocessors.AnnotationAggregator

ParameterTypeDefaultWhat it does
labelsSequence'hatespeech', 'normal', 'offensive'
highlightsstr'majority'
tiesstr'drop'
preprocessoraggregatorhatexplainhighlights=unionties=keep

pyhighlights.components.preprocessors.AnnotationAggregator

ParameterTypeDefaultWhat it does
labelsSequence'hatespeech', 'normal', 'offensive'
highlightsstr'majority'
tiesstr'drop'
preprocessoraggregatorhatexplainties=keep

pyhighlights.components.preprocessors.AnnotationAggregator

ParameterTypeDefaultWhat it does
labelsSequence'hatespeech', 'normal', 'offensive'
highlightsstr'majority'
tiesstr'drop'
preprocessorclass_weights

pyhighlights.components.preprocessors.ClassWeights

ParameterTypeDefaultWhat it does
splitstr'train'
classesint | None
preprocessorhatexplainpipeline

pyhighlights.components.preprocessors.Pipeline

ParameterTypeDefaultWhat it does
stepsListpreprocessor[aggregator, hatexplain], preprocessor[leakage]
preprocessorknowledge_weights

pyhighlights.components.preprocessors.KnowledgeWeights

ParameterTypeDefaultWhat it does
splitstr'train'
entriesint2
preprocessorleakage

pyhighlights.components.preprocessors.LeakageRemover

ParameterTypeDefaultWhat it does
prioritySequence'test', 'val', 'train'
keystr'text'
normalize_keysboolTrue
preprocessorpipeline

pyhighlights.components.preprocessors.Pipeline

ParameterTypeDefaultWhat it does
stepsListpreprocessor[leakage]
selectormlp

pyhighlights.components.models.spp.implementations.MLPSelector

ParameterTypeDefaultWhat it does
hidden_sizesList
taskclass_weightstoyrunnable

pyhighlights.components.tasks.ClassWeightsTask

ParameterTypeDefaultWhat it does
namestr'toy-class-weights'
loaderRegistrationKeydataset[toy]
weightsRegistrationKeypreprocessor[class_weights]
preprocessorOptional
save_pathstr | None
taskgenspptoyrunnable

pyhighlights.components.tasks.GenSPPTask

ParameterTypeDefaultWhat it does
namestr'toy-genspp'
save_pathstr | None
seedsSequence42
batch_sizeint8
max_lengthint | None
vocabulary_sizeint10000
pretrained_model_cardstr | None
add_special_tokensboolTrue
embeddingsstr | None
pretrained_tokens_onlyboolTrue
vocabulary_fromstr'corpus'
requires_embeddingsboolFalse
one_hot_embeddingsint | None
callbacksListcallback[early_stopping, loss], callback[checkpoint, loss]
store_predictionsboolFalse
keep_checkpointsboolTrue
save_weights_onlyboolFalse
faithfulnessboolFalse
diagnosticsboolFalse
highlight_supervisionboolFalse
highlight_lossRegistrationKeyloss[highlight]
highlight_coefficientfloat1.0
trainer_argsDict{'accelerator': 'cpu'}
loaderRegistrationKeydataset[toy]
searchRegistrationKeytrainer[genspp, gru, toy]
preprocessorOptional
val_metricsListmetric[accuracy], metric[f1], metric[f1, highlight], metric[highlight, iou], metric[highlight, precision], metric[highlight, recall], metric[selection_rate], metric[selection_size], metric[selection_spans]
test_metricsListmetric[accuracy], metric[f1], metric[f1, highlight], metric[highlight, iou], metric[highlight, precision], metric[highlight, recall], metric[selection_rate], metric[selection_size], metric[selection_spans]
tasktoyrunnable

pyhighlights.components.tasks.SPPTask

ParameterTypeDefaultWhat it does
namestr'toy'
save_pathstr | None
seedsSequence42
batch_sizeint8
max_lengthint | None
vocabulary_sizeint10000
pretrained_model_cardstr | None
add_special_tokensboolTrue
embeddingsstr | None
pretrained_tokens_onlyboolTrue
vocabulary_fromstr'corpus'
requires_embeddingsboolFalse
one_hot_embeddingsint | None
callbacksListcallback[early_stopping, loss], callback[checkpoint, loss]
store_predictionsboolFalse
keep_checkpointsboolTrue
save_weights_onlyboolFalse
faithfulnessboolFalse
diagnosticsboolFalse
highlight_supervisionboolFalse
highlight_lossRegistrationKeyloss[highlight]
highlight_coefficientfloat1.0
trainer_argsDict{'accelerator': 'cpu', 'max_epochs': 2}
loaderRegistrationKeydataset[toy]
modelRegistrationKeymodel[fr, gru]
preprocessorOptional
train_metricsListmetric[accuracy], metric[f1], metric[f1, highlight], metric[highlight, iou], metric[highlight, precision], metric[highlight, recall], metric[selection_rate], metric[selection_size], metric[selection_spans]
val_metricsListmetric[accuracy], metric[f1], metric[f1, highlight], metric[highlight, iou], metric[highlight, precision], metric[highlight, recall], metric[selection_rate], metric[selection_size], metric[selection_spans]
test_metricsListmetric[accuracy], metric[f1], metric[f1, highlight], metric[highlight, iou], metric[highlight, precision], metric[highlight, recall], metric[selection_rate], metric[selection_size], metric[selection_spans]
torchmetricaccuracy

torchmetrics.Accuracy

ParameterTypeDefaultWhat it does
taskstr'multiclass'
num_classesint2
averagestr'micro'
torchmetricaccuracymulticlass

torchmetrics.Accuracy

ParameterTypeDefaultWhat it does
taskstr'multiclass'
num_classesint3
averagestr'micro'
torchmetricclassf1

pyhighlights.utility.metrics.ClassF1Score

ParameterTypeDefaultWhat it does
pos_labelint1
num_classesint2
torchmetricempty_set

pyhighlights.utility.metrics.EmptySetAccuracy

ParameterTypeDefaultWhat it does
thresholdfloat0.5
ignore_indexint-1
torchmetricexact_set

pyhighlights.utility.metrics.ExactSetMatch

ParameterTypeDefaultWhat it does
thresholdfloat0.5
ignore_indexint-1
torchmetricf1

torchmetrics.F1Score

ParameterTypeDefaultWhat it does
taskstr'multiclass'
num_classesint2
averagestr'macro'
torchmetricf1highlight

pyhighlights.utility.metrics.BinaryHighlightF1Score

ParameterTypeDefaultWhat it does
pos_labelint1
ignore_indexint-1
torchmetricf1link

torchmetrics.classification.MultilabelF1Score

ParameterTypeDefaultWhat it does
num_labelsint2
averagestr'micro'
ignore_indexint-1
torchmetricf1linkmacro

torchmetrics.classification.MultilabelF1Score

ParameterTypeDefaultWhat it does
num_labelsint2
averagestr'macro'
ignore_indexint-1
torchmetricf1multiclass

torchmetrics.F1Score

ParameterTypeDefaultWhat it does
taskstr'multiclass'
num_classesint3
averagestr'macro'
torchmetrichighlightiou

pyhighlights.utility.metrics.BinaryHighlightIoU

ParameterTypeDefaultWhat it does
pos_labelint1
ignore_indexint-1
torchmetrichighlightprecision

pyhighlights.utility.metrics.BinaryHighlightPrecision

ParameterTypeDefaultWhat it does
pos_labelint1
ignore_indexint-1
torchmetrichighlightrecall

pyhighlights.utility.metrics.BinaryHighlightRecall

ParameterTypeDefaultWhat it does
pos_labelint1
ignore_indexint-1
torchmetricselection_rate

pyhighlights.utility.metrics.SelectionRate

No parameters.

torchmetricselection_size

pyhighlights.utility.metrics.SelectionSize

No parameters.

torchmetricselection_spans

pyhighlights.utility.metrics.SelectionSpans

No parameters.

trainergensppgru

pyhighlights.components.models.spp.genspp.GenSPPTrainer

ParameterTypeDefaultWhat it does
modelRegistrationKeymodel[genspp, gru]
n_generationsint100
population_sizeint50
selection_ratefloat0.5
mutation_probabilityfloat1.0
mutation_stdfloat0.05
predictor_epochsint3
task_loss_limitfloat0.1
stop_thresholdfloat0.01
seedint | None
devicesList'cpu'
trainergensppgrutoy

pyhighlights.components.models.spp.genspp.GenSPPTrainer

ParameterTypeDefaultWhat it does
modelRegistrationKeymodel[genspp, gru]
n_generationsint1
population_sizeint2
selection_ratefloat0.5
mutation_probabilityfloat1.0
mutation_stdfloat0.05
predictor_epochsint1
task_loss_limitfloat10.0
stop_thresholdfloat0.01
seedint | None
devicesList'cpu'
trainergenspptransformer

pyhighlights.components.models.spp.genspp.GenSPPTrainer

ParameterTypeDefaultWhat it does
modelRegistrationKeymodel[genspp, transformer]
n_generationsint100
population_sizeint50
selection_ratefloat0.5
mutation_probabilityfloat1.0
mutation_stdfloat0.05
predictor_epochsint3
task_loss_limitfloat0.1
stop_thresholdfloat0.01
seedint | None
devicesList'cpu'

Reading a card#

A key is a name and a set of tags in a namespace, and the pair is what Registry.from_key resolves. A card titled model with the tags fr and gru is the key pyhighlights.configurations.keys.GRU_FR, and the constant is what a script should import rather than the string, since a typo in a constant is an ImportError and a typo in a string is a key that does not exist.

A parameter whose default reads as name[tag, tag] is itself a registration key, which is how a configuration names another one. Writing registrations of your own, in a package beside the library rather than inside it, is Writing your own method. A card marked runnable carries a run_method, so cmn-run offers it from the command line.

API#

Registration keys every pyhighlights configuration is addressed by.

pyhighlights.configurations.keys.LOSS_EARLY_STOPPING = RegistrationKey(name=callback, namespace=pyhighlights, tags=frozenset({'loss', 'early_stopping'}), description=None)#

What a run is monitored by. A study configures these the way it configures its losses: early stopping and checkpointing on the same quantity, and a criterion that turns two quantities into the one they have to agree on.

Field sets shared by model configurations.

Nothing here registers: every model would otherwise inherit a registration it never asked for. The fields each model overrides, such as the backbones, the losses and the optimizer, are named once here and pinned per model in fr, mgr, mcd, mrd, dr, dar, grat and genspp.

Three bases rather than one, because the architectures disagree about what a loss is: SPPModelConfig scores a flat list, PhasedSPPModelConfig scores three lists tied to training phases, and SPPShapeConfig is what the two share.

class pyhighlights.configurations.base.PhasedSPPModelConfig(**data)[source]#

An SPP model whose criteria belong to a training phase.

MCD and MRD are the same shape and differ only in which criteria go in which list, so the lists are declared once. Both refuse supervise_highlights for the same reason: a supervision loss appended to a flat list would be dropped before the first batch, and the phase it belongs to has to be named.

Parameters:
  • name (str)

  • selector_backbones (RegistrationKey[SPPBackbone])

  • selectors (RegistrationKey[SPPSelector])

  • predictor (RegistrationKey[SPPPredictor])

  • predictor_backbone (RegistrationKey[SPPBackbone] | None)

  • temperature (float)

  • compact (bool)

  • select_over (str)

  • encoder_lr (float | None)

  • optimizer (RegistrationKey[Optimizer])

  • train_metrics (List[RegistrationKey[BoundMetric]] | None)

  • val_metrics (List[RegistrationKey[BoundMetric]] | None)

  • test_metrics (List[RegistrationKey[BoundMetric]] | None)

  • shared_losses (List[RegistrationKey[Loss]])

  • predictor_losses (List[RegistrationKey[Loss]])

  • generator_losses (List[RegistrationKey[Loss]])

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.base.SPPModelConfig(**data)[source]#

An SPP model scoring one flat list of criteria.

Parameters:
  • name (str)

  • selector_backbones (RegistrationKey[SPPBackbone])

  • selectors (RegistrationKey[SPPSelector])

  • predictor (RegistrationKey[SPPPredictor])

  • predictor_backbone (RegistrationKey[SPPBackbone] | None)

  • temperature (float)

  • compact (bool)

  • select_over (str)

  • encoder_lr (float | None)

  • optimizer (RegistrationKey[Optimizer])

  • train_metrics (List[RegistrationKey[BoundMetric]] | None)

  • val_metrics (List[RegistrationKey[BoundMetric]] | None)

  • test_metrics (List[RegistrationKey[BoundMetric]] | None)

  • losses (List[RegistrationKey[Loss]])

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.base.SPPShapeConfig(**data)[source]#

The parts every SPP model has, without saying what it optimises.

Split from SPPModelConfig because not every architecture has a flat losses list: MCD and MRD score their criteria per training phase and take three lists instead. The shape lives in one place, so a field added here reaches every model, and each model says only what differs.

Parameters:
  • name (str)

  • selector_backbones (RegistrationKey[SPPBackbone])

  • selectors (RegistrationKey[SPPSelector])

  • predictor (RegistrationKey[SPPPredictor])

  • predictor_backbone (RegistrationKey[SPPBackbone] | None)

  • temperature (float)

  • compact (bool)

  • select_over (str)

  • encoder_lr (float | None)

  • optimizer (RegistrationKey[Optimizer])

  • train_metrics (List[RegistrationKey[BoundMetric]] | None)

  • val_metrics (List[RegistrationKey[BoundMetric]] | None)

  • test_metrics (List[RegistrationKey[BoundMetric]] | None)

compact: bool#

Gather the kept positions into a shorter sequence instead of zeroing the dropped ones in place. Off by default, because it changes what the predictor is trained on rather than fixing what it reads.

Zeroing leaves a dropped position in the sequence, and the predictor can read the fact of it. A recurrent encoder steps over it, and a transformer gives the next kept token a different position embedding. So a selector can signal a label through the shape of the mask rather than the words in it. Compaction closes both mechanisms; it does not close the count, which a shorter sequence still shows.

See SPP.compacted for what it costs.

encoder_lr: float | None#

One rate for the encoders, another for everything above them. None trains the whole model at the optimizer’s own rate, which suits a GRU over a frozen table, where nothing pretrained is fine-tuned. Set it when a pretrained encoder is being fine-tuned: one rate cannot serve both a transformer and a selector initialized from scratch.

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

select_over: str#

the unit the corpus annotates, that a sparsity target is a fraction of, and that an export shows. Words are the same unit whichever backbone read the text. "subtoken" selects over the backbone’s own tokens instead.

Type:

A selection is made over words

Criterion registrations and the bindings that feed them named fields.

class pyhighlights.configurations.losses.AlignmentClassificationLossConfig(**data)[source]#

What the label looks like to a module that only ever read full text.

DAR scores its aligner with this twice: while that module is pretrained on the full input, and afterwards on the highlight, where the term is the generator’s alone.

Parameters:
  • name (str)

  • loss (RegistrationKey[Module])

  • inputs (List[str])

  • coefficient (float)

  • enabled (bool)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.losses.ComplementClassificationLossConfig(**data)[source]#

What the label looks like from everything the highlight left behind.

Parameters:
  • name (str)

  • loss (RegistrationKey[Module])

  • inputs (List[str])

  • coefficient (float)

  • enabled (bool)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.losses.CrossEntropyConfig(**data)[source]#

Cross entropy, weighted per class where a corpus needs it.

Parameters:

weight (List[float] | None)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

weight: List[float] | None#

One weight per class, or nothing for an unweighted loss. A heavily imbalanced corpus is answered correctly by a model that never predicts the rare class, so the numbers such a corpus needs are part of its configuration rather than a detail of its training.

class pyhighlights.configurations.losses.KnowledgeBCEConfig(**data)[source]#

A binary criterion carrying one positive weight per knowledge entry.

Parameters:
  • ignore_index (int)

  • pos_weight (List[float] | None)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

pos_weight: List[float] | None#

One factor per knowledge entry, multiplying the cost of missing a positive there. ClassWeightsTask with KNOWLEDGE_WEIGHTS computes them from a split and writes them to its results, and a study copies them here.

class pyhighlights.configurations.losses.KnowledgeLossConfig(**data)[source]#

Which knowledge base entries explain this example, against the gold links.

The one place in a grounded run where a highlight-side claim meets a gold standard. knowledge_true is -1 on an example the corpus does not annotate and 0 where it annotates that an entry does not apply, so the criterion skips the first and scores the second: an empty knowledge set is an answer, not a missing label.

Parameters:
  • name (str)

  • loss (RegistrationKey[Module])

  • inputs (List[str])

  • coefficient (float)

  • enabled (bool)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.losses.KnowledgeSparsityLossConfig(**data)[source]#

How much of the knowledge base an example is allowed to instantiate.

The same criterion the token axis uses, bound to the knowledge axis: it reads the field names it is given and does not care which axis they are.

Parameters:
  • name (str)

  • loss (RegistrationKey[Module])

  • inputs (List[str])

  • coefficient (float)

  • enabled (bool)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.losses.KnowledgeSupervisionLossConfig(**data)[source]#

Told which entries to name, with a weight per entry.

The same supervision as the term below and a different criterion, for one reason: two classes under a cross entropy carry a single positive weight, and the knowledge axis needs one per entry. The entry that decides a case is frequently the rare one, and a shared weight cannot tell it from the entry that fires on half the corpus.

It scores knowledge_score, the difference of the comparer’s two logits. The gate is taken from the same quantity, so the term and the gate cannot disagree.

Parameters:
  • name (str)

  • loss (RegistrationKey[Module])

  • inputs (List[str])

  • coefficient (float)

  • enabled (bool)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.losses.LossConfig(**data)[source]#

Binds a criterion to the namespace fields it scores.

Parameters:
  • name (str)

  • loss (RegistrationKey[Module])

  • inputs (List[str])

  • coefficient (float)

  • enabled (bool)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.losses.MaskedBCEConfig(**data)[source]#

One independent decision per position of the axis being scored.

Parameters:
  • ignore_index (int)

  • pos_weight (List[float] | None)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

pos_weight: List[float] | None#

One factor per position, multiplying the cost of missing a positive there.

class pyhighlights.configurations.losses.MaskedCrossEntropyConfig(**data)[source]#

Cross entropy over valid, labelled positions of any axis.

Parameters:

weight (List[float] | None)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

weight: List[float] | None#

One weight per class, for an axis whose classes are imbalanced. The knowledge axis typically is, since an example instantiates few entries of a base and many examples instantiate none.

class pyhighlights.configurations.losses.RemainingDiscrepancyLossConfig(**data)[source]#

MRD’s criterion: the complement should stop looking like the whole input.

Scored between the complement and the full input rather than between the highlight and the full input, and maximized. It is the one term in the library a model wants large, which is what the negative coefficient says. A coefficient is otherwise non-negative, since a loss is otherwise something to minimize.

Parameters:
  • name (str)

  • loss (RegistrationKey[Module])

  • inputs (List[str])

  • coefficient (float)

  • enabled (bool)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

Optimizer registrations.

Corpus loader registrations.

class pyhighlights.configurations.datasets.BeerConfig(**data)[source]#

Beer aspects; one variant per aspect.

Parameters:
  • directory (str | None)

  • splits (Dict[str, str] | None)

  • url (str)

  • sha256 (str | None)

  • task (str)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.datasets.HateXplainConfig(**data)[source]#

Hate-speech posts, every annotator judgement kept.

Parameters:
  • directory (str | None)

  • url (str)

  • divisions_url (str)

  • sha256 (str | None)

  • divisions_sha256 (str | None)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

sha256: str | None#

the loader’s defaults, named again where the registered key can be read off them. Two digests because they are two downloads.

Type:

As on R2AConfig

class pyhighlights.configurations.datasets.HotelConfig(**data)[source]#

Hotel aspects; one variant per aspect.

Parameters:
  • directory (str | None)

  • splits (Dict[str, str] | None)

  • url (str)

  • sha256 (str | None)

  • task (str)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.datasets.LoaderConfig(**data)[source]#

Fields every loader shares.

Parameters:

directory (str | None)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.datasets.MoviesConfig(**data)[source]#

ERASER movies: evidence spans become highlights.

Parameters:
  • directory (str | None)

  • task (str)

  • splits (Dict[str, str] | None)

  • url (str)

  • sha256 (str | None)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

sha256: str | None#

the loader’s default, said again where the registered key can be read off it.

Type:

As on R2AConfig

class pyhighlights.configurations.datasets.R2AConfig(**data)[source]#

Fields the Beer and Hotel loaders share; the archive holds both.

Parameters:
  • directory (str | None)

  • splits (Dict[str, str] | None)

  • url (str)

  • sha256 (str | None)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

sha256: str | None#

The digest the loader defaults to, named again here because a registered key is what a run actually builds, and its download is verified against this. A fixture archive passes sha256=None explicitly.

class pyhighlights.configurations.datasets.ToyConfig(**data)[source]#

Synthetic corpus for smoke tests and demos; no download.

Parameters:
  • directory (str | None)

  • sizes (Dict[str, int] | None)

  • triggers (List[str | List[str]])

  • length (int)

  • contaminations (int)

  • min_chunk (int)

  • vocabulary_size (int)

  • seed (int)

contaminations: int#

Chunks of the patterns scattered through the filler. Without them a fragment of a pattern classifies as well as the pattern does. Off here: the registered corpus is a smoke test and a cheap one is the point.

min_chunk: int#

The shortest piece of a pattern worth cutting out as a chunk.

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

triggers: List[str | List[str]]#

a pattern, or a list of patterns all of which have to appear. Tokens are characters here. The default is the cheap smoke-test corpus, one pattern per class; a conjunction whose patterns are shared between classes is the one no single n-gram can solve.

Type:

What each class is

vocabulary_size: int#

How many filler characters, drawn from the letters no trigger uses.

Corpus audits and preprocessing registrations.

class pyhighlights.configurations.preprocessors.ClassWeightsConfig(**data)[source]#

Reads one split’s class frequencies; changes no row.

Parameters:
  • split (str)

  • classes (int | None)

classes: int | None#

Left unset, the number of classes is the largest label seen plus one. Set it wherever a split might not hold every class.

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.preprocessors.HateXplainAggregatorConfig(**data)[source]#

Three annotators to one label and one highlight vector.

Parameters:
  • labels (Sequence[str])

  • highlights (str)

  • ties (str)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.preprocessors.HateXplainPipelineConfig(**data)[source]#

Aggregate the annotators first: repairing leakage before the tie-drop would measure overlap over rows the aggregation then removes.

Parameters:

steps (List[RegistrationKey[Preprocessor]])

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.preprocessors.KnowledgeWeightsConfig(**data)[source]#

Reads one split’s example-to-entry links; changes no row.

ClassWeightsTask runs it unchanged, so the weights land in a results.json a configuration can be copied from rather than being computed inside training and left nowhere.

Parameters:
  • split (str)

  • entries (int)

entries: int#

the largest index a split happens to use is not how many entries there are.

Type:

The size of the knowledge base, required rather than inferred

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.preprocessors.LeakageDetectorConfig(**data)[source]#

Reports what splits share, and refuses a corpus that shares anything.

Parameters:
  • key (str)

  • normalize_keys (bool)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.preprocessors.LeakageRemoverConfig(**data)[source]#

Drops shared and repeated rows, protecting the annotated split first.

Parameters:
  • priority (Sequence[str])

  • key (str)

  • normalize_keys (bool)

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.preprocessors.PipelineConfig(**data)[source]#

Preprocessors run in order, each over what the last returned.

Parameters:

steps (List[RegistrationKey[Preprocessor]])

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.preprocessors.ShortcutDetectorConfig(**data)[source]#

Reports what predicts the label, and refuses a corpus solvable without its evidence.

Parameters:
  • max_length (int)

  • separator (str)

  • tokens (str)

  • label (str)

  • seed (int)

  • permutations (int)

max_length: int#

Longest n-gram scanned. Every shorter one is scanned too.

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

permutations: int#

Label shuffles the refusal threshold is the best of, which makes the check a permutation test at a level of about 1 / permutations.

separator: str#

empty for a corpus of characters, a space for a corpus of words.

Type:

How an n-gram’s tokens are joined for display