Configurations#
pyhighlights.configurations.keys holds every RegistrationKey the other configuration modules register against, so model keys such as GRU_GENSPP and external trainer keys such as GRU_GENSPP_TRAINER are imported from one place.
Every registration#
The cards below are built from the registry itself, once per documentation build, so a key here is a key that exists and its defaults are the ones a run would get.
Each card carries the registration name, its tags as coloured chips, the component it builds, and every parameter with its type and default.
The search box filters by any of those, so transformer mcd narrows to the MCD registrations on a Transformer backbone and loader narrows to the corpora.
123 registrations
analyzerhighlightpositionrunnable
pyhighlights.components.analyzers.HighlightPositionAnalyzer
| Parameter | Type | Default | What it does |
|---|---|---|---|
directory | str | None | | |
pattern | str | 'predictions-seed=*.pkl' | |
bins | int | 10 | |
absolute | bool | False | |
latest | bool | True |
analyzerlabel-studiorunnable
pyhighlights.components.analyzers.LabelStudioExporter
| Parameter | Type | Default | What it does |
|---|---|---|---|
directory | str | None | | |
pattern | str | 'predictions-seed=*.pkl' | |
split | str | 'test' | |
latest | bool | True | |
model_version | str | 'pyhighlights' | |
labels | Sequence | 'highlight' | |
only | int | None | | |
column | str | 'label' | |
stem | str | 'label-studio' |
analyzermetricsrunnable
pyhighlights.components.analyzers.MetricsAnalyzer
| Parameter | Type | Default | What it does |
|---|---|---|---|
directory | str | None | | |
metrics | Sequence | | |
split | str | 'test' | |
pairs | bool | False | |
latest | bool | True |
analyzerpredictionrunnable
pyhighlights.components.analyzers.PredictionAnalyzer
| Parameter | Type | Default | What it does |
|---|---|---|---|
directory | str | None | | |
pattern | str | 'predictions-seed=*.pkl' | |
split | str | 'test' | |
latest | bool | True |
backbonefrozentransformer
pyhighlights.components.models.spp.implementations.TransformerBackbone
| Parameter | Type | Default | What it does |
|---|---|---|---|
pretrained_model_card | str | 'distilbert-base-uncased' | |
num_features | int | None | | |
freeze_transformer | bool | True |
backbonegensppgru
pyhighlights.components.models.spp.implementations.GRUBackbone
| Parameter | Type | Default | What it does |
|---|---|---|---|
vocab_size | int | 10000 | |
embedding_dim | int | 128 | |
hidden_size | int | 16 | |
freeze_embeddings | bool | True | |
num_layers | int | 1 | |
bidirectional | bool | False | |
dropout_rate | float | 0.0 |
backbonegenspptransformer
pyhighlights.components.models.spp.implementations.TransformerBackbone
| Parameter | Type | Default | What it does |
|---|---|---|---|
pretrained_model_card | str | 'distilbert-base-uncased' | |
num_features | int | None | | |
freeze_transformer | bool | True |
backbonegru
pyhighlights.components.models.spp.implementations.GRUBackbone
| Parameter | Type | Default | What it does |
|---|---|---|---|
vocab_size | int | 10000 | |
embedding_dim | int | 128 | |
hidden_size | int | 128 | |
freeze_embeddings | bool | None | | |
num_layers | int | 1 | |
bidirectional | bool | True | |
dropout_rate | float | 0.0 |
backbonegrutransformer
pyhighlights.components.models.spp.implementations.StackedBackbone
| Parameter | Type | Default | What it does |
|---|---|---|---|
pretrained_model_card | str | 'distilbert-base-uncased' | |
hidden_size | int | 128 | |
num_features | int | None | | |
freeze_transformer | bool | True | |
num_layers | int | 1 | |
bidirectional | bool | True | |
dropout_rate | float | 0.0 |
backbonetransformer
pyhighlights.components.models.spp.implementations.TransformerBackbone
| Parameter | Type | Default | What it does |
|---|---|---|---|
pretrained_model_card | str | 'distilbert-base-uncased' | |
num_features | int | None | | |
freeze_transformer | bool | False |
benchmarktoyrunnable
pyhighlights.components.benchmarks.Benchmark
| Parameter | Type | Default | What it does |
|---|---|---|---|
tasks | List | task[toy] | |
name | str | 'toy-benchmark' | |
save_path | str | None | | |
strict | bool | False | |
task_args | Dict | {} |
callbackcheckpointloss
pyhighlights.components.callbacks.WarmupModelCheckpoint
| Parameter | Type | Default | What it does |
|---|---|---|---|
monitor | str | 'val_loss' | |
mode | str | 'min' |
callbackcheckpointscore
pyhighlights.components.callbacks.WarmupModelCheckpoint
| Parameter | Type | Default | What it does |
|---|---|---|---|
monitor | str | 'val_score' | |
mode | str | 'max' |
callbackearly_stoppingloss
pyhighlights.components.callbacks.WarmupEarlyStopping
| Parameter | Type | Default | What it does |
|---|---|---|---|
monitor | str | 'val_loss' | |
mode | str | 'min' | |
patience | int | 5 |
callbackearly_stoppingscore
pyhighlights.components.callbacks.WarmupEarlyStopping
| Parameter | Type | Default | What it does |
|---|---|---|---|
monitor | str | 'val_score' | |
mode | str | 'max' | |
patience | int | 5 |
callbackgeneralization_lossscore
pyhighlights.components.callbacks.GeneralizationLossScore
| Parameter | Type | Default | What it does |
|---|---|---|---|
quality | str | 'val_f1' | |
loss | str | 'val_loss' | |
coefficient | float | 2.0 | |
name | str | 'val_score' |
comparerentailment
pyhighlights.components.models.spp.grounded.EntailmentComparer
| Parameter | Type | Default | What it does |
|---|---|---|---|
hidden_sizes | List | 128 |
criterioncontiguity
pyhighlights.utility.losses.ContiguityPenalty
No parameters.
criterioncross_entropy
pyhighlights.utility.losses.CrossEntropy
| Parameter | Type | Default | What it does |
|---|---|---|---|
weight | Optional | |
criterionjs_div
pyhighlights.utility.losses.JSDiv
No parameters.
criterionkl_div
pyhighlights.utility.losses.KLDiv
No parameters.
criterionknowledgemasked_bce
pyhighlights.utility.losses.MaskedBinaryCrossEntropy
| Parameter | Type | Default | What it does |
|---|---|---|---|
ignore_index | int | -1 | |
pos_weight | Optional | |
criterionmasked_bce
pyhighlights.utility.losses.MaskedBinaryCrossEntropy
| Parameter | Type | Default | What it does |
|---|---|---|---|
ignore_index | int | -1 | |
pos_weight | Optional | |
criterionmasked_cross_entropy
pyhighlights.utility.losses.MaskedCrossEntropy
| Parameter | Type | Default | What it does |
|---|---|---|---|
weight | Optional | |
criterionsparsity
pyhighlights.utility.losses.SparsityPenalty
| Parameter | Type | Default | What it does |
|---|---|---|---|
threshold | float | 0.15 |
datasetbeer
pyhighlights.components.loaders.BeerLoader
| Parameter | Type | Default | What it does |
|---|---|---|---|
directory | str | None | | |
splits | Optional | | |
url | str | 'https://people.csail.mit.edu/yujia/files/r2a/data.zip' | |
sha256 | str | None | '23fcb4cac883ec1de86d83a7747294d7fdae10061d3803fd4c34c930e66f25de' | |
task | str | 'beer0' |
datasetbeertask=beer1
pyhighlights.components.loaders.BeerLoader
| Parameter | Type | Default | What it does |
|---|---|---|---|
directory | str | None | | |
splits | Optional | | |
url | str | 'https://people.csail.mit.edu/yujia/files/r2a/data.zip' | |
sha256 | str | None | '23fcb4cac883ec1de86d83a7747294d7fdae10061d3803fd4c34c930e66f25de' | |
task | str | 'beer0' |
datasetbeertask=beer2
pyhighlights.components.loaders.BeerLoader
| Parameter | Type | Default | What it does |
|---|---|---|---|
directory | str | None | | |
splits | Optional | | |
url | str | 'https://people.csail.mit.edu/yujia/files/r2a/data.zip' | |
sha256 | str | None | '23fcb4cac883ec1de86d83a7747294d7fdae10061d3803fd4c34c930e66f25de' | |
task | str | 'beer0' |
datasethatexplain
pyhighlights.components.loaders.HateXplainLoader
| Parameter | Type | Default | What it does |
|---|---|---|---|
directory | str | None | | |
url | str | 'https://raw.githubusercontent.com/hate-alert/HateXplain/01d742279dac941981f53806154481c0e15ee686/Data/dataset.json' | |
divisions_url | str | 'https://raw.githubusercontent.com/hate-alert/HateXplain/01d742279dac941981f53806154481c0e15ee686/Data/post_id_divisions.json' | |
sha256 | str | None | '63bb3340fee0ec469b09690d04cb68f7c187787dd8b83807f071892c084967fb' | |
divisions_sha256 | str | None | 'c2fb0d89862e7897b11ea3e9380753f15a793482b4b70ad0532dfb1212212835' |
datasethotel
pyhighlights.components.loaders.HotelLoader
| Parameter | Type | Default | What it does |
|---|---|---|---|
directory | str | None | | |
splits | Optional | | |
url | str | 'https://people.csail.mit.edu/yujia/files/r2a/data.zip' | |
sha256 | str | None | '23fcb4cac883ec1de86d83a7747294d7fdae10061d3803fd4c34c930e66f25de' | |
task | str | 'hotel_Location' |
datasethoteltask=hotel_Cleanliness
pyhighlights.components.loaders.HotelLoader
| Parameter | Type | Default | What it does |
|---|---|---|---|
directory | str | None | | |
splits | Optional | | |
url | str | 'https://people.csail.mit.edu/yujia/files/r2a/data.zip' | |
sha256 | str | None | '23fcb4cac883ec1de86d83a7747294d7fdae10061d3803fd4c34c930e66f25de' | |
task | str | 'hotel_Location' |
datasethoteltask=hotel_Service
pyhighlights.components.loaders.HotelLoader
| Parameter | Type | Default | What it does |
|---|---|---|---|
directory | str | None | | |
splits | Optional | | |
url | str | 'https://people.csail.mit.edu/yujia/files/r2a/data.zip' | |
sha256 | str | None | '23fcb4cac883ec1de86d83a7747294d7fdae10061d3803fd4c34c930e66f25de' | |
task | str | 'hotel_Location' |
datasetmovies
pyhighlights.components.loaders.MoviesLoader
| Parameter | Type | Default | What it does |
|---|---|---|---|
directory | str | None | | |
task | str | 'movies' | |
splits | Optional | | |
url | str | 'https://www.eraserbenchmark.com/zipped/{task}.tar.gz' | |
sha256 | str | None | '66e18d4e6c9df9e9f5544572b0bfe92a39673f74ecbfc3859b46cedb2f5b2dee' |
datasettoy
pyhighlights.components.loaders.ToyLoader
| Parameter | Type | Default | What it does |
|---|---|---|---|
directory | str | None | | |
sizes | Optional | | |
triggers | List | 'aa', 'bc' | |
length | int | 20 | |
contaminations | int | 0 | |
min_chunk | int | 2 | |
vocabulary_size | int | 20 | |
seed | int | 0 |
detectorleakage
pyhighlights.components.leakage.LeakageDetector
| Parameter | Type | Default | What it does |
|---|---|---|---|
key | str | 'text' | |
normalize_keys | bool | True |
detectorshortcut
pyhighlights.components.shortcuts.ShortcutDetector
| Parameter | Type | Default | What it does |
|---|---|---|---|
max_length | int | 4 | |
separator | str | '' | |
tokens | str | 'tokens' | |
label | str | 'label' | |
seed | int | 0 | |
permutations | int | 30 |
guiderattentiongru
pyhighlights.components.models.spp.grat.AttentionGuider
| Parameter | Type | Default | What it does |
|---|---|---|---|
backbone | RegistrationKey | backbone[gru] | |
predictor | RegistrationKey | predictor[mlp] | |
noise_sigma | float | 1.0 |
guiderattentiontransformer
pyhighlights.components.models.spp.grat.AttentionGuider
| Parameter | Type | Default | What it does |
|---|---|---|---|
backbone | RegistrationKey | backbone[transformer] | |
predictor | RegistrationKey | predictor[mlp] | |
noise_sigma | float | 1.0 |
lossalignmentclassification
pyhighlights.utility.losses.Loss
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'alignment_classification' | |
loss | RegistrationKey | criterion[cross_entropy] | |
inputs | List | 'aligner_class_logits', 'y_true' | |
coefficient | float | 1.0 | |
enabled | bool | True |
lossclassification
pyhighlights.utility.losses.Loss
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'classification' | |
loss | RegistrationKey | criterion[cross_entropy] | |
inputs | List | 'class_logits', 'y_true' | |
coefficient | float | 1.0 | |
enabled | bool | True |
lossclassificationcomplement
pyhighlights.utility.losses.Loss
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'complement_classification' | |
loss | RegistrationKey | criterion[cross_entropy] | |
inputs | List | 'complement_class_logits', 'y_true' | |
coefficient | float | 1.0 | |
enabled | bool | True |
lossclassificationfull
pyhighlights.utility.losses.Loss
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'full_classification' | |
loss | RegistrationKey | criterion[cross_entropy] | |
inputs | List | 'full_class_logits', 'y_true' | |
coefficient | float | 1.0 | |
enabled | bool | True |
losscontiguity
pyhighlights.utility.losses.Loss
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'contiguity' | |
loss | RegistrationKey | criterion[contiguity] | |
inputs | List | 'highlight_mask', 'mask' | |
coefficient | float | 2.0 | |
enabled | bool | True |
lossdiscrepancy
pyhighlights.utility.losses.Loss
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'discrepancy' | |
loss | RegistrationKey | criterion[kl_div] | |
inputs | List | 'class_logits', 'full_class_logits' | |
coefficient | float | 1.0 | |
enabled | bool | True |
lossdiscrepancyremaining
pyhighlights.utility.losses.Loss
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'remaining_discrepancy' | |
loss | RegistrationKey | criterion[kl_div] | |
inputs | List | 'complement_class_logits', 'full_class_logits' | |
coefficient | float | -1.0 | |
enabled | bool | True |
lossguide
pyhighlights.utility.losses.Loss
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'guide' | |
loss | RegistrationKey | criterion[masked_bce] | |
inputs | List | 'selection_logits', 'guide_target', 'mask' | |
coefficient | float | 1.0 | |
enabled | bool | True |
losshighlight
pyhighlights.utility.losses.Loss
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'highlight' | |
loss | RegistrationKey | criterion[masked_cross_entropy] | |
inputs | List | 'highlight_logits', 'highlight_true', 'mask' | |
coefficient | float | 1.0 | |
enabled | bool | True |
lossjsd
pyhighlights.utility.losses.Loss
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'jsd' | |
loss | RegistrationKey | criterion[js_div] | |
inputs | List | 'class_logits', 'guider_class_logits' | |
coefficient | float | 1.0 | |
enabled | bool | True |
lossknowledge
pyhighlights.utility.losses.Loss
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'knowledge' | |
loss | RegistrationKey | criterion[masked_cross_entropy] | |
inputs | List | 'knowledge_logits', 'knowledge_true', 'knowledge_valid' | |
coefficient | float | 1.0 | |
enabled | bool | True |
lossknowledgesparsity
pyhighlights.utility.losses.Loss
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'knowledge_sparsity' | |
loss | RegistrationKey | criterion[sparsity] | |
inputs | List | 'knowledge_mask', 'knowledge_valid' | |
coefficient | float | 1.0 | |
enabled | bool | True |
lossknowledgesupervised
pyhighlights.utility.losses.Loss
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'knowledge' | |
loss | RegistrationKey | criterion[knowledge, masked_bce] | |
inputs | List | 'knowledge_score', 'knowledge_true', 'knowledge_valid' | |
coefficient | float | 1.0 | |
enabled | bool | True |
losssparsity
pyhighlights.utility.losses.Loss
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'sparsity' | |
loss | RegistrationKey | criterion[sparsity] | |
inputs | List | 'highlight_mask', 'mask' | |
coefficient | float | 1.0 | |
enabled | bool | True |
metricaccuracy
pyhighlights.utility.metrics.BoundMetric
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'accuracy' | |
metric | RegistrationKey | torchmetric[accuracy] | |
inputs | Sequence | 'class_logits', 'y_true' |
metricaccuracymulticlass
pyhighlights.utility.metrics.BoundMetric
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'accuracy' | |
metric | RegistrationKey | torchmetric[accuracy, multiclass] | |
inputs | Sequence | 'class_logits', 'y_true' |
metricclassf1
pyhighlights.utility.metrics.BoundMetric
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'f1' | |
metric | RegistrationKey | torchmetric[class, f1] | |
inputs | Sequence | 'class_logits', 'y_true' |
metricempty_set
pyhighlights.utility.metrics.BoundMetric
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'empty_set_accuracy' | |
metric | RegistrationKey | torchmetric[empty_set] | |
inputs | Sequence | 'knowledge_mask', 'knowledge_true' |
metricexact_set
pyhighlights.utility.metrics.BoundMetric
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'exact_set_match' | |
metric | RegistrationKey | torchmetric[exact_set] | |
inputs | Sequence | 'knowledge_mask', 'knowledge_true' |
metricf1
pyhighlights.utility.metrics.BoundMetric
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'f1' | |
metric | RegistrationKey | torchmetric[f1] | |
inputs | Sequence | 'class_logits', 'y_true' |
metricf1highlight
pyhighlights.utility.metrics.BoundMetric
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'highlight_f1' | |
metric | RegistrationKey | torchmetric[f1, highlight] | |
inputs | Sequence | 'highlight_mask', 'highlight_true' |
metricf1link
pyhighlights.utility.metrics.BoundMetric
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'link_f1' | |
metric | RegistrationKey | torchmetric[f1, link] | |
inputs | Sequence | 'knowledge_mask', 'knowledge_true' |
metricf1linkmacro
pyhighlights.utility.metrics.BoundMetric
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'link_macro_f1' | |
metric | RegistrationKey | torchmetric[f1, link, macro] | |
inputs | Sequence | 'knowledge_mask', 'knowledge_true' |
metricf1multiclass
pyhighlights.utility.metrics.BoundMetric
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'f1' | |
metric | RegistrationKey | torchmetric[f1, multiclass] | |
inputs | Sequence | 'class_logits', 'y_true' |
metrichighlightiou
pyhighlights.utility.metrics.BoundMetric
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'highlight_iou' | |
metric | RegistrationKey | torchmetric[highlight, iou] | |
inputs | Sequence | 'highlight_mask', 'highlight_true' |
metrichighlightprecision
pyhighlights.utility.metrics.BoundMetric
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'highlight_precision' | |
metric | RegistrationKey | torchmetric[highlight, precision] | |
inputs | Sequence | 'highlight_mask', 'highlight_true' |
metrichighlightrecall
pyhighlights.utility.metrics.BoundMetric
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'highlight_recall' | |
metric | RegistrationKey | torchmetric[highlight, recall] | |
inputs | Sequence | 'highlight_mask', 'highlight_true' |
metricselection_rate
pyhighlights.utility.metrics.BoundMetric
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'selection_rate' | |
metric | RegistrationKey | torchmetric[selection_rate] | |
inputs | Sequence | 'highlight_mask', 'mask' |
metricselection_size
pyhighlights.utility.metrics.BoundMetric
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'selection_size' | |
metric | RegistrationKey | torchmetric[selection_size] | |
inputs | Sequence | 'highlight_mask', 'mask' |
metricselection_spans
pyhighlights.utility.metrics.BoundMetric
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'selection_spans' | |
metric | RegistrationKey | torchmetric[selection_spans] | |
inputs | Sequence | 'highlight_mask', 'mask' |
modeldargru
pyhighlights.components.models.spp.dar.DAR
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'dar' | |
selector_backbones | RegistrationKey | backbone[gru] | |
selectors | RegistrationKey | selector[mlp] | |
predictor | RegistrationKey | predictor[mlp] | |
predictor_backbone | RegistrationKey | backbone[gru] | |
temperature | float | 1.0 | |
compact | bool | False | |
select_over | str | 'word' | |
encoder_lr | float | None | | |
optimizer | RegistrationKey | optimizer[adam] | |
train_metrics | Optional | | |
val_metrics | Optional | | |
test_metrics | Optional | | |
losses | List | loss[classification], loss[sparsity], loss[contiguity] | |
aligner_backbone | RegistrationKey | backbone[gru] | |
aligner_loss | RegistrationKey | loss[alignment, classification] | |
pretrain_epochs | int | 20 |
modeldartransformer
pyhighlights.components.models.spp.dar.DAR
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'dar' | |
selector_backbones | RegistrationKey | backbone[transformer] | |
selectors | RegistrationKey | selector[mlp] | |
predictor | RegistrationKey | predictor[mlp] | |
predictor_backbone | RegistrationKey | backbone[transformer] | |
temperature | float | 1.0 | |
compact | bool | False | |
select_over | str | 'word' | |
encoder_lr | float | None | | |
optimizer | RegistrationKey | optimizer[adam] | |
train_metrics | Optional | | |
val_metrics | Optional | | |
test_metrics | Optional | | |
losses | List | loss[classification], loss[sparsity], loss[contiguity] | |
aligner_backbone | RegistrationKey | backbone[transformer] | |
aligner_loss | RegistrationKey | loss[alignment, classification] | |
pretrain_epochs | int | 20 |
modeldrgru
pyhighlights.components.models.spp.dr.DR
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'dr' | |
selector_backbones | RegistrationKey | backbone[gru] | |
selectors | RegistrationKey | selector[mlp] | |
predictor | RegistrationKey | predictor[mlp] | |
predictor_backbone | RegistrationKey | backbone[gru] | |
temperature | float | 1.0 | |
compact | bool | False | |
select_over | str | 'word' | |
encoder_lr | float | None | | |
optimizer | RegistrationKey | optimizer[adam] | |
train_metrics | Optional | | |
val_metrics | Optional | | |
test_metrics | Optional | | |
losses | List | loss[classification], loss[sparsity], loss[contiguity] | |
scale_floor | float | 0.05 |
modeldrtransformer
pyhighlights.components.models.spp.dr.DR
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'dr' | |
selector_backbones | RegistrationKey | backbone[transformer] | |
selectors | RegistrationKey | selector[mlp] | |
predictor | RegistrationKey | predictor[mlp] | |
predictor_backbone | RegistrationKey | backbone[transformer] | |
temperature | float | 1.0 | |
compact | bool | False | |
select_over | str | 'word' | |
encoder_lr | float | None | | |
optimizer | RegistrationKey | optimizer[adam] | |
train_metrics | Optional | | |
val_metrics | Optional | | |
test_metrics | Optional | | |
losses | List | loss[classification], loss[sparsity], loss[contiguity] | |
scale_floor | float | 0.05 |
modelfrgru
pyhighlights.components.models.spp.fr.FR
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'fr' | |
selector_backbones | RegistrationKey | backbone[gru] | |
selectors | RegistrationKey | selector[mlp] | |
predictor | RegistrationKey | predictor[mlp] | |
predictor_backbone | Optional | | |
temperature | float | 1.0 | |
compact | bool | False | |
select_over | str | 'word' | |
encoder_lr | float | None | | |
optimizer | RegistrationKey | optimizer[adam] | |
train_metrics | Optional | | |
val_metrics | Optional | | |
test_metrics | Optional | | |
losses | List | loss[classification], loss[sparsity], loss[contiguity] |
modelfrtransformer
pyhighlights.components.models.spp.fr.FR
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'fr' | |
selector_backbones | RegistrationKey | backbone[transformer] | |
selectors | RegistrationKey | selector[mlp] | |
predictor | RegistrationKey | predictor[mlp] | |
predictor_backbone | Optional | | |
temperature | float | 1.0 | |
compact | bool | False | |
select_over | str | 'word' | |
encoder_lr | float | None | | |
optimizer | RegistrationKey | optimizer[adam] | |
train_metrics | Optional | | |
val_metrics | Optional | | |
test_metrics | Optional | | |
losses | List | loss[classification], loss[sparsity], loss[contiguity] |
modelgensppgru
pyhighlights.components.models.spp.genspp.GenSPP
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'genspp' | |
selector_backbones | RegistrationKey | backbone[genspp, gru] | |
selectors | RegistrationKey | selector[mlp] | |
predictor | RegistrationKey | predictor[mlp] | |
predictor_backbone | RegistrationKey | backbone[genspp, gru] | |
temperature | float | 1.0 | |
compact | bool | False | |
select_over | str | 'word' | |
encoder_lr | float | None | | |
optimizer | RegistrationKey | optimizer[adam, genspp] | |
train_metrics | Optional | | |
val_metrics | Optional | | |
test_metrics | Optional | | |
losses | List | loss[classification] |
modelgenspptransformer
pyhighlights.components.models.spp.genspp.GenSPP
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'genspp' | |
selector_backbones | RegistrationKey | backbone[genspp, transformer] | |
selectors | RegistrationKey | selector[mlp] | |
predictor | RegistrationKey | predictor[mlp] | |
predictor_backbone | RegistrationKey | backbone[genspp, transformer] | |
temperature | float | 1.0 | |
compact | bool | False | |
select_over | str | 'word' | |
encoder_lr | float | None | | |
optimizer | RegistrationKey | optimizer[adam, genspp] | |
train_metrics | Optional | | |
val_metrics | Optional | | |
test_metrics | Optional | | |
losses | List | loss[classification] |
modelgratgru
pyhighlights.components.models.spp.grat.GRAT
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'grat' | |
selector_backbones | RegistrationKey | backbone[gru] | |
selectors | RegistrationKey | selector[mlp] | |
predictor | RegistrationKey | predictor[mlp] | |
predictor_backbone | RegistrationKey | backbone[gru] | |
temperature | float | 1.0 | |
compact | bool | False | |
select_over | str | 'word' | |
encoder_lr | float | None | | |
optimizer | RegistrationKey | optimizer[adam] | |
train_metrics | Optional | | |
val_metrics | Optional | | |
test_metrics | Optional | | |
losses | List | loss[classification], loss[sparsity], loss[contiguity], loss[guide], loss[jsd] | |
guider | RegistrationKey | guider[attention, gru] | |
guider_losses | List | loss[classification] | |
pretrain_epochs | int | 10 | |
guide_decay | float | 0.0001 | |
guide_loss | str | 'guide' | |
jsd_loss | str | 'jsd' |
modelgrattransformer
pyhighlights.components.models.spp.grat.GRAT
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'grat' | |
selector_backbones | RegistrationKey | backbone[transformer] | |
selectors | RegistrationKey | selector[mlp] | |
predictor | RegistrationKey | predictor[mlp] | |
predictor_backbone | RegistrationKey | backbone[transformer] | |
temperature | float | 1.0 | |
compact | bool | False | |
select_over | str | 'word' | |
encoder_lr | float | None | | |
optimizer | RegistrationKey | optimizer[adam] | |
train_metrics | Optional | | |
val_metrics | Optional | | |
test_metrics | Optional | | |
losses | List | loss[classification], loss[sparsity], loss[contiguity], loss[guide], loss[jsd] | |
guider | RegistrationKey | guider[attention, transformer] | |
guider_losses | List | loss[classification] | |
pretrain_epochs | int | 10 | |
guide_decay | float | 0.0001 | |
guide_loss | str | 'guide' | |
jsd_loss | str | 'jsd' |
modelgroundedgru
pyhighlights.components.models.spp.grounded.GroundedSPP
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'grounded' | |
selector_backbones | RegistrationKey | backbone[gru] | |
selectors | RegistrationKey | selector[mlp] | |
predictor | RegistrationKey | predictor[mlp] | |
predictor_backbone | Optional | | |
temperature | float | 1.0 | |
compact | bool | False | |
select_over | str | 'word' | |
encoder_lr | float | None | | |
optimizer | RegistrationKey | optimizer[adam] | |
train_metrics | Optional | | |
val_metrics | Optional | | |
test_metrics | Optional | | |
losses | List | loss[classification], loss[sparsity], loss[contiguity], loss[knowledge] | |
comparer | RegistrationKey | comparer[entailment] |
modelgroundedtransformer
pyhighlights.components.models.spp.grounded.GroundedSPP
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'grounded' | |
selector_backbones | RegistrationKey | backbone[transformer] | |
selectors | RegistrationKey | selector[mlp] | |
predictor | RegistrationKey | predictor[mlp] | |
predictor_backbone | Optional | | |
temperature | float | 1.0 | |
compact | bool | False | |
select_over | str | 'word' | |
encoder_lr | float | None | | |
optimizer | RegistrationKey | optimizer[adam] | |
train_metrics | Optional | | |
val_metrics | Optional | | |
test_metrics | Optional | | |
losses | List | loss[classification], loss[sparsity], loss[contiguity], loss[knowledge] | |
comparer | RegistrationKey | comparer[entailment] |
modelgrumcd
pyhighlights.components.models.spp.mcd.MCD
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'mcd' | |
selector_backbones | RegistrationKey | backbone[gru] | |
selectors | RegistrationKey | selector[mlp] | |
predictor | RegistrationKey | predictor[mlp] | |
predictor_backbone | RegistrationKey | backbone[gru] | |
temperature | float | 1.0 | |
compact | bool | False | |
select_over | str | 'word' | |
encoder_lr | float | None | | |
optimizer | RegistrationKey | optimizer[adam] | |
train_metrics | Optional | | |
val_metrics | Optional | | |
test_metrics | Optional | | |
shared_losses | List | loss[sparsity], loss[contiguity] | |
predictor_losses | List | loss[classification], loss[classification, full] | |
generator_losses | List | loss[discrepancy] |
modelgrumgr
pyhighlights.components.models.spp.mgr.MGR
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'mgr' | |
selector_backbones | List | backbone[gru], backbone[gru], backbone[gru] | |
selectors | List | selector[mlp], selector[mlp], selector[mlp] | |
predictor | RegistrationKey | predictor[mlp] | |
predictor_backbone | RegistrationKey | backbone[gru] | |
temperature | float | 1.0 | |
compact | bool | False | |
select_over | str | 'word' | |
encoder_lr | float | None | | |
optimizer | RegistrationKey | optimizer[adam] | |
train_metrics | Optional | | |
val_metrics | Optional | | |
test_metrics | Optional | | |
losses | List | loss[classification], loss[sparsity], loss[contiguity] | |
inference_head | int | 0 | |
loss_reduction | Literal | 'sum' |
modelgrumrd
pyhighlights.components.models.spp.mrd.MRD
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'mrd' | |
selector_backbones | RegistrationKey | backbone[gru] | |
selectors | RegistrationKey | selector[mlp] | |
predictor | RegistrationKey | predictor[mlp] | |
predictor_backbone | RegistrationKey | backbone[gru] | |
temperature | float | 1.0 | |
compact | bool | False | |
select_over | str | 'word' | |
encoder_lr | float | None | | |
optimizer | RegistrationKey | optimizer[adam] | |
train_metrics | Optional | | |
val_metrics | Optional | | |
test_metrics | Optional | | |
shared_losses | List | loss[sparsity], loss[contiguity] | |
predictor_losses | List | loss[classification, complement], loss[classification, full] | |
generator_losses | List | loss[discrepancy, remaining] |
modelmcdtransformer
pyhighlights.components.models.spp.mcd.MCD
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'mcd' | |
selector_backbones | RegistrationKey | backbone[transformer] | |
selectors | RegistrationKey | selector[mlp] | |
predictor | RegistrationKey | predictor[mlp] | |
predictor_backbone | RegistrationKey | backbone[transformer] | |
temperature | float | 1.0 | |
compact | bool | False | |
select_over | str | 'word' | |
encoder_lr | float | None | | |
optimizer | RegistrationKey | optimizer[adam] | |
train_metrics | Optional | | |
val_metrics | Optional | | |
test_metrics | Optional | | |
shared_losses | List | loss[sparsity], loss[contiguity] | |
predictor_losses | List | loss[classification], loss[classification, full] | |
generator_losses | List | loss[discrepancy] |
modelmgrtransformer
pyhighlights.components.models.spp.mgr.MGR
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'mgr' | |
selector_backbones | List | backbone[transformer], backbone[transformer], backbone[transformer] | |
selectors | List | selector[mlp], selector[mlp], selector[mlp] | |
predictor | RegistrationKey | predictor[mlp] | |
predictor_backbone | RegistrationKey | backbone[transformer] | |
temperature | float | 1.0 | |
compact | bool | False | |
select_over | str | 'word' | |
encoder_lr | float | None | | |
optimizer | RegistrationKey | optimizer[adam] | |
train_metrics | Optional | | |
val_metrics | Optional | | |
test_metrics | Optional | | |
losses | List | loss[classification], loss[sparsity], loss[contiguity] | |
inference_head | int | 0 | |
loss_reduction | Literal | 'sum' |
modelmrdtransformer
pyhighlights.components.models.spp.mrd.MRD
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'mrd' | |
selector_backbones | RegistrationKey | backbone[transformer] | |
selectors | RegistrationKey | selector[mlp] | |
predictor | RegistrationKey | predictor[mlp] | |
predictor_backbone | RegistrationKey | backbone[transformer] | |
temperature | float | 1.0 | |
compact | bool | False | |
select_over | str | 'word' | |
encoder_lr | float | None | | |
optimizer | RegistrationKey | optimizer[adam] | |
train_metrics | Optional | | |
val_metrics | Optional | | |
test_metrics | Optional | | |
shared_losses | List | loss[sparsity], loss[contiguity] | |
predictor_losses | List | loss[classification, complement], loss[classification, full] | |
generator_losses | List | loss[discrepancy, remaining] |
optimizeradam
torch.optim.Adam
| Parameter | Type | Default | What it does |
|---|---|---|---|
lr | float | 0.001 | |
weight_decay | float | 0.0 |
optimizeradamgenspp
torch.optim.Adam
| Parameter | Type | Default | What it does |
|---|---|---|---|
lr | float | 0.01 | |
weight_decay | float | 0.0 |
predictormlp
pyhighlights.components.models.spp.implementations.MLPPredictor
| Parameter | Type | Default | What it does |
|---|---|---|---|
hidden_sizes | List | | |
num_classes | int | 2 |
preprocessoraggregatorhatexplain
pyhighlights.components.preprocessors.AnnotationAggregator
| Parameter | Type | Default | What it does |
|---|---|---|---|
labels | Sequence | 'hatespeech', 'normal', 'offensive' | |
highlights | str | 'majority' | |
ties | str | 'drop' |
preprocessoraggregatorhatexplainhighlights=intersection
pyhighlights.components.preprocessors.AnnotationAggregator
| Parameter | Type | Default | What it does |
|---|---|---|---|
labels | Sequence | 'hatespeech', 'normal', 'offensive' | |
highlights | str | 'majority' | |
ties | str | 'drop' |
preprocessoraggregatorhatexplainhighlights=intersectionties=keep
pyhighlights.components.preprocessors.AnnotationAggregator
| Parameter | Type | Default | What it does |
|---|---|---|---|
labels | Sequence | 'hatespeech', 'normal', 'offensive' | |
highlights | str | 'majority' | |
ties | str | 'drop' |
preprocessoraggregatorhatexplainhighlights=union
pyhighlights.components.preprocessors.AnnotationAggregator
| Parameter | Type | Default | What it does |
|---|---|---|---|
labels | Sequence | 'hatespeech', 'normal', 'offensive' | |
highlights | str | 'majority' | |
ties | str | 'drop' |
preprocessoraggregatorhatexplainhighlights=unionties=keep
pyhighlights.components.preprocessors.AnnotationAggregator
| Parameter | Type | Default | What it does |
|---|---|---|---|
labels | Sequence | 'hatespeech', 'normal', 'offensive' | |
highlights | str | 'majority' | |
ties | str | 'drop' |
preprocessoraggregatorhatexplainties=keep
pyhighlights.components.preprocessors.AnnotationAggregator
| Parameter | Type | Default | What it does |
|---|---|---|---|
labels | Sequence | 'hatespeech', 'normal', 'offensive' | |
highlights | str | 'majority' | |
ties | str | 'drop' |
preprocessorclass_weights
pyhighlights.components.preprocessors.ClassWeights
| Parameter | Type | Default | What it does |
|---|---|---|---|
split | str | 'train' | |
classes | int | None | |
preprocessorhatexplainpipeline
pyhighlights.components.preprocessors.Pipeline
| Parameter | Type | Default | What it does |
|---|---|---|---|
steps | List | preprocessor[aggregator, hatexplain], preprocessor[leakage] |
preprocessorknowledge_weights
pyhighlights.components.preprocessors.KnowledgeWeights
| Parameter | Type | Default | What it does |
|---|---|---|---|
split | str | 'train' | |
entries | int | 2 |
preprocessorleakage
pyhighlights.components.preprocessors.LeakageRemover
| Parameter | Type | Default | What it does |
|---|---|---|---|
priority | Sequence | 'test', 'val', 'train' | |
key | str | 'text' | |
normalize_keys | bool | True |
preprocessorpipeline
pyhighlights.components.preprocessors.Pipeline
| Parameter | Type | Default | What it does |
|---|---|---|---|
steps | List | preprocessor[leakage] |
selectormlp
pyhighlights.components.models.spp.implementations.MLPSelector
| Parameter | Type | Default | What it does |
|---|---|---|---|
hidden_sizes | List | |
taskclass_weightstoyrunnable
pyhighlights.components.tasks.ClassWeightsTask
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'toy-class-weights' | |
loader | RegistrationKey | dataset[toy] | |
weights | RegistrationKey | preprocessor[class_weights] | |
preprocessor | Optional | | |
save_path | str | None | |
taskgenspptoyrunnable
pyhighlights.components.tasks.GenSPPTask
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'toy-genspp' | |
save_path | str | None | | |
seeds | Sequence | 42 | |
batch_size | int | 8 | |
max_length | int | None | | |
vocabulary_size | int | 10000 | |
pretrained_model_card | str | None | | |
add_special_tokens | bool | True | |
embeddings | str | None | | |
pretrained_tokens_only | bool | True | |
vocabulary_from | str | 'corpus' | |
requires_embeddings | bool | False | |
one_hot_embeddings | int | None | | |
callbacks | List | callback[early_stopping, loss], callback[checkpoint, loss] | |
store_predictions | bool | False | |
keep_checkpoints | bool | True | |
save_weights_only | bool | False | |
faithfulness | bool | False | |
diagnostics | bool | False | |
highlight_supervision | bool | False | |
highlight_loss | RegistrationKey | loss[highlight] | |
highlight_coefficient | float | 1.0 | |
trainer_args | Dict | {'accelerator': 'cpu'} | |
loader | RegistrationKey | dataset[toy] | |
search | RegistrationKey | trainer[genspp, gru, toy] | |
preprocessor | Optional | | |
val_metrics | List | metric[accuracy], metric[f1], metric[f1, highlight], metric[highlight, iou], metric[highlight, precision], metric[highlight, recall], metric[selection_rate], metric[selection_size], metric[selection_spans] | |
test_metrics | List | metric[accuracy], metric[f1], metric[f1, highlight], metric[highlight, iou], metric[highlight, precision], metric[highlight, recall], metric[selection_rate], metric[selection_size], metric[selection_spans] |
tasktoyrunnable
pyhighlights.components.tasks.SPPTask
| Parameter | Type | Default | What it does |
|---|---|---|---|
name | str | 'toy' | |
save_path | str | None | | |
seeds | Sequence | 42 | |
batch_size | int | 8 | |
max_length | int | None | | |
vocabulary_size | int | 10000 | |
pretrained_model_card | str | None | | |
add_special_tokens | bool | True | |
embeddings | str | None | | |
pretrained_tokens_only | bool | True | |
vocabulary_from | str | 'corpus' | |
requires_embeddings | bool | False | |
one_hot_embeddings | int | None | | |
callbacks | List | callback[early_stopping, loss], callback[checkpoint, loss] | |
store_predictions | bool | False | |
keep_checkpoints | bool | True | |
save_weights_only | bool | False | |
faithfulness | bool | False | |
diagnostics | bool | False | |
highlight_supervision | bool | False | |
highlight_loss | RegistrationKey | loss[highlight] | |
highlight_coefficient | float | 1.0 | |
trainer_args | Dict | {'accelerator': 'cpu', 'max_epochs': 2} | |
loader | RegistrationKey | dataset[toy] | |
model | RegistrationKey | model[fr, gru] | |
preprocessor | Optional | | |
train_metrics | List | metric[accuracy], metric[f1], metric[f1, highlight], metric[highlight, iou], metric[highlight, precision], metric[highlight, recall], metric[selection_rate], metric[selection_size], metric[selection_spans] | |
val_metrics | List | metric[accuracy], metric[f1], metric[f1, highlight], metric[highlight, iou], metric[highlight, precision], metric[highlight, recall], metric[selection_rate], metric[selection_size], metric[selection_spans] | |
test_metrics | List | metric[accuracy], metric[f1], metric[f1, highlight], metric[highlight, iou], metric[highlight, precision], metric[highlight, recall], metric[selection_rate], metric[selection_size], metric[selection_spans] |
torchmetricaccuracy
torchmetrics.Accuracy
| Parameter | Type | Default | What it does |
|---|---|---|---|
task | str | 'multiclass' | |
num_classes | int | 2 | |
average | str | 'micro' |
torchmetricaccuracymulticlass
torchmetrics.Accuracy
| Parameter | Type | Default | What it does |
|---|---|---|---|
task | str | 'multiclass' | |
num_classes | int | 3 | |
average | str | 'micro' |
torchmetricclassf1
pyhighlights.utility.metrics.ClassF1Score
| Parameter | Type | Default | What it does |
|---|---|---|---|
pos_label | int | 1 | |
num_classes | int | 2 |
torchmetricempty_set
pyhighlights.utility.metrics.EmptySetAccuracy
| Parameter | Type | Default | What it does |
|---|---|---|---|
threshold | float | 0.5 | |
ignore_index | int | -1 |
torchmetricexact_set
pyhighlights.utility.metrics.ExactSetMatch
| Parameter | Type | Default | What it does |
|---|---|---|---|
threshold | float | 0.5 | |
ignore_index | int | -1 |
torchmetricf1
torchmetrics.F1Score
| Parameter | Type | Default | What it does |
|---|---|---|---|
task | str | 'multiclass' | |
num_classes | int | 2 | |
average | str | 'macro' |
torchmetricf1highlight
pyhighlights.utility.metrics.BinaryHighlightF1Score
| Parameter | Type | Default | What it does |
|---|---|---|---|
pos_label | int | 1 | |
ignore_index | int | -1 |
torchmetricf1link
torchmetrics.classification.MultilabelF1Score
| Parameter | Type | Default | What it does |
|---|---|---|---|
num_labels | int | 2 | |
average | str | 'micro' | |
ignore_index | int | -1 |
torchmetricf1linkmacro
torchmetrics.classification.MultilabelF1Score
| Parameter | Type | Default | What it does |
|---|---|---|---|
num_labels | int | 2 | |
average | str | 'macro' | |
ignore_index | int | -1 |
torchmetricf1multiclass
torchmetrics.F1Score
| Parameter | Type | Default | What it does |
|---|---|---|---|
task | str | 'multiclass' | |
num_classes | int | 3 | |
average | str | 'macro' |
torchmetrichighlightiou
pyhighlights.utility.metrics.BinaryHighlightIoU
| Parameter | Type | Default | What it does |
|---|---|---|---|
pos_label | int | 1 | |
ignore_index | int | -1 |
torchmetrichighlightprecision
pyhighlights.utility.metrics.BinaryHighlightPrecision
| Parameter | Type | Default | What it does |
|---|---|---|---|
pos_label | int | 1 | |
ignore_index | int | -1 |
torchmetrichighlightrecall
pyhighlights.utility.metrics.BinaryHighlightRecall
| Parameter | Type | Default | What it does |
|---|---|---|---|
pos_label | int | 1 | |
ignore_index | int | -1 |
torchmetricselection_rate
pyhighlights.utility.metrics.SelectionRate
No parameters.
torchmetricselection_size
pyhighlights.utility.metrics.SelectionSize
No parameters.
torchmetricselection_spans
pyhighlights.utility.metrics.SelectionSpans
No parameters.
trainergensppgru
pyhighlights.components.models.spp.genspp.GenSPPTrainer
| Parameter | Type | Default | What it does |
|---|---|---|---|
model | RegistrationKey | model[genspp, gru] | |
n_generations | int | 100 | |
population_size | int | 50 | |
selection_rate | float | 0.5 | |
mutation_probability | float | 1.0 | |
mutation_std | float | 0.05 | |
predictor_epochs | int | 3 | |
task_loss_limit | float | 0.1 | |
stop_threshold | float | 0.01 | |
seed | int | None | | |
devices | List | 'cpu' |
trainergensppgrutoy
pyhighlights.components.models.spp.genspp.GenSPPTrainer
| Parameter | Type | Default | What it does |
|---|---|---|---|
model | RegistrationKey | model[genspp, gru] | |
n_generations | int | 1 | |
population_size | int | 2 | |
selection_rate | float | 0.5 | |
mutation_probability | float | 1.0 | |
mutation_std | float | 0.05 | |
predictor_epochs | int | 1 | |
task_loss_limit | float | 10.0 | |
stop_threshold | float | 0.01 | |
seed | int | None | | |
devices | List | 'cpu' |
trainergenspptransformer
pyhighlights.components.models.spp.genspp.GenSPPTrainer
| Parameter | Type | Default | What it does |
|---|---|---|---|
model | RegistrationKey | model[genspp, transformer] | |
n_generations | int | 100 | |
population_size | int | 50 | |
selection_rate | float | 0.5 | |
mutation_probability | float | 1.0 | |
mutation_std | float | 0.05 | |
predictor_epochs | int | 3 | |
task_loss_limit | float | 0.1 | |
stop_threshold | float | 0.01 | |
seed | int | None | | |
devices | List | 'cpu' |
Reading a card#
A key is a name and a set of tags in a namespace, and the pair is what Registry.from_key resolves.
A card titled model with the tags fr and gru is the key pyhighlights.configurations.keys.GRU_FR, and the constant is what a script should import rather than the string, since a typo in a constant is an ImportError and a typo in a string is a key that does not exist.
A parameter whose default reads as name[tag, tag] is itself a registration key, which is how a configuration names another one.
Writing registrations of your own, in a package beside the library rather than inside it, is Writing your own method.
A card marked runnable carries a run_method, so cmn-run offers it from the command line.
API#
Registration keys every pyhighlights configuration is addressed by.
- pyhighlights.configurations.keys.LOSS_EARLY_STOPPING = RegistrationKey(name=callback, namespace=pyhighlights, tags=frozenset({'loss', 'early_stopping'}), description=None)#
What a run is monitored by. A study configures these the way it configures its losses: early stopping and checkpointing on the same quantity, and a criterion that turns two quantities into the one they have to agree on.
Field sets shared by model configurations.
Nothing here registers: every model would otherwise inherit a registration it
never asked for. The fields each model overrides, such as the backbones, the
losses and the optimizer, are named once here and pinned per model in fr, mgr,
mcd, mrd, dr, dar, grat and genspp.
Three bases rather than one, because the architectures disagree about what a
loss is: SPPModelConfig scores a flat list,
PhasedSPPModelConfig scores three lists tied to training phases, and
SPPShapeConfig is what the two share.
- class pyhighlights.configurations.base.PhasedSPPModelConfig(**data)[source]#
An SPP model whose criteria belong to a training phase.
MCD and MRD are the same shape and differ only in which criteria go in which list, so the lists are declared once. Both refuse
supervise_highlightsfor the same reason: a supervision loss appended to a flat list would be dropped before the first batch, and the phase it belongs to has to be named.- Parameters:
name (str)
selector_backbones (RegistrationKey[SPPBackbone])
selectors (RegistrationKey[SPPSelector])
predictor (RegistrationKey[SPPPredictor])
predictor_backbone (RegistrationKey[SPPBackbone] | None)
temperature (float)
compact (bool)
select_over (str)
encoder_lr (float | None)
optimizer (RegistrationKey[Optimizer])
train_metrics (List[RegistrationKey[BoundMetric]] | None)
val_metrics (List[RegistrationKey[BoundMetric]] | None)
test_metrics (List[RegistrationKey[BoundMetric]] | None)
shared_losses (List[RegistrationKey[Loss]])
predictor_losses (List[RegistrationKey[Loss]])
generator_losses (List[RegistrationKey[Loss]])
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.base.SPPModelConfig(**data)[source]#
An SPP model scoring one flat list of criteria.
- Parameters:
name (str)
selector_backbones (RegistrationKey[SPPBackbone])
selectors (RegistrationKey[SPPSelector])
predictor (RegistrationKey[SPPPredictor])
predictor_backbone (RegistrationKey[SPPBackbone] | None)
temperature (float)
compact (bool)
select_over (str)
encoder_lr (float | None)
optimizer (RegistrationKey[Optimizer])
train_metrics (List[RegistrationKey[BoundMetric]] | None)
val_metrics (List[RegistrationKey[BoundMetric]] | None)
test_metrics (List[RegistrationKey[BoundMetric]] | None)
losses (List[RegistrationKey[Loss]])
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.base.SPPShapeConfig(**data)[source]#
The parts every SPP model has, without saying what it optimises.
Split from
SPPModelConfigbecause not every architecture has a flatlosseslist: MCD and MRD score their criteria per training phase and take three lists instead. The shape lives in one place, so a field added here reaches every model, and each model says only what differs.- Parameters:
name (str)
selector_backbones (RegistrationKey[SPPBackbone])
selectors (RegistrationKey[SPPSelector])
predictor (RegistrationKey[SPPPredictor])
predictor_backbone (RegistrationKey[SPPBackbone] | None)
temperature (float)
compact (bool)
select_over (str)
encoder_lr (float | None)
optimizer (RegistrationKey[Optimizer])
train_metrics (List[RegistrationKey[BoundMetric]] | None)
val_metrics (List[RegistrationKey[BoundMetric]] | None)
test_metrics (List[RegistrationKey[BoundMetric]] | None)
- compact: bool#
Gather the kept positions into a shorter sequence instead of zeroing the dropped ones in place. Off by default, because it changes what the predictor is trained on rather than fixing what it reads.
Zeroing leaves a dropped position in the sequence, and the predictor can read the fact of it. A recurrent encoder steps over it, and a transformer gives the next kept token a different position embedding. So a selector can signal a label through the shape of the mask rather than the words in it. Compaction closes both mechanisms; it does not close the count, which a shorter sequence still shows.
See
SPP.compactedfor what it costs.
- encoder_lr: float | None#
One rate for the encoders, another for everything above them.
Nonetrains the whole model at the optimizer’s own rate, which suits a GRU over a frozen table, where nothing pretrained is fine-tuned. Set it when a pretrained encoder is being fine-tuned: one rate cannot serve both a transformer and a selector initialized from scratch.
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- select_over: str#
the unit the corpus annotates, that a sparsity target is a fraction of, and that an export shows. Words are the same unit whichever backbone read the text.
"subtoken"selects over the backbone’s own tokens instead.- Type:
A selection is made over words
Criterion registrations and the bindings that feed them named fields.
- class pyhighlights.configurations.losses.AlignmentClassificationLossConfig(**data)[source]#
What the label looks like to a module that only ever read full text.
DAR scores its aligner with this twice: while that module is pretrained on the full input, and afterwards on the highlight, where the term is the generator’s alone.
- Parameters:
name (str)
loss (RegistrationKey[Module])
inputs (List[str])
coefficient (float)
enabled (bool)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.losses.ComplementClassificationLossConfig(**data)[source]#
What the label looks like from everything the highlight left behind.
- Parameters:
name (str)
loss (RegistrationKey[Module])
inputs (List[str])
coefficient (float)
enabled (bool)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.losses.CrossEntropyConfig(**data)[source]#
Cross entropy, weighted per class where a corpus needs it.
- Parameters:
weight (List[float] | None)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- weight: List[float] | None#
One weight per class, or nothing for an unweighted loss. A heavily imbalanced corpus is answered correctly by a model that never predicts the rare class, so the numbers such a corpus needs are part of its configuration rather than a detail of its training.
- class pyhighlights.configurations.losses.KnowledgeBCEConfig(**data)[source]#
A binary criterion carrying one positive weight per knowledge entry.
- Parameters:
ignore_index (int)
pos_weight (List[float] | None)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- pos_weight: List[float] | None#
One factor per knowledge entry, multiplying the cost of missing a positive there.
ClassWeightsTaskwithKNOWLEDGE_WEIGHTScomputes them from a split and writes them to its results, and a study copies them here.
- class pyhighlights.configurations.losses.KnowledgeLossConfig(**data)[source]#
Which knowledge base entries explain this example, against the gold links.
The one place in a grounded run where a highlight-side claim meets a gold standard.
knowledge_trueis-1on an example the corpus does not annotate and0where it annotates that an entry does not apply, so the criterion skips the first and scores the second: an empty knowledge set is an answer, not a missing label.- Parameters:
name (str)
loss (RegistrationKey[Module])
inputs (List[str])
coefficient (float)
enabled (bool)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.losses.KnowledgeSparsityLossConfig(**data)[source]#
How much of the knowledge base an example is allowed to instantiate.
The same criterion the token axis uses, bound to the knowledge axis: it reads the field names it is given and does not care which axis they are.
- Parameters:
name (str)
loss (RegistrationKey[Module])
inputs (List[str])
coefficient (float)
enabled (bool)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.losses.KnowledgeSupervisionLossConfig(**data)[source]#
Told which entries to name, with a weight per entry.
The same supervision as the term below and a different criterion, for one reason: two classes under a cross entropy carry a single positive weight, and the knowledge axis needs one per entry. The entry that decides a case is frequently the rare one, and a shared weight cannot tell it from the entry that fires on half the corpus.
It scores
knowledge_score, the difference of the comparer’s two logits. The gate is taken from the same quantity, so the term and the gate cannot disagree.- Parameters:
name (str)
loss (RegistrationKey[Module])
inputs (List[str])
coefficient (float)
enabled (bool)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.losses.LossConfig(**data)[source]#
Binds a criterion to the namespace fields it scores.
- Parameters:
name (str)
loss (RegistrationKey[Module])
inputs (List[str])
coefficient (float)
enabled (bool)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.losses.MaskedBCEConfig(**data)[source]#
One independent decision per position of the axis being scored.
- Parameters:
ignore_index (int)
pos_weight (List[float] | None)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- pos_weight: List[float] | None#
One factor per position, multiplying the cost of missing a positive there.
- class pyhighlights.configurations.losses.MaskedCrossEntropyConfig(**data)[source]#
Cross entropy over valid, labelled positions of any axis.
- Parameters:
weight (List[float] | None)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- weight: List[float] | None#
One weight per class, for an axis whose classes are imbalanced. The knowledge axis typically is, since an example instantiates few entries of a base and many examples instantiate none.
- class pyhighlights.configurations.losses.RemainingDiscrepancyLossConfig(**data)[source]#
MRD’s criterion: the complement should stop looking like the whole input.
Scored between the complement and the full input rather than between the highlight and the full input, and maximized. It is the one term in the library a model wants large, which is what the negative coefficient says. A coefficient is otherwise non-negative, since a loss is otherwise something to minimize.
- Parameters:
name (str)
loss (RegistrationKey[Module])
inputs (List[str])
coefficient (float)
enabled (bool)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
Optimizer registrations.
Corpus loader registrations.
- class pyhighlights.configurations.datasets.BeerConfig(**data)[source]#
Beer aspects; one variant per aspect.
- Parameters:
directory (str | None)
splits (Dict[str, str] | None)
url (str)
sha256 (str | None)
task (str)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.datasets.HateXplainConfig(**data)[source]#
Hate-speech posts, every annotator judgement kept.
- Parameters:
directory (str | None)
url (str)
divisions_url (str)
sha256 (str | None)
divisions_sha256 (str | None)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.datasets.HotelConfig(**data)[source]#
Hotel aspects; one variant per aspect.
- Parameters:
directory (str | None)
splits (Dict[str, str] | None)
url (str)
sha256 (str | None)
task (str)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.datasets.LoaderConfig(**data)[source]#
Fields every loader shares.
- Parameters:
directory (str | None)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.datasets.MoviesConfig(**data)[source]#
ERASER movies: evidence spans become highlights.
- Parameters:
directory (str | None)
task (str)
splits (Dict[str, str] | None)
url (str)
sha256 (str | None)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.datasets.R2AConfig(**data)[source]#
Fields the Beer and Hotel loaders share; the archive holds both.
- Parameters:
directory (str | None)
splits (Dict[str, str] | None)
url (str)
sha256 (str | None)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- sha256: str | None#
The digest the loader defaults to, named again here because a registered key is what a run actually builds, and its download is verified against this. A fixture archive passes
sha256=Noneexplicitly.
- class pyhighlights.configurations.datasets.ToyConfig(**data)[source]#
Synthetic corpus for smoke tests and demos; no download.
- Parameters:
directory (str | None)
sizes (Dict[str, int] | None)
triggers (List[str | List[str]])
length (int)
contaminations (int)
min_chunk (int)
vocabulary_size (int)
seed (int)
- contaminations: int#
Chunks of the patterns scattered through the filler. Without them a fragment of a pattern classifies as well as the pattern does. Off here: the registered corpus is a smoke test and a cheap one is the point.
- min_chunk: int#
The shortest piece of a pattern worth cutting out as a chunk.
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- triggers: List[str | List[str]]#
a pattern, or a list of patterns all of which have to appear. Tokens are characters here. The default is the cheap smoke-test corpus, one pattern per class; a conjunction whose patterns are shared between classes is the one no single n-gram can solve.
- Type:
What each class is
- vocabulary_size: int#
How many filler characters, drawn from the letters no trigger uses.
Corpus audits and preprocessing registrations.
- class pyhighlights.configurations.preprocessors.ClassWeightsConfig(**data)[source]#
Reads one split’s class frequencies; changes no row.
- Parameters:
split (str)
classes (int | None)
- classes: int | None#
Left unset, the number of classes is the largest label seen plus one. Set it wherever a split might not hold every class.
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.preprocessors.HateXplainAggregatorConfig(**data)[source]#
Three annotators to one label and one highlight vector.
- Parameters:
labels (Sequence[str])
highlights (str)
ties (str)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.preprocessors.HateXplainPipelineConfig(**data)[source]#
Aggregate the annotators first: repairing leakage before the tie-drop would measure overlap over rows the aggregation then removes.
- Parameters:
steps (List[RegistrationKey[Preprocessor]])
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.preprocessors.KnowledgeWeightsConfig(**data)[source]#
Reads one split’s example-to-entry links; changes no row.
ClassWeightsTaskruns it unchanged, so the weights land in aresults.jsona configuration can be copied from rather than being computed inside training and left nowhere.- Parameters:
split (str)
entries (int)
- entries: int#
the largest index a split happens to use is not how many entries there are.
- Type:
The size of the knowledge base, required rather than inferred
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.preprocessors.LeakageDetectorConfig(**data)[source]#
Reports what splits share, and refuses a corpus that shares anything.
- Parameters:
key (str)
normalize_keys (bool)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.preprocessors.LeakageRemoverConfig(**data)[source]#
Drops shared and repeated rows, protecting the annotated split first.
- Parameters:
priority (Sequence[str])
key (str)
normalize_keys (bool)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.preprocessors.PipelineConfig(**data)[source]#
Preprocessors run in order, each over what the last returned.
- Parameters:
steps (List[RegistrationKey[Preprocessor]])
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.preprocessors.ShortcutDetectorConfig(**data)[source]#
Reports what predicts the label, and refuses a corpus solvable without its evidence.
- Parameters:
max_length (int)
separator (str)
tokens (str)
label (str)
seed (int)
permutations (int)
- max_length: int#
Longest n-gram scanned. Every shorter one is scanned too.
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- permutations: int#
Label shuffles the refusal threshold is the best of, which makes the check a permutation test at a level of about
1 / permutations.
- separator: str#
empty for a corpus of characters, a space for a corpus of words.
- Type:
How an n-gram’s tokens are joined for display