Tasks#
A task is one experiment, start to finish: a corpus, its preprocessing, a model, the metrics to score it with, and a list of seeds. It is what a row of a results table is made of, so reproducing a number means running one key rather than remembering which loader went with which checkpoint.
from cinnamon.registry import Registry
from pyhighlights.configurations.keys import TOY_TASK
task = Registry.from_key(TOY_TASK, seeds=[42, 1337, 2024])
results = task.run()
results["summary"]["test_highlight_f1"] # {"mean": ..., "std": ..., "values": [...]}
Registered with run_method="run", so cmn-run drives the same task from the command line.
Anything decided before a run is a parameter of
TaskConfig: what monitors the run, which vector file to read, whether predictions and faithfulness are reported.
A key therefore records what was asked for rather than the part of it somebody remembered to register.
The kwargs above override a key at build time; they are not the only way to set a value.
A value a run computes cannot be a parameter at all: the embedding matrix is fitted against the training split and reaches the model as a tensor, and a vocabulary size measured by a tokenizer is known only once that split has been read.
Constraints between parameters are declared as cinnamon conditions rather than checked in the component, so an invalid combination is rejected while the registry expands keys.
A task that names both pretrained_model_card and embeddings fails the one_embedding_source condition, and a grid varying the embedding source drops that combination before anything trains.
What a run does#
For each seed, in order:
seed_everything, then build a fresh model from its key, since a select-then-predict model trained through a discrete choice lands somewhere different every time, and the spread across seeds is part of the result.Train under the
callbacksthe task names, early stopping and a checkpoint of the best epoch, both onval_lossby default. They have to monitor the same quantity, and a task that is given two refuses to build rather than reporting a model its own stopping rule did not choose;GeneralizationLossScoreis how two quantities become the one they can agree on.Restore that checkpoint before scoring. Early stopping returns after
patienceworse epochs, so the weights still in memory are not the ones anybody would keep.Score validation and test, and store the test predictions when
store_predictionsis set, onepredictions-seed=<seed>.pklper seed, in the run directory.
Then the seeds are summarised and written out, as the mean, the standard deviation and the individual values of every metric.
What lands on disk#
results/<name>/<started>/
├── results.json # every seed's metrics and costs, and their summary
├── manifest.json # the whole configuration tree, and the versions
├── predictions-seed=42.pkl # when ``store_predictions`` is set
├── predictions-seed=1337.pkl
├── seed=42/
│ └── epoch=3-step=128.ckpt
└── seed=1337/…
The weights are the one part of that tree nothing downstream reads: the task restores the best checkpoint itself before scoring, and an analyzer reads results.json and the stored predictions.
On a grid of fine-tuned transformer cells they are also most of the bytes, since a checkpoint holds every encoder the model trains.
keep_checkpoints=False writes the checkpoint, restores it, scores, and then deletes it; save_weights_only keeps it but drops the optimizer state, which is only needed to resume training and no task resumes.
Both are off by default, and the trade the first one makes is that a number cannot be re-scored without training again.
<started> is the moment the run began, 2026-09-09T16-13-00.
A run never overwrites an earlier one: two runs of the same task are two results to compare, and the second quietly replacing the first is a measurement lost to a re-run somebody forgot they had already done.
Two runs inside one second get …-2 appended rather than sharing a directory.
manifest.json is what makes the directory worth keeping.
A task’s own attributes are not enough, since a task holds keys: recording them writes name=model--tags=['fr','gru'] and leaves the hidden size, the sparsity threshold and the learning rate behind that key nowhere in the record.
describe() replaces every key with the
configuration it names, recursively, so the file states the numbers the run used:
{
"started": "2026-09-09T16-13-00",
"component": "pyhighlights.components.tasks.SPPTask",
"key": "name=task--tags=['fr', 'gru', 'movies']--namespace=pyhighlights",
"build_args": {"seeds": [0, 1, 2]},
"versions": {"python": "3.13.15", "pyhighlights": "0.2.0",
"cinnamon-core": "2.0.3", "torch": "2.14.0",
"lightning": "2.6.5"},
"settings": {
"callbacks": [{"@key": "name=callback--tags=['early_stopping','loss']--namespace=pyhighlights",
"monitor": "val_loss", "patience": 5}],
"model": {
"@key": "name=model--tags=['fr', 'gru']--namespace=pyhighlights",
"selector_backbones": {"hidden_size": 128, "bidirectional": true},
"losses": [{"name": "sparsity", "loss": {"threshold": 0.15}}],
"optimizer": {"lr": 0.001}
}
}
}
Predictions sit beside the run rather than inside a checkpoint directory. They belong to the run: a reader that finds them next to a checkpoint can say which seed produced them and not which run, and the analyzers report both.
The versions are there because a metric that moved between two runs of the same configuration is a version difference or nothing at all. Private attributes are absent: they are what the run built, the embedding matrix fitted against the training split among them, and no more a setting than the trained weights are.
key and build_args are what make the file replayable rather than merely readable.
The key alone rebuilds the registered defaults, not the run that was launched, so the overrides go down beside it: a manifest saying seeds: [0, 1, 2] is a run somebody can re-run, and one saying only the key is not. cinnamon annotates every component it builds with both, so a task constructed directly rather than through a key reports null for each.
Build args are resolved like everything else, so an override that names a key, as in Registry.from_key(TASK, model=GRU_MGR), is written out as the settings behind it rather than as a name.
Corpus and model#
loader, preprocessor and model are registration keys, so a task definition swaps its corpus without touching code.
preprocessor is optional only where a corpus needs none:
HateXplain has no label until an
AnnotationAggregator has run,
and the loader says so rather than guessing one.
Text becomes ids in one of two ways.
Name a pretrained_model_card and the matching subword tokenizer is used; leave it unset and a vocabulary is fitted on the training split alone, since fitting it on evaluation text would leak quietly and nothing downstream can tell where an id came from.
Its vocabulary_size has to match the backbone’s vocab_size: an id the embedding has no row for is a crash at the first batch.
vocabulary_size counts ids rather than words, and two of them are reserved: id 0 is padding and id 1 is the unknown token, so a size of 10_000 maps the 9_998 most frequent training tokens.
The two ids are separate so that a highlight over an out-of-vocabulary word is distinguishable from one over padding, which is what an exported explanation has to state.
Everything else a run configures has a page of its own: Token vectors for the token vectors a model starts from, Metrics for what is scored, Highlight supervision for training against a highlight annotation, Faithfulness for the sufficiency and comprehensiveness diagnostics, Diagnostics for reading the inside of a forward pass, Computational cost for what a run spent, Analyzers for reading the results back, and Reproducing published results for running a grid of tasks.
API#
Tasks: one experiment, start to finish.
A task is what a paper’s table row is made of. It names the corpus, the preprocessing, the model and the metrics as registration keys, runs the whole thing over a list of seeds, and writes down what happened. Reproducing a number therefore means running one key, not remembering which loader went with which checkpoint.
Seeds are a list rather than a number because a single run of a select-then-predict model says very little: the selector is trained through a discrete choice, and the spread across seeds is part of the result.
- class pyhighlights.components.tasks.ClassWeightsTask(loader, weights, preprocessor=None, **kwargs)[source]#
Bases:
TaskLoads a corpus, weighs its classes, and writes the numbers down.
A weighted loss needs one number per class, and where that number comes from decides whether a run can be repeated. Computing it inside training leaves it nowhere afterwards; typing it into a configuration by hand leaves it nowhere it can be checked. This is the third way: a run of its own, whose result is the weights and the counts they came from, in the same
results.jsonandmanifest.jsonevery other task writes.So the numbers are readable by a person, persist after the process that computed them exits, and carry the key of the corpus and the preprocessing that produced them. That is what makes them worth copying into a configuration, where every training run’s manifest then records them.
It trains nothing and takes no seeds.
- Parameters:
loader (RegistrationKey[HighlightLoader])
weights (RegistrationKey[ClassWeights])
preprocessor (RegistrationKey[Preprocessor] | None)
- class pyhighlights.components.tasks.GenSPPTask(search, **kwargs)[source]#
Bases:
SPPTaskGenSPP: a genetic search over generators, scored like any other task.
The corpus, the preprocessing, the metrics and the seeds are an
SPPTask’s. What differs is the training: no gradient reaches the generator, so the model is not named directly but by theGenSPPTrainerthat searches for it: one search per seed, seeded with it, since a genetic search over a population of two dozen is the noisiest part of the run.Each seed leaves behind the weights the search settled on and
search.json, one entry per generation undertraining_progress: the best objective the search reached in it, which is1 / fitnessand so falls as the search improves. A search that stopped improving in its tenth generation and one that was still descending when the budget ran out report the same number otherwise.- Parameters:
search (RegistrationKey[GenSPPTrainer])
- SMOKE_CANDIDATES = 8#
How many candidates a diagnosed search may evaluate. Each one trains a predictor over the whole training split, and every batch of that is a page of the record, so a smoke test is a handful of them. A full search evaluates thousands.
- static candidates(search, generations=None)[source]#
How many models a search of this shape trains.
The founders, plus the children every generation draws: couples at
selection_rateof the population, each crossing into two.generationsis how many actually ran, for a search that has already stopped; left out, it is how many were budgeted.- Return type:
int- Parameters:
search (GenSPPTrainer)
generations (int | None)
- check_diagnostics()[source]#
A search is not bounded by what bounds a trainer.
SPPTask.check_diagnostics()readstrainer_args, which here reaches only the throwaway trainer each candidate’s predictor is fitted with. The search itself runs outside Lightning and over the whole split, as many times as it has candidates, so bounding a diagnosed search means bounding the search:population_size,n_generationsand the rate that decides how many children a generation draws.- Return type:
None
- train(seed, loaders, directory, meter)[source]#
Search for a generator, and hand back the model the search settled on.
No epoch of it is ever trained by Lightning, so the trainer built here exists for the scoring pass alone and carries no callbacks.
- Return type:
Tuple[Model,Trainer]- Parameters:
seed (int)
loaders (Mapping[str, DataLoader])
directory (Path)
meter (Meter)
- class pyhighlights.components.tasks.SPPTask(loader, model, preprocessor=None, train_metrics=None, val_metrics=None, test_metrics=None, seeds=(42,), batch_size=32, max_length=None, vocabulary_size=10000, pretrained_model_card=None, add_special_tokens=True, embeddings=None, pretrained_tokens_only=True, vocabulary_from='corpus', requires_embeddings=False, one_hot_embeddings=None, callbacks=None, store_predictions=False, keep_checkpoints=True, save_weights_only=False, faithfulness=False, diagnostics=False, highlight_supervision=False, highlight_loss=None, highlight_coefficient=1.0, trainer_args=None, **kwargs)[source]#
Bases:
TaskTrains a select-then-predict model over a corpus, once per seed.
The pieces are registration keys, so the same task definition swaps its corpus or its model without touching code.
preprocessoris optional only for a corpus that needs none; HateXplain has no label until one has run, and the loader says so rather than guessing.Each seed trains from scratch, restores the checkpoint that scored best on validation, and is evaluated on validation and test.
faithfulnessadds the terms ofpyhighlights.components.faithfulnessover the test split, and is off by default: they are two more columns rather than a correction, and a registered reproduction should report what its paper reports. What lands on disk isresults.json(every seed’s metrics, plus their mean and standard deviation),manifest.json, and, when asked, onepredictions-seed=<seed>.pklper seed.- Parameters:
loader (RegistrationKey[HighlightLoader])
model (RegistrationKey[Model])
preprocessor (RegistrationKey[Preprocessor] | None)
train_metrics (List[RegistrationKey[BoundMetric]] | None)
val_metrics (List[RegistrationKey[BoundMetric]] | None)
test_metrics (List[RegistrationKey[BoundMetric]] | None)
seeds (Sequence[int])
batch_size (int)
max_length (int | None)
vocabulary_size (int)
pretrained_model_card (str | None)
add_special_tokens (bool)
embeddings (str | Path | None)
pretrained_tokens_only (bool)
vocabulary_from (Literal['corpus', 'vectors'])
requires_embeddings (bool)
one_hot_embeddings (int | None)
callbacks (List[RegistrationKey[Callback]] | None)
store_predictions (bool)
keep_checkpoints (bool)
save_weights_only (bool)
faithfulness (bool)
diagnostics (bool)
highlight_supervision (bool)
highlight_loss (RegistrationKey[Loss] | None)
highlight_coefficient (float)
trainer_args (Mapping[str, Any] | None)
- attach_metrics(model)[source]#
Give a model the metrics this task scores it with.
build_model()passes them to the constructor, which is where a model belonging to this task gets them. A searched model is built by the search from the model key alone, so it arrives without any, and this is the one place that is repaired. All three splits are repaired, so a model is not left holding metrics from one path and none from another.- Return type:
None- Parameters:
model (Model)
- bounds_training_batches()[source]#
Whether
trainer_argscuts the training batches down.fast_dev_runruns one batch per split.limit_train_batchesandoverfit_batchesare counts as integers and fractions as floats, so an integer bounds and a float bounds only below one.- Return type:
bool
- build_callbacks(checkpoints)[source]#
What monitors this run, built from the keys the task was given.
A checkpoint callback is told where to write and whether to store weights only, because those are the task’s business rather than the study’s. The run deletes them again once it has scored the epoch they hold.
Nothing is monitored when no keys are given, and the run is scored on the weights it ended on. The registered configuration names the default pair, early stopping and checkpointing on
val_loss, rather than this class naming it. Which callbacks are the default is therefore a decision of the configuration layer and is stated once.- Return type:
List[Callback]- Parameters:
checkpoints (Path)
- build_model(embeddings=None, knowledge=None)[source]#
The model this task names, holding the data it has to be given.
embeddingsandknowledgeare data rather than configuration, so they reach the model here rather than through a registration. Both default to whattokenizer()andloaders()read off the corpus, and both may be passed outright, by a caller that built them itself or by a test.A task that embeds from a vector file and is handed no matrix is refused: it would otherwise build a model with a randomly initialised table and report numbers for it.
- Return type:
Model- Parameters:
embeddings (Tensor | None)
knowledge (InputData | None)
- check_diagnostics()[source]#
Refuse to diagnose a run whose batches nobody bounded.
The record is per batch, so a full run writes gigabytes of it and pays the formatting on every step. It is for a smoke test: one or two epochs over a handful of batches, which is what Lightning’s
fast_dev_runand itslimit_*_batchesalready express and whattrainer_argsalready forwards. A task given neither is a mistake rather than a choice, and it is cheaper to say so before the run than after it.The batches are what this checks, not the epochs:
max_epochscarries a default, so an epoch bound is always present and a rule about it would never fire.The training batches, and a bound that bounds: a
limit_val_batchesleaves training unbounded, and Lightning reads a float as a fraction, solimit_train_batches=1.0is its own default and means every batch. Both are refused.- Return type:
None
- check_supervision(splits)[source]#
Refuse to call a run supervised when nothing supervises it.
The collator pads unannotated positions with
-1and the criterion skips them, so supervising a corpus annotated on test alone trains exactly as an unsupervised run does, and reports itself as the ceiling that run was measured against. Checked where the loaders are built, which is the one thing every path tofitgoes through.- Return type:
None- Parameters:
splits (Mapping[str, DataFrame])
- diagnostics#
Write every stage of the pipeline into this run’s directory. Off by default and only for a run that has been bounded: see
pyhighlights.utility.diagnosticsfor what is recorded andcheck_diagnostics()for why an unbounded run is refused.
- discard_checkpoints(directory)[source]#
Delete this seed’s checkpoints, now that they have been scored.
What a run is read from survives: the metrics, the manifest, the stored predictions and, for a search,
search.json. The weights do not, so a number cannot be re-derived without training again. That is the trade a grid of transformer cells makes to fit on a filesystem, and why this is off by default.- Return type:
None- Parameters:
directory (Path)
- fit(seed, loaders)[source]#
Train one model and score it, leaving its checkpoint behind.
What every seed does whatever produced its model: seed the run, time it, give it a directory of its own, score it on the evaluation splits, and report what it cost. How the model is produced is
train(), which a search replaces.- Return type:
Dict[str,float]- Parameters:
seed (int)
loaders (Mapping[str, DataLoader])
- knowledge()[source]#
The corpus’s knowledge base, one entry as a list of tokens.
A second loader instance, built only to ask: a loader parses and downloads in
read(), never in__init__, so this costs a constructor and the corpus is not read twice.- Return type:
Optional[Sequence[Sequence[str]]]
- one_hot_embeddings#
Width of a one-hot table, or
Nonefor a learned one. A corpus of symbols has nothing to pretrain and nothing to learn: seeone_hot_table().
- requires_embeddings#
Whether the task is meaningless without
embeddings. A reproduction of a run that embeds from a released vector file is: the file is a download the task is given rather than fetches, and forgetting it otherwise trains on whatever vocabulary the corpus happens to produce and reports numbers for it.
- results(runs)[source]#
What the seeds so far reported, and their spread.
seedsis every seed the task was asked for andrunsis what has finished, so a partial result says which of the two it is rather than looking like a task configured with fewer seeds.- Return type:
Dict[str,Any]- Parameters:
runs (Sequence[Mapping[str, float]])
- score(trainer, model, loaders, predictions)[source]#
Score a trained model on whichever evaluation splits exist.
predictionsis where this seed’s predictions go, named rather than a directory: they belong to the run, not to the checkpoint, and a reader that finds them beside a checkpoint can only say which seed produced them, not which run.- Return type:
Dict[str,float]- Parameters:
trainer (Trainer)
model (Model)
loaders (Mapping[str, DataLoader])
predictions (Path)
- scored(trainer, model, loaders, predictions, timer)[source]#
The scoring pass itself, with the timer already installed.
- Return type:
Dict[str,float]- Parameters:
trainer (Trainer)
model (Model)
loaders (Mapping[str, DataLoader])
predictions (Path)
timer (InferenceTimer)
- tokenizer(splits)[source]#
A subword tokenizer when a model card is named, else a vocabulary.
The vocabulary is fitted on
trainonly, and its size has to match the backbone’svocab_size: an id the embedding has no row for is a crash at the first batch. Namingembeddingsfits it against a vector file instead, and the matrix that comes back is handed to the model, which sizes its table to it.one_hot_embeddingsbuilds that matrix rather than reading one, for a corpus whose tokens are symbols.vocabulary_fromdecides what “against a vector file” means, and the two answers are different experiments:"corpus"keeps the training vocabulary and drops the tokens the file has no vector for. A token the training split never saw is unknown at evaluation whatever the file covers."vectors"takes the file’s vocabulary whole. Nothing is unknown that the file covers, which is what a released implementation embedding from a fixed pretrained vocabulary does. It is not a leak, because the file is external and says nothing about the splits.
The choice changes the inputs. An evaluation token absent from the training split is embedded as zero under
"corpus"and as its vector under"vectors".- Return type:
- Parameters:
splits (Mapping[str, DataFrame])
- train(seed, loaders, directory, meter)[source]#
The trained model and the trainer that will score it.
Gradient descent under Lightning, which is what every architecture but GenSPP is trained by.
directoryis this seed’s own, and the checkpoints land in it.meteris handed over rather than read afterwards, because a seed that trains more than one model knows how many only once it has stopped: seeGenSPPTask.train().- Return type:
Tuple[Model,Trainer]- Parameters:
seed (int)
loaders (Mapping[str, DataLoader])
directory (Path)
meter (Meter)
- vocabulary_from#
Where the token ids come from when a vector file is read.
corpuskeeps the training vocabulary and lets the file cover it;vectorstakes the file’s own vocabulary whole, so a token absent from training still has its vector. Seetokenizer().
- class pyhighlights.components.tasks.Task(name='task', save_path=None)[source]#
Bases:
ABCOne experiment, run over a list of seeds and written down.
- Parameters:
name (str)
save_path (str | Path | None)
- property directory: Path#
Where this run writes, stamped with the moment it started.
One directory per run, never reused: two runs of the same task are two results to compare, and the second quietly replacing the first is a measurement lost to a re-run somebody forgot they had already done. The stamp is taken once and kept, so every seed, the metrics and the manifest land together.
- pyhighlights.components.tasks.load_splits(loader, preprocessor=None)[source]#
The corpus a key names, preprocessed by the key that names how.
- Return type:
Dict[str,DataFrame]- Parameters:
loader (RegistrationKey[HighlightLoader])
preprocessor (RegistrationKey[Preprocessor] | None)
- pyhighlights.components.tasks.summarize(runs)[source]#
Mean and standard deviation of each metric across seeds.
The sample standard deviation,
ddof=1. The seeds are a sample of the runs the configuration could produce rather than the whole of them, which is what a table reportingmean +/- stdclaims. The sample form is larger than the population form by a factor ofsqrt(n / (n - 1)): 12% at five seeds and 41% at two. One seed has no spread to report and gives0.0rather than anan.- Return type:
Dict[str,Dict[str,float]]- Parameters:
runs (Sequence[Mapping[str, float]])
- pyhighlights.components.tasks.vocabulary(frames, size)[source]#
Token ids
2tosize - 1, most frequent first.sizecounts the ids in use rather than the tokens named, because it is the embedding’s width that has to hold them: id0is padding and id1is the unknown token, sosize - 2tokens are mapped and the rest fall back to the unknown id. The two are separate ids so that a highlight over an out-of-vocabulary word is distinguishable from one over nothing.Built from the training split alone. A vocabulary fitted on evaluation text would leak it, and quietly, since nothing downstream can tell where an id came from.
- Return type:
Dict[str,int]- Parameters:
frames (Iterable[DataFrame])
size (int)
Callbacks a run is monitored by, and the criterion that selects its epoch.
Early stopping and checkpointing are registered components, so a study configures them the way it configures its losses and its metrics. A study can maximise a metric, monitor a combination of logged quantities, and say that a model spends its first epochs pretraining something the monitor knows nothing about.
One quantity decides both. The callback that stops a run and the callback
that keeps an epoch have to agree, or a run reports a model its own stopping
rule did not choose: stop when two quantities are exhausted and the restored
epoch is whichever one the checkpoint happened to monitor.
GeneralizationLossScore exists so that a study combining two
quantities still monitors one.
- class pyhighlights.components.callbacks.GeneralizationLossScore(quality='val_f1', loss='val_loss', coefficient=2.0, name='val_score')[source]#
One monitored quantity out of a metric and a one-sided loss penalty.
Maximises
qualitywhile charging for validation loss that has risen above its own best:\[score = quality - \lambda \cdot \max(0, loss / loss_{opt} - 1)\]The penalty is Prechelt’s generalization loss (Prechelt, 1998, Early Stopping – But When?), with
loss_optthe lowest validation loss seen so far. A ratio rather than a difference, and that matters: a heavily weighted cross entropy is unbounded while an F1 is not, so a difference would makecoefficienta guess about scale. A relative regression is dimensionless, and the coefficient means one thing: how muchqualityan epoch forfeits per unit of relative loss regression.Monitoring a rare class’s F1 alone can accept a large relative loss regression, and monitoring the loss alone can give up F1 to avoid a small one. The coefficient is what chooses between those, and it belongs to an architecture rather than to the library: a coefficient of 0 monitors
qualityalone and a large one monitors the loss alone, so both of the criteria it replaces are special cases of it.- Parameters:
quality (str)
loss (str)
coefficient (float)
name (str)
- on_validation_end(trainer, pl_module)[source]#
Logged on the hook
EarlyStoppingchecks on.on_validation_endruns after the epoch’s values are logged. It writes the score beforeEarlyStoppingchecks it on the same hook, provided this callback comes first in the list.build_callbacks()puts it there rather than trusting the order it was given.- Return type:
None- Parameters:
trainer (Trainer)
pl_module (LightningModule)
- state_dict()[source]#
The floor the run is charged against, and only that.
loss_optis the whole memory of the criterion: a resumed run that forgot it would charge nothing for a regression it had already seen. The warmup window is not stored, because it is the model’s. A run resumed against a model whosewarmup_epochshas changed therefore keeps a floor set under the old window, which is a different model’s loss.- Return type:
Mapping[str,Any]
- class pyhighlights.components.callbacks.MonitoredScore[source]#
A callback that writes the quantity a run is monitored by.
Marked as a class rather than recognised by name, so that
build_callbacks()can order every criterion before the callbacks that read one.
- class pyhighlights.components.callbacks.WarmupEarlyStopping(monitor, min_delta=0.0, patience=3, verbose=False, mode='min', strict=True, check_finite=True, stopping_threshold=None, divergence_threshold=None, check_on_train_epoch_end=None, log_rank_zero_only=False)[source]#
Early stopping that starts counting when the model starts learning.
Patience is about epochs the monitored model failed to improve on, and an epoch it did not train in is not one of those.
- Parameters:
monitor (str)
min_delta (float)
patience (int)
verbose (bool)
mode (str)
strict (bool)
check_finite (bool)
stopping_threshold (float | None)
divergence_threshold (float | None)
check_on_train_epoch_end (bool | None)
log_rank_zero_only (bool)
- class pyhighlights.components.callbacks.WarmupModelCheckpoint(dirpath=None, filename=None, monitor=None, verbose=False, save_last=None, save_top_k=1, save_on_exception=False, save_weights_only=False, mode='min', auto_insert_metric_name=True, every_n_train_steps=None, train_time_interval=None, every_n_epochs=None, save_on_train_epoch_end=None, enable_version_counter=True)[source]#
Checkpointing that ignores the epochs before the model trains.
Without this the best epoch of a pretraining phase is a candidate for the epoch a run is scored on, and it holds a rationalizer that has taken no gradient step at all.
- Parameters:
dirpath (str | Path | None)
filename (str | None)
monitor (str | None)
verbose (bool)
save_last (bool | Literal['link'] | None)
save_top_k (int)
save_on_exception (bool)
save_weights_only (bool)
mode (str)
auto_insert_metric_name (bool)
every_n_train_steps (int | None)
train_time_interval (timedelta | None)
every_n_epochs (int | None)
save_on_train_epoch_end (bool | None)
enable_version_counter (bool)
- pyhighlights.components.callbacks.warmup_epochs(module)[source]#
Training epochs before the monitored model starts learning.
A model that pretrains a component over the first epochs of the training loop is not improving the thing a monitor watches, so a patience counted from epoch zero can exhaust before the model has taken a single step. G-RAT is that case, since it gates its optimizer on
current_epoch >= pretrain_epochs, so an early-stopping patience shorter than that window stops a run before its model has trained.The number comes from the model rather than from the task’s configuration, because a task repeating it is a second place for it to be wrong. A model that pretrains outside the training loop, as DAR does in
on_train_start, costs no epochs and needs nothing here.- Return type:
int- Parameters:
module (LightningModule)
What a run was: the settings it was given, and the versions that ran it.
A results directory is worth what can be rebuilt from it. Metrics say what
happened; they do not say what produced them, and a task’s own attributes are
not enough either. A task holds keys, so model reads
name=model--tags=['fr','gru'] and the hidden size, the sparsity threshold
and the learning rate behind that key are nowhere in the record.
describe() writes down the whole tree instead. Every key is replaced by
the configuration it names, recursively, so a manifest states the numbers the
run actually used rather than the names of the places they came from. Alongside
them go the versions of the packages that did the computing, since a metric
that moved between two runs of the same configuration is a version difference
or nothing at all.
The record also names the key that built the component and the arguments the
caller overrode, which is what makes a run replayable rather than merely
readable: the key alone rebuilds the registered defaults, not the run that was
launched. Both come from cinnamon, which annotates every component it builds;
a component built by hand has neither, and says so with null.
- pyhighlights.utility.manifest.ANNOTATIONS = ('registration_key', 'build_args')#
What cinnamon puts on a component it builds. Reported as a record of its own rather than left among the settings, where it would read as something the task was configured with.
- pyhighlights.utility.manifest.PACKAGES = ('pyhighlights', 'cinnamon-core', 'torch', 'lightning', 'torchmetrics', 'transformers')#
The packages whose version can change a number. Anything else installed alongside them is noise in a file somebody has to read.
- pyhighlights.utility.manifest.describe(component)[source]#
The manifest for one built component.
Private attributes are what the run built rather than what it was asked for, such as an embedding matrix fitted against the training split. They are no more a setting than the trained weights are.
keyandbuild_argsare what cinnamon wrote on the component when it built it, and arenullfor a component nobody built through a registry.- Return type:
Dict[str,Any]- Parameters:
component (Any)
- pyhighlights.utility.manifest.registration_key(entry)[source]#
The registration key of a resolved entry, old manifests included.
Manifests written before
KEY_FIELDexisted record it askey, and a results tree outlives the release that wrote it.- Return type:
str- Parameters:
entry (Mapping[str, Any])
- pyhighlights.utility.manifest.resolve(value)[source]#
Replace every registration key with the values behind it.
Containers keep their shape, so a list of metric keys becomes a list of metric settings in the same order. A key that appears twice, such as the same backbone under a selector and a predictor, is written out twice, which reads better than a file of cross-references.
The key itself is recorded under
KEY_FIELDrather thankey, so that a configuration with a parameter of that name keeps both.- Return type:
Any- Parameters:
value (Any)
- pyhighlights.utility.manifest.versions()[source]#
The interpreter and the packages that do the computing.
A package that is not installed is left out rather than reported as
None: transformers is optional, and a run that never imported it is not a run whose transformers version was unknown.- Return type:
Dict[str,str]
Task registrations: a corpus, a model, and what to score them with.
The metric set follows the corpus. Beer, Hotel, Movies and Toy have two
classes and take the two-class accuracy and F1. A task over HateXplain, which
has three, names MULTICLASS_ACCURACY_METRIC and MULTICLASS_F1_METRIC
instead. Highlight and selection metrics are the same everywhere, since a
token is either selected or it is not.
- pyhighlights.configurations.tasks.BINARY_METRICS = [RegistrationKey(name=metric, namespace=pyhighlights, tags=frozenset({'accuracy'}), description=None), RegistrationKey(name=metric, namespace=pyhighlights, tags=frozenset({'f1'}), description=None), RegistrationKey(name=metric, namespace=pyhighlights, tags=frozenset({'highlight', 'f1'}), description=None), RegistrationKey(name=metric, namespace=pyhighlights, tags=frozenset({'highlight', 'iou'}), description=None), RegistrationKey(name=metric, namespace=pyhighlights, tags=frozenset({'precision', 'highlight'}), description=None), RegistrationKey(name=metric, namespace=pyhighlights, tags=frozenset({'recall', 'highlight'}), description=None), RegistrationKey(name=metric, namespace=pyhighlights, tags=frozenset({'selection_rate'}), description=None), RegistrationKey(name=metric, namespace=pyhighlights, tags=frozenset({'selection_size'}), description=None), RegistrationKey(name=metric, namespace=pyhighlights, tags=frozenset({'selection_spans'}), description=None)]#
Beer, Hotel, Movies, Toy.
- Type:
Binary corpora
- class pyhighlights.configurations.tasks.ClassWeightsTaskConfig(**data)[source]#
A run whose whole result is the class weights of a corpus.
Its own configuration rather than a field of
TaskConfig: it trains nothing, so seeds, metrics, batches and trainer arguments would all be fields nobody sets. A study registers one per corpus it weights, points it at the same loader and preprocessor its training tasks use, and copies the numbers the run writes into the loss the training tasks name.- Parameters:
name (str)
loader (RegistrationKey[HighlightLoader])
weights (RegistrationKey[ClassWeights])
preprocessor (RegistrationKey[Preprocessor] | None)
save_path (str | None)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.tasks.GenSPPTaskConfig(**data)[source]#
Fields a GenSPP task adds, and the one it drops.
A GenSPP task names a search rather than a model: the model key is the search’s own, since the two disagreeing about which model was evolved is a result nobody could read. Nothing scores the training split, because no epoch of the winning model is ever trained, so only validation and test carry metrics.
- Parameters:
name (str)
save_path (str | None)
seeds (Sequence[int])
batch_size (int)
max_length (int | None)
vocabulary_size (int)
pretrained_model_card (str | None)
add_special_tokens (bool)
embeddings (str | None)
pretrained_tokens_only (bool)
vocabulary_from (str)
requires_embeddings (bool)
one_hot_embeddings (int | None)
callbacks (List[RegistrationKey[Callback]])
store_predictions (bool)
keep_checkpoints (bool)
save_weights_only (bool)
faithfulness (bool)
diagnostics (bool)
highlight_supervision (bool)
highlight_loss (RegistrationKey[Loss])
highlight_coefficient (float)
trainer_args (Dict[str, Any])
loader (RegistrationKey[HighlightLoader])
search (RegistrationKey[GenSPPTrainer])
preprocessor (RegistrationKey[Preprocessor] | None)
val_metrics (List[RegistrationKey[BoundMetric]])
test_metrics (List[RegistrationKey[BoundMetric]])
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- pyhighlights.configurations.tasks.HIGHLIGHT_METRICS = [RegistrationKey(name=metric, namespace=pyhighlights, tags=frozenset({'highlight', 'f1'}), description=None), RegistrationKey(name=metric, namespace=pyhighlights, tags=frozenset({'highlight', 'iou'}), description=None), RegistrationKey(name=metric, namespace=pyhighlights, tags=frozenset({'precision', 'highlight'}), description=None), RegistrationKey(name=metric, namespace=pyhighlights, tags=frozenset({'recall', 'highlight'}), description=None), RegistrationKey(name=metric, namespace=pyhighlights, tags=frozenset({'selection_rate'}), description=None), RegistrationKey(name=metric, namespace=pyhighlights, tags=frozenset({'selection_size'}), description=None), RegistrationKey(name=metric, namespace=pyhighlights, tags=frozenset({'selection_spans'}), description=None)]#
how much of the document was kept, where it sits, and how well what was kept matches the annotation. Precision and recall are reported beside their F1 because a sparsity target moves them apart, and spans because a size alone does not say whether the kept words sit together.
- Type:
What every select-then-predict run reports, whatever the corpus
- class pyhighlights.configurations.tasks.TaskConfig(**data)[source]#
Fields every task shares.
Anything decided before a run is declared here, so a key records what was asked for rather than the part of it somebody remembered to register. A value a run computes cannot be declared at all: the embedding matrix a vector file is read into is fitted against the training split at run time, so it reaches the model as a tensor and never as a parameter.
- Parameters:
name (str)
save_path (str | None)
seeds (Sequence[int])
batch_size (int)
max_length (int | None)
vocabulary_size (int)
pretrained_model_card (str | None)
add_special_tokens (bool)
embeddings (str | None)
pretrained_tokens_only (bool)
vocabulary_from (str)
requires_embeddings (bool)
one_hot_embeddings (int | None)
callbacks (List[RegistrationKey[Callback]])
store_predictions (bool)
keep_checkpoints (bool)
save_weights_only (bool)
faithfulness (bool)
diagnostics (bool)
highlight_supervision (bool)
highlight_loss (RegistrationKey[Loss])
highlight_coefficient (float)
trainer_args (Dict[str, Any])
- add_special_tokens: bool#
Keep
[CLS]and[SEP]. A pretrained encoder was trained reading them; they carry no word, so a selector never sees them either way.
- callbacks: List[RegistrationKey[Callback]]#
early stopping, checkpointing, and any criterion they read. The default pair stops and checkpoints on
val_loss.The stopping callback and the checkpoint callback have to monitor the same quantity. Mix two and a run reports a model its own stopping rule did not choose, silently.
SCORE_*monitors the combinationGeneralizationLossScorewrites, and wants that criterion in this list too.- Type:
What monitors the run
- diagnostics: bool#
Record what each pipeline stage held, for a bounded run. See
diagnostics.
- keep_checkpoints: bool#
Whether the weights survive the run. A checkpoint holds the whole model, and nothing downstream reads one: the task restores the best one itself before scoring, and an analyzer reads
results.jsonand the stored predictions. Turning this off trades the ability to re-score without retraining for disk space, which a grid of transformer cells runs out of.
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- one_hot_embeddings: int | None#
Build a one-hot table of this width instead of reading or learning one, for a corpus whose tokens are symbols. See
one_hot_table().
- requires_embeddings: bool#
Refuse to build without
embeddings, for a reproduction whose numbers are a released vector file’s. Implied byvocabulary_from="vectors".
- save_weights_only: bool#
most of a fine-tuned encoder’s file, and only needed to resume training, which no task does.
- Type:
Weights without the optimizer state
- vocabulary_from: str#
"corpus"fits the vocabulary on the training split and looks its tokens up in the vector file;"vectors"takes the file’s own vocabulary whole, so nothing the file covers is ever unknown. The second needsembeddingsand refuses to build without it.
- class pyhighlights.configurations.tasks.ToyClassWeightsTaskConfig(**data)[source]#
The synthetic corpus, weighed: the end-to-end check of the above.
- Parameters:
name (str)
loader (RegistrationKey[HighlightLoader])
weights (RegistrationKey[ClassWeights])
preprocessor (RegistrationKey[Preprocessor] | None)
save_path (str | None)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.tasks.ToyGenSPPTaskConfig(**data)[source]#
The synthetic corpus against GenSPP, searched two candidates wide.
- Parameters:
name (str)
save_path (str | None)
seeds (Sequence[int])
batch_size (int)
max_length (int | None)
vocabulary_size (int)
pretrained_model_card (str | None)
add_special_tokens (bool)
embeddings (str | None)
pretrained_tokens_only (bool)
vocabulary_from (str)
requires_embeddings (bool)
one_hot_embeddings (int | None)
callbacks (List[RegistrationKey[Callback]])
store_predictions (bool)
keep_checkpoints (bool)
save_weights_only (bool)
faithfulness (bool)
diagnostics (bool)
highlight_supervision (bool)
highlight_loss (RegistrationKey[Loss])
highlight_coefficient (float)
trainer_args (Dict[str, Any])
loader (RegistrationKey[HighlightLoader])
search (RegistrationKey[GenSPPTrainer])
preprocessor (RegistrationKey[Preprocessor] | None)
val_metrics (List[RegistrationKey[BoundMetric]])
test_metrics (List[RegistrationKey[BoundMetric]])
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.tasks.ToyTaskConfig(**data)[source]#
The synthetic corpus against FR: a smoke test that trains in seconds.
- Parameters:
name (str)
save_path (str | None)
seeds (Sequence[int])
batch_size (int)
max_length (int | None)
vocabulary_size (int)
pretrained_model_card (str | None)
add_special_tokens (bool)
embeddings (str | None)
pretrained_tokens_only (bool)
vocabulary_from (str)
requires_embeddings (bool)
one_hot_embeddings (int | None)
callbacks (List[RegistrationKey[Callback]])
store_predictions (bool)
keep_checkpoints (bool)
save_weights_only (bool)
faithfulness (bool)
diagnostics (bool)
highlight_supervision (bool)
highlight_loss (RegistrationKey[Loss])
highlight_coefficient (float)
trainer_args (Dict[str, Any])
loader (RegistrationKey[HighlightLoader])
model (RegistrationKey[Model])
preprocessor (RegistrationKey[Preprocessor] | None)
train_metrics (List[RegistrationKey[BoundMetric]])
val_metrics (List[RegistrationKey[BoundMetric]])
test_metrics (List[RegistrationKey[BoundMetric]])
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
What monitors a run, as configurations rather than as task parameters.
Two sets are registered. The loss pair stops and checkpoints on
val_loss, minimised, and is the task’s default. The score pair monitors
GeneralizationLossScore’s
combination and maximises it.
Both callbacks of a pair monitor the same quantity, and that is the point of pairing them: the epoch that stops a run has to be the epoch the run is scored on. A study is free to mix them, and gets a run whose reported model is not the one its stopping rule chose.
- class pyhighlights.configurations.callbacks.GeneralizationLossScoreConfig(**data)[source]#
A metric, charged for validation loss risen above its own best.
- Parameters:
quality (str)
loss (str)
coefficient (float)
name (str)
- coefficient: float#
How much
qualityan epoch forfeits per unit of relative loss regression. Set it per architecture, because the loss curves differ. 0 monitors the metric alone and a large value monitors the loss alone, so both are special cases of this.
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.callbacks.LossCheckpointConfig(**data)[source]#
Keep the epoch with the lowest validation loss.
- Parameters:
monitor (str)
mode (str)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.callbacks.LossEarlyStoppingConfig(**data)[source]#
Stop when the validation loss stops falling.
- Parameters:
monitor (str)
mode (str)
patience (int)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.callbacks.ScoreCheckpointConfig(**data)[source]#
Keep the epoch with the highest combined score.
- Parameters:
monitor (str)
mode (str)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.callbacks.ScoreEarlyStoppingConfig(**data)[source]#
Stop when the combined score stops rising.
- Parameters:
monitor (str)
mode (str)
patience (int)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)