Backbones#
A backbone is an encoder, and it is the one piece of a model this library does not write an algorithm against.
An architecture asks a backbone for three things, which are encode, pool and output_size, and nothing in FR, MGR, MCD, DR, DAR, MRD, G-RAT or GenSPP mentions a GRU or a Transformer anywhere.
Running an architecture on a different encoder is therefore a registration key rather than a code change.
The contract#
Member |
What it returns |
|---|---|
|
Token states shaped |
|
One vector per sequence, shaped |
|
The width of both, so a head can be built before a batch exists. |
|
Optional. A backbone whose tokens are already embedded by something else says so rather than ignoring the tensor. |
What ships#
Three implementations, and each answers the discrete-choice problem in the same way while reading text differently.
GRUBackbone embeds tokens from a table and encodes them with a GRU, packing the batch so padding stays out of the recurrence.
It is the architecture the reference implementations of FR, MCD, MRD, DR, MGR and DAR use, over a frozen GloVe table.
TransformerBackbone encodes with a pretrained transformer read through transformers, which is an optional dependency, and freeze_transformer decides whether its weights move.
Selection happens over words rather than subtokens, so a word’s subtoken states are folded into one state after the encoder and never before it, which keeps the encoder on the distribution it was pretrained on.
StackedBackbone puts the two together: a frozen pretrained encoder underneath, and a GRU trained from scratch on top.
It is the shape the papers use with a better frozen representation in place of GloVe, and it avoids two failure cases at once.
A frozen transformer read by a linear selector trains a few thousand parameters, which is a probe rather than any of these architectures, while fine-tuning the transformer makes one learning rate wrong for the model, since 1e-3 destroys a pretrained encoder and 2e-5 barely moves a selector initialised from scratch.
Where a study does fine-tune a pretrained encoder, encoder_lr is the field for it: the parameters inside a backbone train at that rate and everything above them at the optimizer’s own.
A rate is defined by where a parameter sits rather than by whether it arrived pretrained, so a backbone is the encoder and a selector, predictor or guider head is not.
Keys#
Key |
What it builds |
|---|---|
|
A bidirectional GRU over an embedding table. |
|
A pretrained transformer, fine-tuned. |
|
The same transformer with its weights held. |
|
A frozen transformer read by a GRU. |
|
The two above with their embeddings frozen, which is what a genetic search over generator parameters needs. |
Pooling is the backbone’s own operation and models never reimplement it, since a recurrent encoder pools by maximum and a transformer by masked mean, and a reimplementation would hand the comparer a different summary than the predictor sees. A row with nothing unmasked pools to zeros rather than to negative infinity, which is an honest reading of an empty complement rather than a number to propagate.
What the bottleneck does and does not guarantee is measured on its own page: Interlocking and what the bottleneck guarantees.
API#
- class pyhighlights.components.models.spp.base.SPPBackbone(*args, **kwargs)[source]#
Bases:
Module,ABCBackend-specific token encoder and pooler.
The same backbone class encodes for the selector and for the predictor. When it encodes for the predictor, it receives a
selection_mask, and no information from a dropped position may reach the state of a kept one.- Parameters:
args (Any)
kwargs (Any)
- abstractmethod encode(features, mask, selection_mask=None)[source]#
Return token states shaped [B, T, D].
maskmarks the positions the encoder attends over:1for content and special tokens,0for padding.selection_maskisNonefor the selector pass. For the predictor pass it marks the kept positions. The content of a dropped position must not reach any kept state through attention, recurrence or a residual path. The state at a dropped position is unconstrained, because the predictor pools kept positions only. Without this guarantee the predictor reads more than the highlight, and the highlight is no longer its input.- Return type:
Tensor- Parameters:
features (Tensor)
mask (Tensor)
selection_mask (Tensor | None)
- load_embeddings(matrix)[source]#
Adopt a pretrained token embedding table.
Optional. A backbone whose tokens are already embedded by something else, such as a pretrained Transformer, has nothing to load. It raises rather than silently ignoring the tensor it was handed.
- Return type:
None- Parameters:
matrix (Tensor)
- abstract property output_size: int#
Token and pooled state width.
- class pyhighlights.components.models.spp.base.SPPPredictor(*args, **kwargs)[source]#
Bases:
Module,ABCMaps a pooled state to class logits.
- Parameters:
args (Any)
kwargs (Any)
- class pyhighlights.components.models.spp.base.SPPSelector(*args, **kwargs)[source]#
Bases:
Module,ABCScores each position of the selection axis as dropped or kept.
- Parameters:
args (Any)
kwargs (Any)
- abstractmethod forward(states)[source]#
Return selection logits shaped [B, T, 2].
- Return type:
Tensor- Parameters:
states (Tensor)
- threshold_parameters()[source]#
The parameters that move the selection decision, if any.
A selector whose logits end in an affine map has one: the output bias, which shifts the boundary between selecting a token and leaving it.
GenSPPTrainermutates these genes with a deviation of their own when it is given one. The default is empty, so a selector that cannot say which of its parameters carry the decision is searched with a single deviation throughout rather than with a guess about its parameter order.- Return type:
List[Parameter]
- pyhighlights.components.models.spp.base.probability(logits, of)[source]#
How much of the class distribution sits on the class
ofnames.One number per row, which is what a faithfulness term is made of: every such term compares this quantity across two inputs to the same predictor.
- Return type:
Tensor- Parameters:
logits (Tensor)
of (Tensor)
- class pyhighlights.components.models.spp.implementations.GRUBackbone(vocab_size, embedding_dim, hidden_size, freeze_embeddings=None, num_layers=1, bidirectional=True, dropout_rate=0.0)[source]#
A recurrent encoder over a token embedding table, max-pooled.
A pretrained table arrives through
load_embeddings(). For the predictor pass, the embeddings of dropped positions are zeroed before the recurrence, so their content reaches no kept state.freeze_embeddingsleft atNonetrains a randomly initialised table and freezes a loaded one, since pretrained vectors are kept fixed unless a run asks otherwise.Truefreezes either table, andFalsetrains either.- Parameters:
vocab_size (int)
embedding_dim (int)
hidden_size (int)
freeze_embeddings (bool | None)
num_layers (int)
bidirectional (bool)
dropout_rate (float)
- load_embeddings(matrix)[source]#
Replace the embedding table with
matrix, frozen unless asked not to be.The table is replaced rather than copied into, because a pretrained vocabulary is as wide as the release covers: requiring the configuration to have guessed that number in advance would make
vocab_sizea value nobody can know before the corpus is read.- Return type:
None- Parameters:
matrix (Tensor)
- class pyhighlights.components.models.spp.implementations.MLPPredictor(input_size, hidden_sizes, num_classes)[source]#
A feed-forward head mapping a pooled state to class logits.
- Parameters:
input_size (int)
hidden_sizes (List[int])
num_classes (int)
- class pyhighlights.components.models.spp.implementations.MLPSelector(input_size, hidden_sizes)[source]#
A feed-forward head scoring each position as dropped or kept.
- Parameters:
input_size (int)
hidden_sizes (List[int])
- class pyhighlights.components.models.spp.implementations.StackedBackbone(pretrained_model_card, hidden_size=128, num_features=None, freeze_transformer=True, num_layers=1, bidirectional=True, dropout_rate=0.0)[source]#
A pretrained encoder read by a recurrent one trained from scratch.
The shape of a bidirectional GRU over a frozen embedding table, with the table replaced by a pretrained transformer. The transformer is the frozen lookup and the GRU is the encoder, so everything trainable starts from scratch and one learning rate serves it. A frozen transformer read by a linear selector would instead train a few thousand parameters, which is a probe rather than a select-then-predict model.
freeze_transformerdefaults to holding the weights. A trainable encoder underneath a trainable GRU needs two learning rates, which is whatencoder_lris for.- Parameters:
pretrained_model_card (str)
hidden_size (int)
num_features (int | None)
freeze_transformer (bool)
num_layers (int)
bidirectional (bool)
dropout_rate (float)
- class pyhighlights.components.models.spp.implementations.TransformerBackbone(pretrained_model_card, num_features=None, freeze_transformer=False)[source]#
A pretrained Hugging Face encoder, mean-pooled.
For the predictor pass, dropped positions are removed from the attention mask, so no kept position attends to them. A frozen encoder stays in evaluation mode, so its dropout never runs and the same input always gives the same states.
- Parameters:
pretrained_model_card (str)
num_features (int | None)
freeze_transformer (bool)
- pyhighlights.components.models.spp.implementations.max_pool(states, mask)[source]#
The largest value each dimension takes over the unmasked positions.
A row with nothing unmasked pools to zeros rather than to
-inf. That row is the empty complement of a model that kept everything, and-infwould propagate through the predictor.- Return type:
Tensor- Parameters:
states (Tensor)
mask (Tensor)
- pyhighlights.components.models.spp.implementations.mlp(sizes)[source]#
Linear layers with GELU between them and none after the last.
- Return type:
Sequential- Parameters:
sizes (List[int])
- pyhighlights.components.models.spp.implementations.recurrent_states(encoder, layer_norm, dropout, inputs, valid)[source]#
Run
encoderoverinputs, back on the width it came in at.Shared by the two backbones that encode recurrently.
GRUBackbonehands the recurrence an embedding table, andStackedBackbonea pretrained encoder’s states. Packing keeps padding out of the recurrence.validmust mark a prefix of each row, with padding only at the end. Packing reads the firstvalid.sum()positions, so a row with a zero between ones would be packed wrongly without an error. Every mask the library builds pads at the end, including the compacted one. A row that is padding throughout gets a length of one, because a length of zero is not a sequence, and it is masked to zeros on the way out.- Return type:
Tensor- Parameters:
encoder (GRU)
layer_norm (LayerNorm)
dropout (Dropout)
inputs (Tensor)
valid (Tensor)
Backbone, selector and predictor registrations shared by the SPP models.
- class pyhighlights.configurations.backbones.FrozenTransformerBackboneConfig(**data)[source]#
The same encoder with its weights held, as a key rather than an argument.
A model names a backbone key, and build arguments reach the component a caller builds, not the ones built underneath it.
freeze_transformeris therefore set by this key rather than passed from the outside.A frozen encoder is a different experiment, not a cheaper approximation of the same one: the selector reads representations nothing adapted to its task. Two places it is the right one: an ablation, where a fine-tuned encoder would absorb the difference being measured, and a run whose size is bounded by memory rather than by patience. A frozen encoder holds no gradient, no gradient buffer and no optimizer state for its parameters.
- Parameters:
pretrained_model_card (str)
num_features (int | None)
freeze_transformer (bool)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)
- class pyhighlights.configurations.backbones.StackedBackboneConfig(**data)[source]#
A frozen transformer read by a GRU trained from scratch.
The reference implementations of FR, MCD, MRD, DR, MGR and DAR use a bidirectional GRU over a frozen pretrained table. This backbone replaces the table with a transformer. While
freeze_transformerholds, one learning rate serves it, because everything trainable starts from scratch. Fine-tuning the transformer needsencoder_lr.hidden_sizeis the library’s GRU default rather than the transformer’s width: the GRU is the encoder here, and 768 units of it is a large layer to train on a corpus of a few thousand documents.- Parameters:
pretrained_model_card (str)
hidden_size (int)
num_features (int | None)
freeze_transformer (bool)
num_layers (int)
bidirectional (bool)
dropout_rate (float)
- model_post_init(_Configuration__context)#
Runs automatically right after Pydantic instantiates an object.
- Return type:
None- Parameters:
_Configuration__context (Any)