Data#
A batch lives on two axes, and which one a field is on is the whole of this page.
The subtoken axis [B, T] is what the encoder reads, features, attention_mask and word_ids.
A backbone tokenizes however it likes, so T is its business: a vocabulary tokenizer gives one position per word, a subword tokenizer gives more.
The word axis [B, W] is what a person reads and what a selection is made over, mask and highlight_true.
A word is the unit the corpus annotates, the unit a sparsity target is a fraction of, and the unit an export shows.
It is also the same unit whichever backbone read the text, which is what lets two backbones share a table.
word_ids maps between them, -1 where a position belongs to no word.
For a vocabulary tokenizer the axes coincide and it is the identity.
Selecting over words#
A selection is made over words by default.
The alternative, select_over="subtoken" on any SPP model, lets a model keep un and drop ##fair, and an export then reports the word unfair: a highlight that is not what the predictor read, in a library whose claim is that it is.
Over words it splits none, by construction.
Pooling happens after the encoder, never before, so a pretrained backbone still attends over its own subtokens. What the predictor is then given is the selection spread back over every piece of each word it kept.
subtoken is kept so the difference can be measured rather than argued about.
Special tokens#
[CLS] and [SEP] are added by default.
They carry no word, so word_ids is -1 there and a selector never sees them, but the encoder attends over them always, whatever the selection, because that is what it was pretrained to read.
Dropping them moves the encoder’s token states away from what it would otherwise produce, and with a frozen backbone nothing can adapt to the difference.
This is what attention_mask is for, and why it is not mask: one says what is read, the other what may be chosen.
Corpora and their loaders have their own page, Datasets, and what a model does with the two axes is Backbones.
API#
- pyhighlights.components.data.COLUMNS = ('sample_id', 'text', 'tokens', 'label', 'highlights')#
Columns every corpus frame carries, in order. A loader parses into them and a preprocessor returns them, so a split is the same shape wherever it came from.
highlightsisNonewhere a split has no annotation, andlabelisNoneonly until a preprocessor has resolved one.
- class pyhighlights.components.data.HighlightCollator(tokenizer, max_length=None, knowledge_size=None)[source]#
Bases:
objectTokenize, align highlights to subtokens, and dynamically pad a batch.
- Parameters:
tokenizer (HighlightTokenizer)
max_length (int | None)
knowledge_size (int | None)
- knowledge(examples)[source]#
The links as a multi-hot
[B, M], orNonewithout a base.Bis the batch andMisknowledge_size, one column per knowledge base entry. These are the same two axes every knowledge-side tensor in this library carries, againstTfor the token axis.A row is
-1where the example carries no annotation and0/1where it carries one, so an example annotated with an empty set is a row of zeros rather than a row of-1.- Return type:
Tensor|None- Parameters:
examples (Sequence[HighlightExample])
- class pyhighlights.components.data.HighlightExample(sample_id, tokens, label, highlights=None, knowledge=None)[source]#
Bases:
object- Parameters:
sample_id (int)
tokens (Sequence[str])
label (int)
highlights (Sequence[int] | None)
knowledge (Sequence[int] | None)
- knowledge: Sequence[int] | None = None#
Which entries of a knowledge base explain this example, as indices into it.
Nonewhere the corpus annotates none, and an empty tuple where it annotates that none apply. The two are different claims: an example that instantiates no entry carries a gold label the grounded pipeline is scored on, not a missing one.
- class pyhighlights.components.data.HighlightTokenizer(*args, **kwargs)[source]#
Bases:
ProtocolWhat the collator needs of a tokenizer, as a structural type.
A
Protocolis not a base class anybody inherits: it lists an attribute and a method, and any object carrying both satisfies it as far as a type checker is concerned.VocabularyTokenizerandHuggingFaceTokenizerinherit nothing and both pass, and so would a tokenizer written outside this package. That is the point, since one of the two wraps an object this library does not own.- encode(tokens, max_length=None)[source]#
Encode
tokens, keeping at mostmax_lengthencoded positions.The budget is counted on the axis the encoder reads, special tokens included, rather than on source tokens. The two coincide only for a one-token-to-one-id encoder. A tokenizer that cannot honour the bound exactly may return more:
HighlightCollatorclamps the result.- Return type:
- Parameters:
tokens (Sequence[str])
max_length (int | None)
- class pyhighlights.components.data.HuggingFaceTokenizer(pretrained_model_card, add_special_tokens=True, use_fast=True, **tokenizer_kwargs)[source]#
Bases:
objectFast-tokenizer adapter preserving source-token alignment.
- Parameters:
pretrained_model_card (str)
add_special_tokens (bool)
use_fast (bool)