Token vectors#
A model starts from token vectors, and a task says where they come from in one of three ways, which are mutually exclusive.
Name a pretrained_model_card and the tokens are embedded by that pretrained encoder, which is what Backbones covers.
Name embeddings and a vector file is fitted against the training split, which is the rest of this page.
Name neither and set one_hot_embeddings=<width> and the tokens carry no pretrained representation at all.
Pretrained token vectors#
embeddings names a GloVe-style vector file, one line per token, the token and then its vector, and it is what a reproduction of the published architectures usually needs.
task = SPPTask(
loader=HATEXPLAIN,
model=GRU_FR,
embeddings="glove.twitter.27B.25d.txt",
)
The file’s width has to be the backbone’s embedding_dim.
The ids the vocabulary hands out start at 2: rows 0 and 1 are the padding and unknown ids, both zero, kept apart so that a highlight over an unknown word is not a highlight over padding.
Its length does not have to be anything: the matrix is handed to the model through
load_embeddings(), which
sizes the table to it, so vocab_size is not a number the configuration has to have guessed.
Loaded vectors are frozen by default, and so is a one-hot table, which reaches the model the same way.
The backbone’s freeze_embeddings decides.
None, the default, freezes a loaded table and trains a randomly initialised one.
True freezes either table, and False trains either.
vocabulary_from decides which tokens are read, and the two answers are different experiments rather than a tidiness choice.
"corpus", the default, reads only the tokens the training split uses.
A token the training split never saw is unknown at evaluation whatever the file covers, and pretrained_tokens_only decides what happens to a training token the file has no vector for, dropped by default, so every row is a released vector, since a random row inside a frozen table is noise nothing can learn away.
Set it to False to keep the token with a random row instead.
"vectors" takes the file’s own vocabulary whole.
Nothing the file covers is ever unknown.
This is not a leak: the file is external, and which words it holds says nothing about which split uses them, which is why the choice is about fidelity rather than hygiene.
An evaluation token absent from the training split is embedded as zero under "corpus" and as its vector under "vectors".
"vectors" needs embeddings and refuses to build without it.
requires_embeddings=True asks for the same refusal while keeping the corpus vocabulary, for a reproduction whose numbers are a vector file’s even though its vocabulary is not.
embeddings and pretrained_model_card are mutually exclusive: a subword tokenizer brings its own embeddings.
One-hot inputs#
A corpus whose tokens are symbols rather than words has nothing to pretrain and nothing to learn:
one_hot_embeddings=<width> builds the table instead of reading one, so every token is orthonormal to every other and the padding and unknown ids, rows 0 and 1, are zero.
A frozen random table is not the same corpus to learn from, its rows are neither unit-length nor orthogonal, so the symbols arrive entangled.
The width has to match the backbone’s embedding_dim, and may exceed the vocabulary, which leaves columns that are always zero.
API#
Pretrained token vectors, read off disk into a vocabulary and a matrix.
A backbone freezes a loaded embedding table by default and trains a randomly initialised one. Reproducing a published result usually means a loaded table: the vectors are fixed, frozen, and the vocabulary is whatever the release covers. That is a corpus-and-file question rather than a model one, so it lives here and reaches the model as a tensor.
- pyhighlights.utility.embeddings.load_vectors(path, tokens=None, pretrained_only=True, generator=None)[source]#
Read a GloVe-style text file into a vocabulary and its embedding matrix.
The file is one line per token: the token, then its vector, whitespace separated. GloVe, fastText and word2vec’s text export share this format.
tokensrestricts the result to a corpus, and is usually the training vocabulary: reading every row of a released file costs memory, and a token the training split never saw is one the run has no reason to embed.Nonereads the file whole, which is what a reproduction of a run embedding from a fixed pretrained vocabulary needs. Seetokenizer()and itsvocabulary_from. It leaks nothing: which words a released vector file holds says nothing about which split uses them.pretrained_onlydecides what happens to a corpus token the file has no vector for.Truedrops it, so every row is a released vector and the unknown id absorbs the rest. A reproduction usually wants this, since a randomly initialised row inside a frozen table is noise nothing can learn away.Falsekeeps the token and gives it a random row.Rows
0and1are zeros and belong to the padding and unknown ids, matchingVocabularyTokenizer, so the ids the vocabulary hands out start at2. Both rows are zero, so neither contributes to a state, and they stay separate ids so that a highlight over an unknown word is not a highlight over padding.- Return type:
Tuple[Dict[str,int],Tensor]- Parameters:
path (str | Path)
tokens (Iterable[str] | None)
pretrained_only (bool)
generator (Generator | None)
- pyhighlights.utility.embeddings.one_hot_table(rows, width)[source]#
A one-hot embedding table: rows 0 and 1 zero, row
jthe vectore[j-2].What a corpus small enough to have no vector file wants, when its tokens are symbols rather than words: every token is orthonormal to every other and nothing about them is learned or guessed. A frozen random table is not the same thing. Its rows are neither unit-length nor orthogonal, so the symbols arrive already entangled.
Row
0is padding and row1is the unknown id, matchingVocabularyTokenizer, so a padded position and an out-of-vocabulary one each contribute nothing while remaining different ids.widthmay exceed the vocabulary, leaving columns that are always zero. That is a declared embedding size larger than the alphabet turned out to be, which costs a few unused input weights and nothing else.- Return type:
Tensor- Parameters:
rows (int)
width (int)