Datasets#
A loader downloads a corpus once and hands back one pandas.DataFrame per split, as distributed.
What is done to it next is a preprocessing step, whether that means repairing splits that overlap or reducing several annotators to one judgement, because each of those is a decision about the study rather than about the corpus.
from pyhighlights.components.leakage import LeakageDetector
from pyhighlights.components.loaders import HotelLoader
from pyhighlights.components.preprocessors import LeakageRemover
loader = HotelLoader(task="hotel_Location")
splits = loader.load() # {"train", "val", "test"} -> DataFrame
LeakageDetector().check(splits) # raises: these splits overlap
splits = LeakageRemover().process(splits) # repaired, annotated split whole
Every loader produces the same five columns.
Column |
Type |
Meaning |
|---|---|---|
|
|
Row index within its split, renumbered whenever rows are dropped |
|
|
The document as distributed |
|
|
Whitespace split of |
|
|
Class label |
|
|
Per-token 0/1, |
Preprocessing#
Preprocessor takes the splits a
loader produced and returns splits of the same shape, so any of them chains with any other.
Five ship with pyhighlights, and
Pipeline runs a list of them in
order, its steps are registration keys, so a study states its pipeline in a configuration instead of in code.
LeakageRemoverWalks the splits in priority order,
test, thenval, thentrain, and keeps in each only rows no earlier split claimed and no earlier row of its own repeated. The annotated split comes first because it is the only one carrying highlights, so training and validation are what give rows up.priorityis a parameter: reproducing a training set rather than an evaluation one is a different, equally legitimate choice.removedrecords the count per split.AnnotationAggregatorReduces per-annotator labels and highlights to one of each. See HateXplain below.
LengthFilterDrops rows longer than
max_lengthtokens, rather than truncating them: a truncated row keeps its label and loses the part of the annotation that fell off the end, which is then scored as if the model had missed it.removedrecords the count per split.LabelMapperRewrites label values through a mapping, collapsing two classes into one, say.
columnmay name the per-annotator judgements rather than a resolved label, and it usually should: a post the three annotators callhatespeech,offensiveandnormalhas no majority over three classes and a clear one over two, so folding the classes before the vote is counted and folding them after give different labels.ClassWeightsReads the class frequencies of one split,
trainunless told otherwise, and returns every row untouched. What it found is inweightsandcounts, andClassWeightsTaskis what writes them somewhere they persist. Being a step rather than a calculation inside a task is what makes it run over the split the study trains on, after the filtering and the aggregation that change the frequencies.
LengthFilter and
LabelMapper ship without a
registration:
max_length and a class mapping are study-specific numbers, and inventing one in the library would make it look like a recommendation.
The other three are registered, since repairing leakage, reducing annotators and reading a split’s class frequencies are the same operation whoever asks for them.
Leakage#
Published splits are not always disjoint, and a corpus that shares rows between training and test reports highlight scores on examples the model was trained on.
Nothing downstream can detect that, so
LeakageDetector looks for it
explicitly.
It holds no data: every method takes the splits to analyse, so the same detector serves a loader’s output and a preprocessed copy of it.
detector.report(splits)A frame with one row per ordered split pair:
overlaprows of right found in left, andratio, that count over the size of right. Keys are whitespace- and case-normalised, since raw equality understates real overlap.detector.check(splits)The same report, but raising when any pair shares a row. Call it in a test or before a run. There is no tolerance to set: a shared row is leakage at any rate, and a corpus distributed with one, R2A is, is read with
reportinstead.detector.repeats(splits)Repeated rows inside each split.
checkdoes not read them: it compares splits with each other, and a split holding the same row twice is repaired byLeakageRemoverrather than refused here.
Shortcuts#
Leakage is one way a number stops meaning what it says; a shortcut is the other.
A corpus is a control only while the evidence it annotates is the only thing that solves it, so
ShortcutDetector asks what else predicts the label.
It is registered as detector/shortcut and holds no data, like the leakage detector beside it, and it reads one split’s frame rather than a mapping of them.
detector.report(frame)Every n-gram up to
max_lengthand every length threshold, in one ranking:accuracyis the best single rule over the feature,permutedthe best that feature reaches on shuffled labels, andbaselinethe majority class. An n-gram is a token sequence rather than the string it joins to, so a corpus of multi-character tokens scanned with the default emptyseparatornames its features as token tuples.detector.check(frame)The same ranking over the corpus with every annotated position replaced by a filler token, raising when anything still predicts the label above the permuted best. The gate is the ablation rather than the ranking: the annotated patterns are meant to predict, and what must be empty is what is left once they are gone. A split carrying no annotation has nothing to ablate and is refused; scan it with
report.It is a permutation test on the maximum, at a level of about
1 / permutations, so a small corpus fails it occasionally without anything being wrong. The message carries the margin: a real shortcut clears the threshold by a distance and holds as the corpus grows.
Beer and Hotel (R2A)#
Multi-aspect BeerAdvocate reviews and TripAdvisor hotel reviews, from the R2A release of Bao et al., 2018, Deriving Machine Attention from Human Rationales, the archive the selective-rationalization literature (RNP, FR, MGR, MCD, G-RAT) draws both corpora from.
- Download:
https://people.csail.mit.edu/yujia/files/r2a/data.zip(162 MB), pinned by default at23fcb4cac883ec1de86d83a7747294d7fdae10061d3803fd4c34c930e66f25deso a changed upstream fails loudly rather than being trained on. Passsha256=Noneto skip the check- Splits:
the zero-leakage partition is published as manifests, Beer at 10.5281/zenodo.22703544, Hotel at 10.5281/zenodo.22711382. They are a receipt rather than an input: the digest above plus a deterministic repair already give the same rows
- Tasks:
beer0,beer1,beer2(appearance, aroma, palate);hotel_Location,hotel_Service,hotel_Cleanliness- Labels:
binary
- Loaders:
pyhighlights.components.loaders.BeerLoaderandpyhighlights.components.loaders.HotelLoader, both over the sharedR2ALoaderparsing- Keys:
pyhighlights.configurations.keys.BEERandpyhighlights.configurations.keys.HOTEL, with the aspect as ataskvariant
Annotation lives in a file named train.
Inside the release, data/oracle/<task>.{train,dev} carry labels and text only, and the files named .test carry no rationale column at all.
The only per-token annotation is the 200-row data/target/<task>.train, which is what this line of work reports highlight scores on.
The default split map therefore reads:
{"train": "oracle/<task>.train", # labels only, large
"val": "oracle/<task>.dev", # labels only
"test": "target/<task>.train"} # 200 rows, annotated
splits is a parameter for anyone who reads the release differently.
Leakage in the distributed splits#
Measured on the splits as distributed.
test ⊂ train is the share of the annotated evaluation split found in training; val ⊂ train the share of validation found in training.
Task |
train |
val |
test |
test ⊂ train |
val ⊂ train |
|---|---|---|---|---|---|
|
14472 |
1812 |
200 |
1.000 |
0.184 |
|
101484 |
12688 |
200 |
1.000 |
0.241 |
|
150098 |
18764 |
200 |
1.000 |
0.241 |
|
32276 |
6392 |
200 |
0.445 |
0.656 |
|
28984 |
5720 |
200 |
0.360 |
0.659 |
|
25748 |
4994 |
200 |
0.415 |
0.646 |
Every Hotel aspect leaks its whole annotated evaluation split into training. All 200 annotated examples are also training examples, so any Hotel highlight score published on these splits is measured on seen data.
Beer leaks less there but keeps roughly two thirds of its validation split inside training.
The splits also repeat rows internally, 1418 of the 14472 hotel_Location training rows are duplicates.
A LeakageRemover repairs all of this: the annotated split stays whole and the offending training and validation rows are dropped.
Validation is repaired before training, so it keeps its rows and training pays for the overlap:
Task |
train |
val |
test |
dropped from train |
dropped from val |
|---|---|---|---|---|---|
|
12546 |
1767 |
200 |
1926 |
45 |
|
84735 |
12407 |
200 |
16749 |
281 |
|
125868 |
18390 |
200 |
24230 |
374 |
|
27973 |
6388 |
200 |
4303 |
4 |
|
25139 |
5720 |
200 |
3845 |
0 |
|
22436 |
4994 |
200 |
3312 |
0 |
Training loses 13-17% of its rows, the annotated split loses none, and LeakageDetector().check() passes for every task.
Numbers published on the distributed splits are reproducible by skipping the repair, and are not comparable with numbers from the repaired ones.
One upstream artifact#
Three rows carry one rationale flag more than their text has tokens, hotel_Location row 59, hotel_Cleanliness row 198, beer1 row 119, and the surplus flag is always 0.
An all-zero surplus is trimmed; any other misalignment raises, since a real shift corrupts every label after it.
HateXplain#
Twitter and Gab posts labelled for hate speech, with token-level rationales from three annotators. Mathew et al., 2021, HateXplain: A Benchmark Dataset for Explainable Hate Speech Detection.
- Download:
dataset.json(12 MB) andpost_id_divisions.jsonfrom thehate-alert/HateXplainrepository, read at commit01d7422rather than atmasterand pinned by default at63bb3340...andc2fb0d89..., the benchmark publishes no digest of its own, and a branch name pins nothing: the same key would name different rows after an upstream push. Passsha256=Noneanddivisions_sha256=Noneto skip the checks- Rows:
20148 posts, 3 annotators each; splits come from the published
post_id_divisions.json- Labels:
hatespeech,normal,offensive- Loader:
- Key:
pyhighlights.configurations.keys.HATEXPLAIN
The loader keeps every judgement. label and highlights come back unset and the raw material sits in annotator_labels and annotator_highlights, two columns this corpus carries and the others do not.
Until an
AnnotationAggregator has run,
the splits are not yet examples and datasets() says so rather than guessing.
There is no default aggregation, because reducing three annotators to one judgement is exactly the choice two studies over this corpus make differently:
labelMajority vote. 919 of the 20148 posts have all three annotators disagreeing, so no majority exists;
ties="drop"removes them, as the paper does, andties="keep"resolves them by annotator order.highlightshighlights="majority"keeps a token marked by more than half of the annotators,"union"by any of them,"intersection"by all.
HATEXPLAIN_PIPELINE chains the
aggregator with a
LeakageRemover, in that order:
repairing leakage first would measure overlap over rows the tie handling then removes.
normal posts carry no rationale by design, and 580 non-normal ones carry none either.
Both come back as all-zero highlights, “no token was marked”, not “not annotated”, so a highlight metric sees them as examples with no positive tokens rather than skipping them.
The published splits share no post id, but they do share text: 6 test posts and 3 validation posts also appear in training, and 28 training posts are duplicates of each other. Small, but nonzero, and invisible to an id-based check. Repairing after aggregation removes 37 rows in total.
ERASER#
Document classification with human evidence spans, from DeYoung et al., 2020, ERASER:
A Benchmark to Evaluate Rationalized NLP Models.
A task ships a docs directory of whitespace-tokenized documents and one JSONL file per split whose rows carry a classification and evidences, groups of [start_token, end_token) spans that become the highlights.
- Download:
https://www.eraserbenchmark.com/zipped/<task>.tar.gz, pinned by default at66e18d4e6c9df9e9f5544572b0bfe92a39673f74ecbfc3859b46cedb2f5b2dee, the benchmark publishes no digest of its own. Passsha256=Noneto skip the check- Splits:
the zero-leakage partition is published as a manifest, 10.5281/zenodo.22711411
- Tasks:
movies(1600 / 200 / 199 rows, 3.9 MB)- Labels:
binary (
NEG/POS)- Loader:
pyhighlights.components.loaders.MoviesLoader, over the sharedERASERLoaderparsing- Key:
pyhighlights.configurations.keys.MOVIES
Only single-document tasks are supported. A select-then-predict model takes one token sequence and no query, so boolq, esnli, evidence_inference, fever, multirc and scifact are refused with an explanation rather than silently folded into a document, pyhighlights has nowhere to put a query yet.
movies needs no such compromise.
Every split is annotated.
Rows with an empty evidences list, one in movies, come back as all-zero highlights.
The test split is annotated far more densely than training (a 0.31 highlight rate against 0.09), since its rationales aggregate several annotators; a sparsity target tuned on training data is not tuned for it.
movies has no cross-split leakage: one training row duplicates another, and that is all a LeakageRemover removes.
Toy#
A synthetic corpus, generated in memory. Tokens are characters, as they are in every toy corpus of this line of work: each document is filler characters with one character pattern per class inserted at a random position, and the highlights are exactly that pattern.
- Download:
none
- Loader:
- Key:
pyhighlights.configurations.keys.TOY
ToyLoader(sizes={"train": 64, "val": 16, "test": 16},
triggers=("aa", "bcd"), seed=0)
The filler is drawn from the letters no trigger uses, so a pattern can only appear where the generator put one, two adjacent filler characters can never spell "aa" if a is not a filler character.
Every split is annotated, a seed makes the corpus reproducible, and there is nothing to fetch, which makes it the cheap way to exercise a model, a configuration or a training loop before pointing it at a real corpus.
The registered configuration is a smoke test rather than the GenSPP paper’s toy corpus, which is longer, contaminated, and has three classes.
Both are
ToyLoader: the difference is
settings, and reading the released one rather than generating a new one is url.
save() writes a generated corpus in the same form that url reads, so publishing one needs no loader of its own; a corpus older than these columns is converted by overriding parse.
Corpus statistics#
pyhighlights.utility.statistics.describe() reports, per split, how long
the documents are and how much of them the annotation marks:
from pyhighlights.utility.statistics import describe
describe(BeerLoader().load())
Document length sets a token budget, what
LengthFilter drops rows over,
and highlight_rate is the ratio
SparsityPenalty compares its
threshold against.
So a sparsity target is read off the training split, and only off it.
An evaluation split’s rate is a fact to report once the numbers are in, never a target: the model has not seen that split, and tuning against it is tuning on the test set.
That leaves a real ceiling, and it belongs to the penalty rather than to any corpus.
One threshold names one corpus-level rate, so a corpus whose splits are annotated at different densities, ERASER movies marks test at 0.31 against training’s 0.09, cannot be served by a single target, and no choice of threshold fixes it.
The mismatch is measured rather than hidden:
selection_rate reports what the selector actually keeps at test, beside the highlight scores, as a share of the document’s own tokens, not of the padded batch, so it is comparable between a corpus of short documents and one of long reviews.
API#
- class pyhighlights.components.loaders.BeerLoader(task='beer0', **kwargs)[source]#
Bases:
R2ALoaderThe three Beer aspects of the R2A archive: appearance, aroma, palate.
A corpus of its own rather than a
taskstring, so the aspect is the only thing left to choose and the artefact each aspect is fetched from has somewhere to live.- Parameters:
task (str)
- class pyhighlights.components.loaders.ERASERLoader(task='movies', splits=None, url='https://www.eraserbenchmark.com/zipped/{task}.tar.gz', sha256='66e18d4e6c9df9e9f5544572b0bfe92a39673f74ecbfc3859b46cedb2f5b2dee', **kwargs)[source]#
Bases:
HighlightLoaderDocument classification with evidence spans, from the ERASER benchmark.
DeYoung et al., 2020, ERASER: A Benchmark to Evaluate Rationalized NLP Models. A task ships a
docsdirectory of whitespace-tokenized documents and one JSONL file per split whose rows carry aclassificationandevidences, which are groups of[start_token, end_token)spans into the document. Those spans become the highlights.Only single-document tasks fit the select-then-predict input, which takes one token sequence and no query.
moviesis such a task; the query-based ones would need the question folded into the document, which would change what a model is shown without saying so, and are refused until pyhighlights has a place to put a query.- Parameters:
task (str)
splits (Mapping[str, str] | None)
url (str)
sha256 (str | None)
- SHA256 = '66e18d4e6c9df9e9f5544572b0bfe92a39673f74ecbfc3859b46cedb2f5b2dee'#
The archive the
moviessplit manifest was built against. The benchmark publishes no digest of its own. One constant serves becauseTASKSis one task; a second would want this keyed by task.
- class pyhighlights.components.loaders.HateXplainLoader(url='https://raw.githubusercontent.com/hate-alert/HateXplain/01d742279dac941981f53806154481c0e15ee686/Data/dataset.json', divisions_url='https://raw.githubusercontent.com/hate-alert/HateXplain/01d742279dac941981f53806154481c0e15ee686/Data/post_id_divisions.json', sha256='63bb3340fee0ec469b09690d04cb68f7c187787dd8b83807f071892c084967fb', divisions_sha256='c2fb0d89862e7897b11ea3e9380753f15a793482b4b70ad0532dfb1212212835', **kwargs)[source]#
Bases:
HighlightLoaderHate-speech posts with per-annotator token rationales.
From Mathew et al., 2021, HateXplain: A Benchmark Dataset for Explainable Hate Speech Detection. Three annotators label every post and mark the tokens supporting a non-
normallabel.Every judgement is kept.
labelandhighlightscome back unset, and the raw material sits inannotator_labelsandannotator_highlights; how three annotators become one label and one highlight vector is a choiceAnnotationAggregatormakes, and 919 of the 20148 posts have no majority at all. Until it has run, the splits are not yet examples anddatasets()says so.- Parameters:
url (str)
divisions_url (str)
sha256 (str | None)
divisions_sha256 (str | None)
- COMMIT = '01d742279dac941981f53806154481c0e15ee686'#
the same key would name different rows after an upstream push, and a run made before it could not be told from a run made after. The benchmark publishes no digest of its own, so
SHA256andDIVISIONS_SHA256were computed against this commit, which is whatmasterresolved to as of 2026-09-15, byte for byte.- Type:
The commit both files are read at. A branch name is not a version
- DIVISIONS_SHA256 = 'c2fb0d89862e7897b11ea3e9380753f15a793482b4b70ad0532dfb1212212835'#
Digest of the official split map. Separate from
SHA256because they are separate downloads: one can change without the other.
- SHA256 = '63bb3340fee0ec469b09690d04cb68f7c187787dd8b83807f071892c084967fb'#
Digest of the 12256170-byte
dataset.jsonthat commit holds.
- class pyhighlights.components.loaders.HighlightLoader(directory=None)[source]#
Bases:
ABCCorpus source: downloads once, hands back one data frame per split.
Frames carry
COLUMNS, withhighlightsaligned totokensandNonewhere a split has no annotation. The corpus comes back as distributed, with overlapping splits, per-annotator judgements and all. Repairing or reducing it ispreprocessors, and what a study does there is its own decision; the loader that fetched the files has no business making it, and splits the user built themselves deserve the same choices.- Parameters:
directory (str | Path | None)
- knowledge()[source]#
The corpus’s knowledge base, one entry as a list of tokens.
Optional, like
SPPBackbone.load_embeddings(): most corpora have none and say so by inheriting this. A corpus that has one returns the entries in the order its annotation indexes them, because the links a split carries are positions into this sequence and nothing downstream can check an order it was never told.It is a property of the corpus rather than of a sample: every example of a run shares it, so it is loaded once and never collated into a batch.
- Return type:
Optional[Sequence[Sequence[str]]]
- class pyhighlights.components.loaders.HotelLoader(task='hotel_Location', **kwargs)[source]#
Bases:
R2ALoaderThe three Hotel aspects of the R2A archive: location, service, cleanliness.
- Parameters:
task (str)
- class pyhighlights.components.loaders.MoviesLoader(task='movies', **kwargs)[source]#
Bases:
ERASERLoaderThe ERASER
moviestask: sentiment with evidence spans.The only single-document ERASER task, and the one this line of work reports on.
- Parameters:
task (str)
- class pyhighlights.components.loaders.R2ALoader(task='hotel_Location', splits=None, url='https://people.csail.mit.edu/yujia/files/r2a/data.zip', sha256='23fcb4cac883ec1de86d83a7747294d7fdae10061d3803fd4c34c930e66f25de', archive_name='r2a.zip', **kwargs)[source]#
Bases:
HighlightLoaderBeer and Hotel aspects from the R2A archive of Bao et al., 2018.
data/target/<task>.trainis the only file in the release carrying per-token annotation, so it is the defaulttestsplit despite its name. That is the file this line of work reports highlight scores on.The distributed splits overlap. Every one of the 200 annotated rows of each Hotel aspect also appears in that aspect’s training file, and Beer keeps about two thirds of its validation split inside training. The loader returns them that way, as it returns every corpus: repairing the overlap is
LeakageRemover, whose default priority drops the offending training and validation rows and keeps the annotated split whole. A study that reproduces the release as distributed configures no such step and gets the leakage with it.The repaired splits are published as manifests, so a reader can check a run against the rows it should have seen rather than take the repair on trust: Beer at 10.5281/zenodo.22703544 and Hotel at 10.5281/zenodo.22711382. They are a receipt rather than an input:
sha256pins the upstream archive and the repair is deterministic, so the splits come out the same without fetching either.- Parameters:
task (str)
splits (Mapping[str, str] | None)
url (str)
sha256 (str | None)
archive_name (str)
- pyhighlights.components.loaders.R2A_SHA256 = '23fcb4cac883ec1de86d83a7747294d7fdae10061d3803fd4c34c930e66f25de'#
The release the split manifests on Zenodo were built against. Pinned so a reproduction fails loudly on a changed upstream rather than training on it.
- class pyhighlights.components.loaders.ToyLoader(sizes=None, triggers=('aa', 'bc'), length=20, vocabulary_size=20, contaminations=0, min_chunk=2, seed=0, url=None, sha256=None, member='corpus.pkl', archive_name=None, train_ratio=0.8, val_ratio=0.2, split_seed=0, **kwargs)[source]#
Bases:
HighlightLoaderSynthetic corpus: each class is a set of character patterns to be found.
Tokens are characters, as they are in every toy corpus of this line of work: the generator samples an alphabet, and a released one is read back a character at a time. A trigger is therefore a string of characters like
"aa", not a phrase.A class’s trigger is a conjunction: every pattern in it has to appear for the class to hold. One pattern is the short spelling of a conjunction of one, so both of these are triggers:
triggers = ["aa", "bc"] # a pattern per class triggers = [["aba", "baa"], ["baa", "abb"]] # two, both required
The second form is the one worth having. When no pattern belongs to a single class (
baasits in both classes above), no single n-gram identifies a class, and a model that memorises one cannot pass. It also makes the gold highlight several disjoint spans rather than one run, which is the shape a real highlight has.Patterns overwrite filler at disjoint positions, in a random order and with at least one filler character between them. Overwriting rather than inserting is what keeps every document exactly
lengthtokens long: a class whose patterns are longer would otherwise produce longer documents, and the document’s own length would say which class it is without reading a character of it. The released generator overwrites for the same reason. A generated sample must satisfy its own class and no other; one that does not is drawn again.Every split is annotated and the highlights are exactly the patterns, which makes it the cheap way to exercise a model, a configuration or a training loop end to end.
It also reads. A toy corpus has a life: it is generated, published, and read back by whoever reproduces the numbers it produced. Those are the same corpus at three moments rather than three kinds of thing, so they are one loader.
ToyLoader(...).save(path)writes whatToyLoader(url=...)reads, in this library’s own columns, and a published toy corpus is then a URL and a digest in a configuration rather than a loader somebody has to write.parse()is the hook for a corpus older than those columns.urldecides which half runs, and a configured source is never fallen back on: a loader given aurlit cannot read raises rather than generating, because a corpus of the right shape and the wrong content is the one failure nothing downstream can see.contaminationsscatters proper chunks of the patterns through the filler, which is what stops a fragment from being enough to classify: without itbcoccurs nowhere but insideabc, so detectingbcnames that class exactly as well as detectingabcdoes. Off by default, since the registered corpus is a smoke test and a cheap one is the point. A reproduction of published numbers pointsurlat the corpus those numbers are for, rather than generating one of the same shape.Whatever the settings,
ShortcutDetectoris what says the corpus is a control: it removes the annotated positions and requires that nothing left predicts the label. Run it on a corpus that was read as readily as on one that was generated. Being published is no evidence of being sound.- Parameters:
sizes (Mapping[str, int] | None)
triggers (Sequence[str | Sequence[str]])
length (int)
vocabulary_size (int)
contaminations (int)
min_chunk (int)
seed (int)
url (str | None)
sha256 (str | None)
member (str)
archive_name (str | None)
train_ratio (float)
val_ratio (float)
split_seed (int)
- ALPHABET = 'abcdefghijklmnopqrstuvwxyz'#
Where filler characters come from, minus whatever the triggers use.
- ATTEMPTS = 100#
Draws allowed per sample before the placement is called impossible.
- divide(frame)[source]#
The splits the corpus carries, or the ones the ratios cut into it.
A corpus
save()wrote carries asplitcolumn, so it comes back divided exactly as it was generated. Regenerating a corpus and re-splitting a corpus are different operations and only one of them is reproducible from the file.A flat corpus has no such column. It is divided the way the GenSPP baselines divide theirs: the first
train_ratiois train, the rest test, andval_ratioof train is sampled off it undersplit_seed. The draw israndom.Random, the generator this loader samples everything else with, so one seed family explains a corpus and its division.- Return type:
Dict[str,DataFrame]- Parameters:
frame (DataFrame)
- fetch()[source]#
The corpus file, downloading and unpacking the artifact if needed.
urlmay be an archive, published or local, or a bare corpus file. A record usually holds the archive rather than a loose file, because the archive is what carries the manifest, the licence and the citation beside the data.A downloaded corpus is pinned or refused. What arrives is read with
pandas.read_pickle(), which executes what the file says to execute, so an unpinned URL is arbitrary code running before a row is read. A local path is the caller’s own file and is read as given.- Return type:
Path
- parse(frame)[source]#
A freshly read frame, as this library’s columns.
The hook a corpus older than those columns overrides. Everything
save()writes arrives here already in them, so the default only fills in what is derivable and says which column is missing otherwise. A corpus that has to be guessed at is one nobody can check.- Return type:
DataFrame- Parameters:
frame (DataFrame)
- place(trigger, generator)[source]#
One sample’s tokens and highlight, from a class’s patterns.
The patterns overwrite filler at disjoint positions, in a random order: disjoint with a gap so that each is its own span, random so that the order they were listed in is not a feature.
The positions are drawn uniformly over the arrangements that fit. With widths
wandkruns,length - sum(w) - (k - 1)tokens of filler are free to sit in thek + 1gaps; choosingkcuts out ofslack + kpicks one such arrangement, and each is equally likely. Sampling each start independently and rejecting the overlaps would not be uniform, and where a pattern sits is a channel this corpus exists to keep shut.Contaminating chunks are placed in the same draw as the patterns. Placing them afterwards, or keeping only the samples that came out valid, conditions the arrangement on the class: reject the draws where a chunk completes a second copy of the pattern and the characters beside the highlight stop being class-independent. That dependence survives removing the highlight and is a shortcut. Here nothing is rejected, because a filler character separates every run and no chunk spells a pattern, so no arrangement can be invalid.
- Parameters:
generator (Random)
- read()[source]#
Generate a corpus, or read the one
urlnames.A configured source is never fallen back on. A loader given a
urlit cannot read raises, where generating instead would hand back a corpus of the right shape and different content. That is the one failure a synthetic corpus cannot afford, because nothing downstream can see it.- Return type:
Dict[str,DataFrame]
- satisfied(text)[source]#
Which classes’ conjunctions
textholds, as a set of labels.A valid sample satisfies exactly its own. Public because it is what the corpus means: a shortcut scan asks it about texts this never wrote.
- Return type:
set- Parameters:
text (str)
- save(path)[source]#
Write the corpus to
path, splits and all, for reading back.The file this writes is what
read()reads: the library’s own columns plus asplit, so a generated corpus can be published and then loaded by URL without anyone writing a second loader for it.- Return type:
Path- Parameters:
path (str | Path)
- pyhighlights.components.loaders.fetch_archive(url, archive, root, sha256=None)[source]#
Fetch an archive once, unpack it once, and return where it landed.
The two steps belong together, and both are already idempotent: a file that is there is not fetched again, and a directory that is there is not unpacked again. So a loader asks for the unpacked corpus rather than sequencing a download and an extraction of its own, and a second run over a warm cache touches the network not at all.
sha256is the digest of the archive as distributed, and the reason a reproduction fails loudly on a changed upstream rather than training on it.- Return type:
Path- Parameters:
url (str)
archive (Path)
root (Path)
sha256 (str | None)
Leakage analysis: what the splits of a corpus share with each other.
Published splits overlap more often than their papers admit. Every annotated
row of an R2A Hotel aspect also sits in that aspect’s training file, and
nothing downstream can detect it: a model reports highlight
scores on rows it was trained on and the numbers look ordinary. Detection
lives here; repair is a preprocessing step, in
pyhighlights.components.preprocessors.
- class pyhighlights.components.leakage.LeakageDetector(key='text', normalize_keys=True)[source]#
Bases:
objectReports what a set of splits shares, and refuses splits that share a row.
Holds no data of its own: every method takes the splits to analyse, so one detector serves whichever loader or preprocessing stage is being checked.
- Parameters:
key (str)
normalize_keys (bool)
- check(splits)[source]#
Return the report, raising when any split pair shares a row.
Meant for a test or the top of a run. A corpus that shares rows between train and test reports highlight scores on examples the model was trained on, and every number downstream is quietly wrong.
There is no tolerance to set. A shared row is leakage at any rate, and a corpus distributed with one, as R2A is, is read with
report(), which says how much it shares without refusing it.Between splits only. A split that holds the same row twice passes here:
repeats()is what counts those, andLeakageRemoveris what drops them.- Return type:
DataFrame- Parameters:
splits (Mapping[str, DataFrame])
- pyhighlights.components.leakage.duplicates(splits, key='text', normalize_keys=True)[source]#
Count repeated rows inside each split.
Keys are whitespace- and case-normalized as in
leakage(), andnormalize_keys=Falsecompares them exactly as the corpus spells them. An empty key is not a repeat of another empty one.- Return type:
Dict[str,int]- Parameters:
splits (Mapping[str, DataFrame])
key (str)
normalize_keys (bool)
- pyhighlights.components.leakage.leakage(splits, key='text', normalize_keys=True)[source]#
Report, for each ordered split pair, how much of the right split the left one already contains.
ratiois the share of right rows found in left, so the test row of a train/test pair answers “how much of my evaluation set did I train on”. Keys are whitespace- and case-normalized unless told otherwise, since raw equality understates real overlap. An empty key matches nothing, andsizecounts it all the same: a row with no text is not shared with anything, and hiding it would report a share of a corpus that is smaller than the one on disk.- Return type:
DataFrame- Parameters:
splits (Mapping[str, DataFrame])
key (str)
normalize_keys (bool)
- pyhighlights.components.leakage.normalize(value)[source]#
One comparable key: collapsed whitespace, stripped, lower-cased.
objectrather thanstrbecause the column holds whatever the loader parsed. A missing value normalizes to the empty string, and an empty key is never counted as shared: two rows a corpus left without text are two unusable rows rather than a duplicate. They are dropped byremove_leakage(), which is where a row leaves a corpus.- Return type:
str- Parameters:
value (object)
Shortcut analysis: what predicts the label without reading the evidence.
A corpus is a control only while the evidence it annotates is the only thing that solves it. If some other feature separates the classes, a model can score well on the task and badly on the explanation, and nothing downstream can tell the two apart. That is the failure a synthetic corpus exists to rule out.
The toy corpus is the main case. Scoring string-matching baselines by highlight F1 asks whether a selection matches the annotation rather than whether it solves the task: a competing n-gram can score near zero against the annotation and still predict the class perfectly, and that is exactly the shortcut a control has to exclude. What is asked here is the other question, over every n-gram the corpus actually contains.
The same scan answers it of a real corpus. Whether punctuation predicts a class is this question, asked of words instead of characters.
What a clean report does and does not say. No scan proves the annotated
evidence is the only solution: any feature fine enough to index the sample
separates it. The claim is bounded, and the bound is the feature family:
single n-grams up to max_length, and sequence length. A conjunction of two
n-grams that neither one predicts alone is outside it, and is what
ablated() is for: remove the evidence and re-run, and nothing expressible
in any feature family should be left to find.
Every score is read against a permuted control. With thousands of features the best one beats the majority baseline on chance alone, so the threshold is not the baseline but the best score any feature reaches once the labels are shuffled.
- class pyhighlights.components.shortcuts.ShortcutDetector(max_length=4, separator='', tokens='tokens', label='label', seed=0, permutations=30)[source]#
Bases:
objectReports what predicts a corpus’s label, and refuses an unexplained winner.
Holds no data of its own, like
LeakageDetector: every method takes the frame to analyse.- Parameters:
max_length (int)
separator (str)
tokens (str)
label (str)
seed (int)
permutations (int)
- check(frame)[source]#
Ablate the annotated evidence, then refuse anything that still predicts.
The gate is the ablation, not the ranking. A scan of the corpus itself cannot be a gate: the annotated patterns are meant to predict, and so is anything that co-occurs with them, so the ranking is full of features that are supposed to be there. Worse, a pattern shared by two of three classes still separates the third by its absence. Being shared makes a pattern insufficient, not uninformative.
What the corpus has to guarantee is the other direction: with the evidence gone, nothing is left. That claim is not bounded by a feature family, because there is nothing for any family to find.
The threshold is the best score any feature reaches on shuffled labels, which is the multiple-comparison control: with thousands of features the best of them beats the majority baseline by chance, and a real shortcut is one that beats what chance already offers.
That makes this a permutation test on the maximum, at a level of about
1 / permutations, so a small corpus fails it occasionally without anything being wrong. The message carries the margin for that reason: a real shortcut clears the threshold by a distance and holds as the corpus grows, where noise clears it by a hair and decays towards the baseline. Re-run on more rows before believing a narrow failure.- Return type:
DataFrame- Parameters:
frame (DataFrame)
- permutations#
Label shuffles the threshold is the best of. The threshold is the largest score any feature reaches on any shuffle, which makes
check()a permutation test on the maximum at a level of about1 / permutations, so a small value refuses a clean corpus often. The n-grams are counted once whatever this is, so more shuffles cost little.
- pyhighlights.components.shortcuts.ablated(frame, filler='▮', separator='')[source]#
The corpus with every annotated position replaced by one filler token.
Removing the evidence is the decisive test: whatever is left cannot be the thing the corpus is about, so a scan that finds anything here has found a shortcut. Unlike a scan of the corpus itself, this one is not bounded by a feature family, because there is nothing left for any family to find.
The replacement is a token the alphabet does not contain, so the hole cannot spell anything; its width is preserved, which keeps the document’s length out of the comparison.
separatorrejoinstextthe way the corpus spells it: empty for a character corpus, a space for words. The scan readstokensand nevertext, so this only decides whether the returned frame is legible, but an ablated word corpus that readsthe▮brownfoxis a frame nobody can check by eye.- Return type:
DataFrame- Parameters:
frame (DataFrame)
filler (str)
separator (str)
- pyhighlights.components.shortcuts.incidence(documents, max_length=4, separator='')[source]#
Which documents hold each n-gram, up to
max_lengthtokens.A document counts once for an n-gram it repeats: the question is whether the n-gram is there, and a count is a different feature.
separatorjoins the tokens for display: empty for a character corpus like the toy one, a space for a corpus of words.An n-gram is the token sequence, never the string it joins to. With an empty separator
("ab", "c")and("a", "bc")both spellabc, and they are two patterns: one says the corpus holds a tokenabfollowed by a tokenc. They are counted apart, and where a corpus makes that distinction visible the names are written as token tuples rather than as the joined string.- Return type:
Dict[str,List[int]]- Parameters:
documents (Iterable[Sequence[str]])
max_length (int)
separator (str)
- pyhighlights.components.shortcuts.lengths(frame, tokens='tokens', label='label', seed=0, permutations=30)[source]#
How well
len(tokens) >= tpredicts, over every threshold there is.The channel an n-gram scan cannot see. Patterns of different lengths in a corpus of variable-length documents make the document’s own length say which class it is, and no feature over its content is involved.
- Return type:
DataFrame- Parameters:
frame (DataFrame)
tokens (str)
label (str)
seed (int)
permutations (int)
- pyhighlights.components.shortcuts.ngrams(frame, max_length=4, separator='', tokens='tokens', label='label', seed=0, permutations=30)[source]#
Every n-gram in the corpus, scored by how well its presence predicts.
- Return type:
DataFrame- Parameters:
frame (DataFrame)
max_length (int)
separator (str)
tokens (str)
label (str)
seed (int)
permutations (int)
- pyhighlights.components.shortcuts.scan(features, labels, seed=0, permutations=30)[source]#
Score every feature by the best rule over it, against a permuted control.
featuresmaps a name to the document indices that hold it. Returns one row per feature (support,accuracy,permuted,mutual_information) sorted by accuracy, withbaseline(the majority class) carried on every row so a number is readable on its own.- Return type:
DataFrame- Parameters:
features (Mapping[str, Sequence[int]])
labels (Sequence[int])
seed (int)
permutations (int)
What a corpus looks like before anything trains on it.
Two numbers decide a run’s hyperparameters and neither is in the frame: how long the documents are, which sets a token budget, and how much of a document its annotation marks, which is what a sparsity target aims at. Both are read off the training split. An evaluation split’s rates are describable after the fact, never a target, and the model has not seen them.
- pyhighlights.utility.statistics.COLUMNS = ['split', 'rows', 'annotated', 'tokens_min', 'tokens_mean', 'tokens_median', 'tokens_p95', 'tokens_max', 'highlight_rate', 'highlights_mean']#
Column order of the frame
describe()returns.
- pyhighlights.utility.statistics.describe(splits)[source]#
Report document length and annotation density for each split.
highlight_rateis marked tokens over all tokens of the annotated rows. It is the same ratioSparsityPenaltycompares itsthresholdagainst, so a training split’s rate is where that threshold comes from.highlights_meanis marked tokens per annotated row, which a rate alone hides: the same rate covers three tokens of thirty and thirty of three hundred.A split with no annotated row reports
NaNrather than zero. Zero would read as an annotation that marks nothing.- Return type:
DataFrame- Parameters:
splits (Mapping[str, DataFrame])
Preprocessing: everything done to a corpus after it is parsed.
A loader hands back a corpus as distributed. What happens next, such as
repairing leaking splits or turning per-annotator judgements into one label
and one highlight vector, is an editorial choice, and two studies over the same
corpus routinely make it differently. So it is a component: pick one, or
compose several with Pipeline, and the choice is named in a
configuration rather than buried in the loader that fetched the files.
- class pyhighlights.components.preprocessors.AnnotationAggregator(labels=(), highlights='majority', ties='drop', annotator_labels='annotator_labels', annotator_highlights='annotator_highlights')[source]#
Bases:
PreprocessorCollapses per-annotator labels and highlights into one of each.
A corpus annotated by several people (HateXplain has three per post) is loaded with every judgement kept, because reducing them is a choice the corpus does not make for you:
label: majority vote. A post whose top label is shared with another has no majority;ties="drop"removes it, as the HateXplain paper does, andties="keep"resolves it by annotator order. An even number of annotators splitting evenly is a tie as much as three splitting three ways is, and a single annotator is never one.highlights:"majority"marks a token more than half the annotators marked,"union"any,"intersection"all.
Rows whose annotation vectors are all the wrong width, and classes carrying no annotation by design, come back all-zero: “no token was marked”, not “not annotated”.
- Parameters:
labels (Sequence[str])
highlights (str)
ties (str)
annotator_labels (str)
annotator_highlights (str)
- class pyhighlights.components.preprocessors.ClassWeights(split='train', classes=None)[source]#
Bases:
PreprocessorComputes a split’s class weights and hands every row back unchanged.
A weight is a number about a corpus, so it is read off the corpus rather than typed into a configuration by hand and hoped to still be right. Doing it here rather than inside a task makes the reading a named step: it is in the pipeline, so it is in the manifest, and it runs over the split the study actually trains on. That is the split after whatever filtering and aggregation came before it, which is what changes the frequencies.
Nothing is added to the frames.
weightsandcountshold what it found, andClassWeightsTaskis what writes them down.- Parameters:
split (str)
classes (int | None)
- class pyhighlights.components.preprocessors.KnowledgeWeights(entries, split='train')[source]#
Bases:
ClassWeightsReads a split’s example-to-entry links and hands every row back unchanged.
The knowledge-axis counterpart of
ClassWeights, and aClassWeightsTaskruns it unchanged:weightsis one positive weight per entry rather than one per class, andcountshow many examples link to each.classesis inherited and unused: a knowledge base has entries rather than classes, andentriesis what sizes this one.entriesis the size of the base, and it is required rather than inferred. The largest index a split happens to use is not the size of the knowledge base, and guessing it would hand the loss a weight vector one entry short. Such a vector broadcasts against the wrong axis or silently drops the last entry.- Parameters:
entries (int)
split (str)
- class pyhighlights.components.preprocessors.LabelMapper(mapping, column='label')[source]#
Bases:
PreprocessorRewrites label values through a mapping.
Collapsing classes is an editorial choice like any other: a study may fold HateXplain’s
offensiveintonormaland train on two classes. It has to happen before the votes are counted, not after: a post two annotators callhatespeechand one callsoffensivehas a majority either way, but one where the votes arehatespeech,offensiveandnormalhas one only once the last two are the same class. Socolumnmay name the per-annotator judgements as readily as a resolved label, and a list-valued column is mapped element by element.A value the mapping does not name is left as it is.
- Parameters:
mapping (Mapping[Any, Any])
column (str)
- class pyhighlights.components.preprocessors.LeakageRemover(priority=('test', 'val', 'train'), key='text', normalize_keys=True)[source]#
Bases:
PreprocessorDrops the rows one split shares with another, its repeats and its blanks.
Which split gives a row up is what
prioritydecides, and it is a judgement about the study rather than about the corpus: keeping the annotated split whole is right when highlight scores are the result, and wrong when the training set is what must be reproduced. Hence a preprocessor: the loader has no business making that call, and splits the user built themselves are just as valid an input as the distributed ones.removedrecords how many rows each split lost, blank rows included.- Parameters:
priority (Sequence[str])
key (str)
normalize_keys (bool)
- class pyhighlights.components.preprocessors.LengthFilter(max_length, column='tokens')[source]#
Bases:
PreprocessorDrops rows longer than
max_lengthtokens.Nothing here is any one corpus’s. The policy, how many tokens and why, is a configuration and lives in the benchmark that adopts it.
Truncating would keep the row and lose the tokens, which for a corpus scored on highlights means scoring against an annotation whose tail was cut off. A study that caps length to bound its compute drops the row instead, and says how many it dropped.
- Parameters:
max_length (int)
column (str)
- pyhighlights.components.preprocessors.PRIORITY = ('test', 'val', 'train')#
Priority runs from the split that must stay intact to the one that can afford to lose rows. The annotated split comes first: it is the only one carrying highlights, so it is the one worth protecting.
- class pyhighlights.components.preprocessors.Pipeline(steps=())[source]#
Bases:
PreprocessorRuns preprocessors in order, each over what the last returned.
Steps are registration keys rather than instances, so a pipeline is something a configuration states: aggregate the annotations, then repair the leakage the aggregation left behind. A second study over the same corpus states a different one without touching either step.
- Parameters:
steps (Sequence[RegistrationKey])
- class pyhighlights.components.preprocessors.Preprocessor[source]#
Bases:
ABCTurns one set of splits into another.
Implementations take the frames a loader produced and return frames of the same shape, so any of them can be chained with any other.
- pyhighlights.components.preprocessors.class_weights(labels, classes=None)[source]#
Inverse-frequency weights,
n / (classes * count)per class.The same formula and the same refusals as scikit-learn’s
compute_class_weight(class_weight="balanced", ...), which raises both for a class declared but absent fromyand for one present but not declared. Written out rather than depended on: scikit-learn is thirty megabytes this package does not otherwise need, for six lines, and it returns an array where a configuration wants a list.What a weighted cross entropy needs when the classes are not the same size: the rarer a class, the more a mistake on it costs, so a model cannot score well by never predicting it.
classesis the number of classes the model has, and defaults to the largest label seen plus one. Pass it wherever the split might not contain every class: an inferred count would silently give the model one output fewer than the task has.- Return type:
List[float]- Parameters:
labels (Sequence[Any])
classes (int | None)
- pyhighlights.components.preprocessors.link_weights(links, entries)[source]#
One positive weight per knowledge base entry,
negatives / positives.What a binary cross entropy over the knowledge axis needs. The axis is imbalanced twice over: most examples instantiate nothing, and among those that do, most entries still do not apply. A single positive weight cannot separate the entry that fires on half the annotated examples from the one that fires on three of them. The deciding entry is frequently the rare one, which is the reason this is a vector.
linksis one sequence of entry indices per example, orNonewhere the corpus annotates none; unannotated examples are skipped, since they are evidence of nothing rather than evidence of absence.- Return type:
List[float]- Parameters:
links (Sequence[Any])
entries (int)
- pyhighlights.components.preprocessors.remove_leakage(splits, priority=('test', 'val', 'train'), key='text', normalize_keys=True)[source]#
Return splits sharing no row, walking them in
priorityorder.Each split keeps only rows no earlier split claimed and no earlier row of its own repeated, so the result has neither cross-split leakage nor internal duplicates. Splits missing from
priorityare handled last, in their original order, and the returned mapping keeps the input order.A row whose key is empty or missing is dropped wherever it sits. It is not a duplicate of the next empty row, since comparing them would claim a leak the corpus does not have. It is nothing to train on either: an empty text is an empty token sequence, and a highlight over it names no word. The check reads the normalized key, so a blank row goes whether or not
normalize_keysis set.sample_idis renumbered, since it indexes rows within a split.- Return type:
Dict[str,DataFrame]- Parameters:
splits (Mapping[str, DataFrame])
priority (Sequence[str])
key (str)
normalize_keys (bool)