Analyzers#
An analyzer reads a results directory back and answers one question about it, returning a pandas.DataFrame rather than printing, so the same analyzer serves a notebook, a test and a LaTeX table.
MetricsAnalyzerOne row per task, one column per metric,
mean +/- stdacross seeds. It walks everyresults.jsonbeneath the directory, so it reads one task or a whole benchmark without being told which. A metric a task never measured reads as-: a grid rarely reports the same set everywhere, and an unannotated corpus has no highlight F1 to give.pairs=Truekeeps the(mean, std)tuples, whichlatex_table()renders as$12.34_{\pm 0.56}$, escaping the underscores every metric name carries. Since a task keeps every run it has ever done,latestdecides which the table is about: the newest run of each task by default, every run when asked. Runs are grouped by the name a result reports rather than by its directory, so a task that was renamed or moved is still the task its results say it is.HighlightPositionAnalyzerWhere in the document the selector looked, binned as a share of the document so lengths are comparable, and how much it kept. A selector that has learned nothing still selects something; position is what tells the two apart, since a model keying on the opening tokens of every document scores like one that found the highlight. A model with several selectors stores one mask per head; the analysis reads the reported head, which is the one every reported metric scored.
Positions are word positions on either selection axis: a selection made over subtokens is folded through
word_idsfirst, exactly asPredictionAnalyzerfolds it, and a word split into several pieces is one word however many of its pieces were selected.absoluteasks the other question. A model keying on the first three words of every document does that regardless of how long the document is, and a share hides it, in the first bin of a short document, and in the first tenth of a long one. Columns then cover the firstbinswords, and a selection past them counts in the total without a column of its own, so the shares sum to less than one by however much the tail holds.PredictionAnalyzerWhat the selector kept, in words, one row per sample: the gold label and the predicted one, the words the run selected, and the highlight they spell out. A stored prediction is token ids and masks, enough to score, and unreadable on its own, so the corpus is reloaded and joined back to it. The run’s
manifest.jsonnames the loader and the preprocessor that produced it, and those are the keys the analyzer builds: a corpus loaded from anywhere else is a different corpus. The corpus is not stored beside the predictions because it would be stored once per run and per seed.Selections are folded from token positions back to words through the
word_idsthe batch carries, so a subword model reports words like every other, and a word counts as selected when any of its subtokens was. A sample the corpus no longer holds is skipped rather than failing the split: a corpus that changed under a run is worth reporting around.This one resolves keys, so the registry has to be built before it runs, inside a cinnamon script it already is.
LabelStudioExporterThe same rows, written where a domain expert can read them: one Label Studio file per run, pre-annotated with the words the model selected, so judging a highlight is a review rather than a fresh annotation. A corpus with no highlight annotation is exactly the case this is for, there is no F1 to report, and what the highlights are worth is a question for somebody who knows the domain.
onlynarrows the export to samples of one class, which a corpus annotated for a rare one needs: the negatives are most of it and the interesting highlights are all on the positives.columndecides which class that is, and the two answers ask different questions,"label"takes the samples that carry the class and asks whether the model found the right words in them;"predicted"takes the samples the model called that class and asks whether the words it kept justify the call. The second is the one available on a corpus with no annotation to select by, and the one that surfaces a confident mistake.labelsnames the span label the file carries, andlabel_studio()does the conversion on any frame carrying the columnsPredictionAnalyzerreports.One
label-studio-seed=<seed>.jsonper seed, beside the predictions it came from. A run’s seeds are separate predictions of the same samples, so one merged file would show each sentence once per seed with different words marked, which is not a thing to read. Nothing to export writes no file rather than an empty one.
The run column of both prediction analyzers is a path under the directory they were pointed at, fr/2026-09-09T16-13-00 rather than 2026-09-09T16-13-00, because a benchmark writes <name>/<started> per task, and two tasks that started inside the same second would otherwise report themselves as the same run.
None of them is interactive and none plots. An analyzer that asks which folder you meant cannot run unattended, and a figure is a presentation choice that belongs to whoever is writing the paper.
API#
Analyzers: what to make of a directory full of results.
A run leaves results.json files and pickled predictions behind. An analyzer
reads them back and answers one question about them, such as what the
numbers are or where the selector looked. It returns a
pandas.DataFrame rather than printing, so the same analyzer serves a
notebook, a test and a LaTeX table.
Nothing here is interactive and nothing plots. An analyzer that asks which folder you meant cannot run unattended, and a figure is a choice about presentation that belongs to whoever is writing the paper.
- class pyhighlights.components.analyzers.Analyzer(directory=None)[source]#
Bases:
ABCReads a results directory and reports on it.
- Parameters:
directory (str | Path | None)
- class pyhighlights.components.analyzers.HighlightPositionAnalyzer(directory=None, pattern='predictions-seed=*.pkl', bins=10, absolute=False, latest=True)[source]#
Bases:
AnalyzerWhere in a document the selector looked, and how much it kept.
Reads the predictions a task stored. A selector that has learned nothing still selects something, and position is what tells the two apart: a model keying on the first tokens of every document scores like a model that found the highlight, until you look at where it selected.
Positions are word positions, since that is what a selection is made over.
selection_rateis the mean of the per-document rates, which is whatpyhighlights.utility.metrics.SelectionRatereports and what a study’s tables are built from. Pooling instead (all kept words over all words) gives a different number on documents of different lengths, since it weights a long document more than a short one, and two quantities under one column name is how a table stops being comparable to itself.Positions are reported as a share of the document, so documents of different lengths are comparable.
absolutereports word positions instead, which is the other question: a model keying on the first three words of every document does that regardless of how long the document is, and a share hides it in the first bin of a short document and the first tenth of a long one. Columns then cover the firstbinswords, and a selection past them is counted in the total without a column of its own, so the reported shares sum to less than one by however much the tail holds.latestreads the most recent run of each task, asMetricsAnalyzerandPredictionAnalyzerdo: a re-run task is one block of rows here and one row there, rather than one in some reports and two in others.- Parameters:
directory (str | Path | None)
pattern (str)
bins (int)
absolute (bool)
latest (bool)
- class pyhighlights.components.analyzers.LabelStudioExporter(directory=None, pattern='predictions-seed=*.pkl', split='test', latest=True, model_version='pyhighlights', labels=('highlight',), only=None, column='label', stem='label-studio')[source]#
Bases:
PredictionAnalyzerWrites each run’s predicted highlights where an annotator can read them.
An expert judging whether a highlight is the right one needs it in front of the text, not as a list of word indices. This writes one Label Studio file per run, pre-annotated with what the model selected, so the reading is a review rather than a fresh annotation.
onlynarrows the export to samples of one class, which is what a corpus annotated for a rare one needs: the negatives are 97% of it and the interesting highlights are all on the positives.columndecides which class that is, and the two answers are different questions."label"selects the samples that carry the class, and asks whether the model found the right words in them."predicted"selects the samples the model called that class, and asks whether the words it kept justify the call. It is the only one of the two available on a corpus with no annotation to select by, and the one that surfaces a confident mistake. Any columnPredictionAnalyzerreports may be named.One file per seed, beside the predictions it came from. A run’s seeds are separate predictions of the same samples, so merging them would show the same sentence once per seed with different words marked, which is not a thing to read.
- Parameters:
directory (str | Path | None)
pattern (str)
split (str)
latest (bool)
model_version (str)
labels (Sequence[str])
only (int | None)
column (str)
stem (str)
- class pyhighlights.components.analyzers.MetricsAnalyzer(directory=None, metrics=(), split='test', pairs=False, latest=True)[source]#
Bases:
AnalyzerOne row per task, one column per metric,
mean ± stdacross seeds.Walks every
results.jsonbeneath the directory, so it reads one task or a whole benchmark without being told which.metricsselects and orders the columns; left empty, every metric found is reported.A metric a task did not measure is
-rather than missing, since a grid of models rarely reports exactly the same set: an unannotated corpus has no highlight F1 to give.A task keeps every run it has ever done, one timestamped directory each, so
latestdecides which of them the table is about: the most recent run of each task by default, every run when asked. Reporting all of them by default would grow the table each time a configuration is re-run, and the numbers a paper quotes are the last ones measured.- Parameters:
directory (str | Path | None)
metrics (Sequence[str])
split (str)
pairs (bool)
latest (bool)
- reports()[source]#
Every run found, newest last, one per task when
latest.The name comes out of
results.json, which the run wrote itself.latest_runs()is what does the grouping.- Return type:
List[Dict[str,Any]]
- pyhighlights.components.analyzers.PREDICTIONS = 'predictions-seed=*.pkl'#
a run stores one file per seed, and every analyzer here reads all of them.
- Type:
The predictions one seed left behind. A glob rather than a name
- class pyhighlights.components.analyzers.PredictionAnalyzer(directory=None, pattern='predictions-seed=*.pkl', split='test', latest=True)[source]#
Bases:
AnalyzerWhat the model selected, in words, one row per sample.
A stored prediction is token ids and masks: enough to score, unreadable on its own. This joins it back to the corpus it came from, so a row says which words the selector kept and what the predictor made of them.
The corpus is not stored beside the predictions, where it would be stored once per run and per seed, so it is reloaded. The run’s
manifest.jsonnames the loader and the preprocessor that produced it, and those keys are what get built here: a corpus loaded from anywhere else is a different corpus.A selection is made over words, so it is already in the unit a person reads: a row lists word positions and the words at them. A run that selected over subtokens instead is folded back through the
word_idsthe batch carries, and a word counts as selected when any of its subtokens was. That is why that setting cannot say what the predictor actually read, and why it is not the default.A task builds its loader and its preprocessor from their keys alone, with no overrides, so rebuilding those keys rebuilds exactly the corpus the run trained against. An override changes which key a task holds, and that key is the one the manifest wrote down.
The registry has to be built before this runs, since it resolves the keys the manifest names. Inside a cinnamon script it already is.
A sample the corpus no longer holds, or one whose words it places differently, is left out rather than refused, because the rest of the split is still worth reading.
skippedcounts both per predictions file, and a file that lost anything says so through the module’s logger.- Parameters:
directory (str | Path | None)
pattern (str)
split (str)
latest (bool)
- corpus(run)[source]#
The split these predictions were made on, keyed by sample id.
- Return type:
Dict[int,Series]- Parameters:
run (Path)
- frames()[source]#
One frame per predictions file, with the file that produced it.
Per file rather than one frame for everything, so a caller that writes something back can write it beside the predictions it came from. Every seed of a run has its own file and its own predictions of the same samples, which is a separate thing to read rather than N copies of one.
- Return type:
Iterator[Tuple[Path,DataFrame]]
- runs()[source]#
The run directories to read, newest per task when
latest.The name comes out of the run’s
manifest.json, falling back to the directory the stamp sits in for a run that wrote none.run_directories()is what does the grouping, for every analyzer that reads stored predictions.- Return type:
List[Path]
- pyhighlights.components.analyzers.escape(value)[source]#
Make a name safe to typeset. An unescaped
_is a subscript.- Return type:
str- Parameters:
value (Any)
- pyhighlights.components.analyzers.label_studio(frame, model_version='pyhighlights', labels=('highlight',), score=1.0)[source]#
Predicted highlights as Label Studio pre-annotations.
Reads the columns
PredictionAnalyzerreports (tokens,selected,label,predictedandtext), so it converts any frame carrying them, whatever produced it. A frame withouttextfalls back to joining the tokens with a space.Each selected word becomes one span. The whole document goes in
dataalongside the gold and predicted label, so a reader sees what the model was given and what it made of it, not only what it highlighted.- Return type:
List[Dict[str,Any]]- Parameters:
frame (DataFrame)
model_version (str)
labels (Sequence[str])
score (float)
- pyhighlights.components.analyzers.latest_runs(items, latest=True)[source]#
One entry per task, newest last, or every entry when
latestis off.A task keeps every run it has ever done, one timestamped directory each, so a re-run or a requeued job would otherwise read as extra rows: a second line in a table, or the same sample exported twice for an annotator.
Each item arrives as
(name, stamp, item). Grouping is by the name the run reported rather than by its directory: the stamp is a path component, and a task that has been renamed or moved is still the task its own record says it is. Recency is the stamp rather than the order the walk produced, because those two disagree in exactly that case. A run undernew-name/2026-09-02is walked before one underold-name/2026-09-01, and taking the last walked would report the older one as current.Both analyzers that read a directory of runs group them this way, and they differ only in where the name comes from: a metrics report carries it, and a predictions file has it in the manifest beside it.
- Return type:
List[TypeVar(T)]- Parameters:
items (Iterable[Tuple[str, str, T]])
latest (bool)
- pyhighlights.components.analyzers.latex_table(frame, precision=2, percentage=True)[source]#
The rows of a table body,
&-separated and\\-terminated.A
(mean, std)pair, which is whatMetricsAnalyzer(pairs=True)reports, becomes$12.34_{\pm 0.56}$. Anything else is escaped and written as it is: metric names carry underscores, and LaTeX reads those as subscripts.The header row is included; the
tabularwrapper is not, since its column specification belongs to the table this goes into.- Return type:
str- Parameters:
frame (DataFrame)
precision (int)
percentage (bool)
- pyhighlights.components.analyzers.offsets(tokens, text=None)[source]#
Character span of each token inside
text.Label Studio addresses a span by character offset into the text it shows, so the offsets have to be into the document the corpus wrote rather than into a rejoining of its tokens: a character corpus spells
aabwhere a rejoining with spaces spellsa a b, and every offset after the first would then be wrong.textdefaults to" ".join(tokens), which is what a corpus of words holds. Tokens are located in order, so repeated tokens take successive occurrences rather than the first one every time.- Return type:
List[Tuple[int,int]]- Parameters:
tokens (Sequence[str])
text (str | None)
- pyhighlights.components.analyzers.readability(frame, by='label')[source]#
Selection size, rate and span count, per run and per class.
Over the rows
PredictionAnalyzer.analyze()returns, which is where a selection is already in words and beside the document it came from.Two columns nothing else reports. A span count, because a rate cannot tell two readable phrases from eight scattered fragments. Twenty per cent of a document in two spans is something a person can read, and the same share in eight is not. And the split by class, because domain experts asked whether the highlights of negative examples differ from those of positive ones, and a pooled average over a split that is 97.7% negative reports the negative examples’ number and calls it the model’s.
byis"label"for the annotation and"predicted"for what the model called it. Both are worth reading: the first shows what was missed, the second what was invented.- Return type:
DataFrame- Parameters:
frame (DataFrame)
by (str)
- pyhighlights.components.analyzers.reported_head(masks)[source]#
The head a model with several selectors is scored on.
The reported metrics score
SPP.inference_head. MGR stores only that head at test, and every other model reports head 0, so a stored mask with a head axis is read at head 0.- Return type:
ndarray- Parameters:
masks (ndarray)
- pyhighlights.components.analyzers.run_of(path, directory)[source]#
Which run a predictions file belongs to, as a path under
directory.The stamp alone does not name a run: a benchmark writes
<name>/<started>per task, and two tasks that started in the same second would report the same run while being different runs.- Return type:
str- Parameters:
path (Path)
directory (Path)
- pyhighlights.components.analyzers.seed_of(path)[source]#
The seed in
predictions-seed=42.pkl, or"?"if it says none.- Return type:
str- Parameters:
path (Path)
- pyhighlights.components.analyzers.separator_of(tokens, text)[source]#
What joins this corpus’s tokens, read off the text it wrote.
A corpus of words joins with a space and a corpus of characters joins with nothing (
ToyLoaderwrites"".join(tokens)), so assuming one of them spells the other’s documents wrongly.- Return type:
str- Parameters:
tokens (Sequence[str])
text (str | None)