Reproducing published results#

pyhighlights ships tools; a paper is a set of values. Those values live in a repository of their own, one per paper, so nothing here carries one study’s numbers and the library’s releases are not tied to anyone’s experiments.

A reproduction is a package of configurations and nothing else: every component it names is the library’s. It registers nothing on import, and a run builds the registry over it with the library beside it.

from pathlib import Path

import pyhighlights
import genspp2025
from cinnamon.registry import Registry

Registry.build(
    directory=Path(genspp2025.__file__).parent,
    external_directories=[Path(pyhighlights.__file__).parent],
)

directory is where registrations are executed and external_directories are indexed so their namespaces can be referenced: the reproduction is what runs, the library is what it points at. Building only the library leaves the paper keys absent, which is the point.

The shape that has worked, and what the one below follows: a package per corpus, one module per kind of thing registered, and its keys beside them in <corpus>/keys.py. Neither corpus imports the other; a shared keys.py holds only what both need, the namespace and the seeds.

Each reproduction takes its own namespace, so two papers over one corpus can prepare it differently without either becoming “the” version of it.

Written so far#

GenSPP, ACL 2025. Ruggeri and Signorelli, Interlocking-free Selective Rationalization Through Genetic-based Learning, paper, reference implementation. Two corpora against FR, MGR, MCD, G-RAT and GenSPP, with the container and the Slurm jobs that run them: nlp-unibo/pyhighlights-genspp2025. Its README carries the settings it registers and, more usefully, every place it is not the released implementation.

The corpus that reproduction reads is not a loader of its own. It is ToyLoader with a url, and the released pickle’s older schema is handled by overriding parse in the reproduction, fifteen lines, because a difference in serialisation is not a difference in what the corpus is. See Datasets.

Running a grid#

A paper’s table is a grid, not one experiment: every model over every corpus. Benchmark is that grid, a list of task keys, run in order, each writing inside the benchmark’s own directory.

from pyhighlights.configurations.keys import TOY_BENCHMARK

report = Registry.from_key(TOY_BENCHMARK).run()
report["failed"]   # tasks that raised, by name

A task that raises does not take the rest of the grid with it: the failure is recorded against that task and the run carries on, since an afternoon of training should not be lost to one bad configuration. strict=True turns that off where a run must be all-or-nothing.

task_args is handed to every task the benchmark builds, which is how a registered grid is run differently without registering a second one:

Registry.from_key(
    TOY_BENCHMARK,
    task_args={"seeds": (0,), "trainer_args": {"max_epochs": 1}},
).run()

One batch and one seed to check that every cell of a grid holds together, or a smaller batch for a card that cannot fit the registered one. A task’s manifest records the arguments it was built with, so a run overridden this way says so rather than reading like the registered configuration.

API#

Benchmarks: several tasks, one report.

A paper’s table is not one experiment but a grid of them: every model over every corpus, each with its own seeds. A benchmark is that grid: a list of task keys, run in order, and one report collecting what each of them found.

It owns no training logic. Whatever a task does, the benchmark does not need to know; it runs the key and keeps the result.

class pyhighlights.components.benchmarks.Benchmark(tasks=(), name='benchmark', save_path=None, strict=False, task_args=None)[source]#

Bases: object

Runs a list of tasks and writes down what each of them reported.

A task that raises does not take the rest of the grid with it: the failure is recorded against that task and the benchmark carries on, since an afternoon of runs should not be lost to one bad configuration. strict turns that off for a run that must be all-or-nothing.

Parameters:
  • tasks (Sequence[RegistrationKey[Task]])

  • name (str)

  • save_path (str | Path | None)

  • strict (bool)

  • task_args (Mapping[str, Any] | None)

report(results)[source]#

The grid so far, and the settings it was run with.

settings is here for the same reason a task writes a manifest: a grid run with task_args produces numbers the registered configuration would not, and a report that does not say so cannot be told from one that was never overridden.

Return type:

Dict[str, Any]

Parameters:

results (Sequence[Mapping[str, Any]])

Benchmark and analyzer registrations.

class pyhighlights.configurations.benchmarks.BenchmarkConfig(**data)[source]#

A grid of tasks, run in order.

Parameters:
  • tasks (List[RegistrationKey[Task]])

  • name (str)

  • save_path (str | None)

  • strict (bool)

  • task_args (Dict[str, Any])

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

task_args: Dict[str, Any]#

Passed to every task this benchmark builds, so a registered grid can be run differently without registering a second one: one batch and one seed to check that every cell holds together, or a smaller batch for a GPU that cannot fit the registered one. Each task’s manifest records what it was built with, so an overridden run says so.

class pyhighlights.configurations.benchmarks.HighlightPositionAnalyzerConfig(**data)[source]#

Where in the document the selector looked.

Parameters:
  • directory (str | None)

  • pattern (str)

  • bins (int)

  • absolute (bool)

  • latest (bool)

latest: bool#

a re-run task is one block of rows here and one row in the metrics table.

Type:

The newest run of each task, as every other analyzer reads

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.benchmarks.LabelStudioExporterConfig(**data)[source]#

Predicted highlights, written where a domain expert can read them.

Parameters:
  • directory (str | None)

  • pattern (str)

  • split (str)

  • latest (bool)

  • model_version (str)

  • labels (Sequence[str])

  • only (int | None)

  • column (str)

  • stem (str)

column: str#

the annotated one, or the predicted one.

Type:

Which class only names

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.benchmarks.MetricsAnalyzerConfig(**data)[source]#

One row per task, mean +/- std across seeds.

Parameters:
  • directory (str | None)

  • metrics (Sequence[str])

  • split (str)

  • pairs (bool)

  • latest (bool)

latest: bool#

The most recent run of each task, or every run a task has ever done.

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.benchmarks.PredictionAnalyzerConfig(**data)[source]#

What the selector kept, in words, one row per sample.

Parameters:
  • directory (str | None)

  • pattern (str)

  • split (str)

  • latest (bool)

latest: bool#

The newest run of each task, as MetricsAnalyzerConfig does. False reads every run a task has ever done, which is a report about the history rather than about the model.

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)

class pyhighlights.configurations.benchmarks.ToyBenchmarkConfig(**data)[source]#

The synthetic corpus, as a benchmark of one: the end-to-end smoke test.

Parameters:
  • tasks (List[RegistrationKey[Task]])

  • name (str)

  • save_path (str | None)

  • strict (bool)

  • task_args (Dict[str, Any])

model_post_init(_Configuration__context)#

Runs automatically right after Pydantic instantiates an object.

Return type:

None

Parameters:

_Configuration__context (Any)