Computational cost#
Every seed also reports what it took to produce, under cost_.
A table of F1 says which model is better and nothing about what it takes to get there.
cost_runtime_sThe seed, end to end: the model built, trained and scored.
cost_inference_batch_s,cost_inference_epoch_sThe test pass, per batch and whole. Test rather than validation, because it is the pass the reported numbers come from and it runs once, on a model that has stopped training. The batch figure is the forward passes alone; the epoch figure is what a caller waits for, batch loading included.
cost_memory_mibThe high-water mark, in mebibytes, the unit
nvidia-smiand every process monitor print. On CUDA it is the run’s own peak on its busiest device, since the counters are reset when the seed starts. It counts the memory handed to tensors, without the allocator’s cache or the CUDA context, so it reads lower thannvidia-smi. On CPU it is the process’s, which only ever rises, a second seed in the same process inherits the first’s peak.cost_parameters,cost_trainable_parameters,cost_frozen_parametersEvery parameter of the scored model, and the two halves of it. A frozen encoder is memory and compute at inference however little it learns, and a model that freezes most of itself is a different proposition to train than one that does not. Counted on the model as it was scored, so a component frozen partway through training counts as frozen, and so does a generator a genetic search settled on, which descent never moved.
cost_concurrency,cost_modelsHow many models the seed trained, and how many of them ran at once. One and one for a model trained by descent. A genetic search trains its founders plus the children of every generation that ran, fewer than
n_generationswhen it reachedstop_threshold, scored one per worker.cost_runtime_per_run_sWhat one model cost, which is what makes the rows comparable:
runtime * concurrency / models. Wall clock alone would report a search as cheap as the hours it happened to take on the machine that ran it.There is no per-model memory column to go with it. Most of what a run holds is the interpreter, torch and the corpus, resident before the first candidate exists, so dividing the peak by the workers reports less memory per model than a run holds doing nothing.
cost_memory_mibis one worker’s ceiling: the largest of a search’s workers, which is the figure comparable to a baseline, itself one worker. A search on CPU runs its workers as processes, and a search on CUDA runs them as threads, one per device. A node running the search needs that much again per worker, less whatever fork left shared between processes.cost_concurrencysays how many workers there were. The operating system offers no exact total, since pages shared by fork are counted once per process holding them.
MetricsAnalyzer reads them like any
other column, so the computational table is the same call with a different prefix:
MetricsAnalyzer(directory="results", split="cost").analyze()
and the test_ table a paper quotes is unchanged by any of this.
API#
What a run cost to produce, beside what it scored.
A table of F1 says which model is better and nothing about what it takes to get there. These are the other half: how long a seed ran, how long one inference batch and one inference pass take, how much memory the run reached (in mebibytes), how many parameters the model carries and how many of them gradient descent moves, and, for a genetic search, how many models were trained at once to produce the one that got scored.
Every column is prefixed cost_, so
MetricsAnalyzer reports them with
split="cost" and leaves them out of the test_ table a paper quotes.
Concurrency is what makes these comparable. A baseline trains one model
per seed; GenSPP trains thousands, several at a time, and reports the winner.
Wall clock alone would say a search is as cheap as the hours it happened to
take on the machine that ran it. So a run records how many models it trained
(cost_models) and how many of them were in flight at once
(cost_concurrency), and derives the per-model figures from both:
cost_runtime_per_run_sruntime * concurrency / models. Workers running at once multiply the work done in a second, so the product is worker-seconds and dividing by the models trained gives what one of them cost. A baseline, at one model and one worker, reports its own wall clock.
The workers counted are the ones that had something to do: a pool of eight scoring a population of four runs four at a time.
There is no per-model memory column, deliberately. Most of what a run
holds is the interpreter, torch and the corpus, resident before the first
candidate exists. Dividing the peak by the workers would report less than
that, which is not what any one model costs. cost_memory_mib is the
ceiling a run needs, which is the question a machine is sized by.
It is one worker’s ceiling, not a node’s. On CUDA the figure is the
memory the caching allocator handed to tensors on the busiest device. It
excludes the allocator’s reserved cache, the CUDA context and host memory, so
it reads lower than nvidia-smi. On CPU the figure is the resident memory
of the largest process. A search on CPU scores its candidates in processes of
their own, and a search on CUDA scores them in threads, one per device. In
both cases the figure is comparable to a baseline, which is one worker.
A node running a search needs that much again for each worker, less whatever
fork left shared between processes. cost_concurrency says how many workers
there were. The operating system offers no exact total for processes.
Resident pages shared by fork are counted once per process that holds them.
Additionally, ru_maxrss over children is the largest single child rather
than their sum.
- class pyhighlights.utility.cost.InferenceTimer[source]#
Times the test pass: the whole of it, and each batch of it.
Test rather than validation, because it is the pass the reported numbers come from and it runs once per seed on a model that has stopped training. The epoch figure covers what a caller waits for, loading the batches included; the batch figure covers the forward passes alone, which is what a second model on the same corpus is compared against.
- class pyhighlights.utility.cost.Meter(concurrency=1, models=1)[source]#
Times a seed and reads what it peaked at.
modelsis how many models the seed trained: one for a baseline, and a search’s whole population for GenSPP.concurrencyis how many of them ran at once. Both are settable after construction, because a search only knows how many generations it ran once it has stopped.- Parameters:
concurrency (int)
models (int)
- columns(model)[source]#
What the seed cost, as columns of
results.json.A count stays an integer: a parameter count written as
1234.0reads as a measurement of something, and these five are counts of things rather than quantities that were measured.- Return type:
Dict[str,float|int]- Parameters:
model (Module)
- pyhighlights.utility.cost.parameters(model, trainable=None)[source]#
How many parameters the model carries.
trainableselects:Truecounts what gradient descent moves,Falsewhat it does not, and left out counts both. All three are worth reporting. A frozen encoder is memory and compute at inference however little it learns, and a model that freezes most of itself is a different proposition to train than one that does not.Counted on the model as it was scored, which is the state a reader of the table gets if they load the checkpoint: a component frozen partway through training counts as frozen, however many epochs moved it first, and so does a generator a genetic search settled on rather than descended to. Trainable here means what gradient descent moves.
- Return type:
int- Parameters:
model (Module)
trainable (bool | None)
- pyhighlights.utility.cost.peak_memory()[source]#
Mebibytes at the high-water mark of the busiest worker a run used.
CUDA reports the run’s own peak on its busiest device, since
Meterresets the counters when it starts. The CPU figure is the process’s high-water mark, which only ever rises: a second seed in the same process inherits the first’s peak rather than measuring its own. That is what the operating system offers, and it is still the honest ceiling for a run of one seed.- Return type:
float