spacr.classify_classes

Define classification classes from annotation or plate-metadata values.

Each class maps a display name to either a source-column/value pair or a random-complement rule. For example:

{"infected": {"column": "annot_1", "value": 1},
 "uninfected": {"column": "annot_2", "value": 0}}

This representation supports classes derived from different annotation columns. In metadata mode, the same rules may reference plate, row, column, field, or well identifiers. normalize_settings() converts supported legacy classification settings to this representation without mutating the input mapping.

Exceptions

ClassDefinitionError

A class definition that cannot select objects, and why.

Classes

ClassRule

One class: its name, and what makes an object a member.

Functions

annotation_column_of(→ str)

Return the annotation column represented by class settings.

assign_classes(→ Any)

Label every row with its class name, or NA.

candidate_columns() → Tuple[str, ...])

The columns the Classes dict may be filled from.

class_metadata_of(→ list)

Return legacy-shaped metadata values derived from class rules.

class_names(→ List[str])

The class names, in order -- what settings['classes'] used to be.

class_rules(→ Tuple[ClassRule, ...])

The classes a settings dict defines, in the order they were given.

fold_into_classes(→ dict)

Populate compatibility keys from current class definitions.

folder_names(→ List[str])

Return ordered class-folder names for model training.

normalize_settings(→ Dict[str, Any])

Return settings with CLASSES as a dict. Never mutates.

values_in(→ Tuple[Any, ...])

The distinct values of column -- the keys the dict is populated with.

Module Contents

exception spacr.classify_classes.ClassDefinitionError[source]

Bases: ValueError

A class definition that cannot select objects, and why.

Initialize self. See help(type(self)) for accurate signature.

class spacr.classify_classes.ClassRule[source]

One class: its name, and what makes an object a member.

Parameters:
  • name – nonblank class label written to matched rows and retained as the ordered training and folder name.

  • column – source-table column compared by an explicit rule; leave it blank only for a random-complement rule.

  • value – exact value selected by equality in column for an explicit rule.

  • random_complement – when true, sample unclaimed rows with assign_classes()’s seed, up to the largest explicit class size; it cannot be combined with column or value.

Either a column/value pair, or random_complement – never both. A rule that says both would have two answers for the same object and no way to choose between them.

to_dict() → Dict[str, Any][source]

Return the serializable selector shape stored in classes.

spacr.classify_classes.annotation_column_of(settings) → str[source]

Return the annotation column represented by class settings.

The first non-empty column in the ordered classes mapping is returned. If no class rule supplies a column, the legacy annotation_column value is used.

Parameters:

settings (mapping) – Classification settings in current or legacy form.

Returns:

str – Annotation column name, or an empty string when none is defined.

spacr.classify_classes.assign_classes(frame: Any, settings: Mapping[str, Any], *, seed: int | None = 0) → Any[source]

Label every row with its class name, or NA.

The random complement is drawn from the rows NO rule claimed, sized to match the largest explicit class so the training set is not lopsided by accident – a comparison group ten times the size of the class it is compared against teaches the model the prior, not the difference.

Parameters:
  • frame – object table whose rows are to be labelled. Rule column names are resolved against this table.

  • settings – classification settings containing the ordered CLASSES definitions.

  • seed – fixes the random complement. A training set that changes every time it is built cannot be compared with the run before it.

Returns:

a Series of class names aligned to frame.

Raises:

ClassDefinitionError – a rule naming a column the table lacks.

spacr.classify_classes.candidate_columns(settings: Mapping[str, Any], available: Sequence[str] = ()) → Tuple[str, ...][source]

The columns the Classes dict may be filled from.

Under the metadata basis these are the plate’s coordinates; otherwise they are whatever annotation columns the table actually has. The GUI populates the dict’s keys from the VALUES of the chosen column, so this is the first half of “you set the column then the keys of this dict get populated”.

Parameters:
  • settings – classification settings whose resolved dataset basis decides whether coordinate metadata or annotation columns are offered.

  • available – the table’s columns, used to filter the metadata list – a database with no well column must not offer one.

spacr.classify_classes.class_metadata_of(settings) → list[source]

Return legacy-shaped metadata values derived from class rules.

Parameters:

settings (mapping) – Classification settings in current or legacy form.

Returns:

list – Class values as [[value], ...] in class order. The legacy class_metadata list is returned when no current class rule has a value; otherwise an empty list is returned.

spacr.classify_classes.class_names(settings: Mapping[str, Any]) → List[str][source]

The class names, in order – what settings['classes'] used to be.

Parameters:

settings – classification settings whose class names are requested.

Downstream (deep_spacr, model_zoo, the evaluation code) reads a list of names and should keep doing so. This is what normalize_settings() writes back under CLASS_FOLDER_NAMES so none of that has to learn the dict.

spacr.classify_classes.class_rules(settings: Mapping[str, Any]) → Tuple[ClassRule, ...][source]

The classes a settings dict defines, in the order they were given.

Parameters:

settings – classification settings containing the class definitions.

Order matters: it is the label order the model is trained with, so it has to be stable rather than whatever a set iterates in.

Raises:

ClassDefinitionError – a malformed dict, or more than one random complement – two classes both meaning “everything else” have no boundary between them.

spacr.classify_classes.fold_into_classes(settings) → dict[source]

Populate compatibility keys from current class definitions.

Parameters:

settings (mutable mapping) – Classification settings to update in place.

Returns:

dict – The same mapping, with non-empty annotation_column and class_metadata values derived from classes.

spacr.classify_classes.folder_names(settings: Mapping[str, Any]) → List[str][source]

Return ordered class-folder names for model training.

Current settings derive folder names from the ordered keys of the classes definition. For compatibility, a legacy list-valued classes setting takes precedence. When no class definitions are present, the function uses class_folder_names, which records the folders written by dataset generation. Invalid or absent definitions produce an empty list unless that recorded folder list is available.

Parameters:

settings (mapping) – Classification settings in current or legacy form.

Returns:

list of str – Folder names in class-label order.

spacr.classify_classes.normalize_settings(settings: Mapping[str, Any]) → Dict[str, Any][source]

Return settings with CLASSES as a dict. Never mutates.

Parameters:

settings – current or legacy classification settings to normalize.

The translation happens ONCE, here, so no downstream reader has to know both shapes. A settings CSV written before this produces the same classes it did before – which is the whole requirement, and what the tests assert.

class_names is written alongside, in order, because that is what deep_spacr and the evaluation code read.

spacr.classify_classes.values_in(frame: Any, column: str, *, limit: int = 100) → Tuple[Any, ...][source]

The distinct values of column – the keys the dict is populated with.

Nulls are excluded: “not annotated” is the absence of a class, and offering it as one is how a user ends up training on their own blanks.

Parameters:
  • frame – table containing the candidate class column.

  • column – column whose distinct non-null values define class choices.

  • limit – refuse to enumerate a free-form column. Past this many distinct values it is a measurement, not a label, and the Gate Editor is what turns a measurement into a class.

Raises:

ClassDefinitionError – the column is missing, or has too many distinct values to be a label.