spacr.filters

The filters table: gates, written back to the database as columns.

A gate drawn in the Gate Editor is a shape on two measurements. What a user wants out of it is a LABEL – this object is in my population, that one is not – attached to the objects themselves, so it can be merged with anything else they have measured. That is what this module writes:

one column per gate, named after the gate, 1 inside and 0 outside.

Why a separate table. The columns could be added to cell, but a gate is not a measurement: it is an interpretation, it is re-drawn often, and it belongs to whichever object the user was looking at. Writing into the measurement tables would mix the two, and re-gating would rewrite a table that the measure step owns. filters is written only by this module, so it can be deleted and rebuilt at any time without losing a measurement.

The bootstrap. The first gate exported has to create the table, and the table has to carry enough identity that a filter can be merged back onto ANY object table or onto png_list. spaCR joins those on plate / row / column / field plus the object label (and the timepoint, when the database is a timelapse), so those are exactly the columns build_filters_frame() collects – from every object table present, not just the anchor, because a gate drawn on nucleus measurements has to merge onto nuclei.

A tolerant reader, a strict writer. Databases in the wild carry both the current column names and the ones spaCR wrote years ago, so reading accepts either spelling. Everything written out uses the canonical name, so the filters table itself never needs the alias machinery.

Exceptions

FilterError

A filter that cannot be built or written, with the reason.

Functions

annotate_from_gates(→ pandas.Series)

Label every object from SEVERAL gates at once.

build_filters_frame(→ pandas.DataFrame)

The filters table's identity, before any gate is written to it.

build_filters_from_relationships(→ pandas.DataFrame)

The filters table as a COPY of the relationships table.

build_relationships_frame(→ pandas.DataFrame)

Every object relationship in the database, as one flat table.

choose_anchor(→ str)

The table the metadata is taken from.

column_name_for(→ str)

The column a gate is written to.

column_names(→ Tuple[str, ...])

Every column in table, in the order the database declares them.

combination_label(→ str)

The class name for one combination of gate memberships.

ensure_filters_table(→ pandas.DataFrame)

Return the filters table, building it the first time.

ensure_relationships_table(→ pandas.DataFrame)

The relationships table, built on demand if the mask step never did.

export_annotation(→ Tuple[str, int])

Write a gate-derived annotation to filters as one column.

export_gate(→ Tuple[str, int])

Write one gate to filters as a 1/0 column.

gate_mask_over_table(→ Tuple[pandas.DataFrame, ...)

Apply a gate to EVERY object, reading only the columns it needs.

identity_columns_of(→ Dict[str, str])

Map canonical identity name -> the spelling table uses.

key_columns(→ List[str])

The identity columns present, in join order.

object_tables(→ Tuple[str, ...])

The object tables this database actually has, in preference order.

png_crop_type(→ Optional[str])

WHICH OBJECT the crops in png_list are pictures of.

read_identity(→ pandas.DataFrame)

The identity columns of one table, and nothing else.

read_sampled(→ pandas.DataFrame)

Read table, optionally taking only a fraction of its rows.

require_full_identity(→ None)

Raise unless keys names every identity column AND the object label.

resolve_column(→ Optional[str])

The first alias present in columns, matched case-insensitively.

row_count(→ int)

How many objects the table has -- what a sample is a fraction OF.

rowid_expression(→ Optional[str])

Which spelling of the implicit row id this table leaves usable.

sampling_clause(→ str)

A SQL fragment taking roughly fraction of the rows.

table_names(→ Tuple[str, ...])

Every table in the database, in the order SQLite lists them.

write_filters_table(→ None)

Replace filters with frame.

write_relationships(→ pandas.DataFrame)

Build and store the relationships table. Called after the mask step.

Module Contents

exception spacr.filters.FilterError[source]

Bases: ValueError

A filter that cannot be built or written, with the reason.

Initialize self. See help(type(self)) for accurate signature.

spacr.filters.annotate_from_gates(frame: pandas.DataFrame, gates, names: Sequence[str], *, mode: str = 'binary') → pandas.Series[source]

Label every object from SEVERAL gates at once.

binary

1 when the object is inside EVERY chosen gate, 0 otherwise. The intersection, because that is what “annotate based on all the gates” means when the answer has to be one column.

multiclass

one class per observed combination of memberships. Only combinations that actually occur become classes – enumerating all 2^n would offer classes with no objects in them, which no classifier can learn and every class-balance report would then have to explain.

Parameters:
  • frame – the objects to label. It must contain every measurement column referenced by the selected gates; the returned Series keeps this frame’s index and row order.

  • gates – a GateSet.

  • names – which gates to use, in the order the label reads.

Returns:

a Series aligned to frame – integers for binary, class names for multiclass.

Raises:

FilterError – no gates chosen, or a mode that does not exist.

spacr.filters.build_filters_frame(db_path: str) → pandas.DataFrame[source]

The filters table’s identity, before any gate is written to it.

The three steps asked for, in order:

  1. which tables are in the database (object_tables());

  2. the object numbers from EVERY object table, plus the metadata – plate, row, column, field, object – and the crop paths from png_list when it exists;

  3. one table carrying enough to merge a filter onto any of them.

Every object table contributes rows, not just the anchor. A gate drawn on nucleus measurements has to merge onto nuclei, and a filters table built only from cell could not express that. Each object also carries an in_<table> flag saying which tables it appears in, which is what makes a merge predictable rather than a thing you discover by counting rows.

Parameters:

db_path – the measurement database.

Returns:

one row per distinct object key.

Raises:

FilterError – no object table to build from.

spacr.filters.build_filters_from_relationships(db_path: str) → pandas.DataFrame[source]

The filters table as a COPY of the relationships table.

“to make the filters table this the relationships table should be the base (it should be copied) and filters added.” A copy rather than a merge onto it: a filter then automatically carries every relationship, and there is one definition of what an object is rather than two that can drift.

Parameters:

db_path – the measurement database. Despite being a build step this can WRITE: the relationships table it copies is created on demand when the mask step never wrote one. An existing one is used as it stands – pass through write_relationships() first if the masks have changed since it was written.

spacr.filters.build_relationships_frame(db_path: str) → pandas.DataFrame[source]

Every object relationship in the database, as one flat table.

One row per object of the FINEST kind present, carrying the label of each coarser object it belongs to. Flat rather than a link table per pair because every question asked of it – “cells with more than three pathogens”, “the mean pathogen intensity per cell” – is a group-by on the parent, and a flat table answers those with no joins at all.

The parent link is cell_id on the child, which is what spacr.io._read_and_join_tables() uses and what the measure step writes. A child measured without a parent mask has no link, and is carried with a null parent rather than dropped: the object exists, and saying it has no parent is different from pretending it is not there.

Parameters:

db_path – the measurement database. Read-only: the frame is returned and nothing is stored, so a caller that wants it on disk goes through write_relationships() or ensure_relationships_table().

Raises:

FilterError – no object table to build from.

spacr.filters.choose_anchor(tables: Sequence[str]) → str[source]

The table the metadata is taken from.

“usually be cell, but if cell does not exist then another table should be used”. Preference order, so a database of only nuclei anchors on nuclei and a database of only organelles on organelles – there is no table that has to be present.

Parameters:

tables – the object tables actually present, as returned by object_tables(). Only membership is tested – the preference comes from the order of OBJECT_TABLES, not from the order given here, so putting nucleus first does not make it the anchor while cell is also in the sequence.

Raises:

FilterError – nothing to anchor on. Naming the tables that WERE found is the difference between a fixable message and a shrug.

spacr.filters.column_name_for(gate_name: str) → str[source]

The column a gate is written to.

The gate’s own name, with anything that is not a letter, digit or underscore collapsed to an underscore. The name is what the user reads in both places, so it is kept recognisable rather than hashed.

Parameters:

gate_name – the gate’s display name. It is stripped, every run of characters outside [0-9A-Za-z_] collapses to one underscore, leading and trailing underscores are dropped, and a result starting with a digit is prefixed g_ – legal in quoted SQLite but not in the tools that read the table afterwards. Two gates whose names differ only in punctuation therefore land on the SAME column, and export_gate() replaces rather than suffixes.

Raises:

FilterError – a name with nothing usable left in it, which would otherwise become an anonymous column called _.

spacr.filters.column_names(db_path: str, table: str) → Tuple[str, ...][source]

Every column in table, in the order the database declares them.

Parameters:
  • db_path – path to the SQLite database.

  • table – table to describe.

Returns:

the column names; empty when the table does not exist.

spacr.filters.combination_label(memberships: Sequence[bool], names: Sequence[str]) → str[source]

The class name for one combination of gate memberships.

Named after the gates the object IS in, in gate order, so the label reads as what it means – live+CD8 rather than class_3. An object in no gate is none, which is a real class: “outside everything” is a finding, not a gap.

Parameters:
  • memberships – one truth value per gate, for ONE object, positionally aligned to names.

  • names – the gate names, in the order the label reads. Order is part of the class name – live+CD8 and CD8+live are two different classes – so the same order has to be used for every object, which is why annotate_from_gates() fixes it once. The two sequences are zipped, so a longer one is truncated silently rather than reported.

spacr.filters.ensure_filters_table(db_path: str, *, rebuild: bool = False) → pandas.DataFrame[source]

Return the filters table, building it the first time.

Parameters:
  • db_path – the measurement database. An existing table is read without modification; the first call, or a rebuild, writes the identity table derived from the database’s object relationships.

  • rebuild – discard and rebuild. The identity is derived entirely from the object tables, but any gate columns already written are LOST – which is why this is a parameter and not something the export path does on its own.

Returns:

the table as it now stands on disk.

spacr.filters.ensure_relationships_table(db_path: str, *, rebuild: bool = False) → pandas.DataFrame[source]

The relationships table, built on demand if the mask step never did.

“if the relationships table does not exist when attempting to make the filters table, then the relationships table gets generated first.” A user who gated before ever running the new mask step must not hit an error for it.

Parameters:
  • db_path – the measurement database. Opened for WRITING whenever the table has to be built, so a database that is only readable can be gated on if it was bootstrapped already, and not otherwise.

  • rebuild – derive the table again from the object tables and replace whatever is stored. Relationships are never refreshed on their own, so a stored table built before the masks changed stays stale until this is passed – which is exactly what write_relationships() does.

spacr.filters.export_annotation(db_path: str, frame: pandas.DataFrame, labels: pandas.Series, column: str, *, object_type: str | None = None) → Tuple[str, int][source]

Write a gate-derived annotation to filters as one column.

Through the same path a single gate takes, so an annotation and a filter are the same kind of thing in the database and merge the same way.

Parameters:
  • db_path – the measurement database. filters is built first if it is not there yet, and rewritten whole afterwards.

  • frame – the objects the labels describe. Only its identity columns are read – the measurements are not needed – and it must carry the FULL identity (IDENTITY_COLUMNS plus object_label, in the canonical spellings), because an object label repeats in every field. See require_full_identity().

  • labels – one label per row of frame, taken POSITIONALLY: the Series index is discarded, so labels that were reindexed or sorted away from the frame’s row order would be attached to the wrong objects. A length mismatch is not checked here and surfaces as a pandas ValueError. Where several rows share one object key the first label wins.

  • column – the name to write it under, sanitised by column_name_for(); an existing column of that name is dropped and replaced. Objects outside frame are left NULL rather than filled, since a multiclass annotation has no zero.

  • object_type – which kind of object frame holds, for the same reason export_gate() takes it: without it an annotation of nucleus 2 also lands on cell 2.

Returns:

(column name, objects labelled).

Raises:

FilterError – frame is missing any part of the object identity, or shares no object key with the filters table.

spacr.filters.export_gate(db_path: str, frame: pandas.DataFrame, inside: numpy.ndarray, gate_name: str, *, rebuild: bool = False, object_type: str | None = None) → Tuple[str, int][source]

Write one gate to filters as a 1/0 column.

Parameters:
  • db_path – the measurement database whose filters table is created when absent and then rewritten with the gate column.

  • frame – the objects the gate was evaluated on. Must carry the FULL identity (IDENTITY_COLUMNS plus object_label, in the canonical spellings – see require_full_identity()); the measurements are not needed and not read.

  • inside – boolean mask over frame, True for objects in the gate.

  • gate_name – names the column.

  • rebuild – rebuild the identity table first, discarding gate columns.

  • object_type – which kind of object frame holds – the name of the measurement table it was read from. GIVE IT WHENEVER YOU KNOW IT. filters holds cells, nuclei and pathogens side by side, and object_label is unique within a field only for ONE kind: a merge that leaves the type out writes a gate drawn on nucleus 2 onto cell 2 as well, which is the same wrong answer require_full_identity() exists to prevent, one axis over. None keeps the type-blind merge for a caller that genuinely cannot say, and for a database with no object_type column.

Returns:

(column name, objects marked).

Raises:

FilterError – the frame is missing any part of the object identity, or the mask does not match it.

Objects NOT in frame get 0, not null. A user who gated on a 20% sample and exported it would otherwise get a column that is null for four objects in five, and null is not what “outside the gate” means. The GUI re-reads the full table before calling this precisely so that the 0s are real – see gate_mask_over_table().

spacr.filters.gate_mask_over_table(db_path: str, table: str, gates, gate_name: str) → Tuple[pandas.DataFrame, numpy.ndarray][source]

Apply a gate to EVERY object, reading only the columns it needs.

The point of the sampling setting is that a user gates on a fraction of a large table; the point of this is that the export does not. The gate’s own columns plus the identity columns are read in full – a handful out of hundreds, so this stays cheap even where reading the whole table is not.

Parameters:
  • db_path – path to the SQLite measurement database, opened read-only.

  • table – the object table to evaluate. Only its identity columns and the measurement columns used by the selected gate path are read.

  • gates – a GateSet.

  • gate_name – which gate in it.

Returns:

(identity frame, mask) ready for export_gate().

Raises:

FilterError – a column the gate needs is not in the table.

spacr.filters.identity_columns_of(db_path: str, table: str) → Dict[str, str][source]

Map canonical identity name -> the spelling table uses.

Missing columns are absent from the map. A table without a field column still merges on the keys it does have; refusing outright would rule out databases that are perfectly usable.

Parameters:
  • db_path – opened read-only, so unlike a missing table a database file that is not there is fatal: it raises sqlite3.OperationalError instead of an empty map.

  • table – the table to inspect; it need not be an object table. A name that is not in the database is not an error – PRAGMA table_info returns nothing for it, so the result is an empty map, the same answer as a table that exists and carries no identity at all.

spacr.filters.key_columns(frame: pandas.DataFrame) → List[str][source]

The identity columns present, in join order.

One definition, used by the bootstrap, the merge and the writer, so a filter can never be merged on a different key than it was built with.

Parameters:

frame – a frame whose columns use the CANONICAL identity names – what read_identity() returns, having aliased them on the way out. Membership is tested exactly and case-sensitively, with no alias lookup: a raw SELECT * from a database that spells them plate and row contributes none of the well keys, and such a frame yields ["object_label"] alone. That is why nothing merges on the result of this function without putting it through require_full_identity() first – an object label is unique only WITHIN a field, so merging on it alone collides objects across plates, wells and fields. Only column names are inspected; no values are touched.

spacr.filters.object_tables(db_path: str) → Tuple[str, ...][source]

The object tables this database actually has, in preference order.

Step 1 of the bootstrap: check which tables are in the database. Only tables that exist AND carry an object label count – a table can be present and empty of the identity a filter needs, and discovering that at merge time rather than here would produce a filter that quietly matches nothing.

Parameters:

db_path – the measurement database. Only the names in OBJECT_TABLES are looked for, so a table holding some other kind of object is invisible here however it is keyed – add it to that tuple rather than expecting discovery.

spacr.filters.png_crop_type(db_path: str) → str | None[source]

WHICH OBJECT the crops in png_list are pictures of.

Parameters:

db_path – measurements database whose png_list schema is read.

png_list names its id column after the object it cropped – cell_id, pathogen_id, organelle_id – so the crop mode a run used is recoverable from the schema rather than having to be remembered.

This is what makes a crop path attachable to the right row. Without it the label is just an integer, and matching integers across object types hands nucleus 2 the crop of CELL 2.

Returns:

'cell', 'pathogen', … or None when there is no png_list, no recognised id column, or no crops at all.

spacr.filters.read_identity(db_path: str, table: str) → pandas.DataFrame[source]

The identity columns of one table, and nothing else.

Identity only, because this runs over every object table in the database and a measurement table is wide – hundreds of columns of which four matter. SELECT * here is the difference between a bootstrap that takes a moment and one that reads the whole database.

Parameters:
  • db_path – path to the SQLite measurement database. It is opened read-only, so a missing file raises sqlite3.OperationalError instead of being created as an empty database.

  • table – must carry an object_label column – that name, in any case. Every other identity column is optional and absent from the returned frame; the object label is not, because without it the rows cannot be told apart.

Raises:

FilterError – table has no object label column.

spacr.filters.read_sampled(db_path: str, table: str, *, fraction: float = 1.0, limit: int | None = None) → pandas.DataFrame[source]

Read table, optionally taking only a fraction of its rows.

Sampling happens in SQL where it can, because the point is to not read the rows. A table that shadows every row id alias is read whole and sampled afterwards – slower, but correct, and it says so in the log rather than quietly returning everything.

Parameters:
  • db_path – path to the SQLite measurement database, opened read-only.

  • table – the table to read. Its name is quoted as one SQLite identifier; a missing table raises sqlite3.OperationalError.

  • fraction – how much of the table to read, in (0, 1].

  • limit – a hard row cap applied after the fraction.

spacr.filters.require_full_identity(keys: Sequence[str], what: str) → None[source]

Raise unless keys names every identity column AND the object label.

The invariant a write-back depends on: object_label is unique only within one field of one well of one plate, so a merge keyed on anything less than the full identity silently joins one plate’s object 7 onto another’s. The partial-identity tolerance elsewhere in this module (identity_columns_of(), build_filters_frame()) is about READING a database that is missing a column; writing a per-object column back into it is the one operation where a partial key is a wrong answer rather than a reduced one.

Parameters:
  • keys – what key_columns() returned for the frame.

  • what – names the operation in the error message, e.g. "a gate".

Raises:

FilterError – any identity column, or the object label, is absent.

spacr.filters.resolve_column(columns: Iterable[str], aliases: Sequence[str]) → str | None[source]

The first alias present in columns, matched case-insensitively.

Returns the name AS SPELLED in the table, which is what has to go in the SQL – returning the canonical spelling would produce queries for columns that are not there.

Parameters:
  • columns – the column names as the table spells them, typically from column_names(). Where a table carries two spellings that differ only in case, the one appearing LAST is the one returned.

  • aliases – candidate spellings in preference order; the first one present wins, which is why the canonical name leads every entry of IDENTITY_ALIASES. A single-element tuple is the way to ask “does this exact column exist, whatever its case”.

spacr.filters.row_count(db_path: str, table: str) → int[source]

How many objects the table has – what a sample is a fraction OF.

Parameters:
  • db_path – path to the SQLite measurement database, opened read-only.

  • table – goes into the query as a quoted name, so it must be a table that exists – a typo raises sqlite3.OperationalError rather than counting zero, and a caller sizing a sample should check with table_names() first.

spacr.filters.rowid_expression(columns: Iterable[str]) → str | None[source]

Which spelling of the implicit row id this table leaves usable.

This is not hypothetical. Every spaCR measurement table has a column called rowID – the row of the plate – and SQLite matches column names case-insensitively, so rowid in a query means THAT column. It holds ‘A’..’P’, and 'A' % 5 is 0 in SQLite, so a sampling clause written the obvious way is true for every row and silently samples nothing. The symptom is a sampling setting that appears to do nothing, which is indistinguishable from the read being slow for another reason.

Parameters:

columns – the table’s own column names, typically from column_names(). Compared case-insensitively, which is the whole point: it is rowID that shadows rowid. Pass the columns of the table about to be queried – an alias is only unusable relative to a particular table.

Returns:

the first alias not shadowed by a real column, or None for a table that shadows all three (or has no row id at all).

spacr.filters.sampling_clause(fraction: float, rowid: str = '_rowid_') → str[source]

A SQL fragment taking roughly fraction of the rows.

Systematic on the row id, not ORDER BY RANDOM(): random ordering sorts the whole table before discarding most of it, which costs MORE than reading everything and is the opposite of the point. Modulo on the row id is an index scan and is also reproducible – the same 20% every time, so a gate drawn on Monday sits on the same cloud on Tuesday.

The bias this trades for is that row id order is insertion order, i.e. roughly well by well. For drawing a gate on a cloud of a million objects that is not a distinction that matters; for anything where it might, the export applies the gate to every row regardless of what was sampled.

Parameters:
  • fraction – in (0, 1]. 1 means everything, and returns no clause.

  • rowid – which row id spelling to use – see rowid_expression(), which is what picks it.

Raises:

FilterError – a fraction outside (0, 1].

spacr.filters.table_names(db_path: str) → Tuple[str, ...][source]

Every table in the database, in the order SQLite lists them.

Parameters:

db_path – the measurement database, opened read-only through a SQLite URI. Read-only mode does not create a file, so a path that is not there raises sqlite3.OperationalError rather than returning an empty tuple – “no tables” always means an empty database, never a wrong path.

spacr.filters.write_filters_table(db_path: str, frame: pandas.DataFrame) → None[source]

Replace filters with frame.

Whole-table replace rather than ALTER + UPDATE: the table is small (one row per object, a handful of columns), it is owned entirely by this module, and a partial write that left a gate column half-populated would be indistinguishable from a gate that selected those rows.

Parameters:
  • db_path – the measurement database, opened for writing.

  • frame – the WHOLE table as it should end up on disk. Since the write replaces rather than merges, any gate column missing from frame is dropped from the database – callers read the current table, add their column to it and pass the result back, never the column on its own.

spacr.filters.write_relationships(db_path: str) → pandas.DataFrame[source]

Build and store the relationships table. Called after the mask step.

Separate from ensure_relationships_table() so the mask step can say “rebuild this, the masks just changed” without a caller having to know the flag.

Parameters:

db_path – the measurement database, opened for writing. Any stored relationships table is REPLACED, so this is the call to make after re-masking and the wrong one to make merely to read the table.