spacr.ops_store

Where an OPS run’s tables live, and the gate that says the objects exist.

The storage contract has two halves. In measurements.db, which is authoritative: ops_geometry, ops_phenotype, ops_objects, ops_reads and ops_barcodes. Beside it, as a cache: parquet copies of ops_reads and ops_barcodes, written in the same step with their row counts asserted equal, so the sidecar cannot drift from the authority unnoticed.

SQLITE IS THE AUTHORITY AND PARQUET IS A CACHE, which is a decision and not a detail. The two can disagree, so one of them has to be right by definition – otherwise a reader that picked the faster one would sometimes be reading yesterday’s numbers with no way to tell. write_table() writes both in one call and asserts the counts match before either is visible as complete; read_table() prefers the cache and falls back without complaining, because a missing cache is a performance question and not a correctness one.

READINESS IS A GATE, NOT A STEP. No sequencing or phenotype channel is read until ops_objects is on disk and its row count checked. objects_ready() is that sentence, and it returns a REASON when the answer is no – a gate that only says “not yet” makes the operator guess which half failed.

Other measurement tables use the same object ids, so later steps need no special case. The numbering is deterministic for that reason: those ids join tables written at different times.

Exceptions

StoreError

A store that cannot be trusted, with the way out in the text.

Classes

Readiness

Whether object sampling may start, and why not when it may not.

Functions

objects_ready(→ Readiness)

May a sequencing or phenotype channel be read yet?

read_table(db_path, table, *[, prefer_cache])

Read one OPS table, preferring its parquet cache.

row_count(→ Optional[int])

Rows in one table, or None when the table is not there.

write_table(→ int)

Write one OPS table to sqlite, and its parquet cache if it has one.

Module Contents

exception spacr.ops_store.StoreError[source]

Bases: ValueError

A store that cannot be trusted, with the way out in the text.

Initialize self. See help(type(self)) for accurate signature.

class spacr.ops_store.Readiness[source]

Whether object sampling may start, and why not when it may not.

Parameters:
  • ready – the verdict. Also what bool(readiness) answers, so a caller can write if not objects_ready(db): and still reach reason when it needs to say why.

  • rows – how many objects the table holds, or None when there is no table to count – which is a different failure from a table with no rows in it, and the two are told apart here rather than by the caller.

  • reason – one sentence naming which of the three things went wrong, empty when nothing did.

spacr.ops_store.objects_ready(db_path: str, *, minimum: int = 1) → Readiness[source]

May a sequencing or phenotype channel be read yet?

The gate is one sentence – “no sequencing or phenotype channel is read until ops_objects is on disk and its row count checked” – and both halves matter. On disk without a count check would pass an empty table, and an empty ops_objects means every later phase silently produces nothing while appearing to run.

IT RETURNS A REASON, not just a verdict. “Not yet” leaves the operator guessing which of three things went wrong: no database, no table, or a table with nothing in it. Those have different fixes.

Parameters:
  • minimum – the fewest objects a real well can have. One is the honest floor – a well with a single nucleus is a bad well, not an impossible one – so this exists to be raised by a caller who knows their plate, not to encode a guess here.

  • db_path – the run’s sqlite database. Its ABSENCE is one of the three answers, so this is not required to exist.

spacr.ops_store.read_table(db_path: str, table: str, *, prefer_cache: bool = True)[source]

Read one OPS table, preferring its parquet cache.

A missing or unreadable cache falls back to sqlite silently, because that is a speed question. A cache whose row count disagrees with the database is NOT silent – it is removed and the authority is returned, since a disagreement is the one failure this arrangement can produce.

Parameters:
  • db_path – the run’s sqlite database, which is the authority.

  • table – which of OPS_TABLES to read.

Raises:

StoreError – when the table is not in the database at all.

spacr.ops_store.row_count(db_path: str, table: str) → int | None[source]

Rows in one table, or None when the table is not there.

Parameters:
  • db_path – the run’s sqlite database.

  • table – the table to count.

spacr.ops_store.write_table(db_path: str, table: str, frame, *, if_exists: str = 'replace') → int[source]

Write one OPS table to sqlite, and its parquet cache if it has one.

BOTH IN ONE CALL, because 372 asks for the row counts to be asserted equal and two calls could not do that – a caller who wrote the database and then crashed would leave a sidecar describing a different run, and nothing would notice until a reader silently preferred it.

Parameters:
  • db_path – the measurements database.

  • table – one of OPS_TABLES.

  • frame – a DataFrame.

Returns:

rows written.

Raises:

StoreError – on an unknown table; when the sidecar it just wrote does not have the same number of rows as the database; or when ops_objects or ops_barcodes would hold two rows for one object, inside the frame or between the frame and the table it is appended to. One row per object is a UNIQUE constraint in the schema, and nothing is written when it would be broken.