spacr.tabular

One reader and one writer for every table spaCR opens or saves.

Why this module exists

An audit across spacr/ found 248 tabular reads and writes – pd.read_csv, pd.read_sql*, .to_csv, .to_sql – and thirteen call sites that normalised a column name at all. There was no funnel; there were 248 doors, and which spelling of a key a frame ended up with depended on which door it came through.

The resulting failure has a consistent shape: columnID worked as a filter, because some path downstream renamed on the way to the fit, while the CSV picker read the raw header and offered column_name and column – the names a user must not have to know.

So: one place normalises, and every reader goes through it. Doing it at the funnel is the point. The picker becomes correct for free, because it reads through the same door and therefore sees columnID.

What a read guarantees

Every frame that leaves read_table() / read_database() has

Writing

Decided, not accidental: spaCR writes canonical names. write_table and write_database canonicalise on the way out by default, so a frame assembled by hand cannot re-export column_name and start the cycle again. The header of an exported file therefore changes for anyone whose downstream script reads column_name – that is a release note, and it is the deliberate half of this compatibility trade-off. canonicalise=False is there for a caller who owes an external format an exact header.

The guards that moved here rather than being left behind

  • A ~ path is expanded, once, for every reader – GitHub issue #108, where a src beginning with ~ was resolved against the working directory and refused with FileNotFoundError: ~<DB>. $HOME and %USERPROFILE% too: a settings CSV carried between machines routinely holds one.

  • The measurements schema migration runs on open, exactly as io._read_db does it, so a legacy database is repaired by any reader rather than only by the one that remembered.

Dependencies

pandas, sqlite3 and spacr.schema. Nothing else at module scope – no spacr.utils, no matplotlib, no torch. That is a requirement rather than tidiness: the CSV picker and the SQL column list have to be able to import this to get canonical names, and they cannot pay for a torch import to do it.

Exceptions

TabularFormatError

A path whose suffix names no format this module can read.

Functions

database_tables(→ Tuple[str, ...])

The table names in a database, sorted.

read_database(→ List[pandas.DataFrame])

Read one or more tables out of a SQLite database.

read_table(→ pandas.DataFrame)

Read one table, whatever it is stored in, with canonical column names.

resolve_path(→ str)

Expand ~ and environment variables in a path, once, for everyone.

table_columns(→ Tuple[str, ...])

The column names a table would be read with, without reading it.

table_format(→ str)

Which reader a path needs: 'csv', 'sqlite', 'parquet',

write_database(→ str)

Write one frame into a SQLite table, with canonical column names.

write_table(→ str)

Write one frame, format chosen by suffix, with canonical column names.

Module Contents

exception spacr.tabular.TabularFormatError[source]

Bases: ValueError

A path whose suffix names no format this module can read.

Initialize self. See help(type(self)) for accurate signature.

spacr.tabular.database_tables(db: Any, *, migrate: bool = False) → Tuple[str, ...][source]

The table names in a database, sorted.

Parameters:
  • db – path to the database. ~ and $VARS are expanded.

  • migrate – run the schema migration first. False, because listing what is there must not rewrite it.

Returns:

the table names.

spacr.tabular.read_database(db: Any, tables: Any, *, canonicalise: bool = True, report: Callable[[str], None] | None = print, warn: Callable[[str], None] | None = None, repair_plate_ids: bool = True, migrate: bool = True, read_only: bool = False, limit: int | None = None, chunksize: int = 100000, **kwargs) → List[pandas.DataFrame][source]

Read one or more tables out of a SQLite database.

Expands ~, validates identifiers before opening the database, runs the schema migration when requested, and reads in chunks to limit peak memory. A missing table raises ValueError with its name.

Parameters:
  • db – path to the database.

  • tables – a table name, or a sequence of them.

  • canonicalise – apply the vocabulary to each frame.

  • report – see read_table().

  • warn – see read_table().

  • repair_plate_ids – collapse a doubled pp plate prefix.

  • migrate – run the schema migration on open.

  • read_only – open through file:...?mode=ro, so the read cannot write to the user’s database. Incompatible with migrate.

  • limit – read at most this many rows per table. None reads all of them.

  • chunksize – rows per chunk.

  • kwargs – passed to pandas.read_sql_query().

Returns:

one frame per requested table, in the order asked for.

Raises:

ValueError – when a table is not in the database.

spacr.tabular.read_table(source: Any, *, table: str | None = None, canonicalise: bool = True, report: Callable[[str], None] | None = print, warn: Callable[[str], None] | None = None, repair_plate_ids: bool = True, **kwargs) → pandas.DataFrame[source]

Read one table, whatever it is stored in, with canonical column names.

CSV, TSV, SQLite, Parquet, Feather and Excel, chosen by suffix. A SQLite path needs table; every other format ignores it.

Parameters:
  • source – path to the file. ~ and $VARS are expanded. A cloud address (s3://, gs://, az://, https://) is read from a cached local copy.

  • table – the table name, for a database.

  • canonicalise – apply the vocabulary. There is no reason to turn this off on the ordinary path – it is what makes the picker and the run agree about what a column is called. Off is for a caller inspecting a file exactly as written.

  • report – called with each agreeing-collision message; print by default, None to silence.

  • warn – called with each disagreeing-collision message; None routes to warnings.warn().

  • repair_plate_ids – collapse a doubled pp plate prefix.

  • kwargs – passed to the underlying pandas reader.

Returns:

a pandas.DataFrame.

spacr.tabular.resolve_path(path: Any) → str[source]

Expand ~ and environment variables in a path, once, for everyone.

GitHub issue #108: a src beginning with ~ produced ~/.../measurements.db, which the migration resolved against the working directory and refused with FileNotFoundError: ~<DB>. Fixed at the funnel rather than at the ~99 sites that build a measurements path by string concatenation.

Parameters:

path – a path, a os.PathLike, or anything else (returned unchanged, so a caller may pass an open connection through).

Returns:

the expanded path, or the object it was given.

spacr.tabular.table_columns(source: Any, *, table: str | None = None, canonicalise: bool = True) → Tuple[str, ...][source]

The column names a table would be read with, without reading it.

The CSV button and the SQL column list show what the run will see – columnID, never column_name – and they get that by asking the reader rather than by growing a second copy of the vocabulary.

A collapsed duplicate is not listed twice: a picker that offered both well and wellID would let a user choose a column the run will not find.

Parameters:
  • source – a CSV/Parquet/Excel path, or a database path with table.

  • table – the table name, for a database.

  • canonicalise – apply the vocabulary.

Returns:

the column names, in order.

spacr.tabular.table_format(path: Any) → str[source]

Which reader a path needs: 'csv', 'sqlite', 'parquet', 'feather' or 'excel'.

Parameters:

path – a path.

Returns:

the format name.

Raises:

TabularFormatError – when the suffix names no known format.

spacr.tabular.write_database(frame: pandas.DataFrame, db: Any, table: str, *, if_exists: str = 'append', canonicalise: bool = True, index: bool = False, migrate: bool = False, **kwargs) → str[source]

Write one frame into a SQLite table, with canonical column names.

Parameters:
  • frame – the frame.

  • db – path to the database; created if absent, ~ expanded.

  • table – the table name.

  • if_exists – 'append' (spaCR’s usual), 'replace', 'fail'.

  • canonicalise – rename legacy spellings on the way out. On by default so a frame assembled by hand cannot put column_name back into a database the reader will then have to repair.

  • index – write the index. False.

  • migrate – run the schema migration first. False: writing a scratch table must not migrate the user’s measurements.

  • kwargs – passed to pandas.DataFrame.to_sql().

Returns:

the resolved database path.

spacr.tabular.write_table(frame: pandas.DataFrame, path: Any, *, canonicalise: bool = True, index: bool = False, **kwargs) → str[source]

Write one frame, format chosen by suffix, with canonical column names.

See the module docstring for why writing canonical was chosen over writing back what was read.

Parameters:
  • frame – the frame.

  • path – destination. ~ and $VARS are expanded; the parent directory is created.

  • canonicalise – rename legacy spellings on the way out.

  • index – pandas’ index argument, defaulted to False because every spaCR export that ever wanted the index has a real column for it and an unnamed Unnamed: 0 on re-read is a bug in waiting.

  • kwargs – passed to the underlying pandas writer.

Returns:

the resolved path written.