spacr.tabular¶
One reader and one writer for every table spaCR opens or saves.
Why this module exists¶
An audit across spacr/ found 248
tabular reads and writes – pd.read_csv, pd.read_sql*, .to_csv,
.to_sql – and thirteen call sites that normalised a column name at
all. There was no funnel; there were 248 doors, and which spelling of a key a
frame ended up with depended on which door it came through.
The resulting failure has a consistent shape: columnID
worked as a filter, because some path downstream renamed on the way to the
fit, while the CSV picker read the raw header and offered column_name and
column – the names a user must not have to know.
So: one place normalises, and every reader goes through it. Doing it at the
funnel is the point. The picker becomes correct for free, because it reads
through the same door and therefore sees columnID.
What a read guarantees¶
Every frame that leaves read_table() / read_database() has
canonical metadata column names (
spacr.schema.canonical_column_name(), which folds case and punctuation);one column per metadata key, with the collision reported – printed when the duplicates agreed, warned with a row count when they did not (
spacr.schema.resolve_metadata_collisions());the
pplate1plate-value repair applied to every plate-bearing column (spacr.schema.normalise_plate_columns()).
Writing¶
Decided, not accidental: spaCR writes canonical names. write_table
and write_database canonicalise on the way out by default, so a frame
assembled by hand cannot re-export column_name and start the cycle again.
The header of an exported file therefore changes for anyone whose downstream
script reads column_name – that is a release note, and it is the
deliberate half of this compatibility trade-off.
canonicalise=False is there for a caller who owes an external format an
exact header.
The guards that moved here rather than being left behind¶
A
~path is expanded, once, for every reader – GitHub issue #108, where asrcbeginning with~was resolved against the working directory and refused withFileNotFoundError: ~<DB>.$HOMEand%USERPROFILE%too: a settings CSV carried between machines routinely holds one.The measurements schema migration runs on open, exactly as
io._read_dbdoes it, so a legacy database is repaired by any reader rather than only by the one that remembered.
Dependencies¶
pandas, sqlite3 and spacr.schema. Nothing else at module
scope – no spacr.utils, no matplotlib, no torch. That is a requirement
rather than tidiness: the CSV picker and the SQL column list have to be able
to import this to get canonical names, and they cannot pay for a torch
import to do it.
Exceptions¶
A path whose suffix names no format this module can read. |
Functions¶
|
The table names in a database, sorted. |
|
Read one or more tables out of a SQLite database. |
|
Read one table, whatever it is stored in, with canonical column names. |
|
Expand |
|
The column names a table would be read with, without reading it. |
|
Which reader a path needs: |
|
Write one frame into a SQLite table, with canonical column names. |
|
Write one frame, format chosen by suffix, with canonical column names. |
Module Contents¶
- exception spacr.tabular.TabularFormatError[source]¶
Bases:
ValueErrorA path whose suffix names no format this module can read.
Initialize self. See help(type(self)) for accurate signature.
- spacr.tabular.database_tables(db: Any, *, migrate: bool = False) Tuple[str, ...][source]¶
The table names in a database, sorted.
- Parameters:
db – path to the database.
~and$VARSare expanded.migrate – run the schema migration first.
False, because listing what is there must not rewrite it.
- Returns:
the table names.
- spacr.tabular.read_database(db: Any, tables: Any, *, canonicalise: bool = True, report: Callable[[str], None] | None = print, warn: Callable[[str], None] | None = None, repair_plate_ids: bool = True, migrate: bool = True, read_only: bool = False, limit: int | None = None, chunksize: int = 100000, **kwargs) List[pandas.DataFrame][source]¶
Read one or more tables out of a SQLite database.
Expands
~, validates identifiers before opening the database, runs the schema migration when requested, and reads in chunks to limit peak memory. A missing table raisesValueErrorwith its name.- Parameters:
db – path to the database.
tables – a table name, or a sequence of them.
canonicalise – apply the vocabulary to each frame.
report – see
read_table().warn – see
read_table().repair_plate_ids – collapse a doubled
ppplate prefix.migrate – run the schema migration on open.
read_only – open through
file:...?mode=ro, so the read cannot write to the user’s database. Incompatible withmigrate.limit – read at most this many rows per table.
Nonereads all of them.chunksize – rows per chunk.
kwargs – passed to
pandas.read_sql_query().
- Returns:
one frame per requested table, in the order asked for.
- Raises:
ValueError – when a table is not in the database.
- spacr.tabular.read_table(source: Any, *, table: str | None = None, canonicalise: bool = True, report: Callable[[str], None] | None = print, warn: Callable[[str], None] | None = None, repair_plate_ids: bool = True, **kwargs) pandas.DataFrame[source]¶
Read one table, whatever it is stored in, with canonical column names.
CSV, TSV, SQLite, Parquet, Feather and Excel, chosen by suffix. A SQLite path needs
table; every other format ignores it.- Parameters:
source – path to the file.
~and$VARSare expanded. A cloud address (s3://,gs://,az://,https://) is read from a cached local copy.table – the table name, for a database.
canonicalise – apply the vocabulary. There is no reason to turn this off on the ordinary path – it is what makes the picker and the run agree about what a column is called. Off is for a caller inspecting a file exactly as written.
report – called with each agreeing-collision message;
printby default,Noneto silence.warn – called with each disagreeing-collision message;
Noneroutes towarnings.warn().repair_plate_ids – collapse a doubled
ppplate prefix.kwargs – passed to the underlying pandas reader.
- Returns:
- spacr.tabular.resolve_path(path: Any) str[source]¶
Expand
~and environment variables in a path, once, for everyone.GitHub issue #108: a
srcbeginning with~produced~/.../measurements.db, which the migration resolved against the working directory and refused withFileNotFoundError: ~<DB>. Fixed at the funnel rather than at the ~99 sites that build a measurements path by string concatenation.- Parameters:
path – a path, a
os.PathLike, or anything else (returned unchanged, so a caller may pass an open connection through).- Returns:
the expanded path, or the object it was given.
- spacr.tabular.table_columns(source: Any, *, table: str | None = None, canonicalise: bool = True) Tuple[str, ...][source]¶
The column names a table would be read with, without reading it.
The CSV button and the SQL column list show what the run will see –
columnID, nevercolumn_name– and they get that by asking the reader rather than by growing a second copy of the vocabulary.A collapsed duplicate is not listed twice: a picker that offered both
wellandwellIDwould let a user choose a column the run will not find.- Parameters:
source – a CSV/Parquet/Excel path, or a database path with
table.table – the table name, for a database.
canonicalise – apply the vocabulary.
- Returns:
the column names, in order.
- spacr.tabular.table_format(path: Any) str[source]¶
Which reader a path needs:
'csv','sqlite','parquet','feather'or'excel'.- Parameters:
path – a path.
- Returns:
the format name.
- Raises:
TabularFormatError – when the suffix names no known format.
- spacr.tabular.write_database(frame: pandas.DataFrame, db: Any, table: str, *, if_exists: str = 'append', canonicalise: bool = True, index: bool = False, migrate: bool = False, **kwargs) str[source]¶
Write one frame into a SQLite table, with canonical column names.
- Parameters:
frame – the frame.
db – path to the database; created if absent,
~expanded.table – the table name.
if_exists –
'append'(spaCR’s usual),'replace','fail'.canonicalise – rename legacy spellings on the way out. On by default so a frame assembled by hand cannot put
column_nameback into a database the reader will then have to repair.index – write the index.
False.migrate – run the schema migration first.
False: writing a scratch table must not migrate the user’s measurements.kwargs – passed to
pandas.DataFrame.to_sql().
- Returns:
the resolved database path.
- spacr.tabular.write_table(frame: pandas.DataFrame, path: Any, *, canonicalise: bool = True, index: bool = False, **kwargs) str[source]¶
Write one frame, format chosen by suffix, with canonical column names.
See the module docstring for why writing canonical was chosen over writing back what was read.
- Parameters:
frame – the frame.
path – destination.
~and$VARSare expanded; the parent directory is created.canonicalise – rename legacy spellings on the way out.
index – pandas’
indexargument, defaulted toFalsebecause every spaCR export that ever wanted the index has a real column for it and an unnamedUnnamed: 0on re-read is a bug in waiting.kwargs – passed to the underlying pandas writer.
- Returns:
the resolved path written.