spacr.columns

Report missing columns and offer the columns that are available.

THE FAILURE THIS REPLACES. A misnamed dependent_variable survives every early check and dies inside the merge, after the whole score table has been read, with a message naming a column the file does not have and saying nothing about what it does have. The user then opens the CSV by hand to read its header. On a large screen that is minutes of reading to answer a question the header row answers instantly.

THREE RULES, and each is a different failure this must not blur:

  1. THE HEADER ROW ONLY. nrows=0. This runs before the run, sometimes to populate a GUI dropdown on the GUI thread, and a score CSV is hundreds of megabytes.

  2. A MISSING FILE IS NOT A MISSING COLUMN. The first cannot answer the question; the second answers it with “not that name, one of these”. A caller told “column not found” about a path that does not exist goes looking in the wrong place.

  3. THE SUGGESTION IS SEPARATE FROM THE LIST. Near-misses are offered first because predictions for prediction is the common typo, but the FULL list is always there too – a suggestion that is wrong and a list that is absent is worse than no suggestion at all.

Exceptions

ColumnNotFound

A named column is absent, and the message says what is present.

Functions

available(→ List[str])

Every column across paths, in order, without duplicates.

describe(→ str)

The sentence to print or raise when name is not there.

headers(→ Dict[str, List[str]])

{path: [column, ...]} for every readable CSV in paths.

missing(→ List[str])

The paths that could not be read. The other half of headers().

resolve(→ str)

name if the CSVs have it, else raise with what they do have.

suggest(→ List[str])

Column names close to name, best first.

Module Contents

exception spacr.columns.ColumnNotFound(message: str, *, name: str = '', available: Sequence[str] = ())[source]

Bases: KeyError

A named column is absent, and the message says what is present.

Parameters:
  • message – human-readable explanation returned for the exception.

  • name – the column that was asked for, kept so a caller can recover it without parsing message.

  • available – the columns the file does have, IN FILE ORDER. Not sorted, because file order is what a user reading the header sees, and sorting hides that related columns are grouped together.

A KeyError subclass so existing except KeyError paths still catch it, and so settings['x'] failures and this one read the same way to a caller that does not care which it was.

Initialize a readable error with structured column context.

spacr.columns.available(paths) → List[str][source]

Every column across paths, in order, without duplicates.

Parameters:

paths – one CSV path, a sequence of paths, or None.

Across, not per file: score_data is routinely a list of one CSV per plate with identical headers, and a user choosing a column does not care which plate it came from.

spacr.columns.describe(name, paths, *, what: str = 'column', setting: str = '') → str[source]

The sentence to print or raise when name is not there.

Parameters:
  • name – missing column name to explain.

  • paths – CSV paths whose headers should be reported.

Names the setting, the files that were read, the near-misses and then every column. Long on purpose: this is the message that decides whether the user fixes it in two minutes or re-runs the screen to find out.

spacr.columns.headers(paths) → Dict[str, List[str]][source]

{path: [column, ...]} for every readable CSV in paths.

Parameters:

paths – one path, a sequence of them, or None.

Returns:

a dict in the order given. A path that does not exist or cannot be parsed is ABSENT from the result rather than mapped to an empty list – see rule 2. Use missing() to ask which those were.

spacr.columns.missing(paths) → List[str][source]

The paths that could not be read. The other half of headers().

Parameters:

paths – one CSV path, a sequence of paths, or None.

spacr.columns.resolve(name, paths, *, what: str = 'column', setting: str = '') → str[source]

name if the CSVs have it, else raise with what they do have.

Parameters:
  • name – the column asked for.

  • paths – the CSVs to look in.

  • what – what kind of column, for the message (“response column”, “count column”).

  • setting – the settings key this came from, so the message names the control the user has to change.

Raises:

ColumnNotFound – carrying available, so a GUI can offer the list rather than re-reading the files to build it.

A CASE-INSENSITIVE MATCH IS ACCEPTED AND RETURNED IN THE FILE’S SPELLING. Predictions when the file says predictions is not a different column, and failing on it teaches a user to distrust the message rather than to fix the name.

spacr.columns.suggest(name, columns: Iterable[str]) → List[str][source]

Column names close to name, best first.

Parameters:
  • name – missing column name for which to find near-matches.

  • columns – available column names to search.

Case-insensitive, because Prediction for prediction is a typo nobody should have to see spelled out.