spacr.regex_infer

Infer filename-parsing regexes for microscope image collections.

propose() groups filenames with a shared token structure and returns ranked regular expressions with named metadata groups. Each proposal reports its coverage, unmatched files, sampled values, and the evidence used to suggest roles such as wellID, fieldID, and chanID. Review these suggestions before importing files from a new naming convention.

Use rename_preview() and structure() to inspect the proposed renaming and folder layout. These functions do not move, rename, or write files.

Examples

>>> names = ["WA01F001C1.tif", "WA01F002C1.tif"]
>>> candidate = propose(names)[0]
>>> candidate.matched, candidate.unmatched
(2, ())
>>> rename_preview(candidate, names)[0]["matched"]
True

Classes

FieldEvidence

Describe the observed values and suggested role of one filename slot.

Proposal

Store a candidate filename regex and the evidence for reviewing it.

Functions

hint_for(→ str)

Infer a metadata role from the literal preceding a value slot.

mask_of(→ Tuple[str, ...])

Replace digit and non-digit tokens with coarse family placeholders.

propose(→ List[Proposal])

Infer and rank candidate regexes for a collection of filenames.

rename_preview(→ List[dict])

Preview parsed metadata and destination folders without writing files.

shape_of(→ Tuple[str, ...])

Replace digit tokens with placeholders to identify a filename family.

structure(→ Dict[str, int])

Count matched preview files by proposed destination folder.

tokenise(→ List[str])

Split a filename into alternating digit and non-digit runs.

Module Contents

class spacr.regex_infer.FieldEvidence[source]

Describe the observed values and suggested role of one filename slot.

Parameters:
  • index – token position in the filename family.

  • values – observed values for this slot in input order.

  • numeric – whether every observed value contains only digits.

  • before – literal filename text immediately preceding this slot.

  • role – suggested spaCR metadata field, or an empty string when the evidence does not support a role.

  • fixed_tail – constant suffix absorbed into this slot’s capture group.

  • because – human-readable evidence supporting the suggested role.

samples(limit: int = 4) → Tuple[str, ...][source]

Return unique example values in their first-seen order.

Parameters:

limit – Maximum number of values to return.

Returns:

tuple of str – Up to limit unique values.

property distinct: int[source]

Return the number of distinct observed values.

class spacr.regex_infer.Proposal[source]

Store a candidate filename regex and the evidence for reviewing it.

Parameters:
  • pattern – regular-expression pattern containing the proposed named capture groups.

  • fields – evidence indexed by proposed capture-group name.

  • matched – number of evaluated basenames matched by pattern.

  • total – number of non-empty input basenames evaluated.

  • unmatched – basenames the pattern could not parse, retained for review.

  • suffix – shared filename extension without its leading period.

compiled()[source]

Compile and return pattern.

evidence() → str[source]

Format coverage, field samples, role evidence, and unmatched files.

property coverage: float[source]

Return the fraction of evaluated filenames that match.

spacr.regex_infer.hint_for(before: str) → str[source]

Infer a metadata role from the literal preceding a value slot.

Parameters:

before – Literal text immediately before the variable filename component. Only its trailing alphabetic run is considered.

Returns:

str – A role from KNOWN_ROLES, or an empty string when the literal provides no recognized hint.

spacr.regex_infer.mask_of(tokens: Sequence[str]) → Tuple[str, ...][source]

Replace digit and non-digit tokens with coarse family placeholders.

Parameters:

tokens – Tokens returned by tokenise().

Returns:

tuple of str – "#" for each digit run and "@" for each non-digit run.

Notes

This coarser grouping lets changing well letters remain variable. Families that exceed MAX_VARYING_LITERALS are rejected later to avoid merging unrelated naming conventions.

spacr.regex_infer.propose(names: Iterable[str], limit: int = 4) → List[Proposal][source]

Infer and rank candidate regexes for a collection of filenames.

Parameters:
  • names – Filenames or paths. Only each basename is inspected; empty names are ignored.

  • limit – Maximum number of proposals to return.

Returns:

list of Proposal – Candidates ordered by matched-file count and then by the number of inferred metadata fields. Returns an empty list when the input has no comparable filename family.

Notes

Mixed naming conventions can produce several proposals. Inspect Proposal.unmatched and Proposal.evidence() before choosing one.

spacr.regex_infer.rename_preview(proposal: Proposal, names: Iterable[str], roles: Dict[str, str] | None = None) → List[dict][source]

Preview parsed metadata and destination folders without writing files.

Parameters:
  • proposal – Candidate returned by propose().

  • names – Filenames or paths to preview.

  • roles – Optional mapping from capture-group names to spaCR metadata roles. Values override the roles suggested by proposal.

Returns:

list of dict – One record per input with old, matched, values, and folder keys. Unmatched files remain in the result with matched=False.

Notes

This function performs no file-system writes or renames.

spacr.regex_infer.shape_of(tokens: Sequence[str]) → Tuple[str, ...][source]

Replace digit tokens with placeholders to identify a filename family.

Parameters:

tokens – Tokens returned by tokenise().

Returns:

tuple of str – Tokens with each all-digit run replaced by "#".

spacr.regex_infer.structure(preview: Sequence[dict]) → Dict[str, int][source]

Count matched preview files by proposed destination folder.

Parameters:

preview – Records returned by rename_preview().

Returns:

dict of str to int – Sorted mapping from folder path to matched-file count. Unmatched files and records without a folder are omitted.

spacr.regex_infer.tokenise(name: str) → List[str][source]

Split a filename into alternating digit and non-digit runs.

Parameters:

name – Filename or filename-like string. The extension remains part of the final non-digit run.

Returns:

list of str – Token runs in their original order.