Train/test leakage audit¶
Classifier accuracy is invalid when related crops occur on both sides of an
evaluation boundary. spaCR now audits the permanent train//test/
split, the ordinary train/validation holdout, every outer and inner CV
boundary, and the CV partition as a whole before fitting.
What is verified¶
The audit checks:
the same path or crop identity never crosses a boundary;
exported augmentations such as
_rot90,_flip_hand_aug3remain with their source object;byte-identical crops are detected by streaming SHA-256 even if renamed;
the requested
cv_group_byidentity (field, well or plate) stays intact;every CV sample is held out exactly once;
no related family is assigned to more than one held-out fold; and
related crops do not carry conflicting class labels.
Filename identities follow spaCR’s
<plate>_<well>_<field>_..._<object>.png crop convention. With
leakage_require_identity=True (the default), a filename that cannot prove
its requested identity fails the audit; an unknown relationship is not
reported as independent.
Ordinary validation is now group-aware too. cv_group_by='well' uses the
same group-stratified partitioner as CV and selects the candidate fold closest
to val_split and the full dataset’s class distribution. Augmentation is
applied only after the split and only to training.
Run the audit directly¶
spacr-leakage /data/classifier_dataset
spacr-leakage /data/classifier_dataset --group-by plate \
--output /tmp/leakage.json
The command prints JSON and exits 0 for a verified split, 1 when
leakage or unverifiable identities are found, and 2 when the dataset
cannot be audited. --no-content-hash reduces I/O but cannot detect renamed
copies. --allow-unverifiable is intended for diagnosing legacy datasets;
it does not make their performance estimate trustworthy.
Classify settings¶
leakage_audit_train_testAudit
train/againsttest/before fitting. DefaultTrue.leakage_hash_contentStream SHA-256 for copy/rename detection. Default
True.leakage_require_identityFail when protected identity or content cannot be verified. Default
True.evaluation_fail_on_leakageStop before fitting when an audit fails. Default
True. Setting it toFalserecords a failed audit but cannot make the resulting metric valid.
Reports are written beside the classifier run as
train_test_leakage_audit.json and
train_validation_leakage_audit.json. CV audit records are also included
in the Classifier Evaluation workbench’s leakage.json.
API¶
Classifier evaluation, calibration, and split-leakage diagnostics.
The module is model-agnostic: it consumes labels, probabilities, fold ids, and sample paths. Deep-learning CV and future classical-ML pipelines can therefore write the same evaluation bundle and the Qt workbench can display one stable artifact format.
- class spacr.classifier_evaluation.FoldLeakageAudit(group_by: str, n_samples: int, n_folds: int, validation_membership_missing: ~typing.List[int] = <factory>, validation_membership_duplicate: ~typing.List[int] = <factory>, overlap_counts: ~typing.Dict[str, int] = <factory>, examples: ~typing.Dict[str, ~typing.List[str]] = <factory>, critical_levels: ~typing.List[str] = <factory>, warnings: ~typing.List[str] = <factory>, unverifiable_counts: ~typing.Dict[str, int] = <factory>, hash_errors: ~typing.List[str] = <factory>, split_name: str = 'all_cv_folds')[source]
Whole-CV proof that each related sample family belongs to one fold.
- Parameters:
group_by – canonical identity level required to remain within one validation fold.
n_samples – number of source paths whose fold membership was audited.
n_folds – number of train/validation fold pairs inspected.
validation_membership_missing – up to twenty sample indexes that were never held out for validation.
validation_membership_duplicate – up to twenty sample indexes held out in more than one fold.
overlap_counts – identities assigned to multiple validation folds, counted at each exact, content, family, object, and acquisition level.
examples – up to ten conflicting identities per level, annotated with the folds that contain them.
critical_levels – completeness, overlap, identity, or label failures that make
passedfalse.warnings – non-fatal caveats and explanations accompanying failures.
unverifiable_counts – samples whose requested identity or optional byte-content hash could not be checked, counted by level.
hash_errors – up to twenty file-specific content-hashing failures.
split_name – stable label for this whole-CV audit record.
- class spacr.classifier_evaluation.LeakageReport(group_by: str, train_samples: int, validation_samples: int, overlap_counts: ~typing.Dict[str, int], examples: ~typing.Dict[str, ~typing.List[str]], split_name: str = '', critical_levels: ~typing.List[str] = <factory>, warnings: ~typing.List[str] = <factory>, unverifiable_counts: ~typing.Dict[str, int] = <factory>, hash_errors: ~typing.List[str] = <factory>)[source]
Overlap counts and examples for one train/validation boundary.
- Parameters:
group_by – protected split level (
cell,field,well, orplate); the legacynonespelling is normalized tocell.train_samples – number of training paths.
validation_samples – number of validation paths.
overlap_counts – overlap count at each identity level.
examples – up to ten shared identities per level.
split_name – caller-supplied label for the audited boundary.
critical_levels – levels that invalidate the requested split.
warnings – non-fatal caveats.
unverifiable_counts – samples lacking a requested identity or content hash, counted by the level that could not be verified.
hash_errors – up to twenty file-specific failures from optional byte-content hashing.
- spacr.classifier_evaluation.audit_cv_folds(paths: Sequence[Any], folds: Sequence[Tuple[Sequence[int], Sequence[int]]], *, labels: Sequence[Any] | None = None, group_by: str = 'well', hash_content: bool = False, require_identity: bool = True, raise_on_leakage: bool = False) FoldLeakageAudit[source]
Verify partition coverage and identity isolation across every CV fold.
This checks the fold assignment as one object, rather than trusting a sample of pairwise boundaries. Each index must be validation exactly once; exact paths, byte-identical content, source/augmentation families and the requested plate/well/field group must map to one held-out fold only.
- Parameters:
paths – one source path per sample, indexed by the fold indices.
folds –
(train_indices, validation_indices)per fold. Indices outsiderange(len(paths))are reported rather than ignored.labels – optional class per path, one value per path. Only used to describe the folds; it does not affect leakage detection.
group_by – identity level that may not cross a fold boundary –
none,field,wellorplate. A well-grouped split permits the same plate on both sides but never the same well.hash_content – also compare file CONTENT, so a byte-identical copy under a different name is caught. Costs one read per file.
require_identity – treat UNVERIFIABLE as critical. A filename that does not encode the requested identity, or a file that cannot be hashed, is otherwise only a warning – so leaving this False means a clean report can still hide leakage nobody could check for.
raise_on_leakage – raise
LeakageErrorinstead of returning a report whosepassedis False.
- Returns:
a
FoldLeakageAuditcarrying the per-level overlap counts, examples, and which levels were critical.- Raises:
ValueError – for an unsupported
group_by, orlabelswhose length does not matchpaths.
- spacr.classifier_evaluation.audit_dataset_splits(root: Any, *, group_by: str = 'well', hash_content: bool = True, require_identity: bool = True, raise_on_leakage: bool = False) LeakageReport[source]
Audit the permanent
train/versustest/dataset boundary.- Parameters:
root – dataset directory holding
train/andtest/. Both are searched recursively, and only.png,.jpg,.jpeg,.tif,.tiff,.bmpand.npyfiles are collected, so a folder of any other format audits as empty.group_by – identity level that may not appear on both sides –
none,field,wellorplate. The defaultwellstill permits the same plate in train and test. It is validated only after the images are found, so a bad value over an empty tree reports the missing images instead.hash_content – defaults to True here, unlike
audit_split_leakage(), so a byte-identical copy saved under a different name fails the audit – at the cost of one read per file.require_identity – defaults to True here: filenames that do not encode the
group_bylevel, and files that cannot be hashed, become critical instead of a warning, so an unverifiable split cannot report as clean.raise_on_leakage – raise
LeakageErrorinstead of returning a report whosepassedis False.
- Raises:
FileNotFoundError – when either side collects no image – a missing or unreadable folder is never treated as a passing split.
- Returns:
a
LeakageReportwhosesplit_nameis alwaystrain_vs_test; thetest/side is counted asvalidation_samples.
- spacr.classifier_evaluation.audit_split_leakage(train_paths: Sequence[Any], validation_paths: Sequence[Any], *, group_by: str = 'well', raise_on_leakage: bool = False, split_name: str = '', hash_content: bool = False, require_identity: bool = False) LeakageReport[source]
Detect related images crossing a train/validation boundary.
Exact/object/augmentation-family overlap is always critical. The requested
group_bylevel is also critical: a well-grouped split permits the same plate on both sides but never the same well.- Parameters:
train_paths – source paths used to fit the model.
validation_paths – paths used only for evaluation.
group_by –
cell,field,well, orplate. Legacynonealiasescell.raise_on_leakage – raise
LeakageErroron a critical overlap.split_name – optional fold/split label stored in the report.
hash_content – also compare file CONTENT, so a byte-identical copy under a different name is caught. Costs one read per file.
require_identity – treat UNVERIFIABLE as critical. A filename that does not encode the requested identity, or a file that cannot be hashed, is otherwise only a warning – so leaving this False means a clean report can still hide leakage nobody could check for.
- Returns: