spacr.plot¶
Scientific plotting and statistical-annotation helpers.
Classes¶
Grouped plot + statistical-test helper for spacr experiment DataFrames. |
Functions¶
|
Plot grouped observations and run assumption-aware comparisons. |
|
Compute a two-set gene overlap from CSVs and draw its Venn diagram. |
|
Every colour in |
|
The DPI this figure can actually be written at, and a word if it is not |
Return |
|
|
|
Return a random |
|
|
Aggregate a well-level DataFrame into a plate-shaped heatmap. |
|
Concatenate feature-importance CSVs and hand off to |
|
The data colours a reader will not find on |
|
Read measurements + annotation from a spacr DB and plot a class-balanced jitter plot. |
|
Show the original and the normalised image side by side in grayscale. |
|
The four outline colours for |
|
Overlay |
|
Plot random |
|
Display per-channel images, label mask and flow field for Cellpose v4 outputs. |
|
Plot Jaccard, Dice, boundary-F1 and average-precision distributions per comparison. |
|
Load per-plate CSVs, filter/outlier-clean and render a |
|
Read one or more measurement DBs, annotate conditions and render a |
|
Plot a horizontal bar chart of raw feature importances. |
|
Plot a histogram of |
|
Render a square grid of percentile-normalised images with a black background. |
|
Plot image and mask overlays. |
|
Plot image and mask overlays, outlining each channel's own mask in magenta. |
|
Show side-by-side images and arrays found across multiple folders. |
|
Overlay Lorenz curves from multiple gRNA count CSVs with per-plate Gini coefficients. |
|
Display per-channel images, label masks and flow fields for a batch. |
|
Show multi-channel image stacks with per-object outlines overlaid. |
|
Overlay mask outlines on the matching channel image for each object type. |
|
Plot organelle segmentation results: raw channel, label mask, morphology-specific diagnostic. |
|
Plot a horizontal bar chart of permutation feature importances with error bars. |
|
Render every plate of a screen as ONE panel, wells square, on one colour scale. |
|
Plot stacked proportion bars per group with chi-squared and pairwise stats. |
|
Render mask overlay, cropped PNG grid and activation-map grid for one FOV. |
|
Show original vs. resized image/label pairs in a 2x2 grid. |
|
Show a single image, its label mask (optionally outlined) and flow image. |
|
Repaint |
|
A binomial GLM on the per-object outcome, standard errors clustered by unit. |
|
Compare conditions on their PER-UNIT proportions, one row per bin. |
|
Each unit's share of every bin, one row per unit. |
|
Return a random |
|
Aggregate vision-model test CSVs under |
|
Write |
|
Display several Cellpose-style label masks side by side for a quick visual QC. |
|
Show three masks side by side with random colormaps. |
|
Read a table (CSV/TSV/XLS/XLSX or a DataFrame) and render a volcano plot. |
Module Contents¶
- class spacr.plot.spacrGraph(df, grouping_column, data_column, graph_type='jitter_box', summary_func='mean', order=None, colors=None, output_dir='./output', save=False, y_lim=None, log_y=False, log_x=False, error_bar_type='std', remove_outliers=False, theme='pastel', representation='object', paired=False, all_to_all=True, compare_group=None, graph_name=None, annotate_stats=False)[source]¶
Grouped plot + statistical-test helper for spacr experiment DataFrames.
Wraps preprocessing (aggregation by object / well / plate), normality and variance testing, group-wise pairwise stats, and plot rendering (bar / jitter / box / violin / jitter_box / jitter_bar / line / line_std) in a single object whose output can optionally be persisted alongside a CSV of stats.
- Parameters:
df – Input DataFrame.
grouping_column – Categorical grouping variable.
data_column – Metric column (or list of columns) to summarise.
graph_type – Plot type. Default
'jitter_box'– a box with the points over it. Seecreate_grouped_plot()for why that default is a correction rather than a taste.summary_func – Aggregator for well/plate level. Default
'mean'.order – Explicit ordering of groups.
colors – Optional colour palette.
output_dir – Save location when
save=True.save – If True, persist plot and stats.
y_lim – Two-element y-axis limits.
log_y – Use log scale for y-axis.
log_x – Use log scale for x-axis.
error_bar_type –
'std'or'sem'. Default'std'.remove_outliers – Drop 1.5*IQR outliers per group before plotting.
theme – Seaborn palette name. Default
'pastel'.representation – Aggregation level —
'object','well'or'plate'. Default'object'.paired – Treat groups as paired samples where applicable.
all_to_all – Run every pairwise comparison;
Falsecompares each group tocompare_group.compare_group – Reference group when
all_to_all=False.graph_name – Prefix for saved file names.
annotate_stats – Draw a bracket over each pairwise comparison with its asterisks (or
ns) above it. DefaultFalse: the tests are run and written to the results table on every plot, but withall_to_all=Truean N-group plot has N(N-1)/2 comparisons and a stack of that many brackets buries the data it is about. Ask for them when the comparisons are few enough to read. Only drawn for a singledata_column; see_draw_comparison_lines().
Store configuration, set the theme, and preprocess the DataFrame.
- create_plot(ax=None)[source]¶
Build the plot for the chosen graph type onto
self.fig.Nothing is displayed: retrieve the figure with
get_figure()(and the statistics withget_results()), or callplt.show().- Parameters:
ax – Existing
Axesto draw into, for placing this graph in a panel of a larger figure.self.figis then set to that axes’ parent figure, so a latersave=Truewrites the whole enclosing figure, not this panel alone.Nonecreates a fresh figure sized from the group count andbar_width— and note that with a singledata_columnthe standardisation pass still callsax.figure.set_size_inches, which resizes a shared figure underneath its other panels.
- perform_levene_test(unique_groups)[source]¶
Levene’s test for equal variance on
data_column[0], MEDIAN-centred.Delegates to
spacr.figures.stats.check_equal_variance(). Two things moved when it did, and both change the number a caller writes into a CSV:The centring is the median (Brown-Forsythe), not SciPy’s default mean. Median centring is less sensitive to non-normal data, and this function is called before the normality verdict is known.
Below
spacr.figures.stats.MIN_N_FOR_ASSUMPTIONSobservations in the smallest group the result is(nan, nan). On three replicates Levene has almost no power, so “p = 0.7, variances are equal” means “we could not tell”, and printing 0.7 into a results table invites exactly the reading that publishes a difference that is not there.
- Parameters:
unique_groups – Groups to compare.
- Returns:
(statistic, p_value), both NaN when the check had no power.
- perform_normality_tests()[source]¶
Evaluate normality for each requested column and group.
Shapiro-Wilk results and the overall verdict come from
spacr.figures.stats.check_normality(), including its Bonferroni correction across groups. Groups with fewer than three observations are reported as"Skipped". Checks below the configured information threshold reportInformative=Falserather than treating a failure to reject as evidence of normality.- Returns:
is_normal (bool) –
Trueonly when every requested column passes.results (list of dict) – Per-group test statistics, sample sizes, and verdicts.
- perform_posthoc_tests(is_normal, unique_groups)[source]¶
Perform post-hoc tests for multiple groups based on all_to_all flag.
- Parameters:
is_normal – Outcome of the normality check, which selects the family of test: True runs Tukey HSD, False runs Dunn’s test with an automatically chosen p-adjustment. It only matters when post-hoc testing runs at all — see
unique_groups. It must be the verdictperform_normality_tests()returned, which isspacr.figures.stats.check_normality()’s. A hand-computed one puts the omnibus test and the pairwise tests on different footing — Kruskal-Wallis across the groups followed by Tukey between them is two different assumptions about one dataset — and it is how the power floor gets bypassed: three replicates buy Dunn’s, not Tukey.unique_groups – The distinct group labels. Only its length is read; the comparisons themselves are rebuilt from
self.df[self.grouping_column], so reordering or renaming entries has no effect. Fewer than three groups returns an empty list, as doesself.all_to_allbeing False, because pairwise correction is meaningless for a single comparison.
- Returns:
A list of per-comparison dicts with
Comparison,Test Statistic(alwaysNone— neither test reports one),p-value,Test Nameand then_object/n_wellcounts; empty when no post-hoc test was warranted. Onlyself.data_column[0]is tested, so extra data columns are ignored here.
- perform_statistical_tests(unique_groups, is_normal)[source]¶
Run one supported group comparison per data column.
- Parameters:
unique_groups (sequence) – Groups to compare. Two groups produce a pairwise test; larger sets produce an omnibus test.
is_normal (bool) – External normality verdict.
Falseforces a rank test;Truestill requires informative engine-level assumption checks.
- Returns:
list of dict – Test name, statistic, p-value, sample counts, effect size, and selection rationale for each data column. Untestable comparisons use
Test Name='not testable'and include the reason.
Notes
Test selection is delegated to
spacr.figures.stats.compare(). Two-group comparisons may use Student’s t, Welch’s t, or Mann-Whitney U; larger comparisons may use one-way ANOVA, Welch’s ANOVA, or Kruskal-Wallis. Paired data use the paired t-test or Wilcoxon signed-rank test.
- preprocess_data()[source]¶
Return a new DataFrame aggregated to the configured representation.
Drops rows with NaN in the grouping or data columns, aggregates the data columns with
summary_funcper well ('prc') or per plate ('plateID', split out ofprcwhen needed) — or leaves them per object — and makes the grouping column an ordered Categorical.- Returns:
The preprocessed DataFrame;
__init__assigns it back toself.dfrather than the frame being modified in place.- Raises:
KeyError – if
representation='plate'and neither aplateIDnor aprccolumn is available.ValueError – if
representationis not'object','well'or'plate'.
- spacr.plot.create_grouped_plot(df, grouping_column, data_column, graph_type='jitter_box', summary_func='mean', order=None, colors=None, output_dir='./output', save=False, y_lim=None, error_bar_type='std')[source]¶
Plot grouped observations and run assumption-aware comparisons.
Pairwise tests are chosen independently by
spacr.figures.stats.compare(). Student’s t, Welch’s t, or Mann-Whitney U is used according to the normality and equal-variance checks; an underpowered assumption check selects the rank test. When at least three groups jointly pass normality, Tukey HSD rows are added.The
'jitter_box'default is a STATISTICAL CORRECTION, not a presentation preference: it shows the observations and their distribution instead of reducing each group to a mean bar. Two groups can have the same mean while having different spreads. The box summarizes the distribution; the jitter stays because the points are the evidence.- Parameters:
df (pandas.DataFrame) – Source observations.
grouping_column (str) – Categorical column defining groups.
data_column (str) – Numeric column to plot and compare.
graph_type ({'bar', 'violin', 'jitter', 'box', 'jitter_box'}, optional) – Plot representation. The default shows every observation together with median, quartiles, and whiskers.
summary_func (str or callable, optional) – Aggregation used by the bar representation.
order (sequence of str, optional) – Group order. By default, sort observed group values.
colors (palette-like, optional) – Colours passed to seaborn. By default, use the house data colour.
output_dir (path-like, optional) – Directory for saved output.
save (bool, optional) – Save the figure and
test_results.csvwhen true.y_lim (sequence of float, optional) – Two-element vertical-axis limits.
error_bar_type ({'std', 'sem'}, optional) – Error statistic for bar plots.
- Returns:
figure (matplotlib.figure.Figure) – Displayed figure. It carries the source recipe used by the interactive representation menu.
results_df (pandas.DataFrame) – Normality, pairwise, and optional Tukey HSD results.
- Raises:
ValueError – If a bar plot receives an unsupported
error_bar_type.
- spacr.plot.create_venn_diagram(file1, file2, gene_column='gene', filter_coeff=0.1, save=True, save_path=None)[source]¶
Compute a two-set gene overlap from CSVs and draw its Venn diagram.
- Parameters:
file1 – First CSV file.
file2 – Second CSV file.
gene_column – Column identifying genes. Default
'gene'.filter_coeff – Threshold on the
coefficientcolumn — positive filters> threshold, negative filters< threshold.save – If True, save as PDF; requires
save_path.save_path – Output PDF path when
saveis True.
- Returns:
{'overlap', 'unique_to_file1', 'unique_to_file2'}lists.- Raises:
ValueError – if
saveis True butsave_pathis missing.
- spacr.plot.data_colours(fig)[source]¶
Every colour in
figthat carries the CLAIM rather than the frame.- Parameters:
fig – Matplotlib figure whose data artists are inspected.
Used only to say when one of them stops working on paper (150 D), never to change one. A data line is identified the same way
_chromeidentifies a reference line, from the opposite side of the same test.
- spacr.plot.deliverable_dpi(fig, dpi, path=None)[source]¶
The DPI this figure can actually be written at, and a word if it is not the one that was asked for.
A resolution preference is a request, not a guarantee.
spacrGraphpins its canvas to at least 10 inches square and grows it with the number of groups, so 600 and 1200 DPI are not available for a large grouped figure – the raster would be tens of thousands of pixels on a side.The old behaviour was to hand the number to matplotlib and find out. This returns the DPI that will be used and says, by name, when that is not the DPI that was requested. Appearing to accept a setting and then quietly delivering another one is the failure this avoids.
- Parameters:
fig – the figure about to be written.
dpi – the requested dots per inch.
path – destination, named in the message when there is one.
- Returns:
the DPI to pass to
savefig.
- spacr.plot.figure_output_preferences()[source]¶
Return
(format, dpi)from the user’s preferences.Degrades to
DEFAULT_FIGURE_FORMAT/DEFAULT_FIGURE_DPIrather than raising: the preference store is Qt’s, and the pipelines that call this run headless from the CLI and from notebooks, where importing PySide6 to decide a file extension would be absurd.
- spacr.plot.figure_path(path, fmt=None)[source]¶
pathwith the extension the figure-format preference will write.THE ONE PLACE A NON-MATPLOTLIB RENDERER CAN ASK.
save_figurerewrites the extension itself, which is why every matplotlib save has honoured the preference for months – but pyqtgraph decides what it writes FROM THE NAME (FastPlot.exportbranches on.pdf/.svg/ else), so a renderer that is handedvolcano.pdfwrites a PDF whatever the user chose. The name has to be settled before the export sees it, and settling it twice in two modules is how the two drift apart.- Parameters:
path – a destination, with or without an extension.
fmt – force a format, bypassing the preference. Unknown formats fall back to the preference rather than raising: a run must not lose a figure over a typo.
- Returns:
the path as a
str.
- spacr.plot.generate_mask_random_cmap(mask)[source]¶
Return a random
ListedColormapsized to the labels inmask.- Parameters:
mask – Label mask array (0 = background).
- Returns:
Random colormap where index 0 is black and remaining entries are random opaque RGBA colours.
- spacr.plot.generate_plate_heatmap(df, plate_number, variable, grouping, min_max, min_count)[source]¶
Aggregate a well-level DataFrame into a plate-shaped heatmap.
The grid is read off the data. It used to be pinned to
r1..r16byc1..c27, so every well of a 1536 plate past row P or past column 27 fell outside theCategorical, became NaN, and was dropped by the groupby — measured, in the database, and absent from the figure with nothing said. Rows and columns now go throughspacr.plate_qc.parse_row_label()/parse_column_label(which isspacr.schema’s letter walk, soAA…AFand beyond are real rows), and the axes span exactly the wells present: a 96 plate is still 8x12 and a 384 still 16x24, because nothing is padded out to the largest format that exists.A well that genuinely cannot be placed — a
prcwith too few parts, or a row/column token holding no position — is reported throughspacr.errors.raise_if_strict()(anERRORonspacr.errors, or a raise underSPACR_STRICT_ERRORS) naming the identifiers concerned. Replacing a silent drop with a quieter silent drop would fix nothing.- Parameters:
df – Long-format DataFrame with a
prc(plate_row_column) identifier and the requestedvariablecolumn.plate_number – Plate ID selecting the subset to display.
variable – Column to aggregate. Ignored when
grouping='count'.grouping – Aggregation —
'count','mean'or'sum'.min_max – Colour scale spec —
'all','allq', or a two-element list[vmin, vmax](floats treated as quantiles).min_count – Drop wells with fewer than this many rows.
- Returns:
(plate_map, (vmin, vmax))— the pivoted matrix, indexed'r<N>'by'c<N>', and the colour-limit tuple.- Raises:
ValueError – if
groupingis not one of the accepted values.KeyError – if
variableis missing and required.
- spacr.plot.graph_importance(settings)[source]¶
Concatenate feature-importance CSVs and hand off to
spacrGraphfor plotting.- Parameters:
settings – Settings dict with
csvs(single path or list),grouping_column,data_column,graph_type,save.- Returns:
None (side-effects: plot shown, artefacts saved).
- spacr.plot.illegible_data_colours(fig, ground, floor=None)[source]¶
The data colours a reader will not find on
ground, as hex.- Parameters:
fig – Matplotlib figure whose data colours are checked.
ground – background colour against which contrast is measured.
The data deliberately does NOT flip, so a palette chosen against near-black can be illegible on paper – and the honest answer is to NAME the colour, because a substitution the user did not ask for changes what the picture says. Deduplicated and sorted so the same sentence comes out of the same figure twice.
- spacr.plot.jitterplot_by_annotation(src, x_column, y_column, plot_title='Jitter Plot', output_path=None, filter_column=None, filter_values=None)[source]¶
Read measurements + annotation from a spacr DB and plot a class-balanced jitter plot.
- Parameters:
src – Path to a spacr experiment directory containing
measurements/measurements.db.x_column – Column used as grouping variable (x-axis).
y_column – Numeric column plotted on the y-axis.
plot_title – Title for the plot. Default
'Jitter Plot'.output_path – If set, save the figure to this path; otherwise show it.
filter_column – Optional column (or list of columns) to filter rows on before plotting.
filter_values – Values (or list of value lists) accepted per
filter_column.
- Returns:
Balanced
DataFrameused for the plot.- Raises:
KeyError – if required plate/row/col columns are missing.
- spacr.plot.normalize_and_visualize(image, normalized_image, title='')[source]¶
Show the original and the normalised image side by side in grayscale.
Multi-channel inputs are averaged over their channels for display.
- Parameters:
image – Original image, 2D or
(H, W, C).normalized_image – Normalised counterpart to compare against.
title – Suffix appended to both panel titles. Default
"".
- Returns:
None
- spacr.plot.outline_palette_colours(palette)[source]¶
The four outline colours for
palette.- Parameters:
palette – a key of
OUTLINE_PALETTES. Anything unknown – includingNone– falls back todefaultrather than raising: a figure drawn in the historic colours is a far smaller problem than a pipeline that stops at the plotting step.- Returns:
{object_name: colour}.
- spacr.plot.overlay_masks_on_images(img_folder, normalize=True, resize=True, save=False, plot=False, thickness=2)[source]¶
Overlay
masks/*outlines onto matching images fromimg_folder.- Parameters:
img_folder – Folder containing images; masks live in
img_folder/maskswith matching filenames.normalize – If True, percentile-normalise images before blending. Default
True.resize – If True, resize the blended overlay to 1000x1000. Default
True.save – If True, write PNGs to
img_folder/overlay/. DefaultFalse.plot – If True, show each overlay via matplotlib. Default
False.thickness – Contour line thickness in pixels. Default
2.
- Returns:
{'written': int, 'failed': [(filename, reason)]}. A field that cannot be read is named and skipped rather than ending the run, so a folder holding one truncated TIFF still produces every other overlay – and the caller can tell which ones are missing.
- spacr.plot.plot_arrays(src, figuresize=10, cmap='inferno', nr=1, normalize=True, q1=1, q2=99)[source]¶
Plot random
.npy/.npzarrays fromsrc, one channel per subplot.- Parameters:
src – Directory or single
.npy/.npzpath.figuresize – Base figure size. Default
10.cmap – Matplotlib colormap. Default
'inferno'.nr – Maximum number of arrays to plot. Default
1.normalize – If True, percentile-normalise before display. Default
True.q1 – Lower percentile for normalisation. Default
1.q2 – Upper percentile for normalisation. Default
99.
- Returns:
None
- spacr.plot.plot_cellpose4_output(batch, masks, flows, cmap='inferno', figuresize=10, nr=1, print_object_number=True)[source]¶
Display per-channel images, label mask and flow field for Cellpose v4 outputs.
- Parameters:
batch – Image batch of shape
(N, H, W, C).masks – Label masks, one per image.
flows – Flow arrays, one per image.
cmap – Colormap for image channels. Default
'inferno'.figuresize – Base figure size. Default
10.nr – Maximum number of images to plot. Default
1.print_object_number – If True, annotate each object with its label ID. Default
True.
- Returns:
None
- spacr.plot.plot_comparison_results(comparison_results)[source]¶
Plot Jaccard, Dice, boundary-F1 and average-precision distributions per comparison.
- Parameters:
comparison_results – Iterable of dicts with per-file metrics (each key ending in
jaccard/dice/boundary_f1/average_precision).- Returns:
The generated
Figure.
- spacr.plot.plot_data_from_csv(settings)[source]¶
Load per-plate CSVs, filter/outlier-clean and render a
spacrGraphplot.- Parameters:
settings – Settings dict — see
settings.get_plot_data_from_csv_default_settingsfor keys (src,data_column,grouping_column,keep_groups,remove_outliers,graph_type,graph_name, …).- Returns:
(fig, results_df, df)— the figure, stats DataFrame and plotted DataFrame.- Raises:
ValueError – if
srcis not a string or list.
- spacr.plot.plot_data_from_db(settings)[source]¶
Read one or more measurement DBs, annotate conditions and render a
spacrGraphplot.Concatenates results across source directories, derives the
recruitmentcolumn if requested, drops missing rows, then hands the data tospacrGraphfor statistics + plotting.- Parameters:
settings – Settings dict. See
settings.set_default_plot_data_from_dbfor accepted keys (notablysrc,database,table_names,data_column,grouping_column,graph_type,graph_name).- Returns:
The plotted DataFrame, or
Nonewhen the requested data or grouping column is missing.- Raises:
ValueError – if
srcis neither a string nor a list.
- spacr.plot.plot_feature_importance(feature_importance_df, title='')[source]¶
Plot a horizontal bar chart of raw feature importances.
- Parameters:
feature_importance_df – DataFrame with columns
featureandimportance.title – what the bars MEAN, when it is not the model’s own importances. Four of the classifiers spaCR offers expose no
feature_importances_and are drawn from permutation importance instead – a different quantity, measuring what the fitted model loses when a column is shuffled – and a panel that did not say so would be passing one off as the other.
- Returns:
The generated
Figure.
- spacr.plot.plot_histogram(df, column, dst=None)[source]¶
Plot a histogram of
df[column]and optionally save it as PDF.- Parameters:
df – DataFrame containing
column.column – Column to plot.
dst – If set, save under
<dst>/<column>_histogram.pdf.
- Returns:
None
- spacr.plot.plot_image_grid(image_paths, percentiles)[source]¶
Render a square grid of percentile-normalised images with a black background.
Each tile carries its source file and the per-channel display range it was stretched to, so a checked export (see
save_figure()) can say when tiles are scaled differently and write a provenance sidecar that rebuilds every tile from its file.- Parameters:
image_paths – Image files to display; extra tiles are filled black.
percentiles – Two-element percentile pair used to normalise each channel.
- Returns:
The generated
Figure.
- spacr.plot.plot_image_mask_overlay(file, channels, cell_channel, nucleus_channel, pathogen_channel, organelle_channel=None, figuresize=10, percentiles=(2, 98), thickness=3, save_pdf=True, mode='outlines', export_tiffs=False, all_on_all=False, all_outlines=False, filter_dict=None, outline_palette='default')[source]¶
Plot image and mask overlays.
Loads the merged
.npystack, draws one panel per requested channel with the object masks applied as contours or filled labels, and closes with a panel showing every object combined.- Parameters:
file – Path to the merged
.npystack for one field of view.channels – Indices of the image channels to draw, one panel each.
cell_channel – Intensity channel the cell mask belongs to, or
Nonewhen there is no cell mask.nucleus_channel – Intensity channel the nucleus mask belongs to, or
None.pathogen_channel – Intensity channel the pathogen mask belongs to, or
None.organelle_channel – Intensity channel the organelle mask belongs to, or
None. DefaultNone.figuresize – Figure height in inches; the figure is drawn four times as wide. Default
10.percentiles – Two-element percentile pair used to normalise each channel. Default
(2, 98).thickness – Contour line width in pixels. Default
3.save_pdf – If True, save the figure into
results/overlay/two directories abovefile, in the configured figure format rather than always as PDF. DefaultTrue.mode –
'outlines'draws mask contours; any other value overlays filled, randomly coloured labels. Default'outlines'.export_tiffs – If True, also write every stack plane as a grayscale TIFF into
results/<stem>/tiff/alongside it. DefaultFalse.all_on_all – If True, draw every mask on every channel. Default
False.all_outlines – If True, draw every mask on the channels that own no mask themselves. Default
False.filter_dict – Optional per-object limits keyed by
'cell','nucleus','pathogen'or'organelle', each holding((min_area, max_area), (min_intensity, max_intensity)); objects outside the limits are dropped before plotting.outline_palette – which outline colours to draw, a key of
OUTLINE_PALETTES.'default'is what spaCR has always drawn;'colourblind'is legible under red-green and blue-yellow deficiency, where the default’s worst pair scores 27 out of 255 – cell is drawn red and pathogen green, the one pair the commonest deficiency removes. Default'default', because changing every figure a user has already made would be worse than the defect.
- Returns:
The generated matplotlib
Figure.
- spacr.plot.plot_image_mask_overlay_magenta_outlines(file, channels, cell_channel, nucleus_channel, pathogen_channel, figuresize=10, percentiles=(2, 98), thickness=3, save_pdf=True, mode='outlines', export_tiffs=False, all_on_all=False, all_outlines=False, filter_dict=None)[source]¶
Plot image and mask overlays, outlining each channel’s own mask in magenta.
Variant of
plot_image_mask_overlay()with noorganelle_channel: whenmodeis'outlines'andall_on_allis False, the mask belonging to a channel is outlined in magenta rather than in that object’s colour. In every other mode it falls back to filled, randomly coloured labels as that function does, but seeded per call rather than per object, so the colours differ between runs and between panels.- Parameters:
file – Path to the merged
.npystack for one field of view.channels – Indices of the image channels to draw, one panel each.
cell_channel – Intensity channel the cell mask belongs to, or
Nonewhen there is no cell mask.nucleus_channel – Intensity channel the nucleus mask belongs to, or
None.pathogen_channel – Intensity channel the pathogen mask belongs to, or
None.figuresize – Figure height in inches; the figure is drawn four times as wide. Default
10.percentiles – Two-element percentile pair used to normalise each channel. Default
(2, 98).thickness – Contour line width in pixels. Default
3.save_pdf – If True, save the figure into
results/overlay/two directories abovefile, in the configured figure format rather than always as PDF. DefaultTrue.mode –
'outlines'draws mask contours; any other value overlays filled, randomly coloured labels. Default'outlines'.export_tiffs – If True, also write every stack plane as a grayscale TIFF into
results/<stem>/tiff/alongside it. DefaultFalse.all_on_all – If True, draw every mask on every channel in its own colour. Default
False.all_outlines – If True, draw every mask on the channels that own no mask themselves. Default
False.filter_dict – Optional per-object limits with a
'cell','nucleus'and'pathogen'entry, each holding((min_area, max_area), (min_intensity, max_intensity)); objects outside the limits are dropped before plotting.
- Returns:
The generated matplotlib
Figure.
- spacr.plot.plot_images_and_arrays(folders, lower_percentile=1, upper_percentile=99, threshold=1000, extensions=None, overlay=False, max_nr=None, randomize=True)[source]¶
Show side-by-side images and arrays found across multiple folders.
Each image is either percentile-normalised (values below
threshold) or shown as a label mask. Optionally overlays object outlines from a matching mask file.- Parameters:
folders – Folders to scan for image/array files.
lower_percentile – Lower percentile clip. Default
1.upper_percentile – Upper percentile clip. Default
99.threshold – Values <= threshold are treated as label data instead of intensity. Default
1000.extensions – File extensions to include. Default
['.npy', '.tif', '.tiff', '.png'].overlay – If True, overlay object outlines. Default
False.max_nr – Maximum number of key groups to plot.
randomize – If True, shuffle key order before plotting. Default
True.
- Returns:
None
- spacr.plot.plot_lorenz_curves(csv_files, name_column='grna_name', value_column='count', remove_keys=None, x_lim=None, y_lim=None, remove_outliers=False, save=True)[source]¶
Overlay Lorenz curves from multiple gRNA count CSVs with per-plate Gini coefficients.
- Parameters:
csv_files – Paths to per-plate CSVs, each with columns
name_columnandvalue_column.name_column – Identifier column used for outlier filtering. Default
'grna_name'.value_column – Column whose distribution is analysed. Default
'count'.remove_keys – Names to exclude before analysis. Default
[](exclude nothing).x_lim – X-axis limits
[lo, hi]. Default[0.0, 1].y_lim – Y-axis limits
[lo, hi]. Default[0, 1].remove_outliers – If True, drop names whose per-well count falls outside a 1.5*IQR window. Default
False.save – If True, save the figure alongside the first CSV under
results/lorenz_curve_with_gini.pdf. DefaultTrue.
- Returns:
None
- spacr.plot.plot_masks(batch, masks, flows, cmap='inferno', figuresize=10, nr=1, file_type='.npz', print_object_number=True)[source]¶
Display per-channel images, label masks and flow fields for a batch.
- Parameters:
batch – Image batch — shape
(N, H, W, C)or a single image of shape(H, W, C).masks – Label masks, one per image (list or ndarray).
flows – Flow arrays, one per image.
cmap – Colormap for image channels. Default
'inferno'.figuresize – Base figure size. Default
10.nr – Maximum number of images to plot. Default
1.file_type – Source file type of
flows—'png'selects the first element of each flow entry. Default'.npz'.print_object_number – If True, annotate each object with its label ID. Default
True.
- Returns:
None
- spacr.plot.plot_merged(src, settings)[source]¶
Show multi-channel image stacks with per-object outlines overlaid.
- Parameters:
src – Folder containing
.npymerged stacks.settings – Plot settings dict — includes channel/mask dims, overlay colours, normalisation, filter and object-count keys.
- Returns:
The last generated
Figurewhensettings['nr']is exceeded; otherwiseNone.
- spacr.plot.plot_object_outlines(src, objects=None, channels=None, max_nr=10)[source]¶
Overlay mask outlines on the matching channel image for each object type.
- Parameters:
src – Experiment root;
masks/<object>_mask_stackand channel folders live directly under it.objects – Object types to plot. Default
['nucleus', 'cell', 'pathogen'].channels – Channel indices paired with
objects(channel folders are named<channel + 1>). Default[0, 1, 2].max_nr – Maximum number of images to plot per object. Default
10.
- Returns:
None
- spacr.plot.plot_organelle_output(img_batch, masks, settings, cmap='inferno', figuresize=10, nr=1, print_object_number=True)[source]¶
Plot organelle segmentation results: raw channel, label mask, morphology-specific diagnostic.
- Parameters:
img_batch – Single-channel image batch of shape
(N, H, W).masks – Label masks, one per image.
settings – Organelle settings dict;
organelle_morphologyandorganelle_methoddrive the diagnostic panel.cmap – Colormap for the raw channel. Default
'inferno'.figuresize – Base figure size. Default
10.nr – Maximum number of images to plot. Default
1.print_object_number – If True, annotate each object with its label ID. Default
True.
- Returns:
None
- spacr.plot.plot_permutation(permutation_df)[source]¶
Plot a horizontal bar chart of permutation feature importances with error bars.
- Parameters:
permutation_df – DataFrame with columns
feature,importance_meanandimportance_std.- Returns:
The generated
Figure.
- spacr.plot.plot_plates(df, variable, grouping, min_max, cmap, min_count=0, verbose=True, dst=None)[source]¶
Render every plate of a screen as ONE panel, wells square, on one colour scale.
The layout, the colour scale and the treatment of unmeasured wells live in
spacr.figures.plates; this function is the call the pipeline already makes, kept at its own signature.WHAT CHANGED, AND WHY (the design – “the lpates look super small on the collected figure”): the plates were laid out four-per-row on a 40 x 5 inch figure, an 8:1 strip that uses about an eighth of a square tile in the figure grid, with wells 1.14:1 rather than square. They are now a SMALL MULTIPLE – 2 x 2 for a four-plate screen – which is a 1.3:1 composite, and the figure is sized from the well grid so the wells come out exactly square.
Two things that were wrong with the picture and not only with its size: each plate carried its OWN colour scale, so the same blue meant a different number on the plate beside it; and a well that was never measured was drawn as a measurement of zero, which on a screen with 155 of 384 wells used is more than half the panel — and set the bottom of the scale. One scale is now shared across the plates, and an unmeasured well is drawn as a neutral wash and left out of the scale.
- Parameters:
df – Long-format DataFrame with a
prccolumn of the formplateID_rowID_columnIDand the column named byvariable.variable – Column to aggregate (see
generate_plate_heatmap()).grouping – Aggregation mode —
'count','mean'or'sum'.min_max – Color-scale spec (
'all','allq',[vmin, vmax]), applied ONCE over every plate rather than once per plate.cmap – Matplotlib colormap name or object.
None— or the legacy'viridis'literal — uses the house single-hue ramp.min_count – Drop wells with fewer than this many rows before plotting. Default
0.verbose – If True, call
plt.show()after building the figure. DefaultTrue.dst – If given, save the figure as
<dst>/plate_heatmap_<variable>.pdf.
- Returns:
The generated matplotlib
Figure.
Example
from spacr.plot import plot_plates fig = plot_plates( df, variable='recruitment', grouping='mean', min_max='allq', cmap=None, min_count=20, )
See also
spacr.figures.plates.build_plates()— the panel itself, which also returns the legend sentence for it.spacr.ml.generate_ml_scores()— produces score dataframes typically fed to this plotter.
- spacr.plot.plot_proportion_stacked_bars(settings, df, group_column, bin_column, prc_column='prc', level='object', cmap='viridis')[source]¶
Plot stacked proportion bars per group with chi-squared and pairwise stats.
- Parameters:
settings – Settings dict —
verbosetoggles pairwise chi-squared verbosity.df – Long-format DataFrame with categorical
group_columnandbin_column.group_column – Group axis of the stacked bars.
bin_column – Categorical column stacked within each bar.
prc_column – Per-well identifier used when aggregating at the well or plate level. Default
'prc'.level – Aggregation level —
'object'for direct counts, or'well'/'plateID'for per-well means with SD bars.cmap – Matplotlib colormap. Default
'viridis'.
- Returns:
(results_df, pairwise_results, fig)— chi-squared summary, pairwise comparison table and the plot figure.
- spacr.plot.plot_region(settings)[source]¶
Render mask overlay, cropped PNG grid and activation-map grid for one FOV.
Reads the FOV’s merged NPY, resolves its PNG crops and activation maps from the measurements and activation DBs, and writes the three figures under
<src>/results/<name>/when possible — in the configured figure format, so PDF only while that is the preference.- Parameters:
settings – Settings dict with
src,name,channels,cell_channel,nucleus_channel,pathogen_channel,percentiles,activation_mode,activation_db,mode,export_tiffs.- Returns:
Tuple
(fig_mask_overlay, fig_png_grid, fig_activation_grid)— any element may beNonewhen the corresponding assets were not found.
- spacr.plot.plot_resize(images, resized_images, labels, resized_labels)[source]¶
Show original vs. resized image/label pairs in a 2x2 grid.
- Parameters:
images – Sequence of original images (first element shown).
resized_images – Sequence of resized images.
labels – Sequence of original label arrays.
resized_labels – Sequence of resized label arrays.
- Returns:
None
- spacr.plot.print_mask_and_flows(stack, mask, flows, overlay=True, max_size=1000, thickness=2)[source]¶
Show a single image, its label mask (optionally outlined) and flow image.
- Parameters:
stack – Original 2D image or
(H, W, C)stack.mask – Label mask matching
stackspatially.flows – Optional list of flow arrays; skipped when
None.overlay – If True, draw mask contours over the image instead of showing the mask alone. Default
True.max_size – Downsample any dimension exceeding this size. Default
1000.thickness – Contour line thickness in pixels. Default
2.
- Returns:
None
- spacr.plot.print_ready(fig, mode=None, announce=True)[source]¶
Repaint
fig’s chrome for paper for the length of the block.THE CONTRACT IS THAT NOTHING SURVIVES IT. Every artist touched is restored in a
finally, so a user watching a plot while it saves must not see it flash and the figure is byte-identical afterwards – the own acceptance, and the reason this is a context manager rather than a function that fixes a figure up.WHAT MOVES: an illegible ground becomes the page; illegible chrome becomes the ink; an illegible GRID becomes a faint print grey rather than the ink, because a grid repainted in the ink is a cage over the data.
WHAT DOES NOT MOVE: every data colour, and any chrome that was already legible on the page. A light-mode save therefore changes nothing at all, which is the property that makes this safe to switch on by default.
- Parameters:
fig – matplotlib figure whose page, chrome and grid artists are repainted temporarily and restored when the context exits.
mode – one of
spacr.figure_style.SAVE_MODES; None asks the preference.'screen'is a no-op by construction.announce – print the 150 D sentence when a data colour has stopped working on the page. Off for a caller that saves in a loop.
- spacr.plot.proportion_mixed_model(df, group_column, bin_column, unit_column)[source]¶
A binomial GLM on the per-object outcome, standard errors clustered by unit.
- Parameters:
df – object-level observations carrying group, bin, and unit fields.
group_column – column naming the conditions to compare.
bin_column – categorical outcome column modelled one bin at a time.
unit_column – column naming clusters used for robust standard errors.
The proportions test throws away how many objects each well contributed; this keeps them while still charging the degrees of freedom the DESIGN supports, by clustering on the unit. Reported beside the other two because when it disagrees with them, the disagreement is the finding.
- spacr.plot.proportion_test_by_unit(df, group_column, bin_column, unit_column)[source]¶
Compare conditions on their PER-UNIT proportions, one row per bin.
- Parameters:
df – object-level observations carrying group, bin, and unit fields.
group_column – column naming the conditions to compare.
bin_column – categorical outcome column whose bins are tested.
unit_column – column naming independent replication units.
The object-level chi-squared asks whether 20,000 objects came from one distribution. Objects in a well share a treatment, a transfection, an imaging session and a monolayer, so that is not the question anyone asked, and its p-value is smaller than the experiment supports by orders of magnitude. This asks the question the design supports: do the WELLS differ, with n = the number of wells.
- spacr.plot.proportions_per_unit(df, group_column, bin_column, unit_column)[source]¶
Each unit’s share of every bin, one row per unit.
- Parameters:
df – object-level observations carrying group, bin, and unit fields.
group_column – column naming the conditions to compare.
bin_column – categorical outcome column whose shares are computed.
unit_column – column naming independent wells, plates, or other replication units.
- Returns:
a frame with
group_column,unit_columnand one column per bin holding a proportion in [0, 1]. Units contributing no objects do not appear.
- spacr.plot.random_cmap(num_objects=100)[source]¶
Return a random
ListedColormapwithnum_objects + 1colours.- Parameters:
num_objects – Number of foreground colours to generate. Default
100.- Returns:
Colormap with index 0 = black and remaining indices random opaque RGBA colours.
- spacr.plot.read_and_plot__vision_results(base_dir, y_axis='accuracy', name_split='_time', y_lim=None)[source]¶
Aggregate vision-model test CSVs under
base_dirand plot mean score per model.- Parameters:
base_dir – Root directory containing
*_test_result.csvfiles nested per epoch.y_axis – Metric column to average. Default
'accuracy'.name_split – Substring that splits filename into model name and epoch info. Default
'_time'.y_lim – Y-axis limits
[lo, hi]. Default[0.8, 0.9].
- Returns:
None
- spacr.plot.save_figure(fig, path, *, fmt=None, dpi=None, close=False, save_mode=None, announce_colours=True, integrity=None, **kwargs)[source]¶
Write
figtopath, honouring the figure preferences.The single place a spaCR figure the user keeps gets written. Before this existed there were sixty-odd
savefigcalls, each with its own hard-coded format and DPI, and the “Figure format” and “Resolution” preferences reached exactly two of them – both writing to a temp directory. Everything a pipeline saved into its results folder ignored both settings entirely.Three things are decided here rather than left to the caller or to matplotlib.
The format follows the preference, and the file NAME follows the format: a PNG written to
figure.pdfis a file no viewer opens. An explicitfmt=still wins, for the few callers that genuinely need one particular format.Fonts are embedded as TrueType (
pdf.fonttype = 42) for the length of the save. matplotlib’s default is Type 3, which draws every glyph as its own content stream: the file is still vector, but Illustrator and Inkscape open the text as unselectable outlines, and the preference that selects this path is labelled “PDF (vector, editable)”. Scoped withrc_contextso a caller that has deliberately chosen otherwise is not changed underneath it.The DPI is passed, always. A PDF page is resolution-independent, but spaCR figures are full of
imshowpanels – cell montages, mask overlays, plate heatmaps – and those are rasterised at the figure’s own 100 DPI unless told otherwise. Without it, 100, 300 and 600 produced byte-identical files. What is passed isdeliverable_dpi(), which says so out loud when the requested number is not achievable here.The chrome is repainted for paper, and only for the length of the write. A figure saved from a dark-themed session was white ink on a white page –
spacr.qt.preferences.get_figure_colorshands both renderers a white foreground on a dark theme and nothing inverted it at export time, so the file was a blank rectangle with some coloured dots in it, and a transparent PNG even looked right in a dark file manager and disappeared when it was pasted into a manuscript.print_ready()moves the furniture and puts it back; the DATA is never touched, because a white point turned black is, on a volcano, the colour of “not a hit”.- Parameters:
fig – a matplotlib
Figure.path – destination; its extension is corrected to the format.
fmt – force a format, bypassing the preference.
dpi – force a DPI, bypassing the preference.
close – close the figure once written.
save_mode – force one of
spacr.figure_style.SAVE_MODES–'print'(light page, dark chrome),'screen'(exactly what is on screen, the old behaviour) or'transparent'. None asks the preference, which defaults to'print'.announce_colours – say when a data colour has stopped working on the light page. It is NAMED, never substituted.
integrity – check the figure’s image panels before writing – display ranges that differ between panels meant for comparison, saturated or clipped pixels, repeated panels, a lossy format, and panels written with fewer pixels than they hold – print any warning, stamp the provenance (source files, display settings, processing steps, spaCR version) into the PNG or PDF metadata and write it to a
<file>.provenance.jsonsidecar beside the figure. None followsSPACR_FIGURE_INTEGRITYand then the Preferences toggle, which is off by default. A figure without image panels is written unchanged either way.kwargs – passed through to
savefig(bbox_inchesetc.). An explicitfacecolorstill wins over the print ground.
- Returns:
the path actually written, as a
str.
- spacr.plot.visualize_cellpose_masks(masks, titles=None, filename=None, save=False, src=None)[source]¶
Display several Cellpose-style label masks side by side for a quick visual QC.
Handy for sanity-checking the masks produced by
spacr.core.preprocess_generate_masks()(e.g. compare the cell, nucleus and pathogen masks of the same field, or two runs against each other). Each mask is rendered with a random-color palette so neighbouring objects stay distinguishable.- Parameters:
masks – Sequence of 2D label mask arrays.
titles – Titles paired positionally with
masks. Falls back to'Mask 1','Mask 2', …filename – Used in the figure suptitle and, when
save, as the output PDF filename.save – If True, save the figure under
<src>/results/<filename>.pdf. DefaultFalse.src – Root folder for saving. Defaults to the current working directory.
- Returns:
None. Displays (and optionally writes) the figure.
- Raises:
AssertionError – if
titlesandmaskshave different lengths.
Example
from spacr.plot import visualize_cellpose_masks visualize_cellpose_masks( [cell_mask, nucleus_mask, pathogen_mask], titles=['cell','nucleus','pathogen'], filename='field_001', save=True, src='/data/plate01', )
See also
spacr.core.preprocess_generate_masks()— produces the masks visualized here.
- spacr.plot.visualize_masks(mask1, mask2, mask3, title='Masks Comparison')[source]¶
Show three masks side by side with random colormaps.
- Parameters:
mask1 – First label mask.
mask2 – Second label mask.
mask3 – Third label mask.
title – Figure suptitle. Default
"Masks Comparison".
- Returns:
None
- spacr.plot.volcano_plot(data: str | pandas.DataFrame, *, fold_change_col: str, p_value_col: str, name_col: str | None = None, x_transform: str = 'none', y_transform: str = '-log10', fold_change_threshold: float | None = None, p_value_threshold: float | None = None, annotate: bool = True, annotate_max: int | None = None, point_size: float = 20.0, alpha: float = 0.7, figsize: Tuple[float, float] = (8.0, 6.0), title: str | None = None, xlim: Tuple[float, float] | None = None, ylim: Tuple[float, float] | None = None, threshold_line_kwargs: dict | None = None, scatter_kwargs: dict | None = None, text_kwargs: dict | None = None, save_path: str | None = None, show: bool = True, ax: matplotlib.pyplot.Axes | None = None, sheet_name: int | str = 0) Tuple[matplotlib.pyplot.Figure, matplotlib.pyplot.Axes, list][source]¶
Read a table (CSV/TSV/XLS/XLSX or a DataFrame) and render a volcano plot.
Auto-detects file type from extension (.csv, .tsv/.tab, .xls/.xlsx) and applies the requested x/y transforms before drawing.
- Parameters:
data – Path to table file or a pandas
DataFrame.fold_change_col – Column of raw fold change (or logFC when
x_transform='none').p_value_col – Column of p-values.
name_col – Optional column supplying point labels.
x_transform – One of
'none','log2','log10','ln'. Use'none'when the column already stores logFC (may be negative).y_transform – One of
'none','-log10','-ln','log10','ln'. Default'-log10'.fold_change_threshold – Threshold on x — in plotted units when
x_transform='none', otherwise in raw FC units.p_value_threshold – Threshold on raw p; drawn as a dashed horizontal line in plotted units.
annotate – Annotate significant points when a name column is supplied.
annotate_max – Cap on the number of annotated points (highest y first).
point_size – Scatter marker size.
alpha – Scatter marker alpha.
figsize – Figure size in inches.
title – Optional figure title.
xlim – Optional x-axis limits.
ylim – Optional y-axis limits.
threshold_line_kwargs – Extra kwargs for threshold lines.
scatter_kwargs – Extra kwargs for the scatter call.
text_kwargs – Extra kwargs for label texts.
save_path – If given, save the figure to this path.
show – Call
plt.show()at the end. DefaultTrue.ax – Existing axes to draw on; a new figure is created if None.
sheet_name – Excel sheet index/name for .xls/.xlsx inputs.
- Returns:
(fig, ax, hits)wherehitsare the labels drawn.- Raises:
ValueError – on unknown transforms, or numeric columns that cannot be coerced.