spacr.methods_export

Workflow inputs and outputs

Methods & Results

Draft methods/results text from recorded settings and results; review every claim before publication.

Open: Regression → Methods & Results.

Inputs and outputs below include conditional alternatives. The guidance and handoff notes say which route applies.

Inputs

  • Run history and artifacts — Project run records, settings, output paths, artifact provenance, status and logs.

  • Regression results and hits — Selected run results folder: coefficient/result CSVs, hit tables, settings and diagnostic figures.

Outputs

  • Shareable reports — Exported HTML/PDF or methods/results documents derived from recorded analysis outputs.

Before this module

  • Regression: Review exported prose and every traced result.

API reference.

Module tutorial.

Methods and Results sections, written from a run digest the model cannot leave.

The deliverable is two paragraphs of a paper: a methods section that says what was actually done, and a results section that says what came out. The hard part is not the prose. The hard part is that a language model asked to write about an experiment will produce numbers, and the numbers will be plausible, and nobody reading the paragraph can tell which of them came from the run.

So the model never sees the data. It sees a run digest: a structured record of the modules that ran, the parameters they ran with, the counts, the versions, the QC verdicts and the statistics already computed by the rest of spaCR. It writes prose around those numbers, and then:

  • verify_numbers() extracts every number from what came back and checks each one against the digest. A number that is not in the digest — not as a value, not as a correct rounding of one, not as a verbatim quote of a digest string — is unsupported, and

  • check_draft() refuses a draft that carries one. The refusal is the guarantee: the model cannot introduce a figure, because a draft carrying an invented figure does not get returned as a draft.

Two consequences worth stating plainly. First, the digest is the contract: anything the prose may quote has to be in it, which is why the digest carries the alpha, the confidence level and the seed as numbers rather than leaving the model to supply the obvious ones. Second, there is a deterministic renderer — render_methods() and render_results() — that writes the same sections from the same digest with no model at all. It is what runs when no AI provider is configured, it is what a rejected draft falls back to, and it is what the number-provenance tests assert against, because a test that plants a number in the digest and requires it in the output must not be testing a stub.

The caveats are not optional. A methods section that omits a QC failure, or the fact that illumination correction did not run, or the on_error=skip that dropped eleven fields, is not a shorter methods section — it is a wrong one. build_digest() collects those into DIGEST_CAVEATS and both renderers, and the prompt, are required to state every one.

Public API:

from spacr.methods_export import build_digest, render_methods

digest = build_digest(project="/data/plate7", run_dir=run.dir,
                      results_folder=".../results/pred/ols")
print(render_methods(digest))
print(render_results(digest))

The AI half lives in spacr.qt.ai.manuscript, which reuses the console’s existing provider plumbing rather than opening a second client.

Classes

Verification

The verdict on one generated section's numbers.

Functions

build_digest(, regression_type, hits, model_path, ...)

Assemble everything a methods and results section may quote.

check_draft(→ Tuple[Verification, Verification])

Verify both sections. (methods verdict, results verdict).

digest_numbers(→ Set[float])

Every number the digest actually asserts.

digest_strings(→ List[str])

Every non-trivial string the digest carries, longest first.

extract_numbers() → List[str])

Every number a paragraph asserts, as it was written.

render_methods(→ str)

Write the methods section from the digest, with no model involved.

render_prompt(→ Tuple[str, str])

(system prompt, user message) for one digest.

render_results(→ str)

Write the results section from the digest, with no model involved.

system_prompt(→ str)

The instruction the model is held to. Names the rule it must not break.

verify_numbers(→ Verification)

Check that every number in text came from digest.

Module Contents

class spacr.methods_export.Verification[source]

The verdict on one generated section’s numbers.

Parameters:
  • ok – no unsupported number was found.

  • checked – how many number tokens were examined.

  • supported – the tokens that trace to the digest.

  • unsupported – the tokens that do not. These are the inventions.

  • missing_caveats – caveats the digest carries that the text omits.

problem() → str[source]

One sentence naming what is wrong, or "".

to_dict() → Dict[str, Any][source]

A JSON-serializable copy.

spacr.methods_export.build_digest(*, project: str | os.PathLike | None = None, run_dir: str | os.PathLike | None = None, macro_path: str | os.PathLike | None = None, settings: Mapping[str, Any] | None = None, results_folder: str | os.PathLike | None = None, metadata_files: Sequence[str | os.PathLike] = (), regression_type: str = '', hits: Any = None, model_path: str | os.PathLike | None = None, title: str = '', top_hits: int = 10, extra: Mapping[str, Any] | None = None) → Dict[str, Any][source]

Assemble everything a methods and results section may quote.

Every source is optional and every one of them is read defensively: a subsystem that cannot answer contributes a note rather than an exception, because a digest is written at the END of a long run and losing it to a missing model card would be absurd.

Parameters:
  • project – the project root. Supplies the provenance summary from spacr.pipeline_graph and the segmentation QC from spacr.seg_qc.

  • run_dir – a spacr.run_journal run folder. Supplies the manifest: versions, timings, seeds, warnings, status.

  • macro_path – the emitted macro.py; defaults to the one inside run_dir. Supplies the per-module steps, the settings each ran with and which of them the user actually chose.

  • settings – settings to read directly, when there is no journal.

  • results_folder – a regression results folder, for the statistics.

  • metadata_files – annotation CSVs to join into the hit list.

  • regression_type – the backend, for how the hit list is ranked.

  • hits – an already-built spacr.hits.HitList, instead of reading results_folder.

  • model_path – a classifier checkpoint whose model card carries the held-out metrics.

  • title – what to call the experiment in the prose.

  • top_hits – how many hits to carry into the digest.

  • extra – anything else to record, under "extra".

Returns:

the digest, JSON-serializable throughout.

spacr.methods_export.check_draft(methods: str, results: str, digest: Mapping[str, Any]) → Tuple[Verification, Verification][source]

Verify both sections. (methods verdict, results verdict).

The methods section additionally has to state the run’s caveats.

Parameters:
  • methods – generated methods-section text to verify.

  • results – generated results-section text to verify.

  • digest – source digest against which numbers and caveats are checked.

spacr.methods_export.digest_numbers(digest: Mapping[str, Any]) → Set[float][source]

Every number the digest actually asserts.

Numeric leaves, plus strings that are entirely a number (a gene id such as "233460" is a string in the tables and a number in prose). Digits embedded in a longer string — a path, a run id, a version — are deliberately NOT harvested here: those are handled by digest_strings(), which lets the prose quote them verbatim without also licensing every digit inside them as a free-standing figure.

Parameters:

digest – a digest as build_digest() returns.

Returns:

the numbers, as floats.

spacr.methods_export.digest_strings(digest: Mapping[str, Any]) → List[str][source]

Every non-trivial string the digest carries, longest first.

Quoting one of these verbatim — a path, a run id, a settings digest, a package version, a gene name — is not inventing a number, so verify_numbers() removes them from the text before it looks for figures. Longest first so a version is removed whole rather than having its prefix eaten by a shorter match.

Parameters:

digest – a digest.

Returns:

the strings, longest first.

spacr.methods_export.extract_numbers(text: str, strings: Sequence[str] = ()) → List[str][source]

Every number a paragraph asserts, as it was written.

strings — usually digest_strings() — are removed first, so a sentence that quotes run 8f21c0a3 or /data/plate7 is not accused of asserting 8 and 7. Structural tokens (list markers, dotted versions, ISO dates, long hex ids) go next. What is left is a claim.

Parameters:
  • text – the prose to check.

  • strings – substrings that may be quoted verbatim.

Returns:

the number tokens, in the order they appear.

spacr.methods_export.render_methods(digest: Mapping[str, Any]) → str[source]

Write the methods section from the digest, with no model involved.

What runs when no AI provider is configured, what a rejected draft falls back to, and what the number-provenance tests assert against. Every number in the output comes from digest; every caveat in the digest is stated. No trailing newline.

Parameters:

digest – collected run facts from which to render the section.

spacr.methods_export.render_prompt(digest: Mapping[str, Any]) → Tuple[str, str][source]

(system prompt, user message) for one digest.

A pure function of the digest — which is the point. The model receives this and nothing else, so there is no path by which raw data could reach it, and a test can assert that a number planted in the digest is the only place a number in the prompt can have come from.

Parameters:

digest – the digest to write about.

Returns:

the two prompt halves.

spacr.methods_export.render_results(digest: Mapping[str, Any]) → str[source]

Write the results section from the digest, with no model involved.

Every number comes from digest. No trailing newline.

Parameters:

digest – collected run facts from which to render the section.

spacr.methods_export.system_prompt() → str[source]

The instruction the model is held to. Names the rule it must not break.

spacr.methods_export.verify_numbers(text: str, digest: Mapping[str, Any], *, require_caveats: bool = False) → Verification[source]

Check that every number in text came from digest.

This is the assertion the whole module exists to make. It is applied to what a model returns, and a draft that fails it is not returned as a draft — see check_draft().

Parameters:
  • text – the generated prose.

  • digest – the digest it was generated from.

  • require_caveats – also require every sentence in the digest’s caveats to be represented. Applied to the methods section, which is where the caveats belong.

Returns:

a Verification.