spacr.gene_facts

What spaCR already knows about one gene, ready to put on a tile.

Clicking a gene in an interactive regression opens the information available for that gene. spacr.gene_tile answers the half that can be WRONG – which gene a clicked guide actually names, and whether that mapping is ambiguous. This module answers the other half: given a gene, what does spaCR hold about it.

IT HOLDS NOTHING ITSELF. Every value here comes out of spacr.annotation – the same join that writes the annotated exports – so the tile and the CSV a reader opens beside it cannot disagree about GRA14’s compartment. This module reads no file, parses no gene id and owns no table; it groups spacr.annotation.annotate()’s 23 columns into the order a human reads them and adds the per-segment DeepTMHMM coordinates that annotate deliberately leaves out of a coefficient export.

Four rules, each of them a failure this project has already had:

  1. ONE PARSE. TGGT1_224750, gene_fraction:gene[224750] and the guide 224750_2 are all gene 224750, and the function that says so is spacr.annotation.gene_number(). A second copy of that rule is how the volcano and the hit list start naming different genes.

  2. A GAP IS SAID OUT LOUD. A gene with no annotation row gets GeneFacts.reason – “no row in the bundled Toxoplasma annotation” – and NOT a block of empty fields, which reads as “measured, found nothing” rather than “not available here”. GeneFacts.known is the flag a caller greys a control on.

  3. A GROUP IS DERIVED FROM THE TABLE, NOT LISTED HERE. The fitness and expression rows are every fit_/expr_ column spacr.annotation.columns() reports, so an eighth published screen added to phenotype.csv appears on the tile without this file being touched, and a column this module has never heard of lands in “other annotation” instead of being silently dropped.

  4. NOTHING HERE IS FOR THE GUI THREAD TO DISCOVER. The first call reads five bundled CSVs (360 ms) and the first GeneFacts.segments reads DeepTMHMM’s 8,140 rows. warm() does both off the GUI thread, and takes the screen’s own terms while it is there: annotate re-checks all five right-hand keys for uniqueness on every call, so one gene costs 20 ms and four hundred cost 21. A warmed click is a dict lookup, 0.02 ms.

Public API:

from spacr import gene_facts

known = gene_facts.facts("fraction:grna[239740_3]")
known.known                 # True
known.value("gene_name")    # 'GRA14'
known.sections()            # (('identity', (('gene name', 'GRA14'), ...

Classes

GeneFacts

Everything the bundled annotation holds about one gene.

Segment

One DeepTMHMM segment of a protein, with its residue coordinates.

Functions

available(→ Tuple[str, ...])

Every annotation column this module can show, in reading order.

clear_cache(→ None)

Forget the derived indices. For tests, and for a reinstall mid-session.

facts(→ GeneFacts)

Everything the bundled annotation holds about the gene value names.

facts_for(→ Dict[str, GeneFacts])

The facts for several genes at once, keyed by gene number.

unavailable_reason(→ str)

Why there is no annotation on this install, or "".

warm() → Tuple[str, ...])

Load every table this module reads. CALL THIS OFF THE GUI THREAD.

Module Contents

class spacr.gene_facts.GeneFacts[source]

Everything the bundled annotation holds about one gene.

Parameters:
  • gene – the bare gene number, or "" when the caller’s value named no gene at all.

  • values – the annotation columns that had something in them, keyed by the column name spacr.annotation.columns() uses. A column with a gap is ABSENT here rather than present and empty.

  • segments – the DeepTMHMM signal peptide and transmembrane segments, in residue order. Empty for a soluble protein, and empty for a gene DeepTMHMM never saw – reason is what tells those apart.

  • reason – why there is nothing, in a sentence, or "" when there is something. Never blank and empty at the same time: an empty panel reads as a bug, “no row in the bundled annotation” reads as an answer.

sections() → Tuple[Tuple[str, Tuple[Tuple[str, str], ...]], ...][source]

The facts as ((heading, ((label, value), ...)), ...).

THE ONE ORDERING, so the Qt tile and the text form cannot drift into presenting the same record two ways – the same shape spacr.gene_tile.GeneTile.sections() returns, so one renderer lays both halves of a tile out with one loop.

A group with nothing in it is not emitted. That is rule 2: a heading over a blank space is a claim that something was looked for and not found, which is a different statement from “not bundled here”.

to_html() → str[source]

The facts as HTML, for a Qt rich-text view.

Matches spacr.gene_tile.GeneTile.to_html() mark for mark, so the two halves of one tile do not read as two different documents.

to_text() → str[source]

The facts as plain text – what a test reads and a log records.

value(column: str, default: Any = None) → Any[source]

One annotation column, or default when it had a gap.

Parameters:

column – annotation column whose value is requested.

property known: bool[source]

Did the annotation hold anything at all about this gene?

The flag a caller greys a control on – the design: a control that cannot do anything says why rather than sitting there inert.

class spacr.gene_facts.Segment[source]

One DeepTMHMM segment of a protein, with its residue coordinates.

Parameters:
  • kind – signal peptide or transmembrane.

  • index – 1-based position among the segments of that kind; 1 for a signal peptide, of which there is at most one.

  • start – first residue, 1-based and inclusive, as DeepTMHMM reports.

  • end – last residue, inclusive.

  • length – residues spanned, as the source file records it rather than recomputed – a disagreement between the two is a fact about the run and is not this module’s to hide.

property label: str[source]

signal peptide span or TM 3 – the tile’s left column.

“span”, not “signal peptide”: the summary row two lines above is already labelled “signal peptide”, and two rows with one label reads as the panel having printed the same field twice.

property text: str[source]

residues 62-78 (17 aa).

spacr.gene_facts.available() → Tuple[str, ...][source]

Every annotation column this module can show, in reading order.

Empty when nothing is bundled, which is what unavailable_reason() turns into a sentence.

spacr.gene_facts.clear_cache() → None[source]

Forget the derived indices. For tests, and for a reinstall mid-session.

spacr.annotation.clear_cache() is called too: the layout here is derived from that module’s tables, and dropping one without the other leaves a layout describing columns that are no longer loaded.

spacr.gene_facts.facts(value: Any) → GeneFacts[source]

Everything the bundled annotation holds about the gene value names.

Parameters:

value – a design term, an accession, a guide id or a bare gene number – every spelling spacr.annotation.gene_number() takes.

Returns:

a GeneFacts, ALWAYS. A term that names no gene comes back with GeneFacts.known false and a GeneFacts.reason saying so, because a caller that has to tell None apart from an empty record is a caller that will forget to.

spacr.gene_facts.facts_for(values: Iterable[Any]) → Dict[str, GeneFacts][source]

The facts for several genes at once, keyed by gene number.

Parameters:

values – anything spacr.annotation.gene_number() accepts – design terms, accessions, guide ids, bare numbers, in any mixture.

Returns:

one entry per DISTINCT gene named, in the order first named. A value naming no gene contributes nothing, so an empty result means nothing in values was a gene.

ONE JOIN FOR THE WHOLE SET. The ambiguous case is three genes at once and each of them wants the same 23 columns; three separate merges cost three times as much and, worse, would be three chances for the key to be built differently.

spacr.gene_facts.unavailable_reason() → str[source]

Why there is no annotation on this install, or "".

A sentence rather than a flag, because it is shown to the user: a panel that merely refused would be indistinguishable from one that broke.

spacr.gene_facts.warm(values: Iterable[Any] = ()) → Tuple[str, ...][source]

Load every table this module reads. CALL THIS OFF THE GUI THREAD.

Parameters:

values – the terms a user might click – a whole results table’s feature column is the intended argument. Every gene among them is joined in ONE pass and cached, which is why passing four hundred costs the same 21 ms as passing one.

Returns:

the columns that came out available, so a caller can tell “the tables loaded” from “the tables are not installed” without a second call.

The whole point of the function, measured: cold, the first click pays 360 ms of CSV reading plus a 20 ms join, inside a mouse press. A plot that freezes for a third of a second when clicked reads as broken. Warmed with the screen’s own terms, a click is a dict lookup – 0.02 ms.