spacr.gene_facts¶
What spaCR already knows about one gene, ready to put on a tile.
Clicking a gene in an interactive regression opens the information available
for that gene. spacr.gene_tile answers the half
that can be WRONG – which gene a clicked guide actually names, and whether
that mapping is ambiguous. This module answers the other half: given a gene,
what does spaCR hold about it.
IT HOLDS NOTHING ITSELF. Every value here comes out of spacr.annotation
– the same join that writes the annotated exports – so the tile and the CSV
a reader opens beside it cannot disagree about GRA14’s compartment. This
module reads no file, parses no gene id and owns no table; it groups
spacr.annotation.annotate()’s 23 columns into the order a human reads
them and adds the per-segment DeepTMHMM coordinates that annotate
deliberately leaves out of a coefficient export.
Four rules, each of them a failure this project has already had:
ONE PARSE.
TGGT1_224750,gene_fraction:gene[224750]and the guide224750_2are all gene224750, and the function that says so isspacr.annotation.gene_number(). A second copy of that rule is how the volcano and the hit list start naming different genes.A GAP IS SAID OUT LOUD. A gene with no annotation row gets
GeneFacts.reason– “no row in the bundled Toxoplasma annotation” – and NOT a block of empty fields, which reads as “measured, found nothing” rather than “not available here”.GeneFacts.knownis the flag a caller greys a control on.A GROUP IS DERIVED FROM THE TABLE, NOT LISTED HERE. The fitness and expression rows are every
fit_/expr_columnspacr.annotation.columns()reports, so an eighth published screen added tophenotype.csvappears on the tile without this file being touched, and a column this module has never heard of lands in “other annotation” instead of being silently dropped.NOTHING HERE IS FOR THE GUI THREAD TO DISCOVER. The first call reads five bundled CSVs (360 ms) and the first
GeneFacts.segmentsreads DeepTMHMM’s 8,140 rows.warm()does both off the GUI thread, and takes the screen’s own terms while it is there:annotatere-checks all five right-hand keys for uniqueness on every call, so one gene costs 20 ms and four hundred cost 21. A warmed click is a dict lookup, 0.02 ms.
Public API:
from spacr import gene_facts
known = gene_facts.facts("fraction:grna[239740_3]")
known.known # True
known.value("gene_name") # 'GRA14'
known.sections() # (('identity', (('gene name', 'GRA14'), ...
Classes¶
Functions¶
|
Every annotation column this module can show, in reading order. |
|
Forget the derived indices. For tests, and for a reinstall mid-session. |
|
Everything the bundled annotation holds about the gene |
|
The facts for several genes at once, keyed by gene number. |
|
Why there is no annotation on this install, or |
|
Load every table this module reads. CALL THIS OFF THE GUI THREAD. |
Module Contents¶
- class spacr.gene_facts.GeneFacts[source]¶
Everything the bundled annotation holds about one gene.
- Parameters:
gene – the bare gene number, or
""when the caller’s value named no gene at all.values – the annotation columns that had something in them, keyed by the column name
spacr.annotation.columns()uses. A column with a gap is ABSENT here rather than present and empty.segments – the DeepTMHMM signal peptide and transmembrane segments, in residue order. Empty for a soluble protein, and empty for a gene DeepTMHMM never saw –
reasonis what tells those apart.reason – why there is nothing, in a sentence, or
""when there is something. Never blank and empty at the same time: an empty panel reads as a bug, “no row in the bundled annotation” reads as an answer.
- sections() Tuple[Tuple[str, Tuple[Tuple[str, str], ...]], ...][source]¶
The facts as
((heading, ((label, value), ...)), ...).THE ONE ORDERING, so the Qt tile and the text form cannot drift into presenting the same record two ways – the same shape
spacr.gene_tile.GeneTile.sections()returns, so one renderer lays both halves of a tile out with one loop.A group with nothing in it is not emitted. That is rule 2: a heading over a blank space is a claim that something was looked for and not found, which is a different statement from “not bundled here”.
- to_html() str[source]¶
The facts as HTML, for a Qt rich-text view.
Matches
spacr.gene_tile.GeneTile.to_html()mark for mark, so the two halves of one tile do not read as two different documents.
- class spacr.gene_facts.Segment[source]¶
One DeepTMHMM segment of a protein, with its residue coordinates.
- Parameters:
kind –
signal peptideortransmembrane.index – 1-based position among the segments of that kind;
1for a signal peptide, of which there is at most one.start – first residue, 1-based and inclusive, as DeepTMHMM reports.
end – last residue, inclusive.
length – residues spanned, as the source file records it rather than recomputed – a disagreement between the two is a fact about the run and is not this module’s to hide.
- spacr.gene_facts.available() Tuple[str, ...][source]¶
Every annotation column this module can show, in reading order.
Empty when nothing is bundled, which is what
unavailable_reason()turns into a sentence.
- spacr.gene_facts.clear_cache() None[source]¶
Forget the derived indices. For tests, and for a reinstall mid-session.
spacr.annotation.clear_cache()is called too: the layout here is derived from that module’s tables, and dropping one without the other leaves a layout describing columns that are no longer loaded.
- spacr.gene_facts.facts(value: Any) GeneFacts[source]¶
Everything the bundled annotation holds about the gene
valuenames.- Parameters:
value – a design term, an accession, a guide id or a bare gene number – every spelling
spacr.annotation.gene_number()takes.- Returns:
a
GeneFacts, ALWAYS. A term that names no gene comes back withGeneFacts.knownfalse and aGeneFacts.reasonsaying so, because a caller that has to tellNoneapart from an empty record is a caller that will forget to.
- spacr.gene_facts.facts_for(values: Iterable[Any]) Dict[str, GeneFacts][source]¶
The facts for several genes at once, keyed by gene number.
- Parameters:
values – anything
spacr.annotation.gene_number()accepts – design terms, accessions, guide ids, bare numbers, in any mixture.- Returns:
one entry per DISTINCT gene named, in the order first named. A value naming no gene contributes nothing, so an empty result means nothing in
valueswas a gene.
ONE JOIN FOR THE WHOLE SET. The ambiguous case is three genes at once and each of them wants the same 23 columns; three separate merges cost three times as much and, worse, would be three chances for the key to be built differently.
Why there is no annotation on this install, or
"".A sentence rather than a flag, because it is shown to the user: a panel that merely refused would be indistinguishable from one that broke.
- spacr.gene_facts.warm(values: Iterable[Any] = ()) Tuple[str, ...][source]¶
Load every table this module reads. CALL THIS OFF THE GUI THREAD.
- Parameters:
values – the terms a user might click – a whole results table’s
featurecolumn is the intended argument. Every gene among them is joined in ONE pass and cached, which is why passing four hundred costs the same 21 ms as passing one.- Returns:
the columns that came out available, so a caller can tell “the tables loaded” from “the tables are not installed” without a second call.
The whole point of the function, measured: cold, the first click pays 360 ms of CSV reading plus a 20 ms join, inside a mouse press. A plot that freezes for a third of a second when clicked reads as broken. Warmed with the screen’s own terms, a click is a dict lookup – 0.02 ms.