Strategy cards
Generated by scripts/build_strategy_examples.py --docs-only; do not edit by hand.
Each strategy shows the same four bars, in the same places: Better than chance (skill: 0 is the same procedure on shuffled data, 1 is perfect), Reach (how much of the question it can speak to), and the two metrics of its task a biologist asks about first, in plain words. Values are the mean over the calibration's held-out tests at default settings, with the 95% interval and the chance level.
| Task | Reach means | Bar 3 | Bar 4 |
|---|---|---|---|
| label calls | coverage | Right calls (accuracy) | Fair across classes (macro F1) |
| ranking | recall @ top 10% -- A ranking scores every candidate, so coverage is always complete; its analogue is how many of the true ones a short list reaches. | True ones ranked first (AUROC) | Clean top of the list (AUPRC lift) |
| set retrieval | genes returned -- A set strategy speaks about the genes it returns, so its reach is how many it returned. | Returned genes that are real (precision) | Members found (recall) |
| cluster recovery | 1 - unclustered share -- Clustering has no coverage; its analogue is the share of genes it placed in a cluster at all. | Label falls out as a cluster (weighted F1) | Partition agreement (adjusted Rand index) |
| values | coverage | Order predicted (Spearman rho) | Variance explained (R-squared, out of sample) |
| replication | findings made -- A replication test has no coverage; its reach is how many findings it made to check. | Findings that hold (replication rate) | Beyond chance (replication lift) |
Below the bars, About this test says what the strategy does, how it is evaluated, what failure looks like and what success looks like. The worked examples are real calibration runs: the failure is the target it failed on most often at default settings (or, where it never failed on real data, its self-test on a table of random labels and edges), the success its best pass, re-run once to list what it says about genes without a known label.
01 · Hold out a category and search for a map that finds it
UMAP + HDBSCAN · cluster recovery · T. gondii weak · P. falciparum weak
Is there a combination of measurements and map settings under which a label nobody showed the map falls out as clusters -- and which unlabelled genes land in them?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
█░░░░░░░░░░░ 0.07 [0.03, 0.11] · chance 0.00 |
█░░░░░░░░░░░ 0.09 [0.01, 0.19] · chance 0.00 |
| Reach 1 - unclustered share |
███████░░░░░ 0.55 [0.30, 0.81] |
█████████░░░ 0.78 [0.46, 0.99] |
| Label falls out as a cluster weighted F1 |
█████░░░░░░░ 0.38 [0.19, 0.57] · chance 0.32 |
██████░░░░░░ 0.54 [0.23, 0.81] · chance 0.50 |
| Partition agreement adjusted Rand index |
░░░░░░░░░░░░ 0.03 [0.01, 0.06] · chance 0.00 |
█░░░░░░░░░░░ 0.08 [0.01, 0.18] · chance 0.00 |
About this test
- What it does. It hides one label, such as compartment, plus anything that restates it. It then builds many maps of the remaining measurements with different settings, clusters each, and keeps the map whose clusters best isolate that label. Unlabeled genes in each label's best cluster become candidates.
- How it is evaluated. A quarter of the label is hidden, whole gene families at a time. The map and each label's cluster are picked on visible genes, then scored by F1 (a 0-to-1 match score) on the hidden genes. The same clusters are rescored 100 times with hidden labels shuffled; F1 must beat that by 0.05.
- What failure looks like, and why. The hidden genes' F1 sits at the shuffled level: the chosen clusters hold the label no better than chance. Likely reasons are that the label is not written into these measurements, or that the best map was simply the luckiest of many and does not carry over to new genes.
- What success looks like, and why. Hidden genes land in the cluster chosen for their label clearly more often than shuffled labels would. The label is then a real feature of expression, fitness or modification data, and the unlabeled genes in that cluster are leads, each tagged with its cluster's F1.
T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: F1 of hidden genes in the cluster chosen for their label on known genes was 0.684 against 0.683 on shuffled data (skill 0.00).
Works when compartment: F1 of hidden genes in the cluster chosen for their label on known genes reached 0.12 against 0.04 on shuffled data (skill 0.08, 577 scored). Run once on every gene, it called 63 genes with no known compartment; the top 5 are listed.
TGME49_253870hypothetical protein -- nucleus - chromatin (support 0.388)TGME49_270720hypothetical protein -- nucleus - chromatin (support 0.388)TGME49_242055DEAD/DEAH box helicase domain-containing protein -- nucleus - chromatin (support 0.388)TGME49_242415histone lysine-specific demethylase -- nucleus - chromatin (support 0.388)TGME49_204100eIF2 kinase IF2K-C -- nucleus - chromatin (support 0.388)
P. falciparum. Fails when pbtransferredphenotype: 'pbtransferredphenotype' is not encoded in what this strategy reads: F1 of hidden genes in the cluster chosen for their label on known genes was 0.535 against 0.535 on shuffled data (skill 0.00).
Works when lopitpflocation: F1 of hidden genes in the cluster chosen for their label on known genes reached 0.16 against 0.04 on shuffled data (skill 0.13, 395 scored). Run once on every gene, it called 152 genes with no known lopitpflocation; the top 5 are listed.
PF3D7_021920040S ribosomal protein S30 -- ribosomes (support 0.844)PF3D7_071960060S ribosomal protein L11a, putative -- ribosomes (support 0.844)PF3D7_0102200ring-infected erythrocyte surface antigen -- Maurer's cleft (support 0.63)PF3D7_0201800knob associated heat shock protein 40 -- Maurer's cleft (support 0.63)PF3D7_0202100liver stage associated protein 2 -- Maurer's cleft (support 0.63)
02 · Find the map where your gene list is one cluster
UMAP + HDBSCAN · set retrieval · T. gondii weak · P. falciparum reliable
Under some combination of measurements and settings, do the genes on my list fall into a single cluster -- and what else is in it?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
░░░░░░░░░░░░ 0.03 [0.00, 0.04] · chance 0.00 |
███░░░░░░░░░ 0.21 [0.09, 0.35] · chance 0.00 |
| Reach genes returned |
███████████░ 433 [166, 802] |
████████░░░░ 102 [58, 133] |
| Returned genes that are real precision |
░░░░░░░░░░░░ 0.04 [0.03, 0.05] · chance 0.02 |
██░░░░░░░░░░ 0.16 [0.08, 0.25] · chance 0.01 |
| Members found recall |
███░░░░░░░░░ 0.22 [0.12, 0.34] · chance 0.09 |
█████░░░░░░░ 0.42 [0.17, 0.69] · chance 0.04 |
About this test
- What it does. You paste a gene list, such as screen hits or a complex. It walks many maps and clusterings and finds the single cluster that best captures your list, balancing how much of the cluster is on the list against how much of the list is in it. The cluster's other members are candidates.
- How it is evaluated. With 20 or more genes, 30% of your list is hidden and the best cluster is chosen on the rest. The score is F1, a 0-to-1 match between the hidden genes and that cluster's other members. Twenty random lists of the same size run the same walk; yours must beat their 95th percentile by 0.05.
- What failure looks like, and why. The hidden genes score no better than random lists, which reach about 0.04. Your list is then not a unit this data can see: its genes do not behave alike in these measurements. Across many maps, something always collects part of any list, so a good-looking cluster alone proves nothing.
- What success looks like, and why. The hidden members fall back into the cluster chosen from the rest, far above random lists. Your list is then a coherent biological unit in this data. The genes that cluster with it, nearest the cluster center first, are a ranked shortlist of new members to test.
T. gondii. Fails when tachyzoite (stageenrichedderived): 'tachyzoite (stageenrichedderived)' is not encoded in what this strategy reads: F1 of the hidden members against the best cluster's other genes was 0.038 against 0.047 on shuffled data (skill -0.01). Also, the set was too wide: 890 genes returned, only 2% of them members.
Works when PM - integral (compartment): F1 of the hidden members against the best cluster's other genes reached 0.08 against 0.02 on shuffled data (skill 0.06, 40 scored). Run once on every gene, it placed 85 genes with no place on the list in the cluster that holds the list; the top 5 are listed.
TGME49_271610pyrroline-5-carboxylate reductase -- in the list's cluster (distance to centre 0.771)TGME49_314000peptide methionine sulfoxide reductase msrB, putative -- in the list's cluster (distance to centre 1.08)TGME49_249200Ctr copper transporter family protein -- in the list's cluster (distance to centre 1.13)TGME49_207930phosphatidylethanolamine-binding protein -- in the list's cluster (distance to centre 1.3)TGME49_232630hypothetical protein -- in the list's cluster (distance to centre 1.34)
P. falciparum. Fails when gametocyte (stageenrichedderived): The signal is real but small: 0.040 beat shuffled data (0.013), but by 0.027, short of the 0.050 margin a PASS requires. Also, the set was too wide: 30 genes returned, only 3% of them members; only 20 hidden items could be scored.
Works when cytosol (lopitpflocation): F1 of the hidden members against the best cluster's other genes reached 0.25 against 0.01 on shuffled data (skill 0.25, 37 scored). Run once on every gene, it placed 100 genes with no place on the list in the cluster that holds the list; the top 5 are listed.
PF3D7_1346100protein transport protein SEC61 subunit alpha -- in the list's cluster (distance to centre 1.08)PF3D7_0621200pyridoxine biosynthesis protein PDX1 -- in the list's cluster (distance to centre 1.43)PF3D7_0317000proteasome subunit alpha type-3, putative -- in the list's cluster (distance to centre 1.67)PF3D7_1342400casein kinase II beta chain -- in the list's cluster (distance to centre 1.82)PF3D7_1468700eukaryotic initiation factor 4A -- in the list's cluster (distance to centre 2.16)
03 · Ask which categories the data can rediscover
UMAP + neighbour AUROC · ranking · T. gondii reliable · P. falciparum reliable
Of all the categories of a label, which ones do the measurements actually encode -- and which would no map, however tuned, ever find?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
██████░░░░░░ 0.50 [0.22, 0.68] · chance 0.00 |
████████░░░░ 0.69 [0.56, 0.79] · chance 0.00 |
| Reach recall @ top 10% |
████░░░░░░░░ 0.35 [0.17, 0.54] · chance 0.10 |
█████░░░░░░░ 0.39 [0.18, 0.60] · chance 0.10 |
| True ones ranked first AUROC |
█████████░░░ 0.75 [0.61, 0.84] · chance 0.50 |
██████████░░ 0.84 [0.78, 0.90] · chance 0.50 |
| Clean top of the list AUPRC lift |
█████████░░░ x9.6 [x1.6, x19] · chance x1.0 |
███████░░░░░ x5.2 [x1.5, x9.5] · chance x1.0 |
About this test
- What it does. It builds one map with a label hidden, then asks for each category, such as each compartment, whether a labeled gene's map neighbors share its category. The result is an atlas of which categories the measurements encode and which they do not.
- How it is evaluated. 30% of the label is hidden and the atlas is built from visible genes only. Hidden genes are then scored by their visible neighbors. The number is the mean AUROC (0.5 is random, 1 is perfect) for the atlas's top-half categories, versus 20 shuffled-label runs, needing a 0.05 margin.
- What failure looks like, and why. The top-ranked categories score no better on hidden genes than with shuffled labels. Then the atlas does not generalize: the categories it called encoded were a fluke of the visible genes, or the label lives only in the experiment that defined it and leaves no trace elsewhere.
- What success looks like, and why. The categories the atlas ranks highest are also the ones best recovered for hidden genes. You learn which distinctions leave a real trace in expression, fitness, modification and structure, so you can take those to prediction strategies and skip searches that cannot work.
T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: hidden-gene AUROC of the categories the atlas ranks in its top half was 0.463 against 0.495 on shuffled data (skill -0.06).
Works when compartment: Hidden-gene AUROC of the categories the atlas ranks in its top half reached 0.87 against 0.50 on shuffled data (skill 0.73, 507 scored). Run once on every gene, it ranked 26 categories by how well the data recovers them; the top 5 are listed.
- 60S ribosome is recoverable (AUROC 0.972, lift 37.7)
- 40S ribosome is recoverable (AUROC 0.965, lift 40.8)
- 19S proteasome is recoverable (AUROC 0.944, lift 78.6)
- 20S proteasome is recoverable (AUROC 0.92, lift 61.3)
- PM - peripheral 1 is recoverable (AUROC 0.892, lift 44.8)
P. falciparum. Fails when stageenrichedderived: The signal is too weak to tell from luck: 0.739 did not clear 0.823, what shuffled data reaches one time in twenty (skill 0.45).
Works when lopitpflocation: Hidden-gene AUROC of the categories the atlas ranks in its top half reached 0.85 against 0.51 on shuffled data (skill 0.70, 417 scored). Run once on every gene, it ranked 24 categories by how well the data recovers them; the top 5 are listed.
- erythrocyte membrane is recoverable (AUROC 0.96, lift 19.1)
- 19s proteasome is recoverable (AUROC 0.953, lift 32)
- ribosomes is recoverable (AUROC 0.945, lift 19.8)
- micronemes 1 is recoverable (AUROC 0.905, lift 26.4)
- Maurer's cleft is recoverable (AUROC 0.893, lift 16.6)
04 · Keep only the modules that survive the whole walk
UMAP + HDBSCAN co-clustering · cluster recovery · T. gondii weak · P. falciparum weak
Which groups of genes stay together whatever map settings are chosen -- the structure that is in the data rather than in one lucky configuration?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
█░░░░░░░░░░░ 0.07 [0.01, 0.13] · chance 0.00 |
██░░░░░░░░░░ 0.15 [0.03, 0.28] · chance 0.00 |
| Reach 1 - unclustered share |
███████████░ 0.89 [0.84, 0.93] |
█████████░░░ 0.79 [0.58, 1.00] |
| Label falls out as a cluster weighted F1 |
█████░░░░░░░ 0.39 [0.27, 0.49] · chance 0.34 |
██████░░░░░░ 0.50 [0.23, 0.78] · chance 0.45 |
| Partition agreement adjusted Rand index |
█░░░░░░░░░░░ 0.04 [0.02, 0.07] · chance 0.00 |
█░░░░░░░░░░░ 0.11 [0.01, 0.25] · chance 0.00 |
About this test
- What it does. It builds and clusters many maps with different settings, then counts how often each pair of genes lands in the same cluster. Groups that stay together in at least the chosen share of maps become modules. No label is used to build them, so they reflect the data, not one lucky setting.
- How it is evaluated. Modules are built from a small walk on 1,500 genes without the label. A quarter of the label is hidden, and each label's best module is picked on visible genes. The number is the F1 match of hidden genes to that module, versus 100 label shuffles; it must beat them by 0.05.
- What failure looks like, and why. Hidden genes fit their chosen module no better than shuffled labels do. A very large module scores the same either way, so it earns nothing. Likely causes are modules too coarse to separate categories, or a label not reflected in structure that survives across maps.
- What success looks like, and why. Hidden genes sit in the module chosen for their label well above chance. That means stable modules track real biology. A stable module with no dominant known label is a candidate new unit, such as an unrecognized complex or pathway, worth following up.
T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: F1 of hidden genes in the module chosen for their label on known genes was 0.494 against 0.509 on shuffled data (skill -0.03).
Works when compartment: F1 of hidden genes in the module chosen for their label on known genes reached 0.22 against 0.12 on shuffled data (skill 0.11, 302 scored). Run once on every gene, it placed 487 genes with no known compartment in stable modules; the top 5 are listed.
TGME49_234990hypothetical protein -- module 444 (nucleus - chromatin) (0.462 of its labeled genes, stability 0.536)TGME49_293790hypothetical protein -- module 553 (nucleus - chromatin) (0.367 of its labeled genes, stability 0.515)TGME49_500319hypothetical protein, conserved -- module 553 (nucleus - chromatin) (0.367 of its labeled genes, stability 0.515)TGME49_226705hypothetical protein -- module 553 (nucleus - chromatin) (0.367 of its labeled genes, stability 0.515)TGME49_295105rhoptry protein, putative -- module 553 (nucleus - chromatin) (0.367 of its labeled genes, stability 0.515)
P. falciparum. Fails when stageenrichedderived: 'stageenrichedderived' is not encoded in what this strategy reads: F1 of hidden genes in the module chosen for their label on known genes was 0.607 against 0.607 on shuffled data (skill 0.00).
Works when lopitpflocation: F1 of hidden genes in the module chosen for their label on known genes reached 0.28 against 0.12 on shuffled data (skill 0.19, 255 scored). Run once on every gene, it placed 852 genes with no known lopitpflocation in stable modules; the top 5 are listed.
PF3D7_0100200rifin -- module 4 (erythrocyte membrane) (0.857 of its labeled genes, stability 0.979)PF3D7_0100400rifin -- module 4 (erythrocyte membrane) (0.857 of its labeled genes, stability 0.979)PF3D7_0101000rifin -- module 4 (erythrocyte membrane) (0.857 of its labeled genes, stability 0.979)PF3D7_0101800stevor -- module 4 (erythrocyte membrane) (0.857 of its labeled genes, stability 0.979)PF3D7_0115400stevor -- module 4 (erythrocyte membrane) (0.857 of its labeled genes, stability 0.979)
05 · Tune a map without labels, then read what it encodes
UMAP + HDBSCAN, chi-square / Kruskal-Wallis · replication · T. gondii reliable · P. falciparum reliable
If I build a map from one kind of evidence only -- expression, say -- and tune it for structure alone, which OTHER measurements do its clusters turn out to separate?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
███████████░ 0.95 [0.88, 1.00] · chance 0.00 |
███████████░ 0.88 [0.82, 0.93] · chance 0.00 |
| Reach findings made |
███████░░░░░ 53 [39, 63] |
███████░░░░░ 54 [52, 56] |
| Findings that hold replication rate |
███████████░ 0.95 [0.89, 1.00] · chance 0.06 |
███████████░ 0.88 [0.83, 0.94] · chance 0.05 |
| Beyond chance replication lift |
███████████░ x21 [x13, x28] · chance x1.0 |
██████████░░ x18 [x14, x24] · chance x1.0 |
About this test
- What it does. It builds a map from one kind of evidence only, such as transcription, tuned for clean clusters with no label in view. Then it tests every measurement the map never saw against those clusters, asking which other biology, such as fitness or compartment, the clusters separate.
- How it is evaluated. Associations are found on a random half of up to 2,000 genes and rechecked on the other half. The number is the share of findings that replicate. It is compared with 10 runs where the second half's clusters are shuffled; it must beat that by 0.2, with at least three findings.
- What failure looks like, and why. Few findings replicate, no more than with shuffled clusters, or fewer than three are found. The significant results were then large-sample artifacts: with many genes, tiny differences look significant. Or the chosen evidence simply does not couple to the other measurements.
- What success looks like, and why. Most associations found on one half hold on the other, evidence that two kinds of biology are coupled. For example, co-expressed clusters that differ in fitness suggest transcription programs track essentiality. Another family may encode something different.
T. gondii. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find. It made too little to score, and the test declined to judge (only 0 findings on the first half, 3 needed).
Works when transcription: Share of first-half findings that replicate on the second half reached 1.00 against 0.04 on shuffled data (skill 1.00, 53 scored). Run once on every gene, it tested 131 held-out features the map was never shown; the top 5 are listed.
- map separates rpf245775eif12kotachy_r1 (effect 0.864, q 2.59e-189)
- map separates rpf129869confluentr1 (effect 0.864, q 2.59e-189)
- map separates rpf129869subconfluentr2 (effect 0.863, q 2.59e-189)
- map separates rpf245775eif12koprebrady_r1 (effect 0.86, q 1.00e-188)
- map separates rpf129869subconfluentr3 (effect 0.857, q 3.09e-188)
P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find. It made too little to score, and the test declined to judge (only 0 findings on the first half, 3 needed).
Works when transcription: Share of first-half findings that replicate on the second half reached 0.98 against 0.08 on shuffled data (skill 0.98, 54 scored). Run once on every gene, it tested 128 held-out features the map was never shown; the top 5 are listed.
- map separates exprlatetrophozoite (effect 0.105, q 4.30e-67)
- map separates expr_schizont (effect 0.0975, q 8.84e-63)
- map separates febrilelrr5ko37c (effect 0.0804, q 7.16e-52)
- map separates riboseqrpflate_trophozoite (effect 0.14, q 6.70e-48)
- map separates gene_type (effect 0.197, q 1.43e-47)
06 · Find which kind of evidence carries a label
kNN ablation · label calls · T. gondii reliable · P. falciparum reliable
Which measurements actually carry the information about this label -- and which are redundant with others or irrelevant to it?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
██░░░░░░░░░░ 0.14 [0.06, 0.25] · chance 0.00 |
██░░░░░░░░░░ 0.13 [0.07, 0.19] · chance 0.00 |
| Reach coverage |
████████████ 1.00 [1.00, 1.00] |
████████████ 1.00 [1.00, 1.00] |
| Right calls accuracy |
███████░░░░░ 0.60 [0.44, 0.77] · chance 0.54 |
████████░░░░ 0.64 [0.44, 0.85] · chance 0.57 |
| Fair across classes macro F1 |
█████░░░░░░░ 0.39 [0.28, 0.54] |
█████░░░░░░░ 0.42 [0.31, 0.49] |
About this test
- What it does. For a label, it asks which kind of evidence carries it. Each family of measurements, or each single dataset, predicts the label alone by a 15-nearest-neighbor vote. It also measures how much the full combination loses when that evidence is removed.
- How it is evaluated. A quarter of the label is hidden. Each kind of evidence is ranked by accuracy on visible labels, and the top one predicts the hidden labels. That hidden accuracy is compared with the hidden accuracy of every kind of evidence, as if picked at random; it must top their 80th percentile.
- What failure looks like, and why. The top-ranked evidence does no better on hidden genes than a randomly chosen kind. The ranking is then not real: several kinds of evidence may carry the label about equally, or none of them carries it well, so the order on visible genes was noise.
- What success looks like, and why. The evidence ranked first also predicts hidden labels best. You learn which experiments encode this biology, which are redundant (strong alone, no loss when removed) and which are unique (costly to remove). That guides a focused map and where the next experiment should go.
T. gondii. Fails when dtm_class: The signal is real but small: 0.754 beat shuffled data (0.744), but by 0.010, short of the 0.020 margin a PASS requires.
Works when compartment: Hidden accuracy of the evidence ranked first (transcription) reached 0.34 against 0.23 on shuffled data (skill 0.14, 951 scored). Run once on every gene, it ranked 14 kinds of evidence by what each carries alone; the top 5 are listed.
- transcription carries the label (alone 0.321, loss when removed 0.0131)
- sequence carries the label (alone 0.281, loss when removed 0.00657)
- translation carries the label (alone 0.28, loss when removed 0.00421)
- localization carries the label (alone 0.278, loss when removed 0.00815)
- protein abundance carries the label (alone 0.257, loss when removed 0.00473)
P. falciparum. Fails when isexported: The signal is too weak to tell from luck: 0.970 did not clear 0.970, what shuffled data reaches one time in twenty (skill 0.04).
Works when lopitpf_location: Hidden accuracy of the evidence ranked first (expr) reached 0.38 against 0.21 on shuffled data (skill 0.22, 395 scored). Run once on every gene, it ranked 48 kinds of evidence by what each carries alone; the top 5 are listed.
- expr carries the label (alone 0.363, loss when removed -0.00443)
- riboseq carries the label (alone 0.337, loss when removed 0.0101)
- n carries the label (alone 0.317, loss when removed 0.0177)
- febrile carries the label (alone 0.315, loss when removed 0.00443)
- steady carries the label (alone 0.309, loss when removed -0.00316)
07 · Call a gene by the genes that behave like it
kNN · label calls · T. gondii reliable · P. falciparum weak
For a gene with no label, what label do the genes most similar to it across every permitted measurement carry?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
███░░░░░░░░░ 0.22 [0.08, 0.36] · chance 0.00 |
█░░░░░░░░░░░ 0.08 [-0.37, 0.42] · chance 0.00 |
| Reach coverage |
███████████░ 0.93 [0.84, 1.00] |
███████████░ 0.96 [0.87, 1.00] |
| Right calls accuracy |
███████░░░░░ 0.62 [0.47, 0.76] · chance 0.48 |
████████░░░░ 0.68 [0.53, 0.85] · chance 0.53 |
| Fair across classes macro F1 |
█████░░░░░░░ 0.44 [0.34, 0.57] |
██████░░░░░░ 0.51 [0.47, 0.55] |
About this test
- What it does. For each unlabeled gene, it finds the most similar labeled genes across all permitted measurements at once. Those neighbors vote, closer ones counting more. The gene is called with the winning label if it holds enough of the vote. No map or clustering sits in between.
- How it is evaluated. A quarter of the label is hidden, whole gene families at a time. Each hidden gene is called by its nearest visible genes. The number is the share of hidden genes called correctly, with abstentions as misses. It must beat 10 shuffled-label runs by 0.05.
- What failure looks like, and why. The share called correctly sits at the shuffled level. With hundreds of measurements, many missing and filled with the median, the nearest genes may share gaps in the data rather than biology. Or the label is not encoded in these measurements at all.
- What success looks like, and why. Hidden genes are called correctly well above chance. This sets the baseline other strategies must beat. Each call on an unlabeled gene is transparent: you can trace it to the neighbors that made it, and a stricter vote threshold gives fewer, safer calls.
T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: correct calls per hidden gene was 0.675 against 0.680 on shuffled data (skill -0.02).
Works when compartment: Correct calls per hidden gene reached 0.37 against 0.05 on shuffled data (skill 0.34, 951 scored). Run once on every gene, it called 2,000 genes with no known compartment; the top 5 are listed.
TGME49_310470cytochrome c oxidase subunit COX2b -- mitochondrion - membranes (support 0.937)TGME49_312940hypothetical protein -- mitochondrion - membranes (support 0.877)TGME49_202500glideosome-associated protein with multiple-membrane spans GAPM1A -- IMC (support 0.873)TGME49_225960STE kinase -- nucleus - chromatin (support 0.871)TGME49_295610histone lysine methyltransferase, SET, putative -- nucleus - chromatin (support 0.87)
P. falciparum. Fails when isexported: One class dominates 'isexported', so guessing it on shuffled data already scores 0.936; the strategy's 0.937 is no better than that (skill 0.01), so what it reads does not separate the classes.
Works when lopitpflocation: Correct calls per hidden gene reached 0.52 against 0.06 on shuffled data (skill 0.49, 395 scored). Run once on every gene, it called 2,000 genes with no known lopitpflocation; the top 5 are listed.
PF3D7_0606100RNA-binding protein, putative -- nucleus 1 (support 1)PF3D7_071960060S ribosomal protein L11a, putative -- ribosomes (support 1)PF3D7_100350040S ribosomal protein S20e, putative -- ribosomes (support 1)PF3D7_1017600conserved Plasmodium protein, unknown function -- nucleus 1 (support 1)PF3D7_124270040S ribosomal protein S17, putative -- ribosomes (support 1)
08 · Call a gene by its neighbours on the map
UMAP + kNN · label calls · T. gondii weak · P. falciparum weak
On a map built without the label, which label do a gene's nearest placed neighbours carry?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
██░░░░░░░░░░ 0.13 [0.01, 0.24] · chance 0.00 |
░░░░░░░░░░░░ -0.00 [-0.53, 0.29] · chance 0.00 |
| Reach coverage |
███████████░ 0.91 [0.81, 1.00] |
███████████░ 0.94 [0.83, 1.00] |
| Right calls accuracy |
███████░░░░░ 0.56 [0.38, 0.74] · chance 0.47 |
████████░░░░ 0.64 [0.42, 0.84] · chance 0.53 |
| Fair across classes macro F1 |
████░░░░░░░░ 0.37 [0.26, 0.48] |
██████░░░░░░ 0.47 [0.36, 0.57] |
About this test
- What it does. It builds one map with the label withheld. Each unlabeled gene is then called by the distance-weighted vote of its nearest labeled neighbors on the map. It is what you do by eye when you see a grey point inside a colored cloud, made systematic.
- How it is evaluated. On up to 2,500 genes the map places, a quarter of the label is hidden, whole gene families at a time. Hidden genes are called by their nearest visible map neighbors. The number is the share called correctly, compared with 20 shuffled-label runs, needing a 0.05 margin.
- What failure looks like, and why. Correct calls sit at the shuffled level. The map may have torn the category apart when squeezing hundreds of measurements into three dimensions, or the label is not encoded in the data. Genes left out of a sampled map are never called.
- What success looks like, and why. Hidden genes are called correctly well above chance. Comparing with strategy 07 on the same label tells you whether the map cleans up noise or loses signal. The calls label unlabeled genes, and you can check each in context on the colored map.
T. gondii. Fails when dtmclass: 'dtmclass' is not encoded in what this strategy reads: correct calls per hidden gene was 0.751 against 0.753 on shuffled data (skill -0.01).
Works when compartment: Correct calls per hidden gene reached 0.27 against 0.05 on shuffled data (skill 0.23, 513 scored). Run once on every gene, it called 610 genes with no known compartment; the top 5 are listed.
TGME49_286670hypothetical protein -- PM - peripheral 1 (support 1)TGME49_238210EGF family domain-containing protein -- PM - peripheral 1 (support 1)TGME49_311450zinc finger, c2h2 type domain-containing protein -- nucleus - chromatin (support 0.992)TGME49_278370Toxoplasma gondii family A protein -- PM - peripheral 1 (support 0.958)TGME49_274170protein kinase (incomplete catalytic triad) -- PM - peripheral 1 (support 0.958)
P. falciparum. Fails when isexported: One class dominates 'isexported', so guessing it on shuffled data already scores 0.955; the strategy's 0.953 is no better than that (skill -0.04), so what it reads does not separate the classes.
Works when lopitpflocation: Correct calls per hidden gene reached 0.35 against 0.08 on shuffled data (skill 0.30, 395 scored). Run once on every gene, it called 1,470 genes with no known lopitpflocation; the top 5 are listed.
PF3D7_1101300rifin -- erythrocyte membrane (support 0.988)PF3D7_132310060S ribosomal protein L6, putative -- ribosomes (support 0.966)PF3D7_0222100Pfmc-2TM Maurer's cleft two transmembrane protein -- erythrocyte membrane (support 0.963)PF3D7_0110600phosphatidylinositol-4-phosphate 5-kinase -- nucleus 1 (support 0.962)PF3D7_1032100mRNA-decapping enzyme subunit 1, putative -- nucleus 1 (support 0.955)
09 · Name a cluster by the label it is enriched for
UMAP + HDBSCAN, hypergeometric · label calls · T. gondii reliable · P. falciparum reliable
Which clusters of a label-blind map hold one label far more often than chance, and what does that make of their unlabelled members?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
████░░░░░░░░ 0.33 [0.25, 0.42] · chance 0.00 |
█████░░░░░░░ 0.42 [0.33, 0.60] · chance 0.00 |
| Reach coverage |
██░░░░░░░░░░ 0.18 [0.09, 0.27] |
██░░░░░░░░░░ 0.14 [0.07, 0.22] |
| Right calls accuracy |
█░░░░░░░░░░░ 0.06 [0.03, 0.08] |
█░░░░░░░░░░░ 0.06 [0.02, 0.09] |
| Fair across classes macro F1 |
█░░░░░░░░░░░ 0.10 [0.06, 0.15] |
██░░░░░░░░░░ 0.17 [0.11, 0.22] |
About this test
- What it does. It clusters a map built without the label, then tests every cluster against every label for enrichment, asking whether a label appears more often than chance. Unlabeled members of clusters that are both significant and enriched by the chosen lift are called with that label.
- How it is evaluated. A quarter of the label is hidden on up to 2,500 mapped genes, and enrichment is recomputed from visible labels. The number is precision: how often a call on a hidden gene is right, since the method abstains by design. It must beat 20 shuffled-label runs by 0.1.
- What failure looks like, and why. Calls on hidden genes are right no more often than with shuffled labels. An enriched cluster can still have most labeled members elsewhere, so its calls can be mostly wrong. Clusters too small to hold several labeled genes also give weak enrichment.
- What success looks like, and why. Calls on hidden genes are right clearly more often than chance. A cluster can then be named by its enrichment, such as an apicoplast-flavored cluster, even if it is not pure. Its unlabeled members become candidates, each backed by a q-value and a lift.
T. gondii. Fails when lopit_unified: The signal is real but small: 0.050 beat shuffled data (0.000), but by 0.050, short of the 0.100 margin a PASS requires. Also, it could reach only 44% of the hidden genes, so most were never called.
Works when compartment: Precision of calls on hidden genes reached 0.32 against 0.00 on shuffled data (skill 0.32, 107 scored). Run once on every gene, it called 102 genes with no known compartment; the top 5 are listed.
TGME49_278090Toxoplasma gondii family A protein -- PM - peripheral 1 (support 30)TGME49_237800dense granule protein GRA11B -- PM - peripheral 1 (support 30)TGME49_243170Toxoplasma gondii family A protein -- PM - peripheral 1 (support 30)TGME49_219348SAG-related sequence SRS55M -- PM - peripheral 1 (support 30)TGME49_315380SAG-related sequence SRS53B -- PM - peripheral 1 (support 30)
P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find. It made too little to score, and the test declined to judge.
Works when lopitpflocation: Precision of calls on hidden genes reached 0.51 against 0.01 on shuffled data (skill 0.51, 70 scored). Run once on every gene, it called 400 genes with no known lopitpflocation; the top 5 are listed.
PF3D7_0100600rifin -- erythrocyte membrane (support 20.2)PF3D7_0100800rifin -- erythrocyte membrane (support 20.2)PF3D7_0100900rifin -- erythrocyte membrane (support 20.2)PF3D7_0101800stevor -- erythrocyte membrane (support 20.2)PF3D7_0101900rifin -- erythrocyte membrane (support 20.2)
10 · Find genes whose label their neighbours contradict
kNN + network neighbours · ranking · T. gondii reliable · P. falciparum reliable
Which labelled genes sit among genes that almost all carry a different label -- possible mislabels, dual-localized or moonlighting proteins?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
████████░░░░ 0.64 [0.49, 0.77] · chance 0.00 |
█████████░░░ 0.74 [0.57, 0.92] · chance 0.00 |
| Reach recall @ top 10% |
██████░░░░░░ 0.48 [0.35, 0.61] · chance 0.10 |
███████░░░░░ 0.61 [0.37, 0.82] · chance 0.10 |
| True ones ranked first AUROC |
██████████░░ 0.82 [0.74, 0.88] · chance 0.50 |
██████████░░ 0.87 [0.79, 0.94] · chance 0.50 |
| Clean top of the list AUPRC lift |
███████░░░░░ x5.0 [x3.5, x6.7] · chance x1.0 |
████████░░░░ x8.8 [x3.7, x15] · chance x1.0 |
About this test
- What it does. It turns the question around and asks which existing labels look wrong. For each labeled gene it checks how many of its nearest measurement neighbors and network partners share its label. Genes whose company mostly carries another label rank as most surprising.
- How it is evaluated. It deliberately swaps 5% of labels, at least ten, to a wrong class, then computes surprise. The number is the AUROC (0.5 is random, 1 is perfect) for the swapped genes rising to the top. It is compared with 20 random gene sets of the same size and must beat them by 0.1.
- What failure looks like, and why. Swapped genes rank no higher than random genes. Neighbors then do not agree enough on the label to expose a wrong one, because the label is weakly encoded in the measurements or too few genes have labeled network partners.
- What success looks like, and why. Deliberately swapped labels rise to the top of the list. On real data, the top genes are likely annotation errors, proteins with two locations or functions, or genes with unusual measurements. Each comes with the label its neighbors carry instead.
T. gondii. Fails when screenanyphenotype: The signal is too weak to tell from luck: 0.623 did not clear 0.632, what shuffled data reaches one time in twenty (skill 0.19). Also, only 15 hidden items could be scored.
Works when compartment: AUROC of surprise for the swapped labels reached 0.83 against 0.50 on shuffled data (skill 0.65, 190 scored). Run once on every gene, it flagged 200 labeled genes whose label looks wrong; the top 5 are listed.
TGME49_228690phosphatidylinositol 3- and 4-kinase -- labeled apicoplast, looks nucleus - chromatin (surprise 1)TGME49_304720SWIM zinc finger domain-containing protein -- labeled dense granules, looks nucleus - chromatin (surprise 1)TGME49_500327hypothetical protein, conserved -- labeled apical 1, looks nucleus - chromatin (surprise 1)TGME49_254370guanylyl cyclase -- labeled PM - integral, looks nucleus - chromatin (surprise 1)TGME49_321340membrane protein, putative -- labeled ER, looks nucleus - chromatin (surprise 1)
P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: AUROC of surprise for the swapped labels was 0.51 against 0.48 on shuffled data (skill 0.06). The test calls that a FAIL, which is the failure it exists to catch.
Works when lopitpflocation: AUROC of surprise for the swapped labels reached 0.87 against 0.50 on shuffled data (skill 0.75, 79 scored). Run once on every gene, it flagged 200 labeled genes whose label looks wrong; the top 5 are listed.
PF3D7_0105200RAP protein RAP1 -- labeled rhoptries, looks mitochondrion (surprise 1)PF3D7_0106200EF hand domain-containing protein, putative -- labeled rhoptries, looks mitochondrion (surprise 1)PF3D7_0215000acyl-CoA synthetase -- labeled PVM, looks nucleus 1 (surprise 1)PF3D7_0301400Plasmodium exported protein, unknown function -- labeled Maurer's cleft, looks apicoplast (surprise 1)PF3D7_0403600conserved Plasmodium protein, unknown function -- labeled transport vesicles, looks mitotic apparatus (surprise 1)
11 · Diffuse a label across one measured network
random walk with restart · label calls · T. gondii reliable · P. falciparum reliable
If labels flow along the edges of one kind of measured relationship, where do they end up -- and how much of a label does that relationship carry?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
█░░░░░░░░░░░ 0.12 [0.08, 0.17] · chance 0.00 |
███░░░░░░░░░ 0.22 [0.14, 0.33] · chance 0.00 |
| Reach coverage |
██████████░░ 0.80 [0.78, 0.82] |
███████████░ 0.95 [0.95, 0.96] |
| Right calls accuracy |
███░░░░░░░░░ 0.29 [0.18, 0.40] · chance 0.19 |
█████░░░░░░░ 0.45 [0.17, 0.72] · chance 0.32 |
| Fair across classes macro F1 |
███░░░░░░░░░ 0.28 [0.17, 0.41] |
████░░░░░░░░ 0.38 [0.17, 0.53] |
About this test
- What it does. It seeds each label on the genes that carry it and lets it spread along the edges of one measured network, such as crosslinks or co-fitness, by a random walk. Hubs are down-weighted. Each reached gene is called by the label that arrives most strongly.
- How it is evaluated. A quarter of the label is hidden, whole gene families at a time, and only visible genes seed the spread. The number is the share of hidden genes called correctly, with unreached genes as misses. It is compared with 10 shuffled-label runs and must beat them by 0.05.
- What failure looks like, and why. Correct calls sit at the shuffled level. The chosen network does not link genes that share this label, or it reaches too few genes, since unreached genes count as misses. A layer built from the label itself is refused, as it would be circular.
- What success looks like, and why. Hidden genes are called correctly well above chance, so this kind of relationship carries the label. Testing each layer builds a table of which networks encode which biology. The best layer's calls label the unlabeled genes it reaches.
T. gondii. Fails when screenanyphenotype: The signal is too weak to tell from luck: 0.450 did not clear 0.482, what shuffled data reaches one time in twenty (skill 0.06).
Works when compartment: Correct calls per hidden gene reached 0.14 against 0.03 on shuffled data (skill 0.11, 951 scored). Run once on every gene, it called 2,000 genes with no known compartment; the top 5 are listed.
TGME49_265190Ulp1 protease family, C-terminal catalytic domain-containing protein -- Golgi (support 1)TGME49_319610eIF2 kinase IF2K-D (incomplete catalytic triad) -- nucleus - chromatin (support 1)TGME49_284645hypothetical protein -- nucleus - chromatin (support 1)TGME49_290990HEAT repeat-containing protein -- nucleus - non-chromatin (support 1)TGME49_224320hypothetical protein -- mitochondrion - soluble (support 1)
P. falciparum. Fails when isexported: The signal is real but small: 0.650 beat shuffled data (0.605), but by 0.046, short of the 0.050 margin a PASS requires.
Works when lopitpflocation: Correct calls per hidden gene reached 0.18 against 0.05 on shuffled data (skill 0.14, 395 scored). Run once on every gene, it called 2,000 genes with no known lopitpf_location; the top 5 are listed.
PF3D7_021920040S ribosomal protein S30 -- ribosomes (support 1)PF3D7_0305600DNA-(apurinic or apyrimidinic site) endonuclease -- ER 1 (support 1)PF3D7_0416800small GTP-binding protein sar1 -- transport vesicles (support 1)PF3D7_0525000zinc finger protein, putative -- nucleus 1 (support 1)PF3D7_061170060S ribosomal protein L39 -- ribosomes (support 1)
12 · Let every network vote, weighted by what it has earned
chance-weighted ensemble vote · label calls · T. gondii weak · P. falciparum reliable
If every measured relationship and the measurements themselves vote on a gene's label, each weighted by how good it has proven to be, what is the verdict?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
████░░░░░░░░ 0.30 [-0.03, 0.54] · chance 0.00 |
██████░░░░░░ 0.51 [0.41, 0.65] · chance 0.00 |
| Reach coverage |
███████████░ 0.90 [0.69, 1.00] |
███████████░ 0.94 [0.81, 1.00] |
| Right calls accuracy |
███████░░░░░ 0.58 [0.40, 0.75] · chance 0.38 |
████████░░░░ 0.65 [0.53, 0.78] · chance 0.32 |
| Fair across classes macro F1 |
█████░░░░░░░ 0.44 [0.33, 0.57] |
██████░░░░░░ 0.50 [0.44, 0.56] |
About this test
- What it does. Every permitted network and the measurement neighbors vote on each gene's label. Each source's vote is weighted by how far it beat its own chance level on known labels. The verdict reaches more genes than any single network and leans on sources that have earned trust.
- How it is evaluated. A quarter of the label is hidden. Weights are learned on an inner holdout of visible labels only, then hidden genes are called. The number is the share called correctly, compared with 5 shuffled-label runs where weights are relearned; it must beat them by 0.05.
- What failure looks like, and why. Correct calls sit at the shuffled level, and source weights are near zero. No source knew more than chance about this label, so the label is invisible in these networks and measurements, however many edges they have.
- What success looks like, and why. Hidden genes are called correctly well above chance. The source weights table ranks each kind of evidence by how much it knows about this label, net of chance. The calls label unlabeled genes, with support as the winner's share of the weighted vote.
T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: correct calls per hidden gene was 0.463 against 0.528 on shuffled data (skill -0.14).
Works when compartment: Correct calls per hidden gene reached 0.42 against 0.13 on shuffled data (skill 0.34, 951 scored). Run once on every gene, it called 2,000 genes with no known compartment; the top 5 are listed.
TGME49_209000HECT-domain (ubiquitin-transferase) domain-containing protein -- nucleus - chromatin (support 1)TGME49_500189E3 ubiquitin-protein ligase, putative -- apicoplast (support 1)TGME49_266010phosphatidylinositol 3- and 4-kinase -- nucleus - chromatin (support 1)TGME49_500429hypothetical protein, conserved -- apicoplast (support 1)TGME49_500142serine-protein kinase ATM, putative -- apicoplast (support 1)
P. falciparum. Fails when stageenrichedderived: The signal is too weak to tell from luck: 0.691 did not clear 0.706, what shuffled data reaches one time in twenty (skill 0.47).
Works when lopitpflocation: Correct calls per hidden gene reached 0.55 against 0.10 on shuffled data (skill 0.50, 395 scored). Run once on every gene, it called 2,000 genes with no known lopitpflocation; the top 5 are listed.
PF3D7_0100200rifin -- erythrocyte membrane (support 1)PF3D7_0100600rifin -- erythrocyte membrane (support 1)PF3D7_0100700Plasmodium exported protein, unknown function, fragment -- erythrocyte membrane (support 1)PF3D7_0100800rifin -- erythrocyte membrane (support 1)PF3D7_0100900rifin -- erythrocyte membrane (support 1)
13 · Place a protein by the proteins it physically touches
weighted partner vote · label calls · T. gondii reliable · P. falciparum weak
For a protein crosslinked to or pulled down with labelled proteins, what does its physical company say about where it lives and what it joins?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
███░░░░░░░░░ 0.25 [0.08, 0.43] · chance 0.00 |
███░░░░░░░░░ 0.22 [0.01, 0.47] · chance 0.00 |
| Reach coverage |
███████░░░░░ 0.57 [0.28, 0.84] |
█████████░░░ 0.71 [0.61, 0.80] |
| Right calls accuracy |
█████░░░░░░░ 0.39 [0.17, 0.61] · chance 0.19 |
██████░░░░░░ 0.54 [0.33, 0.75] · chance 0.37 |
| Fair across classes macro F1 |
████░░░░░░░░ 0.34 [0.17, 0.52] |
█████░░░░░░░ 0.45 [0.35, 0.56] |
About this test
- What it does. It calls each protein by the labels of its measured physical partners, from crosslinks and replicated pulldowns, weighted by how often each contact was seen. Crosslinked proteins were close together in the parasite, so they share a compartment. Every call lists its partners.
- How it is evaluated. Only genes with at least one physical partner are scored. A quarter of the label is hidden, whole gene families at a time, and hidden genes are called by their visible partners' weighted vote. The share called correctly must beat 20 shuffled-label runs by 0.05.
- What failure looks like, and why. Correct calls sit at the shuffled level. Few hidden genes may have labeled partners, since the interactomes cover a few thousand proteins at most. The crosslinker's chemistry also favors lysines and soluble proteins, so some proteins are poorly sampled.
- What success looks like, and why. Hidden proteins are placed correctly well above chance. Within its reach, this is expected to be the most precise localization evidence available. Each call names the partners, labels and contact counts behind it, so you can check it by hand.
T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: correct calls per hidden gene was 0.062 against 0.047 on shuffled data (skill 0.02). Also, it could reach only 6% of the hidden genes, so most were never called; only 16 hidden items could be scored.
Works when compartment: Correct calls per hidden gene reached 0.49 against 0.09 on shuffled data (skill 0.45, 330 scored). Run once on every gene, it called 271 genes with no known compartment; the top 5 are listed.
TGME49_294820type I fatty acid synthase, putative -- ER 2 (support 1)TGME49_242625ATPase family associated with various cellular activities (AAA) subfamily protein -- dense granules (support 1)TGME49_286270hypothetical protein -- cytosol (support 1)TGME49_500222GCC2 and GCC3 domain-containing protein -- cytosol (support 1)TGME49_245440hypothetical protein -- cytosol (support 1)
P. falciparum. Fails when pbtransferredphenotype: The signal is too weak to tell from luck: 0.350 did not clear 0.503, what shuffled data reaches one time in twenty (skill 0.08). Also, only 20 hidden items could be scored.
Works when lopitpflocation: Correct calls per hidden gene reached 0.56 against 0.06 on shuffled data (skill 0.53, 18 scored). Run once on every gene, it called 15 genes with no known lopitpflocation; the top 5 are listed.
PF3D7_0108000proteasome subunit beta type-3, putative -- 19s proteasome (support 1)PF3D7_0306600ATP synthase-associated protein, putative -- mitochondrion (support 1)PF3D7_0629200DnaJ protein, putative -- ER 1 (support 1)PF3D7_0718000dynein heavy chain, putative -- PVM (support 1)PF3D7_0722700respiratory chain complex 3 associated protein 1 -- mitochondrion (support 1)
14 · Annotate function through shared fold
TM-score-weighted vote · label calls · T. gondii reliable · P. falciparum reliable
What does a protein's fold -- its structural similarity to annotated proteins -- say about its enzymatic class or domain family, even without sequence homology?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
████████░░░░ 0.64 [0.60, 0.66] · chance 0.00 |
███████░░░░░ 0.59 [0.54, 0.61] · chance 0.00 |
| Reach coverage |
█████████░░░ 0.78 [0.75, 0.80] |
█████████░░░ 0.75 [0.71, 0.77] |
| Right calls accuracy |
█████████░░░ 0.73 [0.70, 0.76] · chance 0.27 |
████████░░░░ 0.69 [0.66, 0.72] · chance 0.26 |
| Fair across classes macro F1 |
█████████░░░ 0.75 [0.75, 0.76] |
███████░░░░░ 0.60 [0.54, 0.67] |
About this test
- What it does. It asks what a protein's 3D fold says about its enzyme class, even when its sequence matches nothing annotated. Predicted structures are compared, and each unannotated protein takes a vote of its annotated look-alikes, weighted by how similar the folds are (TM-score).
- How it is evaluated. Among proteins with a structural neighbor, a quarter of the annotations are hidden, whole gene families at a time, and called back from their neighbors. The same test is run 20 times on shuffled annotations. The verdict rests on the share of hidden proteins called correctly.
- What failure looks like, and why. The correct-call rate sits at the shuffled level. Likely reasons: the chosen depth is too fine, since a fold rarely knows the exact substrate at EC levels 3 or 4, or too few proteins have structural neighbors, or the layer chosen is not the structural one.
- What success looks like, and why. The correct-call rate clears the shuffled runs, which says fold carries this annotation at this depth. You get calls for hypothetical proteins, listed first, with the neighbors' models to check. Testing a deeper level shows how far that trust extends.
T. gondii. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: correct calls per hidden gene was 0.21 against 0.16 on shuffled data (skill 0.07). It still passed: a false alarm on random data, so read its passes with care.
Works when --: Correct calls per hidden gene reached 0.75 against 0.26 on shuffled data (skill 0.67, 165 scored). Run once on every gene, it called 337 genes with no known ec_number; the top 5 are listed.
TGME49_224890hypothetical protein -- 6 (support 1)TGME49_234220hypothetical protein -- 3 (support 1)TGME49_246010hypothetical protein -- 3 (support 1)TGME49_312260hypothetical protein -- 2 (support 1)TGME49_206300hypothetical protein -- 2 (support 1)
P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: correct calls per hidden gene was 0.21 against 0.16 on shuffled data (skill 0.07). It still passed: a false alarm on random data, so read its passes with care.
Works when --: Correct calls per hidden gene reached 0.72 against 0.27 on shuffled data (skill 0.62, 164 scored). Run once on every gene, it called 172 genes with no known ec_number; the top 5 are listed.
PF3D7_0103800actin-related protein ARP1, putative -- 3 (support 1)PF3D7_0107200CCR4 domain-containing protein 3, putative -- 3 (support 1)PF3D7_0109200cleavage and polyadenylation specificity factor subunit 5, putative -- 3 (support 1)PF3D7_0110100selenocysteine-specific elongation factor, putative -- 3 (support 1)PF3D7_0110700chromatin assembly factor 1 subunit C, putative -- 3 (support 1)
15 · Find the communities several networks agree on
modularity + Louvain consensus · cluster recovery · T. gondii weak · P. falciparum weak
Which groups of genes are communities in more than one kind of measured relationship at once -- co-expressed AND co-fit AND crosslinked?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
█░░░░░░░░░░░ 0.05 [0.02, 0.08] · chance 0.00 |
░░░░░░░░░░░░ 0.03 [-0.03, 0.09] · chance 0.00 |
| Reach 1 - unclustered share |
████████████ 1.00 [1.00, 1.00] |
████████████ 1.00 [1.00, 1.00] |
| Label falls out as a cluster weighted F1 |
████░░░░░░░░ 0.35 [0.24, 0.46] · chance 0.30 |
███░░░░░░░░░ 0.25 [0.14, 0.35] · chance 0.22 |
| Partition agreement adjusted Rand index |
░░░░░░░░░░░░ 0.02 [0.01, 0.03] · chance 0.00 |
░░░░░░░░░░░░ 0.02 [-0.00, 0.04] · chance 0.00 |
About this test
- What it does. It finds groups of genes that form communities in several kinds of relationship at once, such as co-expression, co-fitness and crosslinks. Communities are found in each network separately, and two genes share a module only when enough of those networks put them together.
- How it is evaluated. A quarter of a chosen label is hidden. On the visible genes, the module that best matches each label is picked, and the test asks how well it holds that label's hidden genes (an F1 score). The null reshuffles the hidden genes' labels 100 times.
- What failure looks like, and why. The F1 score is no better than with shuffled labels. The modules may reflect a bias the networks share, such as favoring abundant proteins, rather than biology. Or the chosen label simply does not follow these communities, or the resolution makes modules too big or small.
- What success looks like, and why. The hidden genes land in the module chosen for their label, so the agreement between networks carries real signal. The modules become candidate complexes and pathways. A large module with no dominant label is a set of genes worth reading one by one.
T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: F1 of hidden genes in the community chosen for their label on known genes was 0.579 against 0.579 on shuffled data (skill -0.00).
Works when stageenrichedderived: F1 of hidden genes in the community chosen for their label on known genes reached 0.40 against 0.32 on shuffled data (skill 0.12, 452 scored). Run once on every gene, it called 2,000 genes with no known stageenrichedderived; the top 5 are listed.
TGME49_500435HECT-domain (ubiquitin-transferase) -containing protein -- oocyst (support --)TGME49_232080VPS13 domain-containing protein -- oocyst (support --)TGME49_242625ATPase family associated with various cellular activities (AAA) subfamily protein -- oocyst (support --)TGME49_268370non-specific serine/threonine protein kinase -- oocyst (support --)TGME49_253750PLU-1 family protein -- oocyst (support --)
P. falciparum. Fails when isexported: 'isexported' is not encoded in what this strategy reads: F1 of hidden genes in the community chosen for their label on known genes was 0.341 against 0.369 on shuffled data (skill -0.04).
Works when lopitpflocation: F1 of hidden genes in the community chosen for their label on known genes reached 0.23 against 0.14 on shuffled data (skill 0.10, 389 scored). Run once on every gene, it called 2,000 genes with no known lopitpflocation; the top 5 are listed.
PF3D7_0100100erythrocyte membrane protein 1, PfEMP1 -- nucleus 1 (support --)PF3D7_0100200rifin -- mitochondrion (support --)PF3D7_0100300erythrocyte membrane protein 1, PfEMP1 -- nucleus 1 (support --)PF3D7_0100400rifin -- mitochondrion (support --)PF3D7_0100500erythrocyte membrane protein 1 (PfEMP1), exon 1, pseudogene -- nucleus 1 (support --)
16 · Predict the contacts an interactome missed
logistic regression · ranking · T. gondii reliable · P. falciparum reliable
Which pairs of proteins are probably in physical contact although the crosslinking or pulldown experiment never saw them together?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
███████░░░░░ 0.61 [0.59, 0.62] · chance 0.00 |
████████████ 0.96 [0.96, 0.96] · chance 0.00 |
| Reach recall @ top 10% |
█████░░░░░░░ 0.44 [0.43, 0.45] · chance 0.10 |
███████░░░░░ 0.55 [0.55, 0.55] · chance 0.10 |
| True ones ranked first AUROC |
██████████░░ 0.80 [0.80, 0.81] · chance 0.50 |
████████████ 0.98 [0.98, 0.98] · chance 0.50 |
| Clean top of the list AUPRC lift |
██████░░░░░░ x3.6 [x3.5, x3.7] · chance x1.0 |
███████░░░░░ x5.3 [x5.3, x5.4] · chance x1.0 |
About this test
- What it does. It predicts physical contacts that a crosslinking or pulldown experiment missed. A model learns what separates measured contacts from other pairs: shared partners, support in other networks, and similar measurements. It then scores every unobserved pair with any support.
- How it is evaluated. A fifth of the layer's contacts are hidden and the model is retrained without them. It must rank the hidden contacts above fresh pairs of genes with a similar number of partners. The null is 5 models fed scrambled gene evidence; the number is the AUROC (ranking accuracy).
- What failure looks like, and why. The hidden contacts rank no better than with scrambled evidence. That means the other layers do not know about this interactome's contacts beyond each gene's popularity, or the layer has too few edges to learn from. Derived and literature layers are refused as targets.
- What success looks like, and why. The hidden contacts outrank matched non-contacts, so the other evidence can point to real missed interactions. You get a ranked list of likely contacts with the layers backing each. A link from a labeled to an unlabeled protein also hints at its location.
T. gondii. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: AUROC of hidden pairs against degree-matched non-pairs was 0.52 against 0.49 on shuffled data (skill 0.07). The test calls that a FAIL, which is the failure it exists to catch.
Works when xlms: AUROC of hidden pairs against degree-matched non-pairs reached 0.82 against 0.49 on shuffled data (skill 0.64, 568 scored). Run once on every gene, it proposed 300 links not in the layer; the top 5 are listed.
TGME49_261580 + TGME49_261250histone H2AX / histone H2A1 -- predicted link (score 12.5)TGME49_216450 + TGME49_258150proteasome subunit alpha type-3, putative / proteasome subunit alpha type-7, putative -- predicted link (score 12.4)TGME49_261570 + TGME49_232230ribosomal protein RPL7A / ribosomal protein RPL30 -- predicted link (score 12.1)TGME49_244560 + TGME49_288380heat shock protein 90, putative / heat shock protein HSP90 -- predicted link (score 12.1)TGME49_213280 + TGME49_208850SAG-related sequence SRS25 / SAG-related sequence SRS11 -- predicted link (score 12)
P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: AUROC of hidden pairs against degree-matched non-pairs was 0.52 against 0.49 on shuffled data (skill 0.07). The test calls that a FAIL, which is the failure it exists to catch.
Works when struct: AUROC of hidden pairs against degree-matched non-pairs reached 0.98 against 0.50 on shuffled data (skill 0.96, 914 scored). Run once on every gene, it proposed 300 links not in the layer; the top 5 are listed.
PF3D7_0902600 + PF3D7_1200800serine/threonine protein kinase, FIKK family / serine/threonine protein kinase, FIKK family -- predicted link (score 23.2)PF3D7_0102600 + PF3D7_0424500serine/threonine protein kinase, FIKK family / serine/threonine protein kinase, FIKK family -- predicted link (score 21.8)PF3D7_0102600 + PF3D7_0301200serine/threonine protein kinase, FIKK family / serine/threonine protein kinase, FIKK family -- predicted link (score 21.5)PF3D7_0902500 + PF3D7_0902600serine/threonine protein kinase, FIKK family / serine/threonine protein kinase, FIKK family -- predicted link (score 21.4)PF3D7_0102600 + PF3D7_0726200serine/threonine protein kinase, FIKK family / serine/threonine protein kinase, FIKK family -- predicted link (score 21.4)
17 · Read the literature for biology, not fame
publication-count residual · ranking · T. gondii reliable · P. falciparum weak
Which pairs of genes are written about together more than their popularity explains -- and are those pairs biologically related?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
█████░░░░░░░ 0.38 [0.21, 0.54] · chance 0.00 |
███░░░░░░░░░ 0.26 [0.23, 0.29] · chance 0.00 |
| Reach recall @ top 10% |
██░░░░░░░░░░ 0.14 [0.12, 0.16] · chance 0.10 |
██░░░░░░░░░░ 0.14 [0.10, 0.18] · chance 0.10 |
| True ones ranked first AUROC |
███████░░░░░ 0.62 [0.60, 0.64] · chance 0.50 |
███████░░░░░ 0.56 [0.46, 0.67] · chance 0.50 |
| Clean top of the list AUPRC lift |
███░░░░░░░░░ x1.2 [x1.1, x1.3] · chance x1.0 |
██░░░░░░░░░░ x1.2 [x1.0, x1.4] · chance x1.0 |
About this test
- What it does. It ranks gene pairs by how often papers mention them together beyond what each gene's fame predicts. Each co-mention count is replaced by its excess over what the two genes' own publication totals would give, so famous genes stop dominating the top.
- How it is evaluated. The label is never used to build the literature layer. Among pairs with both genes labeled, the top corrected pairs are taken, and the number is the share sharing a label. It is compared with 20 random sets of co-mentioned pairs; the raw ranking is shown beside it.
- What failure looks like, and why. The top corrected pairs share a label no more often than random co-mentioned pairs. Then the literature's links do not track this label, or too few co-mentioned pairs have both genes labeled for the top to be a real top.
- What success looks like, and why. The corrected top pairs share a compartment or complex more often than chance, so the literature links them for a reason. Pairs with a high score but no shared label are the literature's hypotheses that the current labels do not yet explain.
T. gondii. Fails when screenanyphenotype: The signal is too weak to tell from luck: 0.600 did not clear 0.705, what shuffled data reaches one time in twenty (skill 0.18). Also, only 20 hidden items could be scored.
Works when compartment: Share of the top 200 corrected pairs sharing a compartment label reached 0.75 against 0.39 on shuffled data (skill 0.59, 200 scored). Run once on every gene, it re-ranked 200 pairs once fame is taken out; the top 5 are listed.
TGME49_208030 + TGME49_291890microneme protein MIC4 / microneme protein MIC1 -- linked beyond fame (corrected 7.93)TGME49_227290 + TGME49_305340histone deacetylase HDAC3 / ATPase MORC -- linked beyond fame (corrected 7.88)TGME49_257370 + TGME49_274160preconoidal region protein PCR2 / preconoidal ring protein PCR1 -- linked beyond fame (corrected 7.74)TGME49_218960 + TGME49_310900AP2 domain transcription factor AP2XII-1 / AP2 domain transcription factor AP2XI-2 -- linked beyond fame (corrected 7.53)TGME49_208030 + TGME49_218520microneme protein MIC4 / microneme protein MIC6 -- linked beyond fame (corrected 7.49)
P. falciparum. Fails when isexported: The signal is too weak to tell from luck: 0.983 did not clear 0.991, what shuffled data reaches one time in twenty (skill 0.25).
Works when lopitpflocation: Share of the top 48 corrected pairs sharing a lopitpf_location label reached 0.56 against 0.40 on shuffled data (skill 0.27, 48 scored). Run once on every gene, it re-ranked 200 pairs once fame is taken out; the top 5 are listed.
PF3D7_0323400 + PF3D7_0424100Rh5 interacting protein / reticulocyte binding protein homologue 5 -- linked beyond fame (corrected 5.69)PF3D7_0731500 + PF3D7_1301600erythrocyte binding antigen-175 / erythrocyte binding antigen-140 -- linked beyond fame (corrected 5.49)PF3D7_0905400 + PF3D7_0929400high molecular weight rhoptry protein 3 / high molecular weight rhoptry protein 2 -- linked beyond fame (corrected 5.42)PF3D7_0102500 + PF3D7_1301600erythrocyte binding antigen-181 / erythrocyte binding antigen-140 -- linked beyond fame (corrected 5.32)PF3D7_0102500 + PF3D7_0731500erythrocyte binding antigen-181 / erythrocyte binding antigen-175 -- linked beyond fame (corrected 5.29)
18 · List what the data says and the literature has not written
multi-layer support count · ranking · T. gondii reliable · P. falciparum reliable
Which gene pairs do several independent measurements link that no paper has ever mentioned together?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
█░░░░░░░░░░░ 0.11 [0.11, 0.11] · chance 0.00 |
█░░░░░░░░░░░ 0.11 [0.11, 0.11] · chance 0.00 |
| Reach recall @ top 10% |
███████░░░░░ 0.57 [0.57, 0.57] · chance 0.10 |
███████░░░░░ 0.58 [0.58, 0.59] · chance 0.10 |
| True ones ranked first AUROC |
███████░░░░░ 0.56 [0.56, 0.56] · chance 0.50 |
███████░░░░░ 0.55 [0.55, 0.55] · chance 0.50 |
| Clean top of the list AUPRC lift |
███░░░░░░░░░ x1.5 [x1.5, x1.5] · chance x1.0 |
███░░░░░░░░░ x1.5 [x1.4, x1.5] · chance x1.0 |
About this test
- What it does. It counts, for every gene pair, how many independent measurement networks link it, then lists the well-supported pairs that no paper has ever mentioned together. Paralogs, shared domains and the compartment layer are left out, since those links are already written down.
- How it is evaluated. The literature serves only as the answer key. The test asks whether pairs with more measurement support are co-mentioned more often than random pairs, scored as an AUROC (ranking accuracy). The null repeats this 10 times with gene identities scrambled.
- What failure looks like, and why. The AUROC sits at the scrambled level: measurement support does not predict what biologists write about on this table. Then the unwritten pairs are not a forecast of future papers, and the list is just agreement between experiments, not a knowledge gap.
- What success looks like, and why. Better-supported pairs are co-mentioned more often, so the unwritten ones at the top are the most likely to be written next. Each is a concrete hypothesis with its supporting layers named. A pair of one understudied and one well-studied gene is cheapest to follow up.
T. gondii. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: AUROC of hidden pairs against random non-pairs was 0.50 against 0.50 on shuffled data (skill 0.00). The test calls that a FAIL, which is the failure it exists to catch.
Works when --: AUROC of hidden pairs against random non-pairs reached 0.56 against 0.50 on shuffled data (skill 0.11, 7,659 scored). Run once on every gene, it found 500 measured pairs nobody has written about; the top 5 are listed.
TGME49_253110 + TGME49_230020kinesin-8, putative / kinesin motor domain-containing protein -- measured, unwritten (layers 3)TGME49_226830 + TGME49_311720DnaK family protein / chaperonin protein BiP -- measured, unwritten (layers 3)TGME49_217460 + TGME49_263870glutaminyl-tRNA synthetase (GlnRS) / glutamate-tRNA ligase -- measured, unwritten (layers 3)TGME49_273760 + TGME49_209030heat shock protein HSP70 / actin ACT1 -- measured, unwritten (layers 3)TGME49_204400 + TGME49_261950ATPase synthase subunit alpha, putative / ATP synthase beta subunit ATP-B -- measured, unwritten (layers 3)
P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: AUROC of hidden pairs against random non-pairs was 0.50 against 0.50 on shuffled data (skill 0.00). The test calls that a FAIL, which is the failure it exists to catch.
Works when --: AUROC of hidden pairs against random non-pairs reached 0.55 against 0.50 on shuffled data (skill 0.11, 558 scored). Run once on every gene, it found 255 measured pairs nobody has written about; the top 5 are listed.
PF3D7_0101300 + PF3D7_0222100Pfmc-2TM Maurer's cleft two transmembrane protein / Pfmc-2TM Maurer's cleft two transmembrane protein -- measured, unwritten (layers 2)PF3D7_0101600 + PF3D7_0222600rifin / rifin -- measured, unwritten (layers 2)PF3D7_0101600 + PF3D7_1000500rifin / rifin -- measured, unwritten (layers 2)PF3D7_0101600 + PF3D7_1400600rifin / rifin -- measured, unwritten (layers 2)PF3D7_0101900 + PF3D7_1255000rifin / rifin -- measured, unwritten (layers 2)
19 · Train a classifier on the known genes and call the rest
logistic regression · label calls · T. gondii reliable · P. falciparum reliable
Given every permitted measurement, which label does a model trained on the labelled genes assign to each unlabelled one -- and which measurements does it rely on?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
████░░░░░░░░ 0.32 [0.18, 0.46] · chance 0.00 |
█████░░░░░░░ 0.39 [0.35, 0.43] · chance 0.00 |
| Reach coverage |
████████████ 1.00 [1.00, 1.00] |
████████████ 1.00 [1.00, 1.00] |
| Right calls accuracy |
███████░░░░░ 0.54 [0.41, 0.68] · chance 0.33 |
████████░░░░ 0.63 [0.46, 0.82] · chance 0.41 |
| Fair across classes macro F1 |
██████░░░░░░ 0.48 [0.36, 0.61] |
███████░░░░░ 0.55 [0.43, 0.67] |
About this test
- What it does. It trains a model on the labeled genes to learn which measurements recognize each class, then assigns a label with a probability to every unlabeled gene. It is balanced so the most common class does not win by default, and it lists the measurements each class relies on.
- How it is evaluated. A quarter of the label is hidden, whole gene families at a time, so a paralog cannot give the answer away. The model calls the hidden genes. The number is the share called correctly, compared with the chance rate for a guesser with the same mix of answers.
- What failure looks like, and why. The correct-call rate is near chance. The label may not be encoded in the permitted measurements, or the penalty setting makes the model too simple or lets it fit noise. A class recognized mainly by missing data is another warning sign.
- What success looks like, and why. Hidden genes are called well above chance, so the measurements do carry the label. You get probability-ranked calls for unlabeled genes and a readable list of what defines each class. Raising the probability threshold leaves fewer but surer calls.
T. gondii. Fails when screenanyphenotype: The signal is too weak to tell from luck: 0.550 did not clear 0.589, what shuffled data reaches one time in twenty (skill 0.11).
Works when compartment: Correct calls per hidden gene reached 0.33 against 0.06 on shuffled data (skill 0.29, 951 scored). Run once on every gene, it called 2,000 genes with no known compartment; the top 5 are listed.
TGME49_278390Toxoplasma gondii family A protein -- PM - peripheral 1 (support 0.693)TGME49_320250SAG-related sequence SRS15A -- PM - peripheral 1 (support 0.692)TGME49_278370Toxoplasma gondii family A protein -- PM - peripheral 1 (support 0.621)TGME49_273110SAG-related sequence SRS30D -- PM - peripheral 1 (support 0.598)TGME49_243190Toxoplasma gondii family A protein -- PM - peripheral 1 (support 0.576)
P. falciparum. Fails when isexported: The signal is real but small: 0.938 beat shuffled data (0.902), but by 0.036, short of the 0.050 margin a PASS requires.
Works when lopitpflocation: Correct calls per hidden gene reached 0.45 against 0.07 on shuffled data (skill 0.41, 395 scored). Run once on every gene, it called 2,000 genes with no known lopitpf_location; the top 5 are listed.
PF3D7_1479500stevor -- erythrocyte membrane (support 0.382)PF3D7_031280060S ribosomal protein L26, putative -- ribosomes (support 0.378)PF3D7_1200300rifin -- erythrocyte membrane (support 0.358)PF3D7_0221200Plasmodium exported protein (hyp15), unknown function -- erythrocyte membrane (support 0.357)PF3D7_0713100Pfmc-2TM Maurer's cleft two transmembrane protein -- erythrocyte membrane (support 0.352)
20 · Learn what makes your list special, from positives alone
PU bagging, logistic regression · ranking · T. gondii reliable · P. falciparum reliable
Given only genes that ARE something -- no list of genes that are not -- which other genes look most like them?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
█████████░░░ 0.79 [0.72, 0.85] · chance 0.00 |
███████████░ 0.90 [0.80, 0.98] · chance 0.00 |
| Reach recall @ top 10% |
████████░░░░ 0.66 [0.55, 0.77] · chance 0.10 |
██████████░░ 0.84 [0.58, 0.99] · chance 0.10 |
| True ones ranked first AUROC |
███████████░ 0.89 [0.86, 0.92] · chance 0.50 |
███████████░ 0.95 [0.90, 0.99] · chance 0.50 |
| Clean top of the list AUPRC lift |
██████████░░ x15 [x9.8, x22] · chance x1.0 |
████████████ x42 [x12, x83] · chance x1.0 |
About this test
- What it does. It starts from a list of genes that ARE something, with no list of genes that are not. Many small models each compare your list with a small random draw of other genes, and every other gene is scored only by models that did not train on it, then ranked.
- How it is evaluated. With 20 or more genes, 30% of your list is hidden and the rest trains the models; smaller lists are tested on a known category of similar size. The number is the AUROC: how well hidden members outrank every other gene. The null is 10 random lists of the same size.
- What failure looks like, and why. Hidden members rank no better than for random lists. The measurements may not capture what your list shares, or the list may be too mixed to have a common signature. If the list came from a column, failing to name it would instead inflate the result.
- What success looks like, and why. Hidden members rise to the top, so the models learned what sets your list apart, even from weak signals spread over many measurements. The candidates are genes that look like list members; take the top ones to strategy 24 to see what they share.
T. gondii. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: AUROC of hidden members against every other gene was 0.49 against 0.53 on shuffled data (skill -0.08). The test calls that a FAIL, which is the failure it exists to catch.
Works when PM - integral (compartment): AUROC of hidden members against every other gene reached 0.95 against 0.49 on shuffled data (skill 0.90, 40 scored). Run once on every gene, it ranked 300 genes with no place on the list as likely members of the list; the top 5 are listed.
TGME49_253880GNS1/SUR4 family protein -- likely member (score 5.29)TGME49_276910endoplasmic reticulum lumen protein retaining receptor (ERD2) family protein -- likely member (score 4.64)TGME49_299060sodium/hydrogen exchanger NHE2 -- likely member (score 4.53)TGME49_278850palmitoyltransferase DHHC2 -- likely member (score 4.43)TGME49_309560nmda receptor glutamate-binding chain -- likely member (score 4.34)
P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: AUROC of hidden members against every other gene was 0.49 against 0.53 on shuffled data (skill -0.08). The test calls that a FAIL, which is the failure it exists to catch.
Works when cytosol (lopitpflocation): AUROC of hidden members against every other gene reached 0.98 against 0.50 on shuffled data (skill 0.96, 37 scored). Run once on every gene, it ranked 300 genes with no place on the list as likely members of the list; the top 5 are listed.
PF3D7_1438900thioredoxin peroxidase 1 -- likely member (score 4.17)PF3D7_0518300proteasome subunit beta type-1, putative -- likely member (score 4.1)PF3D7_1403900serine/threonine protein phosphatase CPPED1, putative -- likely member (score 4)PF3D7_1104000phenylalanine--tRNA ligase beta subunit -- likely member (score 3.98)PF3D7_1419300glutathione S-transferase -- likely member (score 3.9)
21 · Predict a measurement, and find the genes that defy the prediction
gradient boosting / ridge · values · T. gondii reliable · P. falciparum reliable
How well does everything else predict this measurement -- and which genes are far from what their profile says they should be?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
███████░░░░░ 0.56 [0.05, 0.91] · chance 0.00 |
██████░░░░░░ 0.52 [0.25, 0.85] · chance 0.00 |
| Reach coverage |
████████████ 1.00 [1.00, 1.00] |
████████████ 1.00 [1.00, 1.00] |
| Order predicted Spearman rho |
███████░░░░░ 0.56 [0.05, 0.91] · chance 0.00 |
██████░░░░░░ 0.52 [0.25, 0.85] · chance 0.00 |
| Variance explained R-squared, out of sample |
█████░░░░░░░ 0.43 [-0.09, 0.87] · chance 0.00 |
░░░░░░░░░░░░ -3.29 [-14.04, 0.72] · chance 0.00 |
About this test
- What it does. It predicts one numeric measurement, such as a fitness score, from all other permitted measurements. By default, other measurements of the same kind are left out. Each gene is predicted by a model that never saw it or its paralog.
- How it is evaluated. A fifth of the measured values are hidden and the model is trained on the rest. The number is the rank correlation between predicted and hidden values. The null is 3 models trained on the same values shuffled among genes.
- What failure looks like, and why. The correlation is no better than with shuffled values. The measurement then has no signature in the other data, so predictions for unmeasured genes mean little and the surprising genes may be noise. Including screens of the same kind would mainly show agreement.
- What success looks like, and why. The hidden values are predicted, so the trait has a signature in the other data. Unmeasured genes get an estimated value, trusted as far as the correlation says. Genes far from their prediction are candidates for unusual function, or artifacts worth rechecking.
T. gondii. Fails when fitinvivoPE: The signal is real but small: 0.052 beat shuffled data (0.000), but by 0.052, short of the 0.100 margin a PASS requires.
Works when fitinvitrohff: Rank correlation of predicted and hidden values reached 0.73 against 0.00 on shuffled data (skill 0.73, 1,465 scored). Run once on every gene, it predicted 815 genes with no measured fitinvitrohff; the top 5 are listed.
TGME49_219790pre-mRNA processing factor PRP3 -- predicted -4.44 (furthest from the median, 0.0805)TGME49_310430Hsp90 domain-containing protein -- predicted -4.26 (furthest from the median, 0.0805)TGME49_316400alpha tubulin TUBA1 -- predicted -3.97 (furthest from the median, 0.0805)TGME49_283590mitochondrial import inner membrane translocase subunit TIM50 -- predicted -3.74 (furthest from the median, 0.0805)TGME49_249180bifunctional dihydrofolate reductase-thymidylate synthase -- predicted -3.62 (furthest from the median, 0.0805)
P. falciparum. Fails when exprschizont: The signal is too weak to tell from luck: 0.036 did not clear 0.049, what shuffled data reaches one time in twenty (skill 0.04).
Works when piggybacmis: Rank correlation of predicted and hidden values reached 0.49 against 0.00 on shuffled data (skill 0.49, 1,077 scored). Run once on every gene, it predicted 335 genes with no measured piggybac_mis; the top 5 are listed.
PF3D7_030440060S ribosomal protein L44 -- predicted 0.0957 (furthest from the median, 0.605)PF3D7_1041300erythrocyte membrane protein 1, PfEMP1 -- predicted 1.03 (furthest from the median, 0.605)PF3D7_0632800erythrocyte membrane protein 1, PfEMP1 -- predicted 1.01 (furthest from the median, 0.605)PF3D7_0223500erythrocyte membrane protein 1, PfEMP1 -- predicted 0.99 (furthest from the median, 0.605)PF3D7_1373500erythrocyte membrane protein 1, PfEMP1 -- predicted 0.988 (furthest from the median, 0.605)
22 · Fill in what was never measured, and say where that is honest
soft-impute, low-rank SVD · values · T. gondii reliable · P. falciparum reliable
For each measurement, can its missing values be estimated from the rest of the table -- and for which measurements is that impossible?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
██████████░░ 0.87 [0.86, 0.87] · chance 0.00 |
████████░░░░ 0.68 [0.67, 0.69] · chance 0.00 |
| Reach coverage |
████████████ 1.00 [1.00, 1.00] |
████████████ 1.00 [1.00, 1.00] |
| Order predicted Spearman rho |
██████████░░ 0.87 [0.86, 0.87] · chance 0.00 |
████████░░░░ 0.68 [0.67, 0.69] · chance 0.00 |
| Variance explained R-squared, out of sample |
█████████░░░ 0.75 [0.75, 0.76] · chance 0.00 |
█████░░░░░░░ 0.45 [0.44, 0.46] · chance 0.00 |
About this test
- What it does. It estimates missing measurements from the rest of the table using a model of a few shared patterns (a low-rank model). It first checks, column by column, how well each measurement can be rebuilt, and fills only where that works; elsewhere a gap stays a gap.
- How it is evaluated. A tenth of every column's measured values is hidden and the table is completed. Each column gets a reliability, the rank correlation on its hidden values; the verdict uses the median over columns. The null is 3 tables with each column shuffled on its own.
- What failure looks like, and why. The median reliability is at the shuffled level. The measurements are then not correlated enough to predict each other, or the rank is off: too few patterns blur programs together, too many fit noise. Missing values should stay unknown rather than be filled.
- What success looks like, and why. Columns are rebuilt far better than shuffled ones, so the measurements share structure. Columns with high reliability can be filled for genes never measured, which beats filling with the median. The reliability list also says which columns must stay gaps.
T. gondii. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: median per-column rank correlation on hidden entries was -0.01 against -0.03 on shuffled data (skill 0.01). The test calls that a FAIL, which is the failure it exists to catch.
Works when --: Median per-column rank correlation on hidden entries reached 0.88 against 0.00 on shuffled data (skill 0.88, 225,394 scored). Run once on every gene, it filled 815 missing measurements; the top 5 are listed.
TGME49_202860hypothetical protein -- percentile -0.388 (column reliability 0.903)TGME49_500035hypothetical protein, conserved -- percentile -0.21 (column reliability 0.903)TGME49_500067hypothetical protein, conserved -- percentile -0.107 (column reliability 0.903)TGME49_500435HECT-domain (ubiquitin-transferase) -containing protein -- percentile -0.0907 (column reliability 0.903)TGME49_500301ubiquitin-associated domain-containing protein -- percentile -0.0854 (column reliability 0.903)
P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: median per-column rank correlation on hidden entries was -0.01 against -0.03 on shuffled data (skill 0.01). The test calls that a FAIL, which is the failure it exists to catch.
Works when --: Median per-column rank correlation on hidden entries reached 0.69 against 0.00 on shuffled data (skill 0.69, 58,452 scored). Run once on every gene, it filled 335 missing measurements; the top 5 are listed.
PF3D7_114864028S ribosomal RNA -- percentile -0.68 (column reliability 0.661)PF3D7_030440060S ribosomal protein L44 -- percentile -0.668 (column reliability 0.661)PF3D7_1148500non-coding RNA -- percentile -0.613 (column reliability 0.661)PF3D7_137130028S ribosomal RNA -- percentile -0.581 (column reliability 0.661)PF3D7_011230018S ribosomal RNA -- percentile -0.492 (column reliability 0.661)
23 · Find what matters more in one condition, and why
residual + gradient boosting / ridge · values · T. gondii weak · P. falciparum reliable
Which genes matter more (or less) in one condition than a baseline predicts -- in the mouse rather than the dish, say -- and can the rest of the data explain which?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
█░░░░░░░░░░░ 0.07 [0.06, 0.09] · chance 0.00 |
████████░░░░ 0.64 [0.64, 0.65] · chance 0.00 |
| Reach coverage |
████████████ 1.00 [1.00, 1.00] |
████████████ 1.00 [1.00, 1.00] |
| Order predicted Spearman rho |
█░░░░░░░░░░░ 0.07 [0.06, 0.09] · chance 0.00 |
████████░░░░ 0.64 [0.64, 0.65] · chance 0.00 |
| Variance explained R-squared, out of sample |
░░░░░░░░░░░░ -0.04 [-0.06, -0.03] · chance 0.00 |
█████░░░░░░░ 0.40 [0.40, 0.41] · chance 0.00 |
About this test
- What it does. It finds genes that matter more or less in one condition, such as the mouse, than a baseline like fibroblasts predicts. The condition screen is corrected for the baseline, and the leftover shift is then predicted from other measurements, with both screens withheld.
- How it is evaluated. A fifth of the genes measured in both screens are hidden, and a model trained on the rest predicts their shift. The number is the rank correlation between predicted and actual shifts. The null is 3 models trained on shuffled shifts.
- What failure looks like, and why. The correlation stays at the shuffled level. Then the condition-specific shift is idiosyncratic or noise, with no signature in the other data. The ranked list of shifted genes remains, but predictions for genes never screened in the condition should not be used.
- What success looks like, and why. Hidden shifts are predicted, so the condition-specific need has a signature, such as secretion or host-facing location. Strongly negative genes are candidates for host interaction or in vivo nutrition, and unscreened genes can be ranked by that signature.
T. gondii. Fails when random data: The signal is real but small: 0.069 beat shuffled data (0.000), but by 0.069, short of the 0.100 margin a PASS requires.
Works when --: Rank correlation of predicted and hidden values reached 0.15 against 0.00 on shuffled data (skill 0.15, 1,465 scored). Run once on every gene, it predicted 680 genes with no label; the top 5 are listed.
TGME49_500302Porphobilinogen Synthase (PBGS) / Aminolevulinic Acid Dehydratase (ALAD) + Pitrilysin (apicoplast stromal processing peptidase) -- predicted shift 0.387 (furthest from the median, 0.13)TGME49_214410hypothetical protein -- predicted shift -0.124 (furthest from the median, 0.13)TGME49_500287phosphatidylinositol 4-kinase, putative -- predicted shift -0.0789 (furthest from the median, 0.13)TGME49_264250hypothetical protein -- predicted shift -0.0603 (furthest from the median, 0.13)TGME49_252065KRUF family protein -- predicted shift -0.0548 (furthest from the median, 0.13)
P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: rank correlation of predicted and hidden values was 0.03 against 0.00 on shuffled data (skill 0.03). The test calls that a FAIL, which is the failure it exists to catch.
Works when --: Rank correlation of predicted and hidden values reached 0.65 against 0.00 on shuffled data (skill 0.65, 1,144 scored). Run once on every gene, it measured the condition-specific shift of 5,720 genes; the top 5 are listed.
PF3D7_07258005.8S ribosomal RNA -- condition-specific shift -0.771 (explained -0.179)PF3D7_1240500Plasmodium RNA of unknown function RUF6 -- condition-specific shift -0.766 (explained -0.203)PF3D7_MIT02200small subunit ribosomal RNA fragment A -- condition-specific shift -0.762 (explained -0.194)PF3D7_MIT02500large subunit ribosomal RNA fragment B -- condition-specific shift -0.762 (explained -0.185)PF3D7_MIT01800ribosomal RNA fragment RNA26t -- condition-specific shift 0.756 (explained -0.194)
24 · Describe what your gene list has in common
hypergeometric + rank-sum · ranking · T. gondii reliable · P. falciparum reliable
What distinguishes the genes on my list from the rest -- which categories are they enriched in, which measurements are shifted, which networks are dense among them?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
███████░░░░░ 0.61 [0.51, 0.72] · chance 0.00 |
██████████░░ 0.85 [0.75, 0.94] · chance 0.00 |
| Reach recall @ top 10% |
██████░░░░░░ 0.47 [0.32, 0.61] · chance 0.10 |
█████████░░░ 0.76 [0.58, 0.92] · chance 0.10 |
| True ones ranked first AUROC |
██████████░░ 0.81 [0.75, 0.86] · chance 0.50 |
███████████░ 0.93 [0.88, 0.97] · chance 0.50 |
| Clean top of the list AUPRC lift |
████████░░░░ x8.1 [x3.9, x13] · chance x1.0 |
███████████░ x19 [x13, x26] · chance x1.0 |
About this test
- What it does. It describes what your gene list has in common, testing every category, measurement and network in the table at once with one multiple-testing correction. The significant features form a profile, which is then used to rank every other gene.
- How it is evaluated. The profile is built from 60% of the list, and the other 40% is hidden. The number is the AUROC: how well hidden members outrank every other gene on the profile. The null is 10 random lists of the same size.
- What failure looks like, and why. Hidden members score no better than random lists. The enriched features are then a description of this particular list rather than something its members truly share. Naming the column the list came from matters, or the profile just reads that column back.
- What success looks like, and why. The profile recognizes members it never saw, so it captures what the list truly shares. You get a readable profile, sorted by q-value, and a ranked set of genes that resemble the list, as candidates for membership.
T. gondii. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: AUROC of hidden members against every other gene was 0.50 against 0.50 on shuffled data (skill -0.01). The test calls that a FAIL, which is the failure it exists to catch.
Works when PM - integral (compartment): AUROC of hidden members against every other gene reached 0.91 against 0.50 on shuffled data (skill 0.82, 54 scored). Run once on every gene, it ranked 300 genes with no place on the list by how much they resemble the list; the top 5 are listed.
TGME49_257120sugar transporter ST1 -- resembles the list (score 19)TGME49_288920ATP-binding cassette G family transporter ABCG96 -- resembles the list (score 17.1)TGME49_229170formate/nitrite transporter FNT3 -- resembles the list (score 16.7)TGME49_299060sodium/hydrogen exchanger NHE2 -- resembles the list (score 16.4)TGME49_305180Na+/H+ exchanger NHE3 -- resembles the list (score 15.8)
P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: AUROC of hidden members against every other gene was 0.50 against 0.50 on shuffled data (skill -0.01). The test calls that a FAIL, which is the failure it exists to catch.
Works when cytosol (lopitpflocation): AUROC of hidden members against every other gene reached 0.93 against 0.50 on shuffled data (skill 0.86, 49 scored). Run once on every gene, it ranked 300 genes with no place on the list by how much they resemble the list; the top 5 are listed.
PF3D7_0621200pyridoxine biosynthesis protein PDX1 -- resembles the list (score 15.7)PF3D7_0623500superoxide dismutase [Fe] -- resembles the list (score 14.8)PF3D7_0906300Maf-like protein, putative -- resembles the list (score 13.8)PF3D7_1308900mRNA-decapping enzyme 2, putative -- resembles the list (score 13.3)PF3D7_0318800triosephosphate isomerase, putative -- resembles the list (score 13.1)
25 · Grow your gene list along the networks
random walk with restart · ranking · T. gondii reliable · P. falciparum reliable
Starting from my genes, which others does a walk across every measured network keep returning to?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
███████░░░░░ 0.59 [0.49, 0.68] · chance 0.00 |
██████████░░ 0.86 [0.74, 0.96] · chance 0.00 |
| Reach recall @ top 10% |
██████░░░░░░ 0.47 [0.38, 0.56] · chance 0.10 |
██████████░░ 0.82 [0.67, 0.96] · chance 0.10 |
| True ones ranked first AUROC |
██████████░░ 0.80 [0.74, 0.84] · chance 0.50 |
███████████░ 0.93 [0.88, 0.98] · chance 0.50 |
| Clean top of the list AUPRC lift |
████████░░░░ x7.7 [x4.4, x12] · chance x1.0 |
████████████ x38 [x13, x72] · chance x1.0 |
About this test
- What it does. It grows your gene list along the measured networks. A random walk starts at your genes, steps along edges and jumps back often; genes it visits most are close by many short paths. Layers are averaged, and hubs are down-weighted so they do not attract every walk.
- How it is evaluated. 30% of your list is hidden and the walk starts from the other 70%. The number is the AUROC: how well hidden members outrank every other gene by visit score. The null is 10 random starting lists of the same size.
- What failure looks like, and why. Hidden members are visited no more than for random lists. Your genes may have few network edges, or be linked only through single long paths. The measurement graph option can reach genes no network covers.
- What success looks like, and why. The walk returns to the hidden members, so your list is tightly knit in the data. The candidates are the list's closest relatives, such as other members of a complex, pathway or secretory route, each shown with the layers linking it directly to your genes.
T. gondii. Fails when tachyzoite (stageenrichedderived): The signal is too weak to tell from luck: 0.548 did not clear 0.561, what shuffled data reaches one time in twenty (skill 0.09).
Works when PM - integral (compartment): AUROC of hidden members against every other gene reached 0.88 against 0.50 on shuffled data (skill 0.76, 40 scored). Run once on every gene, it ranked 300 genes with no place on the list as likely members of the list; the top 5 are listed.
TGME49_257120sugar transporter ST1 -- likely member (score 5.15e-04)TGME49_288920ATP-binding cassette G family transporter ABCG96 -- likely member (score 4.62e-04)TGME49_315560ATP-binding cassette G family transporter ABCG77 -- likely member (score 4.22e-04)TGME49_229170formate/nitrite transporter FNT3 -- likely member (score 4.15e-04)TGME49_305180Na+/H+ exchanger NHE3 -- likely member (score 4.01e-04)
P. falciparum. Fails when gametocyte (stageenrichedderived): 'gametocyte (stageenrichedderived)' is not encoded in what this strategy reads: AUROC of hidden members against every other gene was 0.366 against 0.505 on shuffled data (skill -0.28). Also, only 20 hidden items could be scored.
Works when cytosol (lopitpflocation): AUROC of hidden members against every other gene reached 0.94 against 0.49 on shuffled data (skill 0.89, 37 scored). Run once on every gene, it ranked 300 genes with no place on the list as likely members of the list; the top 5 are listed.
PF3D7_0623500superoxide dismutase [Fe] -- likely member (score 8.22e-04)PF3D7_1361400actin-depolymerizing factor 2 -- likely member (score 6.47e-04)PF3D7_1216000serine--tRNA ligase, putative -- likely member (score 5.73e-04)PF3D7_1226100haloacid dehalogenase-like hydrolase, putative -- likely member (score 5.52e-04)PF3D7_1037100pyruvate kinase 2 -- likely member (score 4.77e-04)
26 · Find categories that split in two on another measurement
UMAP + HDBSCAN · replication · T. gondii untestable · P. falciparum untestable
Which clusters agree about one thing -- a compartment -- and split cleanly on another -- a stage, a phase, a fitness level?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
-- | -- |
| Reach findings made |
-- | -- |
| Findings that hold replication rate |
-- | -- |
| Beyond chance replication lift |
-- | -- |
About this test
- What it does. It finds clusters that agree on one label, such as a compartment, but split cleanly on another label or measurement, such as a stage or a fitness level. Both labels are withheld from the map, so each split is a claim that a category holds two kinds of gene.
- How it is evaluated. Splits are found on a random half of the genes and checked on the other half. The number is the share of findings that reappear. The null is 20 runs with the splitting label shuffled in the second half; passing also needs at least three findings.
- What failure looks like, and why. Few findings replicate, no more than with the shuffled label, or there are fewer than three. The category may have no real internal structure, or clusters are too coarse or too small to hold a whole category with both sides. A fine-grained shared label rarely clusters.
- What success looks like, and why. Splits reappear in genes the discovery never saw, so the category truly contains two kinds of gene, and the measurement carrying the split is named. The minority side is listed, a lead on, for example, which proteins in a compartment differ by phase or essentiality.
T. gondii. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find. It made too little to score, and the test declined to judge (only 0 findings on the first half, 3 needed).
Works when --: None of its 15 calibration runs on real data passed; the usual reason: only 0 findings on the first half, 3 needed. There is no success to show.
P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find. It made too little to score, and the test declined to judge (only 0 findings on the first half, 3 needed).
Works when --: None of its 15 calibration runs on real data passed; the usual reason: only 0 findings on the first half, 3 needed. There is no success to show.
27 · Find kinds of gene defined by two labels at once
UMAP + HDBSCAN · replication · T. gondii reliable · P. falciparum untestable
Which clusters are enriched for a COMBINATION of two labels -- more than either label alone would make them?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
██████░░░░░░ 0.49 [0.38, 0.62] · chance 0.00 |
-- |
| Reach findings made |
███░░░░░░░░░ 5 [4, 7] |
-- |
| Findings that hold replication rate |
██████░░░░░░ 0.50 [0.39, 0.64] · chance 0.03 |
-- |
| Beyond chance replication lift |
███████████░ x22 [x15, x31] · chance x1.0 |
-- |
About this test
- What it does. It looks for groups of genes defined by two labels at once, such as secreted AND fitness-conferring. Both labels are hidden, genes are clustered on a map of the measurements, and each cluster's joint enrichment is compared with the stronger of the two single enrichments. A ratio above 1.25 flags a combination.
- How it is evaluated. Combinations are found on a random half of the genes and checked on the other half: among genes with the first label, is the second label still concentrated in that cluster? The number that decides it is the share of findings that replicate, compared with 20 runs where the second label is shuffled.
- What failure looks like, and why. The replicating share sits near the shuffled level, or fewer than three combinations are found. Often the cluster was simply enriched for one label, so the second adds nothing within it. Clusters may also be too coarse to isolate a combination, or the two labels are not encoded in these measurements.
- What success looks like, and why. Several combinations replicate well above the shuffled level on genes the discovery never saw. That means a real kind of gene exists that neither label shows alone. The cluster's unlabeled members become candidates for the combination.
T. gondii. Fails when random data: The signal is too weak to tell from luck: 0.167 did not clear 0.167, what shuffled data reaches one time in twenty (skill 0.15). Also, only 12 hidden items could be scored.
Works when --: Share of first-half findings that replicate on the second half reached 0.75 against 0.04 on shuffled data (skill 0.74, 4 scored). Run once on every gene, it found 5 structures neither label shows alone; the top 5 are listed.
- pelliclePMIMC x SP (lopit_unified) (purity 1, 27 genes)
- pelliclePMIMC x SP (lopit_unified) (purity 0.571, 266 genes)
- mitochondrion x TM (lopit_unified) (purity 0.448, 29 genes)
- apicalsecretory x SP+TM (lopitunified) (purity 0.263, 19 genes)
- apicalsecretory x SP (lopitunified) (purity 0.25, 25 genes)
P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find. It made too little to score, and the test declined to judge (only 0 findings on the first half, 3 needed).
Works when --: None of its 15 calibration runs on real data passed; the usual reason: only 0 findings on the first half, 3 needed. There is no success to show.
28 · Find paralogs that changed jobs
profile correlation · ranking · T. gondii reliable · P. falciparum weak
Which duplicated genes behave differently across the measurements -- evidence that one copy took on a new role?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
██░░░░░░░░░░ 0.16 [0.06, 0.31] · chance 0.00 |
█░░░░░░░░░░░ 0.11 [-0.14, 0.37] · chance 0.00 |
| Reach recall @ top 10% |
██░░░░░░░░░░ 0.13 [0.10, 0.17] · chance 0.10 |
██░░░░░░░░░░ 0.13 [0.10, 0.15] · chance 0.10 |
| True ones ranked first AUROC |
███████░░░░░ 0.58 [0.53, 0.66] · chance 0.50 |
███████░░░░░ 0.56 [0.44, 0.68] · chance 0.50 |
| Clean top of the list AUPRC lift |
███░░░░░░░░░ x1.2 [x1.0, x1.4] · chance x1.0 |
███░░░░░░░░░ x1.2 [x1.1, x1.4] · chance x1.0 |
About this test
- What it does. It asks which duplicated genes (paralogs) behave differently across the measurements, hinting that one copy took a new job. Divergence is one minus the correlation of the two genes' profiles over the measurements both have. Pairs are ranked from most to least diverged.
- How it is evaluated. Pairs where both genes carry a compartment label are used, with that label hidden from the profiles. The test asks whether pairs in different compartments are more diverged than pairs in the same one. The deciding number is that ranking score (AUROC), versus 20 random reassignments of divergence.
- What failure looks like, and why. The score sits at the random-reassignment level: divergence does not track a change of compartment. Pairs may share too few measurements, making divergence noisy. Or copies in different compartments may still be measured alike, so the profiles cannot tell them apart.
- What success looks like, and why. Pairs in different compartments are clearly more diverged than pairs in the same one. Divergence then carries meaning, so a highly diverged pair with no localization is evidence that one copy changed its place or job. It gives a ranked shortlist of paralogs to compare side by side.
T. gondii. Fails when dtm_class: The signal is real but small: 0.514 beat shuffled data (0.500), but by 0.014, short of the 0.050 margin a PASS requires.
Works when compartment: AUROC of profile divergence for paralogs with different compartment reached 0.55 against 0.49 on shuffled data (skill 0.11, 535 scored). Run once on every gene, it ranked 3,452 paralog pairs by how far they diverged; the top 5 are listed.
TGME49_207005 + TGME49_271050SAG-related sequence SRS48Q / SAG-related sequence SRS34A -- diverged paralogs (divergence 1.7)TGME49_207010 + TGME49_271050SAG-related sequence SRS48K / SAG-related sequence SRS34A -- diverged paralogs (divergence 1.69)TGME49_207015 + TGME49_271050SRS domain-containing protein / SAG-related sequence SRS34A -- diverged paralogs (divergence 1.69)TGME49_316190 + TGME49_316310superoxide dismutase SOD3 / superoxide dismutase SOD1 -- diverged paralogs (divergence 1.61)TGME49_204130 + TGME49_272430perforin-like protein PLP1 / perforin-like protein PLP2 -- diverged paralogs (divergence 1.58)
P. falciparum. Fails when pbtransferredphenotype: 'pbtransferredphenotype' is not encoded in what this strategy reads: AUROC of profile divergence for paralogs with different pbtransferredphenotype was 0.435 against 0.505 on shuffled data (skill -0.14).
Works when lopitpflocation: AUROC of profile divergence for paralogs with different lopitpflocation reached 0.68 against 0.48 on shuffled data (skill 0.39, 96 scored). Run once on every gene, it ranked 1,741 paralog pairs by how far they diverged; the top 5 are listed.
PF3D7_0208900 + PF3D7_04049006-cysteine protein P230p / 6-cysteine protein P41 -- diverged paralogs (divergence 1.29)PF3D7_0401100 + PF3D7_0831500Plasmodium exported protein, unknown function, fragment / Plasmodium exported protein (PHIST), unknown function -- diverged paralogs (divergence 1.21)PF3D7_0831500 + PF3D7_1479200Plasmodium exported protein (PHIST), unknown function / Plasmodium exported protein (PHISTa), unknown function -- diverged paralogs (divergence 1.2)PF3D7_0208900 + PF3D7_06127006-cysteine protein P230p / 6-cysteine protein P12 -- diverged paralogs (divergence 1.19)PF3D7_0525100 + PF3D7_1253400acyl-CoA synthetase / acyl-CoA synthetase -- diverged paralogs (divergence 1.17)
29 · Carry what one parasite shows to the other
orthogroup mapping · values · T. gondii reliable · P. falciparum reliable
What does a gene's ortholog in the other parasite say about it -- its essentiality, its stage, its localization?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
████░░░░░░░░ 0.31 [0.30, 0.33] · chance 0.00 |
████░░░░░░░░ 0.31 [0.30, 0.33] · chance 0.00 |
| Reach coverage |
████████████ 1.00 [1.00, 1.00] |
████████████ 1.00 [1.00, 1.00] |
| Order predicted Spearman rho |
████░░░░░░░░ 0.32 [0.30, 0.33] · chance 0.00 |
████░░░░░░░░ 0.32 [0.30, 0.33] · chance 0.00 |
| Variance explained R-squared, out of sample |
░░░░░░░░░░░░ -1.60 [-1.70, -1.50] · chance 0.00 |
░░░░░░░░░░░░ -85.78 [-88.25, -83.72] · chance 0.00 |
About this test
- What it does. It carries a measurement or label from the other parasite onto this one through shared orthogroups (gene families across species). The source is summarized per orthogroup, its relation to the target is learned on genes measured in both, and that is applied to genes measured only on the other side.
- How it is evaluated. A quarter of the genes measured in both species are hidden and the relation is learned on the rest. For a measurement, the deciding number is the rank correlation between transferred and hidden values; for a label, correct calls. Both are compared with runs where ortholog values go to random genes.
- What failure looks like, and why. The transfer does no better than orthologs dealt to random genes. The two species may use the gene differently, so the measurement is not conserved. Too few genes may be measured in both to learn the relation. Lineage-specific genes have no ortholog at all and are never reached.
- What success looks like, and why. Transferred values track the hidden ones well above the shuffled level. The ortholog then acts as a second, independent measurement of the gene. A Plasmodium knockout screen, for example, becomes a prediction for untested Toxoplasma genes, listed in 'transferred'.
T. gondii. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: rank correlation of transferred and hidden values was -0.05 against 0.02 on shuffled data (skill -0.07). The test calls that a FAIL, which is the failure it exists to catch.
Works when --: Rank correlation of transferred and hidden values reached 0.35 against -0.00 on shuffled data (skill 0.35, 681 scored). Run once on every gene, it predicted 106 genes with no measured fitinvitrohff; the top 5 are listed.
TGME49_500148GCC2 and GCC3 domain-containing protein -- transferred 1 (furthest from the median, 0.291)TGME49_500222GCC2 and GCC3 domain-containing protein -- transferred 1 (furthest from the median, 0.291)TGME49_500227dynein heavy chain, putative -- transferred 1 (furthest from the median, 0.291)TGME49_500062dynein heavy chain, putative -- transferred 1 (furthest from the median, 0.291)TGME49_500072dynein heavy chain, putative -- transferred 1 (furthest from the median, 0.291)
P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: rank correlation of transferred and hidden values was -0.05 against 0.02 on shuffled data (skill -0.07). The test calls that a FAIL, which is the failure it exists to catch.
Works when --: Rank correlation of transferred and hidden values reached 0.34 against 0.01 on shuffled data (skill 0.33, 661 scored). Run once on every gene, it predicted 25 genes with no measured piggybac_mis; the top 5 are listed.
PF3D7_0902600serine/threonine protein kinase, FIKK family -- transferred 1.42 (furthest from the median, -4)PF3D7_114430060S ribosomal protein L41 -- transferred 0.52 (furthest from the median, -4)PF3D7_0417500memo-like protein -- transferred -0.01 (furthest from the median, -4)PF3D7_1468500derlin-1 -- transferred -0.56 (furthest from the median, -4)PF3D7_1360900RNA-binding protein, putative -- transferred -0.82 (furthest from the median, -4)
30 · Test inference on the genes orthology cannot reach
kNN · label calls · T. gondii reliable · P. falciparum weak
Can lineage-specific, hypothetical or understudied genes be called as reliably as the rest -- and what are they?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
███░░░░░░░░░ 0.26 [0.15, 0.34] · chance 0.00 |
░░░░░░░░░░░░ -0.01 [-0.01, -0.01] · chance 0.00 |
| Reach coverage |
███████████░ 0.89 [0.76, 1.00] |
████████████ 1.00 [1.00, 1.00] |
| Right calls accuracy |
███████░░░░░ 0.59 [0.39, 0.78] · chance 0.42 |
████████░░░░ 0.66 [0.32, 1.00] · chance 0.66 |
| Fair across classes macro F1 |
█████░░░░░░░ 0.38 [0.27, 0.54] |
███████░░░░░ 0.58 [0.16, 1.00] |
About this test
- What it does. It calls labels for one hard group of genes, such as lineage-specific, hypothetical or understudied ones, from their nearest neighbors in the measurements. All conservation and orthology columns are removed first, so a call cannot just reflect how conserved a gene is.
- How it is evaluated. A quarter of the label is hidden by whole orthogroups, and only hidden genes inside the chosen group are scored. The deciding number is correct calls per hidden gene of that group. It is compared with 10 runs on shuffled labels, and shown beside the accuracy for the rest of the proteome.
- What failure looks like, and why. Accuracy within the group sits at the shuffled level, even if the whole-proteome number looked good. These genes are where methods are weakest: few labeled neighbors resemble them, and the label may not be encoded in their measurements once conservation is removed.
- What success looks like, and why. Calls in the group beat shuffled labels by a clear margin. The error rate then belongs to the genes being predicted, not borrowed from well-studied ones. Unlabeled genes in the group get calls that can be trusted at that group's own rate.
T. gondii. Fails when dtm_class: The signal is real but small: 0.733 beat shuffled data (0.709), but by 0.024, short of the 0.050 margin a PASS requires.
Works when compartment: Correct calls per hidden gene reached 0.35 against 0.04 on shuffled data (skill 0.32, 377 scored). Run once on every gene, it called 1,449 genes with no known compartment; the top 5 are listed.
TGME49_312940hypothetical protein -- mitochondrion - membranes (support 0.876)TGME49_228670zinc finger, C2H2 type domain-containing protein -- nucleus - chromatin (support 0.87)TGME49_295610histone lysine methyltransferase, SET, putative -- nucleus - chromatin (support 0.805)TGME49_221280hypothetical protein -- nucleus - chromatin (support 0.797)TGME49_283520hypothetical protein -- nucleus - chromatin (support 0.79)
P. falciparum. Fails when stageenrichedderived: 'stageenrichedderived' is not encoded in what this strategy reads: correct calls per hidden gene was 0.318 against 0.323 on shuffled data (skill -0.01). Also, only 22 hidden items could be scored.
Works when lopitpflocation: Correct calls per hidden gene reached 0.55 against 0.07 on shuffled data (skill 0.51, 286 scored). Run once on every gene, it called 2,000 genes with no known lopitpflocation; the top 5 are listed.
PF3D7_0606100RNA-binding protein, putative -- nucleus 1 (support 1)PF3D7_071960060S ribosomal protein L11a, putative -- ribosomes (support 1)PF3D7_100350040S ribosomal protein S20e, putative -- ribosomes (support 1)PF3D7_1017600conserved Plasmodium protein, unknown function -- nucleus 1 (support 1)PF3D7_1032100mRNA-decapping enzyme subunit 1, putative -- nucleus 1 (support 1)
31 · Call a gene only when independent strategies agree
kNN + logistic + network vote · label calls · T. gondii reliable · P. falciparum works when tuned
Where do measurement neighbours, a trained classifier and the networks give the same answer -- and how much more often is that answer right?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
████░░░░░░░░ 0.35 [0.17, 0.48] · chance 0.00 |
███░░░░░░░░░ 0.25 [-0.26, 0.55] · chance 0.00 |
| Reach coverage |
██████████░░ 0.87 [0.76, 0.96] |
██████████░░ 0.86 [0.77, 0.96] |
| Right calls accuracy |
████████░░░░ 0.63 [0.48, 0.76] |
████████░░░░ 0.67 [0.54, 0.84] |
| Fair across classes macro F1 |
██████░░░░░░ 0.50 [0.43, 0.61] |
███████░░░░░ 0.55 [0.51, 0.58] |
About this test
- What it does. It runs three methods built on different evidence: measurement neighbors, a trained classifier and network partners. A gene is called only where enough of them give the same label, two of three by default. Their failures are largely independent, so agreement should be right more often.
- How it is evaluated. A quarter of the label is hidden; each method trains on the visible genes. The deciding number is precision, the share of agreed calls on hidden genes that are correct. It is compared with agreement between the same methods trained on shuffled labels, and each method's own precision is shown.
- What failure looks like, and why. Agreed calls are no more precise than chance agreement on shuffled labels. The methods may share a failure, so they agree on wrong answers. Or agreement reaches very few genes, since a call needs every method to speak, and few genes may have network partners.
- What success looks like, and why. Agreed calls are clearly more precise than chance and than single methods. The result is a smaller, surer set of labels for unlabeled genes. Raising the agreement to three of three gives the smallest and most certain set.
T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: precision of calls on hidden genes was 0.662 against 0.668 on shuffled data (skill -0.02).
Works when compartment: Precision of calls on hidden genes reached 0.59 against 0.19 on shuffled data (skill 0.50, 629 scored). Run once on every gene, it called 1,787 genes with no known compartment; the top 5 are listed.
TGME49_225745hypothetical protein -- nucleus - chromatin (support 3)TGME49_266010phosphatidylinositol 3- and 4-kinase -- nucleus - chromatin (support 3)TGME49_286270hypothetical protein -- nucleus - chromatin (support 3)TGME49_271145hypothetical protein -- nucleus - chromatin (support 3)TGME49_253870hypothetical protein -- apicoplast (support 3)
P. falciparum. Fails when isexported: The signal is real but small: 0.938 beat shuffled data (0.936), but by 0.002, short of the 0.100 margin a PASS requires.
Works when lopitpflocation: Precision of calls on hidden genes reached 0.71 against 0.23 on shuffled data (skill 0.62, 303 scored). Run once on every gene, it called 2,000 genes with no known lopitpf_location; the top 5 are listed.
PF3D7_0100100erythrocyte membrane protein 1, PfEMP1 -- erythrocyte membrane (support 3)PF3D7_0100400rifin -- erythrocyte membrane (support 3)PF3D7_0100700Plasmodium exported protein, unknown function, fragment -- erythrocyte membrane (support 3)PF3D7_0101100exported protein family 4 -- erythrocyte membrane (support 3)PF3D7_0101600rifin -- erythrocyte membrane (support 3)
32 · Put the understudied genes first
kNN + logistic + network vote · label calls · T. gondii reliable · P. falciparum works when tuned
Which genes nobody has written about can the data say something trustworthy about?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
████░░░░░░░░ 0.32 [0.13, 0.46] · chance 0.00 |
██░░░░░░░░░░ 0.19 [-0.39, 0.55] · chance 0.00 |
| Reach coverage |
██████████░░ 0.87 [0.76, 0.97] |
██████████░░ 0.87 [0.78, 0.96] |
| Right calls accuracy |
███████░░░░░ 0.62 [0.48, 0.77] |
████████░░░░ 0.67 [0.56, 0.84] |
| Fair across classes macro F1 |
██████░░░░░░ 0.48 [0.41, 0.59] |
██████░░░░░░ 0.51 [0.45, 0.56] |
About this test
- What it does. It makes the agreed calls of strategy 31 only for genes with no focal or substantive paper. Candidates are ranked by how many methods agree times how little the gene has been written about. The top of the list is where the data can say something new.
- How it is evaluated. A quarter of the label is hidden, and agreed calls are scored only on hidden understudied genes. The deciding number is precision, the share of those calls that are right. It is compared with 5 runs on shuffled labels, never borrowing trust from well-studied genes.
- What failure looks like, and why. Precision on understudied genes sits at the shuffled level. The measurements were often designed around well-studied genes, so understudied ones may be poorly covered. The methods may also rarely agree on them, leaving too few calls to trust.
- What success looks like, and why. Agreed calls on understudied genes are clearly more precise than chance. Each top candidate is then a first hypothesis about a gene with no literature. The precision shown beside it is the only evidence for that call there is.
T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: precision of calls on hidden genes was 0.643 against 0.667 on shuffled data (skill -0.07).
Works when compartment: Precision of calls on hidden genes reached 0.57 against 0.16 on shuffled data (skill 0.48, 546 scored). Run once on every gene, it called 1,698 little-studied genes with no known compartment; the top 5 are listed.
TGME49_225745hypothetical protein -- nucleus - chromatin (support 0.546, 3 sources agree)TGME49_266010phosphatidylinositol 3- and 4-kinase -- nucleus - chromatin (support 0.546, 3 sources agree)TGME49_286270hypothetical protein -- nucleus - chromatin (support 0.546, 3 sources agree)TGME49_271145hypothetical protein -- nucleus - chromatin (support 0.546, 3 sources agree)TGME49_253870hypothetical protein -- apicoplast (support 0.546, 3 sources agree)
P. falciparum. Fails when isexported: One class dominates 'isexported', so guessing it on shuffled data already scores 0.972; the strategy's 0.972 is no better than that (skill 0.00), so what it reads does not separate the classes.
Works when lopitpflocation: Precision of calls on hidden genes reached 0.70 against 0.20 on shuffled data (skill 0.63, 215 scored). Run once on every gene, it called 2,000 little-studied genes with no known lopitpflocation; the top 5 are listed.
PF3D7_0100100erythrocyte membrane protein 1, PfEMP1 -- erythrocyte membrane (support 0.603, 3 sources agree)PF3D7_0100400rifin -- erythrocyte membrane (support 0.603, 3 sources agree)PF3D7_0100700Plasmodium exported protein, unknown function, fragment -- erythrocyte membrane (support 0.603, 3 sources agree)PF3D7_0101100exported protein family 4 -- erythrocyte membrane (support 0.603, 3 sources agree)PF3D7_0101600rifin -- erythrocyte membrane (support 0.603, 3 sources agree)
33 · Put every layer into one space and read a gene's neighbourhood
logistic edge model · ranking · T. gondii reliable · P. falciparum reliable
Which genes are the nearest neighbours of this one when every permitted network and the whole measurement table are combined into a single graph, and what evidence puts each of them there?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
███████░░░░░ 0.59 [0.58, 0.59] · chance 0.00 |
█████░░░░░░░ 0.39 [0.37, 0.41] · chance 0.00 |
| Reach recall @ top 10% |
██░░░░░░░░░░ 0.19 [0.18, 0.19] · chance 0.10 |
██░░░░░░░░░░ 0.17 [0.16, 0.19] · chance 0.10 |
| True ones ranked first AUROC |
█████████░░░ 0.79 [0.79, 0.79] · chance 0.50 |
████████░░░░ 0.71 [0.68, 0.73] · chance 0.50 |
| Clean top of the list AUPRC lift |
███░░░░░░░░░ x1.6 [x1.6, x1.6] · chance x1.0 |
███░░░░░░░░░ x1.4 [x1.4, x1.5] · chance x1.0 |
About this test
- What it does. It merges every permitted network layer and the measurement table into one graph and lists a gene's nearest neighbors. Each pair gets one probability, with the sources that support it. Nothing is written back into a layer, and each pair is marked as measured or inferred.
- How it is evaluated. One layer's edges are hidden by whole orthogroups and that layer is dropped from the inputs. The deciding number is how well hidden edges rank above non-pairs of similar connectivity (AUROC). It is compared with 10 rewirings that keep each gene's edge count but scramble which genes are linked.
- What failure looks like, and why. The score does not beat the rewired networks. The graph then mostly knows which genes are well connected, not who pairs with whom. The other layers and measurements may carry little about the hidden layer, especially if it is sparse.
- What success looks like, and why. Hidden edges are found clearly above the rewired level. Then the combined graph knows real pairings, not just busy genes. A gene's neighbor list, with the evidence behind each entry, can suggest partners that no single network records.
T. gondii. Fails when xlms: The signal is real but small: 0.837 beat shuffled data (0.791), but by 0.046, short of the 0.050 margin a PASS requires.
Works when coexpression: AUROC of hidden coexpression edges against degree-matched non-pairs reached 0.79 against 0.49 on shuffled data (skill 0.59, 3,349 scored). Run once on every gene, it ranked already-measured pairs at the top: all 500 pairs it listed are recorded by some layer, so this run proposes no unmeasured pair.
P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: AUROC of hidden coexpression edges against degree-matched non-pairs was 0.54 against 0.54 on shuffled data (skill 0.00). The test calls that a FAIL, which is the failure it exists to catch.
Works when coexpression: AUROC of hidden coexpression edges against degree-matched non-pairs reached 0.75 against 0.56 on shuffled data (skill 0.43, 5,277 scored). Run once on every gene, it ranked already-measured pairs at the top: all 500 pairs it listed are recorded by some layer, so this run proposes no unmeasured pair.
34 · Train on the networks and rank the edges they are missing
logistic / spectral embedding · ranking · T. gondii reliable · P. falciparum reliable
Which pairs of genes does the combined evidence imply although no measured layer records them, how strong is each claim, and how good is the model that makes it when it is scored against a degree-matched null rather than a random one?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
███████░░░░░ 0.59 [0.58, 0.59] · chance 0.00 |
█████░░░░░░░ 0.39 [0.37, 0.42] · chance 0.00 |
| Reach recall @ top 10% |
██░░░░░░░░░░ 0.19 [0.18, 0.19] · chance 0.10 |
██░░░░░░░░░░ 0.17 [0.16, 0.19] · chance 0.10 |
| True ones ranked first AUROC |
█████████░░░ 0.79 [0.79, 0.79] · chance 0.50 |
████████░░░░ 0.71 [0.68, 0.73] · chance 0.50 |
| Clean top of the list AUPRC lift |
███░░░░░░░░░ x1.6 [x1.6, x1.6] · chance x1.0 |
███░░░░░░░░░ x1.4 [x1.4, x1.5] · chance x1.0 |
About this test
- What it does. It trains one model per network layer, each blind to its own layer, and ranks gene pairs that the evidence implies but no layer records. Each gap comes with a calibrated probability and the sources behind it. An interpretable baseline ranks unless a learned embedding clearly beats it.
- How it is evaluated. The chosen layer's edges are hidden by orthogroup and removed from the inputs. The deciding number is how well hidden edges rank above non-pairs of similar connectivity (AUROC), against 10 rewirings that keep edge counts. The easier random-pair score and their gap are shown beside it.
- What failure looks like, and why. A score near 0.5, or at the rewired level, means the model learned only which genes are well connected. Crosslinks show a subtler failure: complexes are recovered, but rewired pairs inside a complex score almost as well, so single missing pairs cannot be named.
- What success looks like, and why. Hidden edges clear the rewired bar. The 'gaps' table then gives a ranked shortlist of pairs no experiment recorded, each with its supporting evidence. These are claims to test, for example with strategy 13 or 16, not measured edges.
T. gondii. Fails when xlms: The signal is real but small: 0.849 beat shuffled data (0.804), but by 0.045, short of the 0.050 margin a PASS requires.
Works when coexpression: AUROC of hidden coexpression edges against degree-matched non-pairs reached 0.79 against 0.49 on shuffled data (skill 0.59, 3,349 scored). Run once on every gene, it proposed 300 unmeasured pairs; the top 5 are listed.
TGME49_328400 + TGME49_324200aminotransferase, class V family protein / hypothetical protein -- predicted pair (probability 0.847)TGME49_200600 + TGME49_315855hypothetical protein / hypothetical protein -- predicted pair (probability 0.842)TGME49_204020 + TGME49_239100ribosomal protein RPL8 / ribosomal protein RPS7 -- predicted pair (probability 0.842)TGME49_210690 + TGME49_310490ribosomal protein RPS6 / ribosomal protein RPL27A -- predicted pair (probability 0.842)TGME49_327800 + TGME49_326700dynein-1-alpha heavy chain, flagellar inner arm I1 complex, putative / DNA-directed RNA polymerase, putative -- predicted pair (probability 0.84)
P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: AUROC of hidden coexpression edges against degree-matched non-pairs was 0.54 against 0.54 on shuffled data (skill 0.00). The test calls that a FAIL, which is the failure it exists to catch.
Works when coexpression: AUROC of hidden coexpression edges against degree-matched non-pairs reached 0.75 against 0.56 on shuffled data (skill 0.43, 5,277 scored). Run once on every gene, it proposed 300 unmeasured pairs; the top 5 are listed.
PF3D7_1479300 + PF3D7_API00200Plasmodium exported protein (PHISTa), unknown function, pseudogene / tRNA Histidine -- predicted pair (probability 1)PF3D7_API06000 + PF3D7_API06100tRNA Alanine / tRNA Asparagine -- predicted pair (probability 1)PF3D7_0413000 + PF3D7_0420800Plasmodium RNA of unknown function RUF6 / Plasmodium RNA of unknown function RUF6 -- predicted pair (probability 1)PF3D7_API06000 + PF3D7_API06200tRNA Alanine / tRNA Leucine -- predicted pair (probability 1)PF3D7_0631700 + PF3D7_0701300rifin, pseudogene / rifin, pseudogene -- predicted pair (probability 1)
35 · Call genes with a stated error rate
split conformal prediction · label calls · T. gondii reliable · P. falciparum reliable
Which genes can be given a label with a guaranteed error rate -- and for which does the data leave two or more labels equally possible?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
██████░░░░░░ 0.50 [0.21, 0.73] · chance 0.00 |
████████░░░░ 0.63 [0.44, 0.78] · chance 0.00 |
| Reach coverage |
██░░░░░░░░░░ 0.21 [0.06, 0.44] |
█████░░░░░░░ 0.39 [0.12, 0.68] |
| Right calls accuracy |
██░░░░░░░░░░ 0.17 [0.05, 0.38] |
████░░░░░░░░ 0.32 [0.10, 0.56] |
| Fair across classes macro F1 |
███░░░░░░░░░ 0.25 [0.09, 0.48] |
████░░░░░░░░ 0.36 [0.17, 0.55] |
About this test
- What it does. It turns a classifier's scores into a set of possible labels for each gene, promised to hold the true label for at least 1 - alpha of new genes (90% by default). A one-label set is a call with that promise behind it. A two-label set shows the data cannot tell those compartments apart.
- How it is evaluated. A quarter of the label is hidden by orthogroup, and the model trains and calibrates on the rest. The deciding number is set efficiency: how far the sets narrow the possible labels. It is compared with 10 shuffled-label runs, and coverage on hidden genes is checked against the promise.
- What failure looks like, and why. Sets are as wide as on shuffled labels, so the measurements barely narrow the answer. Coverage can also fall short if new genes do not resemble the calibration genes. A rare compartment may be under-covered unless per-class thresholds are used.
- What success looks like, and why. Sets are clearly narrower than chance while coverage meets the promise. Many unlabeled genes then get single-label calls with a stated error rate. The multi-label sets are findings too, pointing to compartments these data cannot separate.
T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: set efficiency: 1 - (mean set size - 1) / (classes - 1) was 0.188 against 0.173 on shuffled data (skill 0.02). Also, it could reach only 19% of the hidden genes, so most were never called.
Works when compartment: Set efficiency: 1 - (mean set size - 1) / (classes - 1) reached 0.79 against 0.30 on shuffled data (skill 0.70, 951 scored). Run once on every gene, it gave prediction sets to 14 genes with no known compartment; the top 5 are listed.
TGME49_256830SacI domain-containing protein -- apicoplast (set apicoplast (size 1))TGME49_309110tRNA-specific 2-thiouridylase MNMA -- apicoplast (set apicoplast (size 1))TGME49_2758303'-5' exonuclease, putative -- mitochondrion - soluble (set mitochondrion - soluble (size 1))TGME49_208040aldo-keto reductase -- apicoplast (set apicoplast (size 1))TGME49_271950hypothetical protein -- apicoplast (set apicoplast (size 1))
P. falciparum. Fails when isexported: 'isexported' is not encoded in what this strategy reads: set efficiency: 1 - (mean set size - 1) / (classes - 1) was 0.115 against 0.102 on shuffled data (skill 0.01). Also, it could reach only 11% of the hidden genes, so most were never called.
Works when lopitpflocation: Set efficiency: 1 - (mean set size - 1) / (classes - 1) reached 0.85 against 0.22 on shuffled data (skill 0.81, 395 scored). Run once on every gene, it gave prediction sets to 27 genes with no known lopitpflocation; the top 5 are listed.
PF3D7_0115400stevor -- erythrocyte membrane (set erythrocyte membrane (size 1))PF3D7_0306600ATP synthase-associated protein, putative -- mitochondrion (set mitochondrion (size 1))PF3D7_031280060S ribosomal protein L26, putative -- ribosomes (set ribosomes (size 1))PF3D7_0324100Pfmc-2TM Maurer's cleft two transmembrane protein -- erythrocyte membrane (set erythrocyte membrane (size 1))PF3D7_0400800stevor -- erythrocyte membrane (set erythrocyte membrane (size 1))
36 · Smooth the measurements along the networks, then classify
graph convolution + logistic regression · label calls · T. gondii reliable · P. falciparum reliable
Does a gene's label follow from its own measurements together with those of its network neighbours -- and how much does the model lean on each?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
█████░░░░░░░ 0.40 [0.16, 0.57] · chance 0.00 |
█████░░░░░░░ 0.46 [0.37, 0.54] · chance 0.00 |
| Reach coverage |
████████████ 1.00 [1.00, 1.00] |
████████████ 1.00 [1.00, 1.00] |
| Right calls accuracy |
████████░░░░ 0.63 [0.53, 0.74] · chance 0.36 |
█████████░░░ 0.71 [0.59, 0.84] · chance 0.45 |
| Fair across classes macro F1 |
███████░░░░░ 0.56 [0.48, 0.67] |
███████░░░░░ 0.62 [0.55, 0.69] |
About this test
- What it does. It averages each gene's measurements over its network partners, one and two steps out, and trains a logistic regression on the gene's own profile beside its neighborhood's. Only measurements travel along edges, never labels, so hidden labels cannot leak through neighbors.
- How it is evaluated. A quarter of the label is hidden by whole orthogroups and the model trains on the visible genes. The deciding number is the share of hidden genes called correctly. It is compared with the chance level for the same mix of predicted and true classes.
- What failure looks like, and why. Accuracy sits at the chance level for those class mixes. The label may not be encoded in the measurements or the networks. Genes with no edges fall back to their own profile, so sparse networks add little, while two steps on a dense network blur most genes together.
- What success looks like, and why. Calls clearly beat chance, and ideally beat strategy 19, which uses the gene's own profile only. The 'where the model looks' table then shows which compartments the networks encode better than the gene's own measurements. Unlabeled genes get calls.
T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: correct calls per hidden gene was 0.512 against 0.534 on shuffled data (skill -0.05).
Works when compartment: Correct calls per hidden gene reached 0.51 against 0.09 on shuffled data (skill 0.46, 951 scored). Run once on every gene, it called 2,000 genes with no known compartment; the top 5 are listed.
TGME49_219348SAG-related sequence SRS55M -- PM - peripheral 1 (support 0.993)TGME49_238500SAG-related sequence SRS22F -- PM - peripheral 1 (support 0.992)TGME49_252200palmitoyltransferase DHHC7 -- rhoptries 2 (support 0.991)TGME49_238850SAG-related sequence SRS22I -- PM - peripheral 1 (support 0.991)TGME49_320250SAG-related sequence SRS15A -- PM - peripheral 1 (support 0.99)
P. falciparum. Fails when isexported: The signal is real but small: 0.883 beat shuffled data (0.834), but by 0.048, short of the 0.050 margin a PASS requires.
Works when lopitpflocation: Correct calls per hidden gene reached 0.63 against 0.10 on shuffled data (skill 0.59, 395 scored). Run once on every gene, it called 2,000 genes with no known lopitpf_location; the top 5 are listed.
PF3D7_1040200stevor -- erythrocyte membrane (support 0.993)PF3D7_1400200rifin -- erythrocyte membrane (support 0.992)PF3D7_1200300rifin -- erythrocyte membrane (support 0.992)PF3D7_130280040S ribosomal protein S7, putative -- ribosomes (support 0.992)PF3D7_124270040S ribosomal protein S17, putative -- ribosomes (support 0.99)
37 · Let a random forest find what defines a label
random forest + permutation importance · label calls · T. gondii reliable · P. falciparum reliable
Which measurements, in which combinations and past which thresholds, define a label -- and which unlabelled genes carry that definition?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
█████░░░░░░░ 0.45 [0.22, 0.65] · chance 0.00 |
████░░░░░░░░ 0.36 [0.19, 0.53] · chance 0.00 |
| Reach coverage |
████████████ 1.00 [1.00, 1.00] |
████████████ 1.00 [1.00, 1.00] |
| Right calls accuracy |
████████░░░░ 0.70 [0.57, 0.83] · chance 0.44 |
█████████░░░ 0.74 [0.66, 0.86] · chance 0.51 |
| Fair across classes macro F1 |
███████░░░░░ 0.56 [0.46, 0.68] |
███████░░░░░ 0.55 [0.51, 0.59] |
About this test
- What it does. It trains a random forest, hundreds of voting decision trees, on every permitted measurement to call a label. Trees can capture thresholds and combinations, such as high expression AND a signal peptide. Measurements are ranked by how much accuracy drops when each is shuffled.
- How it is evaluated. A quarter of the label is hidden by whole orthogroups and the forest trains on the rest. The deciding number is the share of hidden genes called correctly. It is compared with the chance level for the same mix of predicted and true classes.
- What failure looks like, and why. Accuracy sits at chance: the measurements do not define the label. If the forest merely ties strategy 19, the signal is linear and the simpler model's weights explain it better. Noisy labels can also let trees memorize training genes when the leaf size is small.
- What success looks like, and why. Calls clearly beat chance and ideally beat strategy 19. Unlabeled genes then get calls, and 'what defines the label' names the measurements that truly matter. Shuffling any of them costs accuracy on genes the forest never saw.
T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: correct calls per hidden gene was 0.688 against 0.685 on shuffled data (skill 0.01).
Works when compartment: Correct calls per hidden gene reached 0.52 against 0.12 on shuffled data (skill 0.45, 951 scored). Run once on every gene, it called 2,000 genes with no known compartment; the top 5 are listed.
TGME49_263610tetratricopeptide repeat-containing protein -- nucleus - chromatin (support 0.672)TGME49_320250SAG-related sequence SRS15A -- PM - peripheral 1 (support 0.668)TGME49_221720hypothetical protein -- nucleus - chromatin (support 0.664)TGME49_286270hypothetical protein -- nucleus - chromatin (support 0.631)TGME49_224760SAG-related sequence SRS40E -- PM - peripheral 1 (support 0.623)
P. falciparum. Fails when isexported: The signal is too weak to tell from luck: 0.971 did not clear 0.976, what shuffled data reaches one time in twenty (skill 0.09).
Works when lopitpflocation: Correct calls per hidden gene reached 0.66 against 0.11 on shuffled data (skill 0.61, 395 scored). Run once on every gene, it called 2,000 genes with no known lopitpf_location; the top 5 are listed.
PF3D7_124270040S ribosomal protein S17, putative -- ribosomes (support 0.964)PF3D7_031280060S ribosomal protein L26, putative -- ribosomes (support 0.918)PF3D7_110540040S ribosomal protein S4, putative -- ribosomes (support 0.883)PF3D7_1300700rifin -- erythrocyte membrane (support 0.843)PF3D7_1400200rifin -- erythrocyte membrane (support 0.838)
38 · Learn how much to trust each kind of evidence
stacked logistic regression · label calls · T. gondii reliable · P. falciparum reliable
Given measurement neighbours, a linear model and the measured networks, how should their answers be combined for THIS label -- and what does the combination call?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
█████░░░░░░░ 0.41 [0.18, 0.57] · chance 0.00 |
█████░░░░░░░ 0.45 [0.38, 0.52] · chance 0.00 |
| Reach coverage |
████████████ 1.00 [1.00, 1.00] |
████████████ 1.00 [1.00, 1.00] |
| Right calls accuracy |
███████░░░░░ 0.62 [0.51, 0.74] · chance 0.34 |
████████░░░░ 0.70 [0.60, 0.82] · chance 0.43 |
| Fair across classes macro F1 |
███████░░░░░ 0.56 [0.47, 0.67] |
███████░░░░░ 0.62 [0.57, 0.67] |
About this test
- What it does. It learns how much to trust each kind of evidence for this label: measurement neighbors, a logistic regression and network partners. Each predicts labeled genes out of fold, and a second model learns which to believe for which class. It then calls unlabeled genes.
- How it is evaluated. A quarter of the label is hidden by whole orthogroups; the base predictions use only visible genes, so no hidden label reaches either level. The deciding number is the share of hidden genes called correctly, compared with the chance level for the same class mixes.
- What failure looks like, and why. Accuracy sits at chance, meaning none of the three evidence types carries the label. If it only ties strategies 07, 19 or 31, combining adds little for this label. Few folds give the second model noisier training data.
- What success looks like, and why. Calls clearly beat chance, and usually beat the single methods. 'Trust by evidence' then answers which experiments matter for this label, for example networks for complexes but not secreted proteins. Unlabeled genes get the combined calls.
T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: correct calls per hidden gene was 0.362 against 0.395 on shuffled data (skill -0.05).
Works when compartment: Correct calls per hidden gene reached 0.50 against 0.08 on shuffled data (skill 0.46, 951 scored). Run once on every gene, it called 2,000 genes with no known compartment; the top 5 are listed.
TGME49_224770SAG-related sequence SRS40D -- PM - peripheral 1 (support 0.998)TGME49_320250SAG-related sequence SRS15A -- PM - peripheral 1 (support 0.997)TGME49_224760SAG-related sequence SRS40E -- PM - peripheral 1 (support 0.996)TGME49_273110SAG-related sequence SRS30D -- PM - peripheral 1 (support 0.996)TGME49_238850SAG-related sequence SRS22I -- PM - peripheral 1 (support 0.995)
P. falciparum. Fails when isexported: The signal is real but small: 0.927 beat shuffled data (0.878), but by 0.049, short of the 0.050 margin a PASS requires.
Works when lopitpflocation: Correct calls per hidden gene reached 0.61 against 0.08 on shuffled data (skill 0.57, 395 scored). Run once on every gene, it called 2,000 genes with no known lopitpf_location; the top 5 are listed.
PF3D7_0306600ATP synthase-associated protein, putative -- mitochondrion (support 0.986)PF3D7_130280040S ribosomal protein S7, putative -- ribosomes (support 0.985)PF3D7_124270040S ribosomal protein S17, putative -- ribosomes (support 0.985)PF3D7_0713100Pfmc-2TM Maurer's cleft two transmembrane protein -- erythrocyte membrane (support 0.981)PF3D7_031280060S ribosomal protein L26, putative -- ribosomes (support 0.978)
39 · Predict a value with an interval that holds
gradient boosting / ridge + split conformal · values · T. gondii weak · P. falciparum reliable
For a gene never measured, what value is expected -- and within what range, with a guaranteed chance of containing the truth?
| T. gondii | P. falciparum | |
|---|---|---|
| Better than chance skill |
███████░░░░░ 0.55 [0.04, 0.91] · chance 0.00 |
██████░░░░░░ 0.50 [0.22, 0.84] · chance 0.00 |
| Reach coverage |
████████████ 1.00 [1.00, 1.00] |
████████████ 1.00 [1.00, 1.00] |
| Order predicted Spearman rho |
███████░░░░░ 0.55 [0.04, 0.91] · chance 0.00 |
██████░░░░░░ 0.50 [0.22, 0.84] · chance 0.00 |
| Variance explained R-squared, out of sample |
█████░░░░░░░ 0.41 [-0.13, 0.87] · chance 0.00 |
░░░░░░░░░░░░ -3.49 [-14.82, 0.71] · chance 0.00 |
About this test
- What it does. It predicts a measurement, such as a fitness score, from the other permitted measurements. Each prediction gets an interval in the measurement's own units, promised to contain the true value for at least 1 - alpha of new genes. Measured genes outside their interval are listed as surprises.
- How it is evaluated. A fifth of the measured values are hidden, and the model trains and calibrates on the rest. The deciding number is the rank correlation between predicted and hidden values, compared with the chance level. The share of hidden values inside their intervals is checked against the promise.
- What failure looks like, and why. Predictions track hidden values no better than chance, and intervals are as wide as the measurement's spread. The other measurements then know little about this one. Leaving out the same kind of screen can remove the only related evidence.
- What success looks like, and why. Predictions track hidden values and intervals are narrower than the spread. Unmeasured genes get a value worth acting on, with a stated range. Measured genes far outside their interval are surprises the rest of the data cannot explain.
T. gondii. Fails when fitinvivoPE: The signal is too weak to tell from luck: 0.038 did not clear 0.043, what shuffled data reaches one time in twenty (skill 0.04).
Works when fitinvitrohff: Rank correlation of predicted and hidden values reached 0.72 against 0.00 on shuffled data (skill 0.72, 1,465 scored). Run once on every gene, it predicted 815 unmeasured genes, each with a 90% interval; the top 5 are listed.
TGME49_219790pre-mRNA processing factor PRP3 -- predicted -3.96 (interval -6.54 to -1.38)TGME49_316400alpha tubulin TUBA1 -- predicted -3.86 (interval -6.44 to -1.27)TGME49_310430Hsp90 domain-containing protein -- predicted -3.6 (interval -6.18 to -1.02)TGME49_224350aminopeptidase N2 -- predicted -3.54 (interval -6.12 to -0.956)TGME49_285660DEAD/DEAH box helicase domain-containing protein -- predicted -3.52 (interval -6.1 to -0.941)
P. falciparum. Fails when exprschizont: The signal is real but small: 0.074 beat shuffled data (0.000), but by 0.074, short of the 0.100 margin a PASS requires.
Works when piggybacmis: Rank correlation of predicted and hidden values reached 0.46 against 0.00 on shuffled data (skill 0.46, 1,077 scored). Run once on every gene, it predicted 335 unmeasured genes, each with a 90% interval; the top 5 are listed.
PF3D7_1342700DNA-directed RNA polymerases I, II, and III subunit RPABC4, putative -- predicted 0.163 (interval -0.41 to 0.736)PF3D7_API04600hypothetical protein -- predicted 0.195 (interval -0.378 to 0.768)PF3D7_0114750Plasmodium exported protein, unknown function, fragment -- predicted 0.244 (interval -0.329 to 0.817)PF3D7_091820050S ribosomal protein L3, apicoplast, putative -- predicted 0.25 (interval -0.323 to 0.823)PF3D7_1105100histone H2B -- predicted 0.261 (interval -0.312 to 0.834)