Strategy cards

Generated by scripts/build_strategy_examples.py --docs-only; do not edit by hand.

Each strategy shows the same four bars, in the same places: Better than chance (skill: 0 is the same procedure on shuffled data, 1 is perfect), Reach (how much of the question it can speak to), and the two metrics of its task a biologist asks about first, in plain words. Values are the mean over the calibration's held-out tests at default settings, with the 95% interval and the chance level.

Task Reach means Bar 3 Bar 4
label calls coverage Right calls (accuracy) Fair across classes (macro F1)
ranking recall @ top 10% -- A ranking scores every candidate, so coverage is always complete; its analogue is how many of the true ones a short list reaches. True ones ranked first (AUROC) Clean top of the list (AUPRC lift)
set retrieval genes returned -- A set strategy speaks about the genes it returns, so its reach is how many it returned. Returned genes that are real (precision) Members found (recall)
cluster recovery 1 - unclustered share -- Clustering has no coverage; its analogue is the share of genes it placed in a cluster at all. Label falls out as a cluster (weighted F1) Partition agreement (adjusted Rand index)
values coverage Order predicted (Spearman rho) Variance explained (R-squared, out of sample)
replication findings made -- A replication test has no coverage; its reach is how many findings it made to check. Findings that hold (replication rate) Beyond chance (replication lift)

Below the bars, About this test says what the strategy does, how it is evaluated, what failure looks like and what success looks like. The worked examples are real calibration runs: the failure is the target it failed on most often at default settings (or, where it never failed on real data, its self-test on a table of random labels and edges), the success its best pass, re-run once to list what it says about genes without a known label.

01 · Hold out a category and search for a map that finds it

UMAP + HDBSCAN · cluster recovery · T. gondii weak · P. falciparum weak

Is there a combination of measurements and map settings under which a label nobody showed the map falls out as clusters -- and which unlabelled genes land in them?

T. gondii P. falciparum
Better than chance
skill
█░░░░░░░░░░░ 0.07 [0.03, 0.11] · chance 0.00 █░░░░░░░░░░░ 0.09 [0.01, 0.19] · chance 0.00
Reach
1 - unclustered share
███████░░░░░ 0.55 [0.30, 0.81] █████████░░░ 0.78 [0.46, 0.99]
Label falls out as a cluster
weighted F1
█████░░░░░░░ 0.38 [0.19, 0.57] · chance 0.32 ██████░░░░░░ 0.54 [0.23, 0.81] · chance 0.50
Partition agreement
adjusted Rand index
░░░░░░░░░░░░ 0.03 [0.01, 0.06] · chance 0.00 █░░░░░░░░░░░ 0.08 [0.01, 0.18] · chance 0.00

About this test

T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: F1 of hidden genes in the cluster chosen for their label on known genes was 0.684 against 0.683 on shuffled data (skill 0.00).
Works when compartment: F1 of hidden genes in the cluster chosen for their label on known genes reached 0.12 against 0.04 on shuffled data (skill 0.08, 577 scored). Run once on every gene, it called 63 genes with no known compartment; the top 5 are listed.

P. falciparum. Fails when pbtransferredphenotype: 'pbtransferredphenotype' is not encoded in what this strategy reads: F1 of hidden genes in the cluster chosen for their label on known genes was 0.535 against 0.535 on shuffled data (skill 0.00).
Works when lopitpflocation: F1 of hidden genes in the cluster chosen for their label on known genes reached 0.16 against 0.04 on shuffled data (skill 0.13, 395 scored). Run once on every gene, it called 152 genes with no known lopitpflocation; the top 5 are listed.

02 · Find the map where your gene list is one cluster

UMAP + HDBSCAN · set retrieval · T. gondii weak · P. falciparum reliable

Under some combination of measurements and settings, do the genes on my list fall into a single cluster -- and what else is in it?

T. gondii P. falciparum
Better than chance
skill
░░░░░░░░░░░░ 0.03 [0.00, 0.04] · chance 0.00 ███░░░░░░░░░ 0.21 [0.09, 0.35] · chance 0.00
Reach
genes returned
███████████░ 433 [166, 802] ████████░░░░ 102 [58, 133]
Returned genes that are real
precision
░░░░░░░░░░░░ 0.04 [0.03, 0.05] · chance 0.02 ██░░░░░░░░░░ 0.16 [0.08, 0.25] · chance 0.01
Members found
recall
███░░░░░░░░░ 0.22 [0.12, 0.34] · chance 0.09 █████░░░░░░░ 0.42 [0.17, 0.69] · chance 0.04

About this test

T. gondii. Fails when tachyzoite (stageenrichedderived): 'tachyzoite (stageenrichedderived)' is not encoded in what this strategy reads: F1 of the hidden members against the best cluster's other genes was 0.038 against 0.047 on shuffled data (skill -0.01). Also, the set was too wide: 890 genes returned, only 2% of them members.
Works when PM - integral (compartment): F1 of the hidden members against the best cluster's other genes reached 0.08 against 0.02 on shuffled data (skill 0.06, 40 scored). Run once on every gene, it placed 85 genes with no place on the list in the cluster that holds the list; the top 5 are listed.

P. falciparum. Fails when gametocyte (stageenrichedderived): The signal is real but small: 0.040 beat shuffled data (0.013), but by 0.027, short of the 0.050 margin a PASS requires. Also, the set was too wide: 30 genes returned, only 3% of them members; only 20 hidden items could be scored.
Works when cytosol (lopitpflocation): F1 of the hidden members against the best cluster's other genes reached 0.25 against 0.01 on shuffled data (skill 0.25, 37 scored). Run once on every gene, it placed 100 genes with no place on the list in the cluster that holds the list; the top 5 are listed.

03 · Ask which categories the data can rediscover

UMAP + neighbour AUROC · ranking · T. gondii reliable · P. falciparum reliable

Of all the categories of a label, which ones do the measurements actually encode -- and which would no map, however tuned, ever find?

T. gondii P. falciparum
Better than chance
skill
██████░░░░░░ 0.50 [0.22, 0.68] · chance 0.00 ████████░░░░ 0.69 [0.56, 0.79] · chance 0.00
Reach
recall @ top 10%
████░░░░░░░░ 0.35 [0.17, 0.54] · chance 0.10 █████░░░░░░░ 0.39 [0.18, 0.60] · chance 0.10
True ones ranked first
AUROC
█████████░░░ 0.75 [0.61, 0.84] · chance 0.50 ██████████░░ 0.84 [0.78, 0.90] · chance 0.50
Clean top of the list
AUPRC lift
█████████░░░ x9.6 [x1.6, x19] · chance x1.0 ███████░░░░░ x5.2 [x1.5, x9.5] · chance x1.0

About this test

T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: hidden-gene AUROC of the categories the atlas ranks in its top half was 0.463 against 0.495 on shuffled data (skill -0.06).
Works when compartment: Hidden-gene AUROC of the categories the atlas ranks in its top half reached 0.87 against 0.50 on shuffled data (skill 0.73, 507 scored). Run once on every gene, it ranked 26 categories by how well the data recovers them; the top 5 are listed.

P. falciparum. Fails when stageenrichedderived: The signal is too weak to tell from luck: 0.739 did not clear 0.823, what shuffled data reaches one time in twenty (skill 0.45).
Works when lopitpflocation: Hidden-gene AUROC of the categories the atlas ranks in its top half reached 0.85 against 0.51 on shuffled data (skill 0.70, 417 scored). Run once on every gene, it ranked 24 categories by how well the data recovers them; the top 5 are listed.

04 · Keep only the modules that survive the whole walk

UMAP + HDBSCAN co-clustering · cluster recovery · T. gondii weak · P. falciparum weak

Which groups of genes stay together whatever map settings are chosen -- the structure that is in the data rather than in one lucky configuration?

T. gondii P. falciparum
Better than chance
skill
█░░░░░░░░░░░ 0.07 [0.01, 0.13] · chance 0.00 ██░░░░░░░░░░ 0.15 [0.03, 0.28] · chance 0.00
Reach
1 - unclustered share
███████████░ 0.89 [0.84, 0.93] █████████░░░ 0.79 [0.58, 1.00]
Label falls out as a cluster
weighted F1
█████░░░░░░░ 0.39 [0.27, 0.49] · chance 0.34 ██████░░░░░░ 0.50 [0.23, 0.78] · chance 0.45
Partition agreement
adjusted Rand index
█░░░░░░░░░░░ 0.04 [0.02, 0.07] · chance 0.00 █░░░░░░░░░░░ 0.11 [0.01, 0.25] · chance 0.00

About this test

T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: F1 of hidden genes in the module chosen for their label on known genes was 0.494 against 0.509 on shuffled data (skill -0.03).
Works when compartment: F1 of hidden genes in the module chosen for their label on known genes reached 0.22 against 0.12 on shuffled data (skill 0.11, 302 scored). Run once on every gene, it placed 487 genes with no known compartment in stable modules; the top 5 are listed.

P. falciparum. Fails when stageenrichedderived: 'stageenrichedderived' is not encoded in what this strategy reads: F1 of hidden genes in the module chosen for their label on known genes was 0.607 against 0.607 on shuffled data (skill 0.00).
Works when lopitpflocation: F1 of hidden genes in the module chosen for their label on known genes reached 0.28 against 0.12 on shuffled data (skill 0.19, 255 scored). Run once on every gene, it placed 852 genes with no known lopitpflocation in stable modules; the top 5 are listed.

05 · Tune a map without labels, then read what it encodes

UMAP + HDBSCAN, chi-square / Kruskal-Wallis · replication · T. gondii reliable · P. falciparum reliable

If I build a map from one kind of evidence only -- expression, say -- and tune it for structure alone, which OTHER measurements do its clusters turn out to separate?

T. gondii P. falciparum
Better than chance
skill
███████████░ 0.95 [0.88, 1.00] · chance 0.00 ███████████░ 0.88 [0.82, 0.93] · chance 0.00
Reach
findings made
███████░░░░░ 53 [39, 63] ███████░░░░░ 54 [52, 56]
Findings that hold
replication rate
███████████░ 0.95 [0.89, 1.00] · chance 0.06 ███████████░ 0.88 [0.83, 0.94] · chance 0.05
Beyond chance
replication lift
███████████░ x21 [x13, x28] · chance x1.0 ██████████░░ x18 [x14, x24] · chance x1.0

About this test

T. gondii. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find. It made too little to score, and the test declined to judge (only 0 findings on the first half, 3 needed).
Works when transcription: Share of first-half findings that replicate on the second half reached 1.00 against 0.04 on shuffled data (skill 1.00, 53 scored). Run once on every gene, it tested 131 held-out features the map was never shown; the top 5 are listed.

P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find. It made too little to score, and the test declined to judge (only 0 findings on the first half, 3 needed).
Works when transcription: Share of first-half findings that replicate on the second half reached 0.98 against 0.08 on shuffled data (skill 0.98, 54 scored). Run once on every gene, it tested 128 held-out features the map was never shown; the top 5 are listed.

06 · Find which kind of evidence carries a label

kNN ablation · label calls · T. gondii reliable · P. falciparum reliable

Which measurements actually carry the information about this label -- and which are redundant with others or irrelevant to it?

T. gondii P. falciparum
Better than chance
skill
██░░░░░░░░░░ 0.14 [0.06, 0.25] · chance 0.00 ██░░░░░░░░░░ 0.13 [0.07, 0.19] · chance 0.00
Reach
coverage
████████████ 1.00 [1.00, 1.00] ████████████ 1.00 [1.00, 1.00]
Right calls
accuracy
███████░░░░░ 0.60 [0.44, 0.77] · chance 0.54 ████████░░░░ 0.64 [0.44, 0.85] · chance 0.57
Fair across classes
macro F1
█████░░░░░░░ 0.39 [0.28, 0.54] █████░░░░░░░ 0.42 [0.31, 0.49]

About this test

T. gondii. Fails when dtm_class: The signal is real but small: 0.754 beat shuffled data (0.744), but by 0.010, short of the 0.020 margin a PASS requires.
Works when compartment: Hidden accuracy of the evidence ranked first (transcription) reached 0.34 against 0.23 on shuffled data (skill 0.14, 951 scored). Run once on every gene, it ranked 14 kinds of evidence by what each carries alone; the top 5 are listed.

P. falciparum. Fails when isexported: The signal is too weak to tell from luck: 0.970 did not clear 0.970, what shuffled data reaches one time in twenty (skill 0.04).
Works when lopit
pf_location: Hidden accuracy of the evidence ranked first (expr) reached 0.38 against 0.21 on shuffled data (skill 0.22, 395 scored). Run once on every gene, it ranked 48 kinds of evidence by what each carries alone; the top 5 are listed.

07 · Call a gene by the genes that behave like it

kNN · label calls · T. gondii reliable · P. falciparum weak

For a gene with no label, what label do the genes most similar to it across every permitted measurement carry?

T. gondii P. falciparum
Better than chance
skill
███░░░░░░░░░ 0.22 [0.08, 0.36] · chance 0.00 █░░░░░░░░░░░ 0.08 [-0.37, 0.42] · chance 0.00
Reach
coverage
███████████░ 0.93 [0.84, 1.00] ███████████░ 0.96 [0.87, 1.00]
Right calls
accuracy
███████░░░░░ 0.62 [0.47, 0.76] · chance 0.48 ████████░░░░ 0.68 [0.53, 0.85] · chance 0.53
Fair across classes
macro F1
█████░░░░░░░ 0.44 [0.34, 0.57] ██████░░░░░░ 0.51 [0.47, 0.55]

About this test

T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: correct calls per hidden gene was 0.675 against 0.680 on shuffled data (skill -0.02).
Works when compartment: Correct calls per hidden gene reached 0.37 against 0.05 on shuffled data (skill 0.34, 951 scored). Run once on every gene, it called 2,000 genes with no known compartment; the top 5 are listed.

P. falciparum. Fails when isexported: One class dominates 'isexported', so guessing it on shuffled data already scores 0.936; the strategy's 0.937 is no better than that (skill 0.01), so what it reads does not separate the classes.
Works when lopitpflocation: Correct calls per hidden gene reached 0.52 against 0.06 on shuffled data (skill 0.49, 395 scored). Run once on every gene, it called 2,000 genes with no known lopitpflocation; the top 5 are listed.

08 · Call a gene by its neighbours on the map

UMAP + kNN · label calls · T. gondii weak · P. falciparum weak

On a map built without the label, which label do a gene's nearest placed neighbours carry?

T. gondii P. falciparum
Better than chance
skill
██░░░░░░░░░░ 0.13 [0.01, 0.24] · chance 0.00 ░░░░░░░░░░░░ -0.00 [-0.53, 0.29] · chance 0.00
Reach
coverage
███████████░ 0.91 [0.81, 1.00] ███████████░ 0.94 [0.83, 1.00]
Right calls
accuracy
███████░░░░░ 0.56 [0.38, 0.74] · chance 0.47 ████████░░░░ 0.64 [0.42, 0.84] · chance 0.53
Fair across classes
macro F1
████░░░░░░░░ 0.37 [0.26, 0.48] ██████░░░░░░ 0.47 [0.36, 0.57]

About this test

T. gondii. Fails when dtmclass: 'dtmclass' is not encoded in what this strategy reads: correct calls per hidden gene was 0.751 against 0.753 on shuffled data (skill -0.01).
Works when compartment: Correct calls per hidden gene reached 0.27 against 0.05 on shuffled data (skill 0.23, 513 scored). Run once on every gene, it called 610 genes with no known compartment; the top 5 are listed.

P. falciparum. Fails when isexported: One class dominates 'isexported', so guessing it on shuffled data already scores 0.955; the strategy's 0.953 is no better than that (skill -0.04), so what it reads does not separate the classes.
Works when lopitpflocation: Correct calls per hidden gene reached 0.35 against 0.08 on shuffled data (skill 0.30, 395 scored). Run once on every gene, it called 1,470 genes with no known lopitpflocation; the top 5 are listed.

09 · Name a cluster by the label it is enriched for

UMAP + HDBSCAN, hypergeometric · label calls · T. gondii reliable · P. falciparum reliable

Which clusters of a label-blind map hold one label far more often than chance, and what does that make of their unlabelled members?

T. gondii P. falciparum
Better than chance
skill
████░░░░░░░░ 0.33 [0.25, 0.42] · chance 0.00 █████░░░░░░░ 0.42 [0.33, 0.60] · chance 0.00
Reach
coverage
██░░░░░░░░░░ 0.18 [0.09, 0.27] ██░░░░░░░░░░ 0.14 [0.07, 0.22]
Right calls
accuracy
█░░░░░░░░░░░ 0.06 [0.03, 0.08] █░░░░░░░░░░░ 0.06 [0.02, 0.09]
Fair across classes
macro F1
█░░░░░░░░░░░ 0.10 [0.06, 0.15] ██░░░░░░░░░░ 0.17 [0.11, 0.22]

About this test

T. gondii. Fails when lopit_unified: The signal is real but small: 0.050 beat shuffled data (0.000), but by 0.050, short of the 0.100 margin a PASS requires. Also, it could reach only 44% of the hidden genes, so most were never called.
Works when compartment: Precision of calls on hidden genes reached 0.32 against 0.00 on shuffled data (skill 0.32, 107 scored). Run once on every gene, it called 102 genes with no known compartment; the top 5 are listed.

P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find. It made too little to score, and the test declined to judge.
Works when lopitpflocation: Precision of calls on hidden genes reached 0.51 against 0.01 on shuffled data (skill 0.51, 70 scored). Run once on every gene, it called 400 genes with no known lopitpflocation; the top 5 are listed.

10 · Find genes whose label their neighbours contradict

kNN + network neighbours · ranking · T. gondii reliable · P. falciparum reliable

Which labelled genes sit among genes that almost all carry a different label -- possible mislabels, dual-localized or moonlighting proteins?

T. gondii P. falciparum
Better than chance
skill
████████░░░░ 0.64 [0.49, 0.77] · chance 0.00 █████████░░░ 0.74 [0.57, 0.92] · chance 0.00
Reach
recall @ top 10%
██████░░░░░░ 0.48 [0.35, 0.61] · chance 0.10 ███████░░░░░ 0.61 [0.37, 0.82] · chance 0.10
True ones ranked first
AUROC
██████████░░ 0.82 [0.74, 0.88] · chance 0.50 ██████████░░ 0.87 [0.79, 0.94] · chance 0.50
Clean top of the list
AUPRC lift
███████░░░░░ x5.0 [x3.5, x6.7] · chance x1.0 ████████░░░░ x8.8 [x3.7, x15] · chance x1.0

About this test

T. gondii. Fails when screenanyphenotype: The signal is too weak to tell from luck: 0.623 did not clear 0.632, what shuffled data reaches one time in twenty (skill 0.19). Also, only 15 hidden items could be scored.
Works when compartment: AUROC of surprise for the swapped labels reached 0.83 against 0.50 on shuffled data (skill 0.65, 190 scored). Run once on every gene, it flagged 200 labeled genes whose label looks wrong; the top 5 are listed.

P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: AUROC of surprise for the swapped labels was 0.51 against 0.48 on shuffled data (skill 0.06). The test calls that a FAIL, which is the failure it exists to catch.
Works when lopitpflocation: AUROC of surprise for the swapped labels reached 0.87 against 0.50 on shuffled data (skill 0.75, 79 scored). Run once on every gene, it flagged 200 labeled genes whose label looks wrong; the top 5 are listed.

11 · Diffuse a label across one measured network

random walk with restart · label calls · T. gondii reliable · P. falciparum reliable

If labels flow along the edges of one kind of measured relationship, where do they end up -- and how much of a label does that relationship carry?

T. gondii P. falciparum
Better than chance
skill
█░░░░░░░░░░░ 0.12 [0.08, 0.17] · chance 0.00 ███░░░░░░░░░ 0.22 [0.14, 0.33] · chance 0.00
Reach
coverage
██████████░░ 0.80 [0.78, 0.82] ███████████░ 0.95 [0.95, 0.96]
Right calls
accuracy
███░░░░░░░░░ 0.29 [0.18, 0.40] · chance 0.19 █████░░░░░░░ 0.45 [0.17, 0.72] · chance 0.32
Fair across classes
macro F1
███░░░░░░░░░ 0.28 [0.17, 0.41] ████░░░░░░░░ 0.38 [0.17, 0.53]

About this test

T. gondii. Fails when screenanyphenotype: The signal is too weak to tell from luck: 0.450 did not clear 0.482, what shuffled data reaches one time in twenty (skill 0.06).
Works when compartment: Correct calls per hidden gene reached 0.14 against 0.03 on shuffled data (skill 0.11, 951 scored). Run once on every gene, it called 2,000 genes with no known compartment; the top 5 are listed.

P. falciparum. Fails when isexported: The signal is real but small: 0.650 beat shuffled data (0.605), but by 0.046, short of the 0.050 margin a PASS requires.
Works when lopit
pflocation: Correct calls per hidden gene reached 0.18 against 0.05 on shuffled data (skill 0.14, 395 scored). Run once on every gene, it called 2,000 genes with no known lopitpf_location; the top 5 are listed.

12 · Let every network vote, weighted by what it has earned

chance-weighted ensemble vote · label calls · T. gondii weak · P. falciparum reliable

If every measured relationship and the measurements themselves vote on a gene's label, each weighted by how good it has proven to be, what is the verdict?

T. gondii P. falciparum
Better than chance
skill
████░░░░░░░░ 0.30 [-0.03, 0.54] · chance 0.00 ██████░░░░░░ 0.51 [0.41, 0.65] · chance 0.00
Reach
coverage
███████████░ 0.90 [0.69, 1.00] ███████████░ 0.94 [0.81, 1.00]
Right calls
accuracy
███████░░░░░ 0.58 [0.40, 0.75] · chance 0.38 ████████░░░░ 0.65 [0.53, 0.78] · chance 0.32
Fair across classes
macro F1
█████░░░░░░░ 0.44 [0.33, 0.57] ██████░░░░░░ 0.50 [0.44, 0.56]

About this test

T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: correct calls per hidden gene was 0.463 against 0.528 on shuffled data (skill -0.14).
Works when compartment: Correct calls per hidden gene reached 0.42 against 0.13 on shuffled data (skill 0.34, 951 scored). Run once on every gene, it called 2,000 genes with no known compartment; the top 5 are listed.

P. falciparum. Fails when stageenrichedderived: The signal is too weak to tell from luck: 0.691 did not clear 0.706, what shuffled data reaches one time in twenty (skill 0.47).
Works when lopitpflocation: Correct calls per hidden gene reached 0.55 against 0.10 on shuffled data (skill 0.50, 395 scored). Run once on every gene, it called 2,000 genes with no known lopitpflocation; the top 5 are listed.

13 · Place a protein by the proteins it physically touches

weighted partner vote · label calls · T. gondii reliable · P. falciparum weak

For a protein crosslinked to or pulled down with labelled proteins, what does its physical company say about where it lives and what it joins?

T. gondii P. falciparum
Better than chance
skill
███░░░░░░░░░ 0.25 [0.08, 0.43] · chance 0.00 ███░░░░░░░░░ 0.22 [0.01, 0.47] · chance 0.00
Reach
coverage
███████░░░░░ 0.57 [0.28, 0.84] █████████░░░ 0.71 [0.61, 0.80]
Right calls
accuracy
█████░░░░░░░ 0.39 [0.17, 0.61] · chance 0.19 ██████░░░░░░ 0.54 [0.33, 0.75] · chance 0.37
Fair across classes
macro F1
████░░░░░░░░ 0.34 [0.17, 0.52] █████░░░░░░░ 0.45 [0.35, 0.56]

About this test

T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: correct calls per hidden gene was 0.062 against 0.047 on shuffled data (skill 0.02). Also, it could reach only 6% of the hidden genes, so most were never called; only 16 hidden items could be scored.
Works when compartment: Correct calls per hidden gene reached 0.49 against 0.09 on shuffled data (skill 0.45, 330 scored). Run once on every gene, it called 271 genes with no known compartment; the top 5 are listed.

P. falciparum. Fails when pbtransferredphenotype: The signal is too weak to tell from luck: 0.350 did not clear 0.503, what shuffled data reaches one time in twenty (skill 0.08). Also, only 20 hidden items could be scored.
Works when lopitpflocation: Correct calls per hidden gene reached 0.56 against 0.06 on shuffled data (skill 0.53, 18 scored). Run once on every gene, it called 15 genes with no known lopitpflocation; the top 5 are listed.

14 · Annotate function through shared fold

TM-score-weighted vote · label calls · T. gondii reliable · P. falciparum reliable

What does a protein's fold -- its structural similarity to annotated proteins -- say about its enzymatic class or domain family, even without sequence homology?

T. gondii P. falciparum
Better than chance
skill
████████░░░░ 0.64 [0.60, 0.66] · chance 0.00 ███████░░░░░ 0.59 [0.54, 0.61] · chance 0.00
Reach
coverage
█████████░░░ 0.78 [0.75, 0.80] █████████░░░ 0.75 [0.71, 0.77]
Right calls
accuracy
█████████░░░ 0.73 [0.70, 0.76] · chance 0.27 ████████░░░░ 0.69 [0.66, 0.72] · chance 0.26
Fair across classes
macro F1
█████████░░░ 0.75 [0.75, 0.76] ███████░░░░░ 0.60 [0.54, 0.67]

About this test

T. gondii. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: correct calls per hidden gene was 0.21 against 0.16 on shuffled data (skill 0.07). It still passed: a false alarm on random data, so read its passes with care.
Works when --: Correct calls per hidden gene reached 0.75 against 0.26 on shuffled data (skill 0.67, 165 scored). Run once on every gene, it called 337 genes with no known ec_number; the top 5 are listed.

P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: correct calls per hidden gene was 0.21 against 0.16 on shuffled data (skill 0.07). It still passed: a false alarm on random data, so read its passes with care.
Works when --: Correct calls per hidden gene reached 0.72 against 0.27 on shuffled data (skill 0.62, 164 scored). Run once on every gene, it called 172 genes with no known ec_number; the top 5 are listed.

15 · Find the communities several networks agree on

modularity + Louvain consensus · cluster recovery · T. gondii weak · P. falciparum weak

Which groups of genes are communities in more than one kind of measured relationship at once -- co-expressed AND co-fit AND crosslinked?

T. gondii P. falciparum
Better than chance
skill
█░░░░░░░░░░░ 0.05 [0.02, 0.08] · chance 0.00 ░░░░░░░░░░░░ 0.03 [-0.03, 0.09] · chance 0.00
Reach
1 - unclustered share
████████████ 1.00 [1.00, 1.00] ████████████ 1.00 [1.00, 1.00]
Label falls out as a cluster
weighted F1
████░░░░░░░░ 0.35 [0.24, 0.46] · chance 0.30 ███░░░░░░░░░ 0.25 [0.14, 0.35] · chance 0.22
Partition agreement
adjusted Rand index
░░░░░░░░░░░░ 0.02 [0.01, 0.03] · chance 0.00 ░░░░░░░░░░░░ 0.02 [-0.00, 0.04] · chance 0.00

About this test

T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: F1 of hidden genes in the community chosen for their label on known genes was 0.579 against 0.579 on shuffled data (skill -0.00).
Works when stageenrichedderived: F1 of hidden genes in the community chosen for their label on known genes reached 0.40 against 0.32 on shuffled data (skill 0.12, 452 scored). Run once on every gene, it called 2,000 genes with no known stageenrichedderived; the top 5 are listed.

P. falciparum. Fails when isexported: 'isexported' is not encoded in what this strategy reads: F1 of hidden genes in the community chosen for their label on known genes was 0.341 against 0.369 on shuffled data (skill -0.04).
Works when lopitpflocation: F1 of hidden genes in the community chosen for their label on known genes reached 0.23 against 0.14 on shuffled data (skill 0.10, 389 scored). Run once on every gene, it called 2,000 genes with no known lopitpflocation; the top 5 are listed.

16 · Predict the contacts an interactome missed

logistic regression · ranking · T. gondii reliable · P. falciparum reliable

Which pairs of proteins are probably in physical contact although the crosslinking or pulldown experiment never saw them together?

T. gondii P. falciparum
Better than chance
skill
███████░░░░░ 0.61 [0.59, 0.62] · chance 0.00 ████████████ 0.96 [0.96, 0.96] · chance 0.00
Reach
recall @ top 10%
█████░░░░░░░ 0.44 [0.43, 0.45] · chance 0.10 ███████░░░░░ 0.55 [0.55, 0.55] · chance 0.10
True ones ranked first
AUROC
██████████░░ 0.80 [0.80, 0.81] · chance 0.50 ████████████ 0.98 [0.98, 0.98] · chance 0.50
Clean top of the list
AUPRC lift
██████░░░░░░ x3.6 [x3.5, x3.7] · chance x1.0 ███████░░░░░ x5.3 [x5.3, x5.4] · chance x1.0

About this test

T. gondii. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: AUROC of hidden pairs against degree-matched non-pairs was 0.52 against 0.49 on shuffled data (skill 0.07). The test calls that a FAIL, which is the failure it exists to catch.
Works when xlms: AUROC of hidden pairs against degree-matched non-pairs reached 0.82 against 0.49 on shuffled data (skill 0.64, 568 scored). Run once on every gene, it proposed 300 links not in the layer; the top 5 are listed.

P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: AUROC of hidden pairs against degree-matched non-pairs was 0.52 against 0.49 on shuffled data (skill 0.07). The test calls that a FAIL, which is the failure it exists to catch.
Works when struct: AUROC of hidden pairs against degree-matched non-pairs reached 0.98 against 0.50 on shuffled data (skill 0.96, 914 scored). Run once on every gene, it proposed 300 links not in the layer; the top 5 are listed.

17 · Read the literature for biology, not fame

publication-count residual · ranking · T. gondii reliable · P. falciparum weak

Which pairs of genes are written about together more than their popularity explains -- and are those pairs biologically related?

T. gondii P. falciparum
Better than chance
skill
█████░░░░░░░ 0.38 [0.21, 0.54] · chance 0.00 ███░░░░░░░░░ 0.26 [0.23, 0.29] · chance 0.00
Reach
recall @ top 10%
██░░░░░░░░░░ 0.14 [0.12, 0.16] · chance 0.10 ██░░░░░░░░░░ 0.14 [0.10, 0.18] · chance 0.10
True ones ranked first
AUROC
███████░░░░░ 0.62 [0.60, 0.64] · chance 0.50 ███████░░░░░ 0.56 [0.46, 0.67] · chance 0.50
Clean top of the list
AUPRC lift
███░░░░░░░░░ x1.2 [x1.1, x1.3] · chance x1.0 ██░░░░░░░░░░ x1.2 [x1.0, x1.4] · chance x1.0

About this test

T. gondii. Fails when screenanyphenotype: The signal is too weak to tell from luck: 0.600 did not clear 0.705, what shuffled data reaches one time in twenty (skill 0.18). Also, only 20 hidden items could be scored.
Works when compartment: Share of the top 200 corrected pairs sharing a compartment label reached 0.75 against 0.39 on shuffled data (skill 0.59, 200 scored). Run once on every gene, it re-ranked 200 pairs once fame is taken out; the top 5 are listed.

P. falciparum. Fails when isexported: The signal is too weak to tell from luck: 0.983 did not clear 0.991, what shuffled data reaches one time in twenty (skill 0.25).
Works when lopit
pflocation: Share of the top 48 corrected pairs sharing a lopitpf_location label reached 0.56 against 0.40 on shuffled data (skill 0.27, 48 scored). Run once on every gene, it re-ranked 200 pairs once fame is taken out; the top 5 are listed.

18 · List what the data says and the literature has not written

multi-layer support count · ranking · T. gondii reliable · P. falciparum reliable

Which gene pairs do several independent measurements link that no paper has ever mentioned together?

T. gondii P. falciparum
Better than chance
skill
█░░░░░░░░░░░ 0.11 [0.11, 0.11] · chance 0.00 █░░░░░░░░░░░ 0.11 [0.11, 0.11] · chance 0.00
Reach
recall @ top 10%
███████░░░░░ 0.57 [0.57, 0.57] · chance 0.10 ███████░░░░░ 0.58 [0.58, 0.59] · chance 0.10
True ones ranked first
AUROC
███████░░░░░ 0.56 [0.56, 0.56] · chance 0.50 ███████░░░░░ 0.55 [0.55, 0.55] · chance 0.50
Clean top of the list
AUPRC lift
███░░░░░░░░░ x1.5 [x1.5, x1.5] · chance x1.0 ███░░░░░░░░░ x1.5 [x1.4, x1.5] · chance x1.0

About this test

T. gondii. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: AUROC of hidden pairs against random non-pairs was 0.50 against 0.50 on shuffled data (skill 0.00). The test calls that a FAIL, which is the failure it exists to catch.
Works when --: AUROC of hidden pairs against random non-pairs reached 0.56 against 0.50 on shuffled data (skill 0.11, 7,659 scored). Run once on every gene, it found 500 measured pairs nobody has written about; the top 5 are listed.

P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: AUROC of hidden pairs against random non-pairs was 0.50 against 0.50 on shuffled data (skill 0.00). The test calls that a FAIL, which is the failure it exists to catch.
Works when --: AUROC of hidden pairs against random non-pairs reached 0.55 against 0.50 on shuffled data (skill 0.11, 558 scored). Run once on every gene, it found 255 measured pairs nobody has written about; the top 5 are listed.

19 · Train a classifier on the known genes and call the rest

logistic regression · label calls · T. gondii reliable · P. falciparum reliable

Given every permitted measurement, which label does a model trained on the labelled genes assign to each unlabelled one -- and which measurements does it rely on?

T. gondii P. falciparum
Better than chance
skill
████░░░░░░░░ 0.32 [0.18, 0.46] · chance 0.00 █████░░░░░░░ 0.39 [0.35, 0.43] · chance 0.00
Reach
coverage
████████████ 1.00 [1.00, 1.00] ████████████ 1.00 [1.00, 1.00]
Right calls
accuracy
███████░░░░░ 0.54 [0.41, 0.68] · chance 0.33 ████████░░░░ 0.63 [0.46, 0.82] · chance 0.41
Fair across classes
macro F1
██████░░░░░░ 0.48 [0.36, 0.61] ███████░░░░░ 0.55 [0.43, 0.67]

About this test

T. gondii. Fails when screenanyphenotype: The signal is too weak to tell from luck: 0.550 did not clear 0.589, what shuffled data reaches one time in twenty (skill 0.11).
Works when compartment: Correct calls per hidden gene reached 0.33 against 0.06 on shuffled data (skill 0.29, 951 scored). Run once on every gene, it called 2,000 genes with no known compartment; the top 5 are listed.

P. falciparum. Fails when isexported: The signal is real but small: 0.938 beat shuffled data (0.902), but by 0.036, short of the 0.050 margin a PASS requires.
Works when lopit
pflocation: Correct calls per hidden gene reached 0.45 against 0.07 on shuffled data (skill 0.41, 395 scored). Run once on every gene, it called 2,000 genes with no known lopitpf_location; the top 5 are listed.

20 · Learn what makes your list special, from positives alone

PU bagging, logistic regression · ranking · T. gondii reliable · P. falciparum reliable

Given only genes that ARE something -- no list of genes that are not -- which other genes look most like them?

T. gondii P. falciparum
Better than chance
skill
█████████░░░ 0.79 [0.72, 0.85] · chance 0.00 ███████████░ 0.90 [0.80, 0.98] · chance 0.00
Reach
recall @ top 10%
████████░░░░ 0.66 [0.55, 0.77] · chance 0.10 ██████████░░ 0.84 [0.58, 0.99] · chance 0.10
True ones ranked first
AUROC
███████████░ 0.89 [0.86, 0.92] · chance 0.50 ███████████░ 0.95 [0.90, 0.99] · chance 0.50
Clean top of the list
AUPRC lift
██████████░░ x15 [x9.8, x22] · chance x1.0 ████████████ x42 [x12, x83] · chance x1.0

About this test

T. gondii. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: AUROC of hidden members against every other gene was 0.49 against 0.53 on shuffled data (skill -0.08). The test calls that a FAIL, which is the failure it exists to catch.
Works when PM - integral (compartment): AUROC of hidden members against every other gene reached 0.95 against 0.49 on shuffled data (skill 0.90, 40 scored). Run once on every gene, it ranked 300 genes with no place on the list as likely members of the list; the top 5 are listed.

P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: AUROC of hidden members against every other gene was 0.49 against 0.53 on shuffled data (skill -0.08). The test calls that a FAIL, which is the failure it exists to catch.
Works when cytosol (lopitpflocation): AUROC of hidden members against every other gene reached 0.98 against 0.50 on shuffled data (skill 0.96, 37 scored). Run once on every gene, it ranked 300 genes with no place on the list as likely members of the list; the top 5 are listed.

21 · Predict a measurement, and find the genes that defy the prediction

gradient boosting / ridge · values · T. gondii reliable · P. falciparum reliable

How well does everything else predict this measurement -- and which genes are far from what their profile says they should be?

T. gondii P. falciparum
Better than chance
skill
███████░░░░░ 0.56 [0.05, 0.91] · chance 0.00 ██████░░░░░░ 0.52 [0.25, 0.85] · chance 0.00
Reach
coverage
████████████ 1.00 [1.00, 1.00] ████████████ 1.00 [1.00, 1.00]
Order predicted
Spearman rho
███████░░░░░ 0.56 [0.05, 0.91] · chance 0.00 ██████░░░░░░ 0.52 [0.25, 0.85] · chance 0.00
Variance explained
R-squared, out of sample
█████░░░░░░░ 0.43 [-0.09, 0.87] · chance 0.00 ░░░░░░░░░░░░ -3.29 [-14.04, 0.72] · chance 0.00

About this test

T. gondii. Fails when fitinvivoPE: The signal is real but small: 0.052 beat shuffled data (0.000), but by 0.052, short of the 0.100 margin a PASS requires.
Works when fitinvitrohff: Rank correlation of predicted and hidden values reached 0.73 against 0.00 on shuffled data (skill 0.73, 1,465 scored). Run once on every gene, it predicted 815 genes with no measured fitinvitrohff; the top 5 are listed.

P. falciparum. Fails when exprschizont: The signal is too weak to tell from luck: 0.036 did not clear 0.049, what shuffled data reaches one time in twenty (skill 0.04).
Works when piggybac
mis: Rank correlation of predicted and hidden values reached 0.49 against 0.00 on shuffled data (skill 0.49, 1,077 scored). Run once on every gene, it predicted 335 genes with no measured piggybac_mis; the top 5 are listed.

22 · Fill in what was never measured, and say where that is honest

soft-impute, low-rank SVD · values · T. gondii reliable · P. falciparum reliable

For each measurement, can its missing values be estimated from the rest of the table -- and for which measurements is that impossible?

T. gondii P. falciparum
Better than chance
skill
██████████░░ 0.87 [0.86, 0.87] · chance 0.00 ████████░░░░ 0.68 [0.67, 0.69] · chance 0.00
Reach
coverage
████████████ 1.00 [1.00, 1.00] ████████████ 1.00 [1.00, 1.00]
Order predicted
Spearman rho
██████████░░ 0.87 [0.86, 0.87] · chance 0.00 ████████░░░░ 0.68 [0.67, 0.69] · chance 0.00
Variance explained
R-squared, out of sample
█████████░░░ 0.75 [0.75, 0.76] · chance 0.00 █████░░░░░░░ 0.45 [0.44, 0.46] · chance 0.00

About this test

T. gondii. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: median per-column rank correlation on hidden entries was -0.01 against -0.03 on shuffled data (skill 0.01). The test calls that a FAIL, which is the failure it exists to catch.
Works when --: Median per-column rank correlation on hidden entries reached 0.88 against 0.00 on shuffled data (skill 0.88, 225,394 scored). Run once on every gene, it filled 815 missing measurements; the top 5 are listed.

P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: median per-column rank correlation on hidden entries was -0.01 against -0.03 on shuffled data (skill 0.01). The test calls that a FAIL, which is the failure it exists to catch.
Works when --: Median per-column rank correlation on hidden entries reached 0.69 against 0.00 on shuffled data (skill 0.69, 58,452 scored). Run once on every gene, it filled 335 missing measurements; the top 5 are listed.

23 · Find what matters more in one condition, and why

residual + gradient boosting / ridge · values · T. gondii weak · P. falciparum reliable

Which genes matter more (or less) in one condition than a baseline predicts -- in the mouse rather than the dish, say -- and can the rest of the data explain which?

T. gondii P. falciparum
Better than chance
skill
█░░░░░░░░░░░ 0.07 [0.06, 0.09] · chance 0.00 ████████░░░░ 0.64 [0.64, 0.65] · chance 0.00
Reach
coverage
████████████ 1.00 [1.00, 1.00] ████████████ 1.00 [1.00, 1.00]
Order predicted
Spearman rho
█░░░░░░░░░░░ 0.07 [0.06, 0.09] · chance 0.00 ████████░░░░ 0.64 [0.64, 0.65] · chance 0.00
Variance explained
R-squared, out of sample
░░░░░░░░░░░░ -0.04 [-0.06, -0.03] · chance 0.00 █████░░░░░░░ 0.40 [0.40, 0.41] · chance 0.00

About this test

T. gondii. Fails when random data: The signal is real but small: 0.069 beat shuffled data (0.000), but by 0.069, short of the 0.100 margin a PASS requires.
Works when --: Rank correlation of predicted and hidden values reached 0.15 against 0.00 on shuffled data (skill 0.15, 1,465 scored). Run once on every gene, it predicted 680 genes with no label; the top 5 are listed.

P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: rank correlation of predicted and hidden values was 0.03 against 0.00 on shuffled data (skill 0.03). The test calls that a FAIL, which is the failure it exists to catch.
Works when --: Rank correlation of predicted and hidden values reached 0.65 against 0.00 on shuffled data (skill 0.65, 1,144 scored). Run once on every gene, it measured the condition-specific shift of 5,720 genes; the top 5 are listed.

24 · Describe what your gene list has in common

hypergeometric + rank-sum · ranking · T. gondii reliable · P. falciparum reliable

What distinguishes the genes on my list from the rest -- which categories are they enriched in, which measurements are shifted, which networks are dense among them?

T. gondii P. falciparum
Better than chance
skill
███████░░░░░ 0.61 [0.51, 0.72] · chance 0.00 ██████████░░ 0.85 [0.75, 0.94] · chance 0.00
Reach
recall @ top 10%
██████░░░░░░ 0.47 [0.32, 0.61] · chance 0.10 █████████░░░ 0.76 [0.58, 0.92] · chance 0.10
True ones ranked first
AUROC
██████████░░ 0.81 [0.75, 0.86] · chance 0.50 ███████████░ 0.93 [0.88, 0.97] · chance 0.50
Clean top of the list
AUPRC lift
████████░░░░ x8.1 [x3.9, x13] · chance x1.0 ███████████░ x19 [x13, x26] · chance x1.0

About this test

T. gondii. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: AUROC of hidden members against every other gene was 0.50 against 0.50 on shuffled data (skill -0.01). The test calls that a FAIL, which is the failure it exists to catch.
Works when PM - integral (compartment): AUROC of hidden members against every other gene reached 0.91 against 0.50 on shuffled data (skill 0.82, 54 scored). Run once on every gene, it ranked 300 genes with no place on the list by how much they resemble the list; the top 5 are listed.

P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: AUROC of hidden members against every other gene was 0.50 against 0.50 on shuffled data (skill -0.01). The test calls that a FAIL, which is the failure it exists to catch.
Works when cytosol (lopitpflocation): AUROC of hidden members against every other gene reached 0.93 against 0.50 on shuffled data (skill 0.86, 49 scored). Run once on every gene, it ranked 300 genes with no place on the list by how much they resemble the list; the top 5 are listed.

25 · Grow your gene list along the networks

random walk with restart · ranking · T. gondii reliable · P. falciparum reliable

Starting from my genes, which others does a walk across every measured network keep returning to?

T. gondii P. falciparum
Better than chance
skill
███████░░░░░ 0.59 [0.49, 0.68] · chance 0.00 ██████████░░ 0.86 [0.74, 0.96] · chance 0.00
Reach
recall @ top 10%
██████░░░░░░ 0.47 [0.38, 0.56] · chance 0.10 ██████████░░ 0.82 [0.67, 0.96] · chance 0.10
True ones ranked first
AUROC
██████████░░ 0.80 [0.74, 0.84] · chance 0.50 ███████████░ 0.93 [0.88, 0.98] · chance 0.50
Clean top of the list
AUPRC lift
████████░░░░ x7.7 [x4.4, x12] · chance x1.0 ████████████ x38 [x13, x72] · chance x1.0

About this test

T. gondii. Fails when tachyzoite (stageenrichedderived): The signal is too weak to tell from luck: 0.548 did not clear 0.561, what shuffled data reaches one time in twenty (skill 0.09).
Works when PM - integral (compartment): AUROC of hidden members against every other gene reached 0.88 against 0.50 on shuffled data (skill 0.76, 40 scored). Run once on every gene, it ranked 300 genes with no place on the list as likely members of the list; the top 5 are listed.

P. falciparum. Fails when gametocyte (stageenrichedderived): 'gametocyte (stageenrichedderived)' is not encoded in what this strategy reads: AUROC of hidden members against every other gene was 0.366 against 0.505 on shuffled data (skill -0.28). Also, only 20 hidden items could be scored.
Works when cytosol (lopitpflocation): AUROC of hidden members against every other gene reached 0.94 against 0.49 on shuffled data (skill 0.89, 37 scored). Run once on every gene, it ranked 300 genes with no place on the list as likely members of the list; the top 5 are listed.

26 · Find categories that split in two on another measurement

UMAP + HDBSCAN · replication · T. gondii untestable · P. falciparum untestable

Which clusters agree about one thing -- a compartment -- and split cleanly on another -- a stage, a phase, a fitness level?

T. gondii P. falciparum
Better than chance
skill
-- --
Reach
findings made
-- --
Findings that hold
replication rate
-- --
Beyond chance
replication lift
-- --

About this test

T. gondii. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find. It made too little to score, and the test declined to judge (only 0 findings on the first half, 3 needed).
Works when --: None of its 15 calibration runs on real data passed; the usual reason: only 0 findings on the first half, 3 needed. There is no success to show.

P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find. It made too little to score, and the test declined to judge (only 0 findings on the first half, 3 needed).
Works when --: None of its 15 calibration runs on real data passed; the usual reason: only 0 findings on the first half, 3 needed. There is no success to show.

27 · Find kinds of gene defined by two labels at once

UMAP + HDBSCAN · replication · T. gondii reliable · P. falciparum untestable

Which clusters are enriched for a COMBINATION of two labels -- more than either label alone would make them?

T. gondii P. falciparum
Better than chance
skill
██████░░░░░░ 0.49 [0.38, 0.62] · chance 0.00 --
Reach
findings made
███░░░░░░░░░ 5 [4, 7] --
Findings that hold
replication rate
██████░░░░░░ 0.50 [0.39, 0.64] · chance 0.03 --
Beyond chance
replication lift
███████████░ x22 [x15, x31] · chance x1.0 --

About this test

T. gondii. Fails when random data: The signal is too weak to tell from luck: 0.167 did not clear 0.167, what shuffled data reaches one time in twenty (skill 0.15). Also, only 12 hidden items could be scored.
Works when --: Share of first-half findings that replicate on the second half reached 0.75 against 0.04 on shuffled data (skill 0.74, 4 scored). Run once on every gene, it found 5 structures neither label shows alone; the top 5 are listed.

P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find. It made too little to score, and the test declined to judge (only 0 findings on the first half, 3 needed).
Works when --: None of its 15 calibration runs on real data passed; the usual reason: only 0 findings on the first half, 3 needed. There is no success to show.

28 · Find paralogs that changed jobs

profile correlation · ranking · T. gondii reliable · P. falciparum weak

Which duplicated genes behave differently across the measurements -- evidence that one copy took on a new role?

T. gondii P. falciparum
Better than chance
skill
██░░░░░░░░░░ 0.16 [0.06, 0.31] · chance 0.00 █░░░░░░░░░░░ 0.11 [-0.14, 0.37] · chance 0.00
Reach
recall @ top 10%
██░░░░░░░░░░ 0.13 [0.10, 0.17] · chance 0.10 ██░░░░░░░░░░ 0.13 [0.10, 0.15] · chance 0.10
True ones ranked first
AUROC
███████░░░░░ 0.58 [0.53, 0.66] · chance 0.50 ███████░░░░░ 0.56 [0.44, 0.68] · chance 0.50
Clean top of the list
AUPRC lift
███░░░░░░░░░ x1.2 [x1.0, x1.4] · chance x1.0 ███░░░░░░░░░ x1.2 [x1.1, x1.4] · chance x1.0

About this test

T. gondii. Fails when dtm_class: The signal is real but small: 0.514 beat shuffled data (0.500), but by 0.014, short of the 0.050 margin a PASS requires.
Works when compartment: AUROC of profile divergence for paralogs with different compartment reached 0.55 against 0.49 on shuffled data (skill 0.11, 535 scored). Run once on every gene, it ranked 3,452 paralog pairs by how far they diverged; the top 5 are listed.

P. falciparum. Fails when pbtransferredphenotype: 'pbtransferredphenotype' is not encoded in what this strategy reads: AUROC of profile divergence for paralogs with different pbtransferredphenotype was 0.435 against 0.505 on shuffled data (skill -0.14).
Works when lopitpflocation: AUROC of profile divergence for paralogs with different lopitpflocation reached 0.68 against 0.48 on shuffled data (skill 0.39, 96 scored). Run once on every gene, it ranked 1,741 paralog pairs by how far they diverged; the top 5 are listed.

29 · Carry what one parasite shows to the other

orthogroup mapping · values · T. gondii reliable · P. falciparum reliable

What does a gene's ortholog in the other parasite say about it -- its essentiality, its stage, its localization?

T. gondii P. falciparum
Better than chance
skill
████░░░░░░░░ 0.31 [0.30, 0.33] · chance 0.00 ████░░░░░░░░ 0.31 [0.30, 0.33] · chance 0.00
Reach
coverage
████████████ 1.00 [1.00, 1.00] ████████████ 1.00 [1.00, 1.00]
Order predicted
Spearman rho
████░░░░░░░░ 0.32 [0.30, 0.33] · chance 0.00 ████░░░░░░░░ 0.32 [0.30, 0.33] · chance 0.00
Variance explained
R-squared, out of sample
░░░░░░░░░░░░ -1.60 [-1.70, -1.50] · chance 0.00 ░░░░░░░░░░░░ -85.78 [-88.25, -83.72] · chance 0.00

About this test

T. gondii. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: rank correlation of transferred and hidden values was -0.05 against 0.02 on shuffled data (skill -0.07). The test calls that a FAIL, which is the failure it exists to catch.
Works when --: Rank correlation of transferred and hidden values reached 0.35 against -0.00 on shuffled data (skill 0.35, 681 scored). Run once on every gene, it predicted 106 genes with no measured fitinvitrohff; the top 5 are listed.

P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: rank correlation of transferred and hidden values was -0.05 against 0.02 on shuffled data (skill -0.07). The test calls that a FAIL, which is the failure it exists to catch.
Works when --: Rank correlation of transferred and hidden values reached 0.34 against 0.01 on shuffled data (skill 0.33, 661 scored). Run once on every gene, it predicted 25 genes with no measured piggybac_mis; the top 5 are listed.

30 · Test inference on the genes orthology cannot reach

kNN · label calls · T. gondii reliable · P. falciparum weak

Can lineage-specific, hypothetical or understudied genes be called as reliably as the rest -- and what are they?

T. gondii P. falciparum
Better than chance
skill
███░░░░░░░░░ 0.26 [0.15, 0.34] · chance 0.00 ░░░░░░░░░░░░ -0.01 [-0.01, -0.01] · chance 0.00
Reach
coverage
███████████░ 0.89 [0.76, 1.00] ████████████ 1.00 [1.00, 1.00]
Right calls
accuracy
███████░░░░░ 0.59 [0.39, 0.78] · chance 0.42 ████████░░░░ 0.66 [0.32, 1.00] · chance 0.66
Fair across classes
macro F1
█████░░░░░░░ 0.38 [0.27, 0.54] ███████░░░░░ 0.58 [0.16, 1.00]

About this test

T. gondii. Fails when dtm_class: The signal is real but small: 0.733 beat shuffled data (0.709), but by 0.024, short of the 0.050 margin a PASS requires.
Works when compartment: Correct calls per hidden gene reached 0.35 against 0.04 on shuffled data (skill 0.32, 377 scored). Run once on every gene, it called 1,449 genes with no known compartment; the top 5 are listed.

P. falciparum. Fails when stageenrichedderived: 'stageenrichedderived' is not encoded in what this strategy reads: correct calls per hidden gene was 0.318 against 0.323 on shuffled data (skill -0.01). Also, only 22 hidden items could be scored.
Works when lopitpflocation: Correct calls per hidden gene reached 0.55 against 0.07 on shuffled data (skill 0.51, 286 scored). Run once on every gene, it called 2,000 genes with no known lopitpflocation; the top 5 are listed.

31 · Call a gene only when independent strategies agree

kNN + logistic + network vote · label calls · T. gondii reliable · P. falciparum works when tuned

Where do measurement neighbours, a trained classifier and the networks give the same answer -- and how much more often is that answer right?

T. gondii P. falciparum
Better than chance
skill
████░░░░░░░░ 0.35 [0.17, 0.48] · chance 0.00 ███░░░░░░░░░ 0.25 [-0.26, 0.55] · chance 0.00
Reach
coverage
██████████░░ 0.87 [0.76, 0.96] ██████████░░ 0.86 [0.77, 0.96]
Right calls
accuracy
████████░░░░ 0.63 [0.48, 0.76] ████████░░░░ 0.67 [0.54, 0.84]
Fair across classes
macro F1
██████░░░░░░ 0.50 [0.43, 0.61] ███████░░░░░ 0.55 [0.51, 0.58]

About this test

T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: precision of calls on hidden genes was 0.662 against 0.668 on shuffled data (skill -0.02).
Works when compartment: Precision of calls on hidden genes reached 0.59 against 0.19 on shuffled data (skill 0.50, 629 scored). Run once on every gene, it called 1,787 genes with no known compartment; the top 5 are listed.

P. falciparum. Fails when isexported: The signal is real but small: 0.938 beat shuffled data (0.936), but by 0.002, short of the 0.100 margin a PASS requires.
Works when lopit
pflocation: Precision of calls on hidden genes reached 0.71 against 0.23 on shuffled data (skill 0.62, 303 scored). Run once on every gene, it called 2,000 genes with no known lopitpf_location; the top 5 are listed.

32 · Put the understudied genes first

kNN + logistic + network vote · label calls · T. gondii reliable · P. falciparum works when tuned

Which genes nobody has written about can the data say something trustworthy about?

T. gondii P. falciparum
Better than chance
skill
████░░░░░░░░ 0.32 [0.13, 0.46] · chance 0.00 ██░░░░░░░░░░ 0.19 [-0.39, 0.55] · chance 0.00
Reach
coverage
██████████░░ 0.87 [0.76, 0.97] ██████████░░ 0.87 [0.78, 0.96]
Right calls
accuracy
███████░░░░░ 0.62 [0.48, 0.77] ████████░░░░ 0.67 [0.56, 0.84]
Fair across classes
macro F1
██████░░░░░░ 0.48 [0.41, 0.59] ██████░░░░░░ 0.51 [0.45, 0.56]

About this test

T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: precision of calls on hidden genes was 0.643 against 0.667 on shuffled data (skill -0.07).
Works when compartment: Precision of calls on hidden genes reached 0.57 against 0.16 on shuffled data (skill 0.48, 546 scored). Run once on every gene, it called 1,698 little-studied genes with no known compartment; the top 5 are listed.

P. falciparum. Fails when isexported: One class dominates 'isexported', so guessing it on shuffled data already scores 0.972; the strategy's 0.972 is no better than that (skill 0.00), so what it reads does not separate the classes.
Works when lopitpflocation: Precision of calls on hidden genes reached 0.70 against 0.20 on shuffled data (skill 0.63, 215 scored). Run once on every gene, it called 2,000 little-studied genes with no known lopitpflocation; the top 5 are listed.

33 · Put every layer into one space and read a gene's neighbourhood

logistic edge model · ranking · T. gondii reliable · P. falciparum reliable

Which genes are the nearest neighbours of this one when every permitted network and the whole measurement table are combined into a single graph, and what evidence puts each of them there?

T. gondii P. falciparum
Better than chance
skill
███████░░░░░ 0.59 [0.58, 0.59] · chance 0.00 █████░░░░░░░ 0.39 [0.37, 0.41] · chance 0.00
Reach
recall @ top 10%
██░░░░░░░░░░ 0.19 [0.18, 0.19] · chance 0.10 ██░░░░░░░░░░ 0.17 [0.16, 0.19] · chance 0.10
True ones ranked first
AUROC
█████████░░░ 0.79 [0.79, 0.79] · chance 0.50 ████████░░░░ 0.71 [0.68, 0.73] · chance 0.50
Clean top of the list
AUPRC lift
███░░░░░░░░░ x1.6 [x1.6, x1.6] · chance x1.0 ███░░░░░░░░░ x1.4 [x1.4, x1.5] · chance x1.0

About this test

T. gondii. Fails when xlms: The signal is real but small: 0.837 beat shuffled data (0.791), but by 0.046, short of the 0.050 margin a PASS requires.
Works when coexpression: AUROC of hidden coexpression edges against degree-matched non-pairs reached 0.79 against 0.49 on shuffled data (skill 0.59, 3,349 scored). Run once on every gene, it ranked already-measured pairs at the top: all 500 pairs it listed are recorded by some layer, so this run proposes no unmeasured pair.

P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: AUROC of hidden coexpression edges against degree-matched non-pairs was 0.54 against 0.54 on shuffled data (skill 0.00). The test calls that a FAIL, which is the failure it exists to catch.
Works when coexpression: AUROC of hidden coexpression edges against degree-matched non-pairs reached 0.75 against 0.56 on shuffled data (skill 0.43, 5,277 scored). Run once on every gene, it ranked already-measured pairs at the top: all 500 pairs it listed are recorded by some layer, so this run proposes no unmeasured pair.

34 · Train on the networks and rank the edges they are missing

logistic / spectral embedding · ranking · T. gondii reliable · P. falciparum reliable

Which pairs of genes does the combined evidence imply although no measured layer records them, how strong is each claim, and how good is the model that makes it when it is scored against a degree-matched null rather than a random one?

T. gondii P. falciparum
Better than chance
skill
███████░░░░░ 0.59 [0.58, 0.59] · chance 0.00 █████░░░░░░░ 0.39 [0.37, 0.42] · chance 0.00
Reach
recall @ top 10%
██░░░░░░░░░░ 0.19 [0.18, 0.19] · chance 0.10 ██░░░░░░░░░░ 0.17 [0.16, 0.19] · chance 0.10
True ones ranked first
AUROC
█████████░░░ 0.79 [0.79, 0.79] · chance 0.50 ████████░░░░ 0.71 [0.68, 0.73] · chance 0.50
Clean top of the list
AUPRC lift
███░░░░░░░░░ x1.6 [x1.6, x1.6] · chance x1.0 ███░░░░░░░░░ x1.4 [x1.4, x1.5] · chance x1.0

About this test

T. gondii. Fails when xlms: The signal is real but small: 0.849 beat shuffled data (0.804), but by 0.045, short of the 0.050 margin a PASS requires.
Works when coexpression: AUROC of hidden coexpression edges against degree-matched non-pairs reached 0.79 against 0.49 on shuffled data (skill 0.59, 3,349 scored). Run once on every gene, it proposed 300 unmeasured pairs; the top 5 are listed.

P. falciparum. Fails when random data (noise table): On the noise table every label and edge is dealt out at random, so there is nothing to find: AUROC of hidden coexpression edges against degree-matched non-pairs was 0.54 against 0.54 on shuffled data (skill 0.00). The test calls that a FAIL, which is the failure it exists to catch.
Works when coexpression: AUROC of hidden coexpression edges against degree-matched non-pairs reached 0.75 against 0.56 on shuffled data (skill 0.43, 5,277 scored). Run once on every gene, it proposed 300 unmeasured pairs; the top 5 are listed.

35 · Call genes with a stated error rate

split conformal prediction · label calls · T. gondii reliable · P. falciparum reliable

Which genes can be given a label with a guaranteed error rate -- and for which does the data leave two or more labels equally possible?

T. gondii P. falciparum
Better than chance
skill
██████░░░░░░ 0.50 [0.21, 0.73] · chance 0.00 ████████░░░░ 0.63 [0.44, 0.78] · chance 0.00
Reach
coverage
██░░░░░░░░░░ 0.21 [0.06, 0.44] █████░░░░░░░ 0.39 [0.12, 0.68]
Right calls
accuracy
██░░░░░░░░░░ 0.17 [0.05, 0.38] ████░░░░░░░░ 0.32 [0.10, 0.56]
Fair across classes
macro F1
███░░░░░░░░░ 0.25 [0.09, 0.48] ████░░░░░░░░ 0.36 [0.17, 0.55]

About this test

T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: set efficiency: 1 - (mean set size - 1) / (classes - 1) was 0.188 against 0.173 on shuffled data (skill 0.02). Also, it could reach only 19% of the hidden genes, so most were never called.
Works when compartment: Set efficiency: 1 - (mean set size - 1) / (classes - 1) reached 0.79 against 0.30 on shuffled data (skill 0.70, 951 scored). Run once on every gene, it gave prediction sets to 14 genes with no known compartment; the top 5 are listed.

P. falciparum. Fails when isexported: 'isexported' is not encoded in what this strategy reads: set efficiency: 1 - (mean set size - 1) / (classes - 1) was 0.115 against 0.102 on shuffled data (skill 0.01). Also, it could reach only 11% of the hidden genes, so most were never called.
Works when lopitpflocation: Set efficiency: 1 - (mean set size - 1) / (classes - 1) reached 0.85 against 0.22 on shuffled data (skill 0.81, 395 scored). Run once on every gene, it gave prediction sets to 27 genes with no known lopitpflocation; the top 5 are listed.

36 · Smooth the measurements along the networks, then classify

graph convolution + logistic regression · label calls · T. gondii reliable · P. falciparum reliable

Does a gene's label follow from its own measurements together with those of its network neighbours -- and how much does the model lean on each?

T. gondii P. falciparum
Better than chance
skill
█████░░░░░░░ 0.40 [0.16, 0.57] · chance 0.00 █████░░░░░░░ 0.46 [0.37, 0.54] · chance 0.00
Reach
coverage
████████████ 1.00 [1.00, 1.00] ████████████ 1.00 [1.00, 1.00]
Right calls
accuracy
████████░░░░ 0.63 [0.53, 0.74] · chance 0.36 █████████░░░ 0.71 [0.59, 0.84] · chance 0.45
Fair across classes
macro F1
███████░░░░░ 0.56 [0.48, 0.67] ███████░░░░░ 0.62 [0.55, 0.69]

About this test

T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: correct calls per hidden gene was 0.512 against 0.534 on shuffled data (skill -0.05).
Works when compartment: Correct calls per hidden gene reached 0.51 against 0.09 on shuffled data (skill 0.46, 951 scored). Run once on every gene, it called 2,000 genes with no known compartment; the top 5 are listed.

P. falciparum. Fails when isexported: The signal is real but small: 0.883 beat shuffled data (0.834), but by 0.048, short of the 0.050 margin a PASS requires.
Works when lopit
pflocation: Correct calls per hidden gene reached 0.63 against 0.10 on shuffled data (skill 0.59, 395 scored). Run once on every gene, it called 2,000 genes with no known lopitpf_location; the top 5 are listed.

37 · Let a random forest find what defines a label

random forest + permutation importance · label calls · T. gondii reliable · P. falciparum reliable

Which measurements, in which combinations and past which thresholds, define a label -- and which unlabelled genes carry that definition?

T. gondii P. falciparum
Better than chance
skill
█████░░░░░░░ 0.45 [0.22, 0.65] · chance 0.00 ████░░░░░░░░ 0.36 [0.19, 0.53] · chance 0.00
Reach
coverage
████████████ 1.00 [1.00, 1.00] ████████████ 1.00 [1.00, 1.00]
Right calls
accuracy
████████░░░░ 0.70 [0.57, 0.83] · chance 0.44 █████████░░░ 0.74 [0.66, 0.86] · chance 0.51
Fair across classes
macro F1
███████░░░░░ 0.56 [0.46, 0.68] ███████░░░░░ 0.55 [0.51, 0.59]

About this test

T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: correct calls per hidden gene was 0.688 against 0.685 on shuffled data (skill 0.01).
Works when compartment: Correct calls per hidden gene reached 0.52 against 0.12 on shuffled data (skill 0.45, 951 scored). Run once on every gene, it called 2,000 genes with no known compartment; the top 5 are listed.

P. falciparum. Fails when isexported: The signal is too weak to tell from luck: 0.971 did not clear 0.976, what shuffled data reaches one time in twenty (skill 0.09).
Works when lopit
pflocation: Correct calls per hidden gene reached 0.66 against 0.11 on shuffled data (skill 0.61, 395 scored). Run once on every gene, it called 2,000 genes with no known lopitpf_location; the top 5 are listed.

38 · Learn how much to trust each kind of evidence

stacked logistic regression · label calls · T. gondii reliable · P. falciparum reliable

Given measurement neighbours, a linear model and the measured networks, how should their answers be combined for THIS label -- and what does the combination call?

T. gondii P. falciparum
Better than chance
skill
█████░░░░░░░ 0.41 [0.18, 0.57] · chance 0.00 █████░░░░░░░ 0.45 [0.38, 0.52] · chance 0.00
Reach
coverage
████████████ 1.00 [1.00, 1.00] ████████████ 1.00 [1.00, 1.00]
Right calls
accuracy
███████░░░░░ 0.62 [0.51, 0.74] · chance 0.34 ████████░░░░ 0.70 [0.60, 0.82] · chance 0.43
Fair across classes
macro F1
███████░░░░░ 0.56 [0.47, 0.67] ███████░░░░░ 0.62 [0.57, 0.67]

About this test

T. gondii. Fails when screenanyphenotype: 'screenanyphenotype' is not encoded in what this strategy reads: correct calls per hidden gene was 0.362 against 0.395 on shuffled data (skill -0.05).
Works when compartment: Correct calls per hidden gene reached 0.50 against 0.08 on shuffled data (skill 0.46, 951 scored). Run once on every gene, it called 2,000 genes with no known compartment; the top 5 are listed.

P. falciparum. Fails when isexported: The signal is real but small: 0.927 beat shuffled data (0.878), but by 0.049, short of the 0.050 margin a PASS requires.
Works when lopit
pflocation: Correct calls per hidden gene reached 0.61 against 0.08 on shuffled data (skill 0.57, 395 scored). Run once on every gene, it called 2,000 genes with no known lopitpf_location; the top 5 are listed.

39 · Predict a value with an interval that holds

gradient boosting / ridge + split conformal · values · T. gondii weak · P. falciparum reliable

For a gene never measured, what value is expected -- and within what range, with a guaranteed chance of containing the truth?

T. gondii P. falciparum
Better than chance
skill
███████░░░░░ 0.55 [0.04, 0.91] · chance 0.00 ██████░░░░░░ 0.50 [0.22, 0.84] · chance 0.00
Reach
coverage
████████████ 1.00 [1.00, 1.00] ████████████ 1.00 [1.00, 1.00]
Order predicted
Spearman rho
███████░░░░░ 0.55 [0.04, 0.91] · chance 0.00 ██████░░░░░░ 0.50 [0.22, 0.84] · chance 0.00
Variance explained
R-squared, out of sample
█████░░░░░░░ 0.41 [-0.13, 0.87] · chance 0.00 ░░░░░░░░░░░░ -3.49 [-14.82, 0.71] · chance 0.00

About this test

T. gondii. Fails when fitinvivoPE: The signal is too weak to tell from luck: 0.038 did not clear 0.043, what shuffled data reaches one time in twenty (skill 0.04).
Works when fitinvitrohff: Rank correlation of predicted and hidden values reached 0.72 against 0.00 on shuffled data (skill 0.72, 1,465 scored). Run once on every gene, it predicted 815 unmeasured genes, each with a 90% interval; the top 5 are listed.

P. falciparum. Fails when exprschizont: The signal is real but small: 0.074 beat shuffled data (0.000), but by 0.074, short of the 0.100 margin a PASS requires.
Works when piggybac
mis: Rank correlation of predicted and hidden values reached 0.46 against 0.00 on shuffled data (skill 0.46, 1,077 scored). Run once on every gene, it predicted 335 unmeasured genes, each with a 90% interval; the top 5 are listed.