Strategy calibration

Generated by scripts/calibrate_strategies.py --publish on 2026-09-26 from 7,640 self-test runs (results/calibration_2026-09-26b).

Every strategy carries a self-test that hides information already known -- labels, set members, edges or values -- asks the strategy for it back, and compares the answer with the same procedure on shuffled data. Calibration runs that test over a grid of the strategy's settings, over several held-out labels and five seeds, so each number below is a mean with an interval rather than one draw.

Grades: reliable (defaults beat chance with the interval above 0.05 and pass at least 60%), works when tuned (only the tuned setting does), weak (above chance on average but not reliably), no skill, untestable (fewer than five conclusive runs).

What a high skill here does and does not mean

Each strategy is scored against ITS OWN null, and a null can be too easy. The clearest case is edge prediction: strategy 18 scores its held-out pairs against non-pairs drawn at random, and a random pair of genes is usually a pair of obscure genes, so much of what such a test measures is that well-connected genes are well connected. Its skill here is therefore an upper bound (strategies 16, 33 and 34 use degree-matched non-pairs instead). docs/graphspace.md re-measures the same question against degree-matched and configuration-model nulls, where the honest figure is an AUROC near 0.67 rather than 0.97 -- and it reports the gap between the two nulls as its own quantity. Read a high number here as a reason to look at how the null was built, not as a result.

Toxoplasma gondii

01 · Hold out a category and search for a map that finds it (UMAP + HDBSCAN) -- weak

25% of the label is hidden (whole orthogroups together, so no gene is recovered through a visible paralog). The walk is built as set -- its genes per map (up to 4,000), feature sets and grids (up to three values each); the configuration, and the one cluster that best isolates each label, are both chosen using visible labels only. Metric: for each label, the F1 of its hidden genes against its chosen cluster, weighted by size -- does the structure found on known genes hold the unknown ones? Null: the same chosen clusters scored after permuting the hidden genes' labels, 100 times. (A second search on shuffled labels was the null once; it picks the largest cluster for every label, which scores F1 near 2p by size alone and made the null beat real labels.) Pass: above the null's 95th percentile by at least 0.05.

Metric: F1 of hidden genes in the cluster chosen for their label on known genes. 150 runs, 6 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults minclustersize=[20, 50]; n_neighbors=[15, 50]; selection=[eom, leaf] 0.071 [0.034, 0.105] 44% [27, 63] 0.376 0.322 25
tuned minclustersize=[10, 25]; n_neighbors=[25, 100]; selection=[eom, leaf] 0.070 [0.026, 0.115] 40% [17, 69] 0.394 0.342 10

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.074 [0.050, 0.103] 60% 5
dtm_class 0.078 [0.061, 0.096] 0% 5
lopit_unified 0.108 [0.095, 0.120] 100% 5
screenanyphenotype -0.005 [-0.023, 0.006] 0% 5
stageenrichedderived 0.159 [0.106, 0.208] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

02 · Find the map where your gene list is one cluster (UMAP + HDBSCAN) -- weak

30% of the set is hidden; the walk as set (genes per map up to 4,000, feature sets, up to three values of each grid) picks the cluster with the best F1 for the other 70%. Metric: F1 of the hidden members against that cluster's other genes -- precision is the share of the cluster's candidates that are hidden members, recall the share of hidden members among them. Null: 20 random sets of the same size through the same walk. Pass: above the null's 95th percentile by at least 0.05 (an F1 margin; random sets score about 0.04).

Metric: F1 of the hidden members against the best cluster's other genes. 100 runs, 4 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults minclustersize=[20, 50]; n_neighbors=[15, 50] 0.025 [0.004, 0.042] 16% [6, 35] 0.058 0.033 25
tuned minclustersize=[20, 50]; n_neighbors=[10, 30] 0.028 [0.002, 0.047] 10% [2, 40] 0.06 0.032 10

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.037 [0.029, 0.047] 20% 5
dtm_class 0.024 [0.008, 0.037] 0% 5
lopit_unified 0.040 [0.028, 0.054] 20% 5
screenanyphenotype 0.047 [0.032, 0.062] 40% 5
stageenrichedderived -0.009 [-0.019, 0.000] 0% 5

Sensitivity (mean skill at each value, all other settings pooled):

03 · Ask which categories the data can rediscover (UMAP + neighbour AUROC) -- reliable

30% of the label is hidden. The atlas is built from visible genes only (leave-one-out neighbour AUROC per category); hidden genes are then scored by their visible neighbours. Metric: mean hidden-gene AUROC over the categories the atlas ranks in its top half. Null: the same with the visible labels shuffled, 20 times. Pass: above the null's 95th percentile by 0.05. The rank agreement between atlas and hidden recovery is reported.

Metric: hidden-gene AUROC of the categories the atlas ranks in its top half. 75 runs, 3 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults k=15 0.499 [0.217, 0.679] 80% [61, 91] 0.75 0.501 25
tuned k=15 0.502 [0.222, 0.680] 80% [49, 94] 0.75 0.501 10

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.617 [0.544, 0.686] 100% 5
dtm_class 0.563 [0.539, 0.587] 100% 5
lopit_unified 0.699 [0.662, 0.734] 100% 5
screenanyphenotype -0.057 [-0.143, 0.024] 0% 5
stageenrichedderived 0.676 [0.632, 0.712] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

04 · Keep only the modules that survive the whole walk (UMAP + HDBSCAN co-clustering) -- weak

Modules are built from a small walk on 1,500 genes without the label. 25% of the label is hidden; the module that best isolates each label is chosen on the visible genes. Metric: the size-weighted F1 of each label's hidden genes against its chosen module. Null: the same modules scored after permuting the hidden genes' labels, 100 times -- a large module scores the same either way and earns nothing. Pass: above the null's 95th percentile by 0.05.

Metric: F1 of hidden genes in the module chosen for their label on known genes. 75 runs, 3 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults threshold=0.5 0.073 [0.014, 0.127] 56% [37, 73] 0.386 0.336 25
tuned threshold=0.3 0.045 [-0.041, 0.104] 50% [24, 76] 0.37 0.336 10

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.063 [0.024, 0.096] 60% 5
dtm_class 0.087 [0.065, 0.109] 40% 5
lopit_unified 0.084 [0.039, 0.123] 80% 5
screenanyphenotype -0.033 [-0.119, 0.024] 0% 5
stageenrichedderived 0.149 [0.116, 0.186] 80% 5

Sensitivity (mean skill at each value, all other settings pooled):

05 · Tune a map without labels, then read what it encodes (UMAP + HDBSCAN, chi-square / Kruskal-Wallis) -- reliable

Pattern 5. One label-free map on up to 2,000 genes. Held-out features significant at q < 0.05 on a random half of the genes are the findings; each is re-tested (p < 0.05) on the other half. Metric: share that replicate. Null: the same with the second half's cluster labels permuted, 10 times. Pass: above the null's 95th percentile by 0.2, with at least three findings.

Metric: share of first-half findings that replicate on the second half. 70 runs, 14 settings, targets: (the strategy chooses its own).

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults map_from=transcription 0.949 [0.884, 0.997] 100% [57, 100] 0.952 0.059 5
tuned map_from=chemistry 0.964 [0.947, 0.982] 100% [34, 100] 0.966 0.046 2

Sensitivity (mean skill at each value, all other settings pooled):

06 · Find which kind of evidence carries a label (kNN ablation) -- reliable

25% of the label is hidden. Each kind of evidence is ranked by cross-validated accuracy on the visible labels; the top-ranked one then predicts the hidden labels. Metric: its hidden accuracy. Null: the hidden accuracy of every kind of evidence, i.e. choosing at random. Pass: above the null's 80th percentile by 0.02 (with ~15 kinds of evidence the 95th would demand the single best, which asks more than a ranking must deliver).

Metric: hidden accuracy of the evidence ranked first (transcription). 75 runs, 3 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults k=15 0.140 [0.060, 0.251] 64% [45, 80] 0.604 0.543 25
tuned k=5 0.159 [0.081, 0.247] 70% [40, 89] 0.59 0.525 10

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.094 [0.074, 0.114] 100% 5
dtm_class 0.078 [0.068, 0.087] 80% 5
lopit_unified 0.139 [0.127, 0.156] 100% 5
screenanyphenotype 0.165 [0.123, 0.227] 0% 5
stageenrichedderived 0.349 [0.316, 0.374] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

07 · Call a gene by the genes that behave like it (kNN) -- reliable

Pattern 1: 25% of the label hidden by whole orthogroups; each hidden gene called by its k nearest visible genes with the same vote threshold. Metric: hidden genes called correctly. Null: 10 runs with visible labels shuffled. Pass: above the null's 95th percentile by 0.05.

Metric: correct calls per hidden gene. 225 runs, 9 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults k=15; min_share=0.3 0.225 [0.079, 0.358] 60% [41, 77] 0.619 0.476 25
tuned k=5; min_share=0.0 0.292 [0.163, 0.428] 80% [49, 94] 0.629 0.468 10

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.297 [0.278, 0.316] 100% 5
dtm_class 0.181 [0.173, 0.190] 80% 5
lopit_unified 0.373 [0.364, 0.383] 100% 5
screenanyphenotype 0.050 [-0.029, 0.111] 0% 5
stageenrichedderived 0.496 [0.454, 0.537] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

08 · Call a gene by its neighbours on the map (UMAP + kNN) -- weak

Pattern 1 on the genes the map places (up to 2,500): 25% of the label hidden by whole orthogroups; hidden genes called by their k nearest visible genes in the map. Null: 20 runs on shuffled visible labels. Pass: above the null's 95th percentile by 0.05.

Metric: correct calls per hidden gene. 225 runs, 9 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults k=15; n_neighbors=25 0.125 [0.008, 0.236] 60% [41, 77] 0.565 0.474 25
tuned k=5; n_neighbors=60 0.114 [-0.037, 0.229] 60% [31, 83] 0.538 0.46 10

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.169 [0.148, 0.188] 100% 5
dtm_class 0.034 [0.004, 0.057] 0% 5
lopit_unified 0.231 [0.217, 0.243] 100% 5
screenanyphenotype -0.056 [-0.180, 0.028] 0% 5
stageenrichedderived 0.296 [0.260, 0.328] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

09 · Name a cluster by the label it is enriched for (UMAP + HDBSCAN, hypergeometric) -- reliable

Pattern 1 scored by precision, on the genes a blind map places (up to 2,500): 25% of the label hidden; the enrichment is recomputed from visible labels, and the calls it makes on hidden genes are scored -- the strategy abstains on noise and unenriched clusters by design, so what matters is how often a call is right. Null: 20 runs on shuffled visible labels. Pass: above the null's 95th percentile by 0.1.

Metric: precision of calls on hidden genes. 300 runs, 12 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults minclustersize=25; min_lift=1.5; selection=leaf 0.332 [0.249, 0.423] 100% [84, 100] 0.333 0.002 20
tuned minclustersize=10; min_lift=3.0; selection=leaf 0.409 [0.303, 0.501] 100% [68, 100] 0.411 0.004 8

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.367 [0.271, 0.460] 100% 5
dtm_class 0.276 [0.207, 0.356] 100% 5
lopit_unified 0.548 [0.493, 0.603] 100% 5
screenanyphenotype nan [nan, nan] nan% 0
stageenrichedderived 0.480 [0.402, 0.557] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

10 · Find genes whose label their neighbours contradict (kNN + network neighbours) -- reliable

5% of the labels (at least ten) are swapped to a wrong class, drawn in proportion to class size. Surprise is computed with the corrupted labels. Metric: AUROC of surprise for the swapped genes among all labelled genes. Null: 20 random sets of the same size. Pass: above the null's 95th percentile by 0.1.

Metric: AUROC of surprise for the swapped labels. 75 runs, 3 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults k=15 0.638 [0.490, 0.769] 100% [87, 100] 0.818 0.498 25
tuned k=15 0.650 [0.505, 0.776] 100% [72, 100] 0.823 0.496 10

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.608 [0.583, 0.631] 100% 5
dtm_class 0.796 [0.789, 0.804] 100% 5
lopit_unified 0.633 [0.610, 0.655] 100% 5
screenanyphenotype 0.342 [0.281, 0.396] 100% 5
stageenrichedderived 0.810 [0.797, 0.824] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

11 · Diffuse a label across one measured network (random walk with restart) -- reliable

Pattern 1: 25% of the label hidden by whole orthogroups; the fields are seeded from visible genes only, so a hidden gene never seeds its own call. Metric: hidden genes called correctly (unreached genes count as misses). Null: 10 runs on shuffled labels. Pass: above the null's 95th percentile by 0.05.

Metric: correct calls per hidden gene. 900 runs, 36 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults layer=coexpression; restart=0.5 0.125 [0.082, 0.169] 75% [53, 89] 0.291 0.186 20
tuned layer=coexpression; restart=0.2 0.136 [0.095, 0.176] 88% [53, 98] 0.288 0.175 8

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.089 [0.082, 0.095] 100% 5
dtm_class 0.159 [0.150, 0.165] 100% 5
lopit_unified 0.188 [0.176, 0.199] 100% 5
screenanyphenotype 0.099 [0.088, 0.114] 40% 5
stageenrichedderived nan [nan, nan] nan% 0

Sensitivity (mean skill at each value, all other settings pooled):

12 · Let every network vote, weighted by what it has earned (chance-weighted ensemble vote) -- weak

Pattern 1: 25% of the label hidden; the weights are learned on an inner holdout of the visible labels only, then hidden genes are called. Null: 5 runs with visible labels shuffled (weights relearned each time). Pass: above the null's 95th percentile by 0.05.

Metric: correct calls per hidden gene. 75 runs, 3 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults k=15 0.300 [-0.034, 0.536] 80% [61, 91] 0.577 0.383 25
tuned k=15 0.246 [-0.208, 0.542] 80% [49, 94] 0.559 0.392 10

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.321 [0.313, 0.330] 100% 5
dtm_class 0.495 [0.406, 0.570] 100% 5
lopit_unified 0.394 [0.362, 0.432] 100% 5
screenanyphenotype -0.334 [-0.780, 0.088] 0% 5
stageenrichedderived 0.625 [0.529, 0.723] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

13 · Place a protein by the proteins it physically touches (weighted partner vote) -- reliable

Pattern 1 restricted to genes with at least one physical partner: 25% of the label hidden by whole orthogroups; hidden genes called by their visible partners' weighted vote. Null: 20 runs on shuffled labels. Pass: above the null's 95th percentile by 0.05.

Metric: correct calls per hidden gene. 25 runs, 1 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults -- 0.249 [0.076, 0.428] 60% [41, 77] 0.385 0.194 25
tuned -- 0.257 [0.075, 0.426] 60% [31, 83] 0.39 0.189 10

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.428 [0.414, 0.443] 100% 5
dtm_class 0.302 [0.280, 0.335] 100% 5
lopit_unified 0.488 [0.462, 0.510] 100% 5
screenanyphenotype 0.008 [-0.013, 0.021] 0% 5
stageenrichedderived 0.021 [0.008, 0.035] 0% 5

14 · Annotate function through shared fold (TM-score-weighted vote) -- reliable

Pattern 1 restricted to proteins with a structural neighbour: 25% of the annotation (at the chosen level) hidden by whole orthogroups; hidden proteins called by their visible structural neighbours. Null: 20 runs on shuffled annotations. Pass: above the null's 95th percentile by 0.05.

Metric: correct calls per hidden gene. 15 runs, 3 settings, targets: (the strategy chooses its own).

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults level=1 0.637 [0.602, 0.664] 100% [57, 100] 0.733 0.266 5
tuned level=3 0.713 [0.706, 0.719] 100% [34, 100] 0.736 0.083 2

Sensitivity (mean skill at each value, all other settings pooled):

15 · Find the communities several networks agree on (modularity + Louvain consensus) -- weak

On genes placed in a community: 25% of the label hidden; the community that best isolates each label is chosen on the visible genes (the communities themselves never see labels). Metric: the size-weighted F1 of each label's hidden genes against its chosen community. Null: the same communities scored after permuting the hidden genes' labels, 100 times. Pass: above the null's 95th percentile by 0.05.

Metric: F1 of hidden genes in the community chosen for their label on known genes. 150 runs, 6 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults agreement=0.5; resolution=1.0 0.054 [0.017, 0.085] 40% [23, 59] 0.345 0.304 25
tuned agreement=0.5; resolution=1.0 0.048 [-0.001, 0.087] 40% [17, 69] 0.338 0.299 10

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.041 [0.037, 0.045] 0% 5
dtm_class 0.071 [0.062, 0.080] 60% 5
lopit_unified 0.098 [0.091, 0.104] 100% 5
screenanyphenotype -0.015 [-0.038, 0.007] 0% 5
stageenrichedderived 0.076 [0.053, 0.101] 40% 5

Sensitivity (mean skill at each value, all other settings pooled):

16 · Predict the contacts an interactome missed (logistic regression) -- reliable

Pattern 3: 20% of the layer's edges hidden; the model is trained on the rest against DEGREE-MATCHED non-edges -- each with a gene of similar degree at both ends -- and scores the hidden edges against fresh degree-matched non-edges. Against random non-edges this test read AUROC 0.99 on every correlation layer, because a random pair is usually two obscure genes and degree separates them. The layer's source measurements leave the similarity feature with it, and derived, annotation and literature layers are refused as targets. Metric: AUROC. Null: 5 models trained with each gene's evidence read from a random other gene (identities permuted). Pass: above the null's 95th percentile by 0.05.

Metric: AUROC of hidden pairs against degree-matched non-pairs. 10 runs, 2 settings, targets: (the strategy chooses its own).

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults layer=xlms 0.606 [0.590, 0.623] 100% [57, 100] 0.805 0.504 5
tuned layer=struct 0.926 [0.925, 0.927] 100% [34, 100] 0.963 0.506 2

Sensitivity (mean skill at each value, all other settings pooled):

17 · Read the literature for biology, not fame (publication-count residual) -- reliable

The label is never used to build the literature layer. Among co-mentioned pairs with both genes labelled, the top k by corrected residual are taken (k = the chosen number, capped at a fifth of the pool). Metric: the share of those pairs sharing a label. Null: 20 random sets of k co-mentioned pairs. Pass: above the null's 95th percentile by 0.05.

Metric: share of the top 50 corrected pairs sharing a compartment label. 75 runs, 3 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults top=200 0.376 [0.209, 0.545] 80% [61, 91] 0.707 0.523 25
tuned top=200 0.379 [0.210, 0.547] 80% [49, 94] 0.707 0.521 10

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.587 [0.585, 0.589] 100% 5
dtm_class 0.155 [0.141, 0.169] 100% 5
lopit_unified 0.609 [0.602, 0.615] 100% 5
screenanyphenotype 0.183 [0.169, 0.197] 0% 5
stageenrichedderived 0.344 [0.322, 0.358] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

18 · List what the data says and the literature has not written (multi-layer support count) -- reliable

Pattern 3 with the literature as truth: co-mentioned pairs among genes in the measurement layers against 5 times as many random pairs. Score: number of measurement layers linking the pair. Metric: AUROC. Null: 10 runs with the measurement layers' gene identities permuted. Pass: above the null's 95th percentile by 0.02.

Metric: AUROC of hidden pairs against random non-pairs. 15 runs, 3 settings, targets: (the strategy chooses its own).

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults min_layers=2 0.112 [0.111, 0.112] 100% [57, 100] 0.556 0.5 5
tuned min_layers=1 0.111 [0.111, 0.111] 100% [34, 100] 0.556 0.5 2

Sensitivity (mean skill at each value, all other settings pooled):

19 · Train a classifier on the known genes and call the rest (logistic regression) -- reliable

Pattern 1: 25% of the label hidden by whole orthogroups; the model is trained on the visible genes and calls the hidden ones. Metric: hidden genes called correctly. Null: the analytic chance level for the same predicted and true class mixes (re-running a multinomial fit on shuffled labels ten times would take longer than it tells). Pass: above that level's 95% bound by 0.05.

Metric: correct calls per hidden gene. 100 runs, 4 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults C=0.01 0.318 [0.180, 0.465] 80% [61, 91] 0.542 0.326 25
tuned C=10.0 0.369 [0.203, 0.498] 80% [49, 94] 0.612 0.362 10

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.407 [0.394, 0.417] 100% 5
dtm_class 0.339 [0.322, 0.360] 100% 5
lopit_unified 0.496 [0.488, 0.505] 100% 5
screenanyphenotype 0.043 [-0.018, 0.111] 0% 5
stageenrichedderived 0.573 [0.542, 0.603] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

20 · Learn what makes your list special, from positives alone (PU bagging, logistic regression) -- reliable

Pattern 2: 30% of the set hidden; the other 70% are the positives. Metric: AUROC of the hidden members against every other non-query gene. Null: 10 random sets of the same size. Pass: above the null's 95th percentile by 0.1.

Metric: AUROC of hidden members against every other gene. 75 runs, 3 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults bags=15 0.787 [0.722, 0.849] 100% [87, 100] 0.895 0.504 25
tuned bags=5 0.800 [0.741, 0.853] 100% [72, 100] 0.9 0.5 10

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.880 [0.869, 0.892] 100% 5
dtm_class 0.663 [0.635, 0.691] 100% 5
lopit_unified 0.840 [0.827, 0.853] 100% 5
screenanyphenotype 0.787 [0.768, 0.805] 100% 5
stageenrichedderived 0.757 [0.745, 0.777] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

21 · Predict a measurement, and find the genes that defy the prediction (gradient boosting / ridge) -- reliable

Pattern 4: 20% of the measured values hidden; the model is trained on the rest. Metric: rank correlation between predicted and hidden values. Null: 3 models trained on the visible values shuffled among the measured genes. Pass: above the null's 95th percentile by 0.1.

Metric: rank correlation of predicted and hidden values. 60 runs, 4 settings, targets: exprtachy, fitinvitrohff, fitinvivo_PE.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults model=boosted; own_kind=leave out 0.560 [0.051, 0.909] 67% [42, 85] 0.56 0.0 15
tuned model=boosted; own_kind=include 0.666 [0.148, 0.968] 100% [61, 100] 0.666 0.0 6

Tuned setting, per held-out target:

target skill [95% CI] pass runs
expr_tachy 0.969 [0.965, 0.972] 100% 5
fitinvitrohff 0.892 [0.890, 0.893] 100% 5
fitinvivoPE 0.128 [0.114, 0.148] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

22 · Fill in what was never measured, and say where that is honest (soft-impute, low-rank SVD) -- reliable

Pattern 4 across the whole table: 10% of every column's measured entries hidden, the table completed at the chosen rank. Metric: median over columns of the rank correlation on hidden entries. Null: 3 completions of a table whose columns were each shuffled independently. Pass: above the null's 95th percentile by 0.1.

Metric: median per-column rank correlation on hidden entries. 15 runs, 3 settings, targets: (the strategy chooses its own).

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults rank=20 0.868 [0.864, 0.873] 100% [57, 100] 0.868 0.001 5
tuned rank=60 0.903 [0.902, 0.904] 100% [34, 100] 0.903 -0.001 2

Sensitivity (mean skill at each value, all other settings pooled):

23 · Find what matters more in one condition, and why (residual + gradient boosting / ridge) -- weak

Pattern 4 on the shift: 20% of genes with both measurements hidden; a model trained on the rest predicts their shift. Metric: rank correlation on hidden genes. Null: 3 models trained on shuffled shifts. Pass: above the null's 95th percentile by 0.1.

Metric: rank correlation of predicted and hidden values. 20 runs, 4 settings, targets: (the strategy chooses its own).

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults model=boosted; own_kind=leave out 0.073 [0.060, 0.085] 0% [0, 43] 0.073 0.0 5
tuned model=ridge; own_kind=include 0.068 [0.055, 0.082] 0% [0, 66] 0.068 0.0 2

Sensitivity (mean skill at each value, all other settings pooled):

24 · Describe what your gene list has in common (hypergeometric + rank-sum) -- reliable

Pattern 2: 40% of the set hidden; the profile is built from the other 60% and scores every gene. Metric: AUROC of the hidden members against every other non-query gene. Null: 10 random sets of the same size. Pass: above the null's 95th percentile by 0.1.

Metric: AUROC of hidden members against every other gene. 25 runs, 1 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults -- 0.615 [0.510, 0.719] 100% [87, 100] 0.807 0.499 25
tuned -- 0.644 [0.518, 0.749] 100% [72, 100] 0.822 0.5 10

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.791 [0.766, 0.814] 100% 5
dtm_class 0.565 [0.476, 0.658] 100% 5
lopit_unified 0.678 [0.643, 0.706] 100% 5
screenanyphenotype 0.598 [0.528, 0.672] 100% 5
stageenrichedderived 0.442 [0.389, 0.498] 100% 5

25 · Grow your gene list along the networks (random walk with restart) -- reliable

Pattern 2: 30% of the set hidden; the walk is seeded from the other 70%. Metric: AUROC of the hidden members against every other non-seed gene. Null: 10 random seed sets of the same size. Pass: above the null's 95th percentile by 0.1.

Metric: AUROC of hidden members against every other gene. 150 runs, 6 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults mode=networks + measurements; restart=0.3 0.594 [0.488, 0.683] 100% [87, 100] 0.799 0.506 25
tuned mode=networks + measurements; restart=0.6 0.608 [0.533, 0.669] 100% [72, 100] 0.806 0.504 10

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.696 [0.655, 0.737] 100% 5
dtm_class 0.604 [0.601, 0.609] 100% 5
lopit_unified 0.698 [0.661, 0.729] 100% 5
screenanyphenotype 0.590 [0.565, 0.624] 100% 5
stageenrichedderived 0.391 [0.339, 0.445] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

26 · Find categories that split in two on another measurement (UMAP + HDBSCAN) -- untestable

Pattern 5: significant findings (q < 0.05) on a random half of the mapped genes; each is checked on the other half -- the minority group enriched again among the cluster's A-matching genes (hypergeometric p < 0.05), or the measurement bimodal again. Metric: share replicating. Null: 20 runs with B shuffled within the second half. Pass: above the null's 95th percentile by 0.2 with at least three findings.

Metric: share of findings that replicate. 15 runs, 3 settings, targets: (the strategy chooses its own).

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults minclustersize=40 nan [nan, nan] nan% [nan, nan] nan nan 0

Sensitivity (mean skill at each value, all other settings pooled):

27 · Find kinds of gene defined by two labels at once (UMAP + HDBSCAN) -- reliable

Pattern 5: conjunctions found on a random half; each is checked on the other half as the second label's enrichment in the cluster among genes carrying the first label (hypergeometric p < 0.05). Metric: share replicating. Null: 20 runs with the second label shuffled within the second half. Pass: above the null's 95th percentile by 0.2 with at least three findings.

Metric: share of first-half findings that replicate on the second half. 15 runs, 3 settings, targets: (the strategy chooses its own).

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults minclustersize=15 0.490 [0.376, 0.623] 100% [57, 100] 0.502 0.027 5
tuned minclustersize=15 0.531 [0.322, 0.740] 100% [34, 100] 0.542 0.027 2

Sensitivity (mean skill at each value, all other settings pooled):

28 · Find paralogs that changed jobs (profile correlation) -- reliable

Paralog pairs with both genes labelled; the labels are withheld from the profiles. Metric: AUROC of divergence for pairs whose labels differ against pairs whose labels match. Null: 20 random reassignments of the divergence values to pairs. Pass: above the null's 95th percentile by 0.05.

Metric: AUROC of profile divergence for paralogs with different compartment. 75 runs, 3 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults min_shared=10 0.164 [0.057, 0.312] 60% [39, 78] 0.582 0.5 20
tuned min_shared=10 0.164 [0.056, 0.312] 62% [31, 86] 0.582 0.5 8

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.094 [0.085, 0.103] 40% 5
dtm_class 0.026 [0.022, 0.030] 0% 5
lopit_unified 0.151 [0.141, 0.160] 100% 5
screenanyphenotype nan [nan, nan] nan% 0
stageenrichedderived 0.384 [0.376, 0.392] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

29 · Carry what one parasite shows to the other (orthogroup mapping) -- reliable

Numeric target: 25% of the genes measured in both species hidden; the relation is learned on the rest. Metric: rank correlation of transferred and hidden values. Null: 20 runs with the ortholog values permuted among genes. Categorical target: Pattern 1 on genes with an ortholog value. Pass: above the null's 95th percentile by 0.1 (numeric) or 0.05 (categorical).

Metric: rank correlation of transferred and hidden values. 5 runs, 1 settings, targets: (the strategy chooses its own).

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults -- 0.314 [0.298, 0.333] 100% [57, 100] 0.316 0.003 5
tuned -- 0.321 [0.293, 0.349] 100% [34, 100] 0.323 0.003 2

30 · Test inference on the genes orthology cannot reach (kNN) -- reliable

Pattern 1 restricted to the stratum: 25% of the label hidden by whole orthogroups; only hidden genes inside the stratum are scored. Null: 10 runs on shuffled labels. Pass: above the null's 95th percentile by 0.05.

Metric: correct calls per hidden gene. 100 runs, 4 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults stratum=lineage-specific 0.264 [0.146, 0.345] 75% [53, 89] 0.586 0.416 20
tuned stratum=lineage-specific 0.281 [0.166, 0.358] 75% [41, 93] 0.591 0.412 8

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.291 [0.251, 0.321] 100% 5
dtm_class 0.086 [0.066, 0.107] 0% 5
lopit_unified 0.349 [0.339, 0.361] 100% 5
screenanyphenotype nan [nan, nan] nan% 0
stageenrichedderived 0.332 [0.290, 0.371] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

31 · Call a gene only when independent strategies agree (kNN + logistic + network vote) -- reliable

Pattern 1 scored by precision: 25% of the label hidden; each method is trained on the visible genes and the agreed calls on hidden genes are scored. Metric: share of agreed calls that are correct. Null: 5 runs with visible labels shuffled. Pass: above the null's 95th percentile by 0.1. Single-method precisions are reported.

Metric: precision of calls on hidden genes. 75 runs, 3 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults min_agree=2 0.348 [0.174, 0.485] 60% [41, 77] 0.705 0.513 25
tuned min_agree=3 0.608 [0.331, 0.774] 80% [49, 94] 0.835 0.523 10

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.723 [0.676, 0.769] 100% 5
dtm_class 0.790 [0.761, 0.819] 100% 5
lopit_unified 0.730 [0.711, 0.747] 100% 5
screenanyphenotype -0.034 [-0.152, 0.070] 0% 5
stageenrichedderived 0.781 [0.752, 0.818] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

32 · Put the understudied genes first (kNN + logistic + network vote) -- reliable

Pattern 1 scored by precision and restricted to understudied genes: 25% of the label hidden; agreed calls are scored only on hidden understudied genes. Null: 5 runs with visible labels shuffled. Pass: above the null's 95th percentile by 0.1.

Metric: precision of calls on hidden genes. 75 runs, 3 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults min_agree=2 0.320 [0.134, 0.456] 60% [41, 77] 0.702 0.53 25
tuned min_agree=3 0.585 [0.318, 0.754] 80% [49, 94] 0.831 0.545 10

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.689 [0.615, 0.760] 100% 5
dtm_class 0.773 [0.741, 0.805] 100% 5
lopit_unified 0.702 [0.680, 0.724] 100% 5
screenanyphenotype -0.060 [-0.208, 0.038] 0% 5
stageenrichedderived 0.754 [0.728, 0.790] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

33 · Put every layer into one space and read a gene's neighbourhood (logistic edge model) -- reliable

A layer's edges are hidden by orthogroup -- whole groups at a time, so no hidden edge survives through a visible paralog -- and the layer is removed from its own features. Metric: AUROC of the hidden edges against DEGREE-MATCHED non-pairs, one per hidden edge, matched on connectivity at both ends. Null: 10 configuration-model rewirings of the hidden edges, which keep their degree sequence and destroy their topology, so a model reading fame cannot beat it. Pass: above the null's 95th percentile by 0.05. The AUROC against random non-pairs and the gap between the two are reported as numbers, not as the verdict.

Metric: AUROC of hidden coexpression edges against degree-matched non-pairs. 150 runs, 30 settings, targets: (the strategy chooses its own).

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults k=10; knn=15; layer=coexpression 0.587 [0.582, 0.592] 100% [57, 100] 0.79 0.491 5
tuned k=10; knn=15; layer=cotranslation 0.901 [0.864, 0.939] 100% [34, 100] 0.949 0.481 2

Sensitivity (mean skill at each value, all other settings pooled):

34 · Train on the networks and rank the edges they are missing (logistic / spectral embedding) -- reliable

The chosen layer's edges are hidden by orthogroup and the layer is removed from its own features. Metric: AUROC of the hidden edges against DEGREE-MATCHED non-pairs. Null: 10 configuration-model rewirings of the hidden edges, which preserve their degree sequence, so fame alone cannot clear the bar. Pass: above the null's 95th percentile by 0.05. The random-null AUROC, the fame gap, precision@k, the Brier score and the reliability gap are all reported as numbers beside the verdict.

Metric: AUROC of hidden coexpression edges against degree-matched non-pairs. 100 runs, 20 settings, targets: (the strategy chooses its own).

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults fraction=0.25; layer=coexpression; model=logistic 0.587 [0.582, 0.592] 100% [57, 100] 0.79 0.491 5
tuned fraction=0.4; layer=cotranslation; model=logistic 0.880 [0.877, 0.884] 100% [34, 100] 0.939 0.492 2

Sensitivity (mean skill at each value, all other settings pooled):

35 · Call genes with a stated error rate (split conformal prediction) -- reliable

25% of the label hidden by whole orthogroups; the model is trained and calibrated on the rest (itself split by orthogroup) and builds a set for every hidden gene. Metric: set efficiency, 1 - (mean set size - 1) / (classes - 1) -- how far the sets narrow the possibilities. Null: 10 runs on shuffled labels, where the model learns nothing and the sets must grow to keep the promise. Pass: above the null's 95th percentile by 0.05. Also reported: set coverage on the hidden genes beside the 1 - alpha promised, and the scorecard of the single-label calls.

Metric: set efficiency: 1 - (mean set size - 1) / (classes - 1). 150 runs, 6 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults alpha=0.1; model=logistic 0.500 [0.211, 0.729] 80% [61, 91] 0.588 0.176 25
tuned alpha=0.2; model=kNN 0.496 [0.203, 0.765] 80% [49, 94] 0.609 0.186 10

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.449 [0.405, 0.495] 100% 5
dtm_class 0.399 [0.378, 0.418] 100% 5
lopit_unified 0.652 [0.574, 0.725] 100% 5
screenanyphenotype 0.071 [-0.040, 0.156] 0% 5
stageenrichedderived 0.884 [0.868, 0.903] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

36 · Smooth the measurements along the networks, then classify (graph convolution + logistic regression) -- reliable

Pattern 1: 25% of the label hidden by whole orthogroups; the smoothed features are computed from measurements only, and the model is trained on the visible genes. Metric: hidden genes called correctly. Null: the analytic chance level for the same predicted and true class mixes. Pass: above that level's 95% bound by 0.05.

Metric: correct calls per hidden gene. 150 runs, 6 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults C=0.5; hops=2 0.399 [0.164, 0.569] 80% [61, 91] 0.629 0.359 25
tuned C=0.5; hops=1 0.401 [0.185, 0.560] 80% [49, 94] 0.629 0.358 10

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.441 [0.431, 0.451] 100% 5
dtm_class 0.441 [0.425, 0.454] 100% 5
lopit_unified 0.522 [0.510, 0.534] 100% 5
screenanyphenotype -0.045 [-0.088, -0.001] 0% 5
stageenrichedderived 0.646 [0.621, 0.673] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

37 · Let a random forest find what defines a label (random forest + permutation importance) -- reliable

Pattern 1: 25% of the label hidden by whole orthogroups; the forest is trained on the visible genes and calls the hidden ones. Metric: hidden genes called correctly. Null: the analytic chance level for the same predicted and true class mixes. Pass: above that level's 95% bound by 0.05.

Metric: correct calls per hidden gene. 150 runs, 6 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults min_leaf=2; trees=300 0.454 [0.216, 0.647] 80% [61, 91] 0.701 0.444 25
tuned min_leaf=5; trees=300 0.451 [0.210, 0.659] 80% [49, 94] 0.696 0.43 10

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.426 [0.406, 0.446] 100% 5
dtm_class 0.770 [0.756, 0.788] 100% 5
lopit_unified 0.502 [0.491, 0.511] 100% 5
screenanyphenotype -0.009 [-0.052, 0.030] 0% 5
stageenrichedderived 0.651 [0.628, 0.673] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

38 · Learn how much to trust each kind of evidence (stacked logistic regression) -- reliable

Pattern 1: 25% of the label hidden by whole orthogroups; base predictions for the meta-model are out of fold within the visible genes only, so no hidden label reaches either level. Metric: hidden genes called correctly. Null: the analytic chance level for the same predicted and true class mixes. Pass: above that level's 95% bound by 0.05.

Metric: correct calls per hidden gene. 75 runs, 3 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults k=15 0.408 [0.183, 0.573] 80% [61, 91] 0.619 0.342 25
tuned k=10 0.406 [0.198, 0.559] 80% [49, 94] 0.613 0.337 10

Tuned setting, per held-out target:

target skill [95% CI] pass runs
compartment 0.435 [0.419, 0.450] 100% 5
dtm_class 0.456 [0.442, 0.470] 100% 5
lopit_unified 0.541 [0.533, 0.550] 100% 5
screenanyphenotype -0.021 [-0.049, 0.007] 0% 5
stageenrichedderived 0.636 [0.613, 0.658] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

39 · Predict a value with an interval that holds (gradient boosting / ridge + split conformal) -- weak

Pattern 4: 20% of the measured values hidden; the model is trained and calibrated on the rest (split by orthogroup) and predicts the hidden ones. Metric: rank correlation of predicted and hidden values. Null: the chance distribution of a rank correlation, with refits on shuffled values reported. Pass: above its 95th percentile by 0.1. Also reported: the share of hidden values inside their intervals, beside the coverage promised.

Metric: rank correlation of predicted and hidden values. 60 runs, 4 settings, targets: exprtachy, fitinvitrohff, fitinvivo_PE.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults alpha=0.1; model=boosted 0.552 [0.036, 0.906] 67% [42, 85] 0.552 0.0 15
tuned alpha=0.2; model=ridge 0.532 [0.029, 0.858] 67% [30, 90] 0.532 0.0 6

Tuned setting, per held-out target:

target skill [95% CI] pass runs
expr_tachy 0.855 [0.841, 0.868] 100% 5
fitinvitrohff 0.701 [0.695, 0.708] 100% 5
fitinvivoPE 0.046 [0.029, 0.063] 0% 5

Sensitivity (mean skill at each value, all other settings pooled):

Plasmodium falciparum

01 · Hold out a category and search for a map that finds it (UMAP + HDBSCAN) -- weak

25% of the label is hidden (whole orthogroups together, so no gene is recovered through a visible paralog). The walk is built as set -- its genes per map (up to 4,000), feature sets and grids (up to three values each); the configuration, and the one cluster that best isolates each label, are both chosen using visible labels only. Metric: for each label, the F1 of its hidden genes against its chosen cluster, weighted by size -- does the structure found on known genes hold the unknown ones? Null: the same chosen clusters scored after permuting the hidden genes' labels, 100 times. (A second search on shuffled labels was the null once; it picks the largest cluster for every label, which scores F1 near 2p by size alone and made the null beat real labels.) Pass: above the null's 95th percentile by at least 0.05.

Metric: F1 of hidden genes in the cluster chosen for their label on known genes. 120 runs, 6 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults minclustersize=[20, 50]; n_neighbors=[15, 50]; selection=[eom, leaf] 0.087 [0.009, 0.195] 25% [11, 47] 0.536 0.503 20
tuned minclustersize=[20, 50]; n_neighbors=[10, 30]; selection=[eom, leaf] 0.114 [0.019, 0.287] 25% [7, 59] 0.566 0.535 8

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported 0.345 [0.214, 0.475] 0% 5
lopitpflocation 0.094 [0.078, 0.109] 100% 5
pbtransferredphenotype -0.000 [-0.000, 0.000] 0% 5
stageenrichedderived 0.036 [0.014, 0.062] 0% 5

Sensitivity (mean skill at each value, all other settings pooled):

02 · Find the map where your gene list is one cluster (UMAP + HDBSCAN) -- reliable

30% of the set is hidden; the walk as set (genes per map up to 4,000, feature sets, up to three values of each grid) picks the cluster with the best F1 for the other 70%. Metric: F1 of the hidden members against that cluster's other genes -- precision is the share of the cluster's candidates that are hidden members, recall the share of hidden members among them. Null: 20 random sets of the same size through the same walk. Pass: above the null's 95th percentile by at least 0.05 (an F1 margin; random sets score about 0.04).

Metric: F1 of the hidden members against the best cluster's other genes. 80 runs, 4 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults minclustersize=[20, 50]; n_neighbors=[15, 50] 0.214 [0.093, 0.352] 90% [70, 97] 0.23 0.022 20
tuned minclustersize=[20, 50]; n_neighbors=[15, 50] 0.217 [0.085, 0.353] 75% [41, 93] 0.234 0.022 8

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported 0.242 [0.211, 0.274] 100% 5
lopitpflocation 0.151 [0.119, 0.200] 100% 5
pbtransferredphenotype 0.411 [0.388, 0.430] 100% 5
stageenrichedderived 0.051 [0.035, 0.067] 60% 5

Sensitivity (mean skill at each value, all other settings pooled):

03 · Ask which categories the data can rediscover (UMAP + neighbour AUROC) -- reliable

30% of the label is hidden. The atlas is built from visible genes only (leave-one-out neighbour AUROC per category); hidden genes are then scored by their visible neighbours. Metric: mean hidden-gene AUROC over the categories the atlas ranks in its top half. Null: the same with the visible labels shuffled, 20 times. Pass: above the null's 95th percentile by 0.05. The rank agreement between atlas and hidden recovery is reported.

Metric: hidden-gene AUROC of the categories the atlas ranks in its top half. 60 runs, 3 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults k=15 0.686 [0.565, 0.793] 100% [84, 100] 0.843 0.503 20
tuned k=15 0.652 [0.555, 0.717] 100% [68, 100] 0.828 0.506 8

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported 0.771 [0.705, 0.856] 100% 5
lopitpflocation 0.668 [0.641, 0.696] 100% 5
pbtransferredphenotype 0.506 [0.452, 0.562] 100% 5
stageenrichedderived 0.798 [0.721, 0.875] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

04 · Keep only the modules that survive the whole walk (UMAP + HDBSCAN co-clustering) -- weak

Modules are built from a small walk on 1,500 genes without the label. 25% of the label is hidden; the module that best isolates each label is chosen on the visible genes. Metric: the size-weighted F1 of each label's hidden genes against its chosen module. Null: the same modules scored after permuting the hidden genes' labels, 100 times -- a large module scores the same either way and earns nothing. Pass: above the null's 95th percentile by 0.05.

Metric: F1 of hidden genes in the module chosen for their label on known genes. 60 runs, 3 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults threshold=0.5 0.145 [0.034, 0.284] 45% [26, 66] 0.503 0.446 20
tuned threshold=0.5 0.186 [0.047, 0.387] 50% [22, 78] 0.519 0.453 8

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported 0.358 [0.222, 0.471] 0% 5
lopitpflocation 0.140 [0.114, 0.166] 100% 5
pbtransferredphenotype 0.076 [0.032, 0.108] 80% 5
stageenrichedderived 0.007 [-0.000, 0.020] 0% 5

Sensitivity (mean skill at each value, all other settings pooled):

05 · Tune a map without labels, then read what it encodes (UMAP + HDBSCAN, chi-square / Kruskal-Wallis) -- reliable

Pattern 5. One label-free map on up to 2,000 genes. Held-out features significant at q < 0.05 on a random half of the genes are the findings; each is re-tested (p < 0.05) on the other half. Metric: share that replicate. Null: the same with the second half's cluster labels permuted, 10 times. Pass: above the null's 95th percentile by 0.2, with at least three findings.

Metric: share of first-half findings that replicate on the second half. 250 runs, 50 settings, targets: (the strategy chooses its own).

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults map_from=transcription 0.876 [0.817, 0.934] 100% [57, 100] 0.882 0.054 5
tuned map_from=isoelectric 0.953 [0.906, 1.000] 100% [34, 100] 0.956 0.061 2

Sensitivity (mean skill at each value, all other settings pooled):

06 · Find which kind of evidence carries a label (kNN ablation) -- reliable

25% of the label is hidden. Each kind of evidence is ranked by cross-validated accuracy on the visible labels; the top-ranked one then predicts the hidden labels. Metric: its hidden accuracy. Null: the hidden accuracy of every kind of evidence, i.e. choosing at random. Pass: above the null's 80th percentile by 0.02 (with ~15 kinds of evidence the 95th would demand the single best, which asks more than a ranking must deliver).

Metric: hidden accuracy of the evidence ranked first (expr). 60 runs, 3 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults k=15 0.133 [0.066, 0.191] 65% [43, 82] 0.643 0.569 20
tuned k=5 0.145 [0.077, 0.203] 75% [41, 93] 0.63 0.553 8

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported 0.106 [0.053, 0.159] 0% 5
lopitpflocation 0.180 [0.167, 0.189] 100% 5
pbtransferredphenotype 0.150 [0.130, 0.173] 100% 5
stageenrichedderived 0.164 [0.094, 0.234] 80% 5

Sensitivity (mean skill at each value, all other settings pooled):

07 · Call a gene by the genes that behave like it (kNN) -- weak

Pattern 1: 25% of the label hidden by whole orthogroups; each hidden gene called by its k nearest visible genes with the same vote threshold. Metric: hidden genes called correctly. Null: 10 runs with visible labels shuffled. Pass: above the null's 95th percentile by 0.05.

Metric: correct calls per hidden gene. 180 runs, 9 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults k=15; min_share=0.3 0.076 [-0.369, 0.417] 50% [30, 70] 0.684 0.532 20
tuned k=15; min_share=0.3 0.219 [0.020, 0.423] 50% [22, 78] 0.699 0.539 8

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported -0.523 [-1.101, 0.041] 0% 5
lopitpflocation 0.453 [0.438, 0.471] 100% 5
pbtransferredphenotype 0.375 [0.359, 0.391] 100% 5
stageenrichedderived -0.000 [-0.055, 0.055] 0% 5

Sensitivity (mean skill at each value, all other settings pooled):

08 · Call a gene by its neighbours on the map (UMAP + kNN) -- weak

Pattern 1 on the genes the map places (up to 2,500): 25% of the label hidden by whole orthogroups; hidden genes called by their k nearest visible genes in the map. Null: 20 runs on shuffled visible labels. Pass: above the null's 95th percentile by 0.05.

Metric: correct calls per hidden gene. 180 runs, 9 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults k=15; n_neighbors=25 -0.003 [-0.531, 0.286] 60% [39, 78] 0.635 0.525 20
tuned k=50; n_neighbors=25 0.148 [0.025, 0.270] 62% [31, 86] 0.643 0.537 8

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported -0.581 [-1.235, -0.007] 0% 5
lopitpflocation 0.262 [0.250, 0.273] 100% 5
pbtransferredphenotype 0.273 [0.255, 0.285] 100% 5
stageenrichedderived 0.084 [0.019, 0.146] 20% 5

Sensitivity (mean skill at each value, all other settings pooled):

09 · Name a cluster by the label it is enriched for (UMAP + HDBSCAN, hypergeometric) -- reliable

Pattern 1 scored by precision, on the genes a blind map places (up to 2,500): 25% of the label hidden; the enrichment is recomputed from visible labels, and the calls it makes on hidden genes are scored -- the strategy abstains on noise and unenriched clusters by design, so what matters is how often a call is right. Null: 20 runs on shuffled visible labels. Pass: above the null's 95th percentile by 0.1.

Metric: precision of calls on hidden genes. 240 runs, 12 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults minclustersize=25; min_lift=3.0; selection=leaf 0.418 [0.326, 0.603] 100% [76, 100] 0.422 0.008 12
tuned minclustersize=25; min_lift=1.5; selection=leaf 0.560 [0.447, 0.703] 100% [68, 100] 0.564 0.011 8

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported 0.358 [0.212, 0.496] 100% 5
lopitpflocation 0.456 [0.408, 0.516] 100% 5
pbtransferredphenotype 0.760 [0.696, 0.814] 100% 5
stageenrichedderived 0.565 [0.417, 0.700] 100% 3

Sensitivity (mean skill at each value, all other settings pooled):

10 · Find genes whose label their neighbours contradict (kNN + network neighbours) -- reliable

5% of the labels (at least ten) are swapped to a wrong class, drawn in proportion to class size. Surprise is computed with the corrupted labels. Metric: AUROC of surprise for the swapped genes among all labelled genes. Null: 20 random sets of the same size. Pass: above the null's 95th percentile by 0.1.

Metric: AUROC of surprise for the swapped labels. 60 runs, 3 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults k=15 0.740 [0.573, 0.918] 100% [84, 100] 0.871 0.501 20
tuned k=15 0.765 [0.612, 0.924] 100% [68, 100] 0.882 0.497 8

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported 0.992 [0.991, 0.993] 100% 5
lopitpflocation 0.720 [0.700, 0.742] 100% 5
pbtransferredphenotype 0.521 [0.492, 0.547] 100% 5
stageenrichedderived 0.727 [0.659, 0.818] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

11 · Diffuse a label across one measured network (random walk with restart) -- reliable

Pattern 1: 25% of the label hidden by whole orthogroups; the fields are seeded from visible genes only, so a hidden gene never seeds its own call. Metric: hidden genes called correctly (unreached genes count as misses). Null: 10 runs on shuffled labels. Pass: above the null's 95th percentile by 0.05.

Metric: correct calls per hidden gene. 480 runs, 24 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults layer=coexpression; restart=0.5 0.217 [0.137, 0.331] 93% [70, 99] 0.446 0.317 15
tuned layer=coexpression; restart=0.2 0.244 [0.135, 0.442] 100% [61, 100] 0.45 0.308 6

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported 0.376 [0.265, 0.469] 100% 5
lopitpflocation 0.142 [0.135, 0.152] 100% 5
pbtransferredphenotype 0.177 [0.152, 0.201] 100% 5
stageenrichedderived nan [nan, nan] nan% 0

Sensitivity (mean skill at each value, all other settings pooled):

12 · Let every network vote, weighted by what it has earned (chance-weighted ensemble vote) -- reliable

Pattern 1: 25% of the label hidden; the weights are learned on an inner holdout of the visible labels only, then hidden genes are called. Null: 5 runs with visible labels shuffled (weights relearned each time). Pass: above the null's 95th percentile by 0.05.

Metric: correct calls per hidden gene. 60 runs, 3 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults k=15 0.512 [0.415, 0.650] 65% [43, 82] 0.646 0.323 20
tuned k=15 0.496 [0.395, 0.612] 62% [31, 86] 0.655 0.334 8

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported 0.705 [0.478, 0.847] 40% 5
lopitpflocation 0.468 [0.445, 0.490] 100% 5
pbtransferredphenotype 0.445 [0.423, 0.476] 100% 5
stageenrichedderived 0.429 [0.284, 0.531] 20% 5

Sensitivity (mean skill at each value, all other settings pooled):

13 · Place a protein by the proteins it physically touches (weighted partner vote) -- weak

Pattern 1 restricted to genes with at least one physical partner: 25% of the label hidden by whole orthogroups; hidden genes called by their visible partners' weighted vote. Null: 20 runs on shuffled labels. Pass: above the null's 95th percentile by 0.05.

Metric: correct calls per hidden gene. 20 runs, 1 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults -- 0.216 [0.015, 0.468] 33% [15, 58] 0.539 0.371 15
tuned -- 0.210 [0.026, 0.524] 33% [10, 70] 0.542 0.355 6

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported 0.149 [0.037, 0.341] 0% 5
lopitpflocation 0.481 [0.401, 0.524] 100% 5
pbtransferredphenotype 0.018 [-0.107, 0.109] 0% 5
stageenrichedderived nan [nan, nan] nan% 0

14 · Annotate function through shared fold (TM-score-weighted vote) -- reliable

Pattern 1 restricted to proteins with a structural neighbour: 25% of the annotation (at the chosen level) hidden by whole orthogroups; hidden proteins called by their visible structural neighbours. Null: 20 runs on shuffled annotations. Pass: above the null's 95th percentile by 0.05.

Metric: correct calls per hidden gene. 15 runs, 3 settings, targets: (the strategy chooses its own).

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults level=1 0.587 [0.544, 0.613] 100% [57, 100] 0.693 0.258 5
tuned level=2 0.588 [0.561, 0.615] 100% [34, 100] 0.644 0.137 2

Sensitivity (mean skill at each value, all other settings pooled):

15 · Find the communities several networks agree on (modularity + Louvain consensus) -- weak

On genes placed in a community: 25% of the label hidden; the community that best isolates each label is chosen on the visible genes (the communities themselves never see labels). Metric: the size-weighted F1 of each label's hidden genes against its chosen community. Null: the same communities scored after permuting the hidden genes' labels, 100 times. Pass: above the null's 95th percentile by 0.05.

Metric: F1 of hidden genes in the community chosen for their label on known genes. 120 runs, 6 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults agreement=0.5; resolution=1.0 0.029 [-0.032, 0.088] 50% [30, 70] 0.249 0.225 20
tuned agreement=0.3; resolution=2.0 0.055 [0.002, 0.100] 50% [22, 78] 0.194 0.148 8

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported 0.025 [0.020, 0.030] 0% 5
lopitpflocation 0.095 [0.085, 0.108] 100% 5
pbtransferredphenotype 0.083 [0.074, 0.094] 100% 5
stageenrichedderived 0.021 [-0.058, 0.100] 20% 5

Sensitivity (mean skill at each value, all other settings pooled):

16 · Predict the contacts an interactome missed (logistic regression) -- reliable

Pattern 3: 20% of the layer's edges hidden; the model is trained on the rest against DEGREE-MATCHED non-edges -- each with a gene of similar degree at both ends -- and scores the hidden edges against fresh degree-matched non-edges. Against random non-edges this test read AUROC 0.99 on every correlation layer, because a random pair is usually two obscure genes and degree separates them. The layer's source measurements leave the similarity feature with it, and derived, annotation and literature layers are refused as targets. Metric: AUROC. Null: 5 models trained with each gene's evidence read from a random other gene (identities permuted). Pass: above the null's 95th percentile by 0.05.

Metric: AUROC of hidden pairs against degree-matched non-pairs. 5 runs, 1 settings, targets: (the strategy chooses its own).

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults layer=struct 0.958 [0.957, 0.961] 100% [57, 100] 0.98 0.508 5
tuned layer=struct 0.961 [0.958, 0.963] 100% [34, 100] 0.98 0.504 2

Sensitivity (mean skill at each value, all other settings pooled):

17 · Read the literature for biology, not fame (publication-count residual) -- weak

The label is never used to build the literature layer. Among co-mentioned pairs with both genes labelled, the top k by corrected residual are taken (k = the chosen number, capped at a fifth of the pool). Metric: the share of those pairs sharing a label. Null: 20 random sets of k co-mentioned pairs. Pass: above the null's 95th percentile by 0.05.

Metric: share of the top 48 corrected pairs sharing a lopit_pf_location label. 60 runs, 3 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults top=200 0.263 [0.229, 0.287] 40% [20, 64] 0.727 0.624 15
tuned top=1000 0.258 [0.213, 0.297] 33% [10, 70] 0.727 0.628 6

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported 0.237 [0.187, 0.289] 0% 5
lopitpflocation 0.267 [0.262, 0.271] 100% 5
pbtransferredphenotype 0.285 [0.258, 0.309] 20% 5
stageenrichedderived nan [nan, nan] nan% 0

Sensitivity (mean skill at each value, all other settings pooled):

18 · List what the data says and the literature has not written (multi-layer support count) -- reliable

Pattern 3 with the literature as truth: co-mentioned pairs among genes in the measurement layers against 5 times as many random pairs. Score: number of measurement layers linking the pair. Metric: AUROC. Null: 10 runs with the measurement layers' gene identities permuted. Pass: above the null's 95th percentile by 0.02.

Metric: AUROC of hidden pairs against random non-pairs. 15 runs, 3 settings, targets: (the strategy chooses its own).

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults min_layers=2 0.108 [0.106, 0.109] 100% [57, 100] 0.554 0.5 5
tuned min_layers=1 0.110 [0.110, 0.110] 100% [34, 100] 0.554 0.499 2

Sensitivity (mean skill at each value, all other settings pooled):

19 · Train a classifier on the known genes and call the rest (logistic regression) -- reliable

Pattern 1: 25% of the label hidden by whole orthogroups; the model is trained on the visible genes and calls the hidden ones. Metric: hidden genes called correctly. Null: the analytic chance level for the same predicted and true class mixes (re-running a multinomial fit on shuffled labels ten times would take longer than it tells). Pass: above that level's 95% bound by 0.05.

Metric: correct calls per hidden gene. 80 runs, 4 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults C=0.01 0.388 [0.351, 0.435] 100% [84, 100] 0.63 0.409 20
tuned C=0.1 0.440 [0.360, 0.509] 100% [68, 100] 0.687 0.445 8

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported 0.418 [0.353, 0.483] 100% 5
lopitpflocation 0.510 [0.498, 0.522] 100% 5
pbtransferredphenotype 0.346 [0.326, 0.367] 100% 5
stageenrichedderived 0.456 [0.411, 0.506] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

20 · Learn what makes your list special, from positives alone (PU bagging, logistic regression) -- reliable

Pattern 2: 30% of the set hidden; the other 70% are the positives. Metric: AUROC of the hidden members against every other non-query gene. Null: 10 random sets of the same size. Pass: above the null's 95th percentile by 0.1.

Metric: AUROC of hidden members against every other gene. 60 runs, 3 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults bags=15 0.903 [0.801, 0.976] 100% [84, 100] 0.951 0.498 20
tuned bags=40 0.906 [0.809, 0.978] 100% [68, 100] 0.953 0.506 8

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported 0.917 [0.904, 0.933] 100% 5
lopitpflocation 0.944 [0.937, 0.952] 100% 5
pbtransferredphenotype 0.995 [0.994, 0.996] 100% 5
stageenrichedderived 0.757 [0.740, 0.777] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

21 · Predict a measurement, and find the genes that defy the prediction (gradient boosting / ridge) -- reliable

Pattern 4: 20% of the measured values hidden; the model is trained on the rest. Metric: rank correlation between predicted and hidden values. Null: 3 models trained on the visible values shuffled among the measured genes. Pass: above the null's 95th percentile by 0.1.

Metric: rank correlation of predicted and hidden values. 60 runs, 4 settings, targets: exprschizont, meanplddt, piggybac_mis.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults model=boosted; own_kind=leave out 0.518 [0.247, 0.849] 100% [80, 100] 0.518 0.0 15
tuned model=boosted; own_kind=include 0.523 [0.246, 0.847] 100% [61, 100] 0.523 0.0 6

Tuned setting, per held-out target:

target skill [95% CI] pass runs
expr_schizont 0.241 [0.195, 0.284] 100% 5
mean_plddt 0.852 [0.843, 0.861] 100% 5
piggybac_mis 0.463 [0.449, 0.476] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

22 · Fill in what was never measured, and say where that is honest (soft-impute, low-rank SVD) -- reliable

Pattern 4 across the whole table: 10% of every column's measured entries hidden, the table completed at the chosen rank. Metric: median over columns of the rank correlation on hidden entries. Null: 3 completions of a table whose columns were each shuffled independently. Pass: above the null's 95th percentile by 0.1.

Metric: median per-column rank correlation on hidden entries. 15 runs, 3 settings, targets: (the strategy chooses its own).

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults rank=20 0.678 [0.672, 0.685] 100% [57, 100] 0.679 0.001 5
tuned rank=20 0.679 [0.678, 0.681] 100% [34, 100] 0.68 0.002 2

Sensitivity (mean skill at each value, all other settings pooled):

23 · Find what matters more in one condition, and why (residual + gradient boosting / ridge) -- reliable

Pattern 4 on the shift: 20% of genes with both measurements hidden; a model trained on the rest predicts their shift. Metric: rank correlation on hidden genes. Null: 3 models trained on shuffled shifts. Pass: above the null's 95th percentile by 0.1.

Metric: rank correlation of predicted and hidden values. 20 runs, 4 settings, targets: (the strategy chooses its own).

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults model=boosted; own_kind=leave out 0.642 [0.637, 0.647] 100% [57, 100] 0.642 0.0 5
tuned model=boosted; own_kind=include 0.642 [0.641, 0.644] 100% [34, 100] 0.642 0.0 2

Sensitivity (mean skill at each value, all other settings pooled):

24 · Describe what your gene list has in common (hypergeometric + rank-sum) -- reliable

Pattern 2: 40% of the set hidden; the profile is built from the other 60% and scores every gene. Metric: AUROC of the hidden members against every other non-query gene. Null: 10 random sets of the same size. Pass: above the null's 95th percentile by 0.1.

Metric: AUROC of hidden members against every other gene. 20 runs, 1 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults -- 0.852 [0.755, 0.937] 100% [84, 100] 0.926 0.501 20
tuned -- 0.854 [0.756, 0.937] 100% [68, 100] 0.927 0.501 8

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported 0.923 [0.917, 0.930] 100% 5
lopitpflocation 0.837 [0.823, 0.852] 100% 5
pbtransferredphenotype 0.949 [0.946, 0.952] 100% 5
stageenrichedderived 0.700 [0.664, 0.730] 100% 5

25 · Grow your gene list along the networks (random walk with restart) -- reliable

Pattern 2: 30% of the set hidden; the walk is seeded from the other 70%. Metric: AUROC of the hidden members against every other non-seed gene. Null: 10 random seed sets of the same size. Pass: above the null's 95th percentile by 0.1.

Metric: AUROC of hidden members against every other gene. 120 runs, 6 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults mode=networks + measurements; restart=0.3 0.859 [0.744, 0.959] 100% [84, 100] 0.93 0.501 20
tuned mode=networks + measurements; restart=0.6 0.854 [0.734, 0.961] 100% [68, 100] 0.93 0.508 8

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported 0.907 [0.884, 0.937] 100% 5
lopitpflocation 0.846 [0.823, 0.870] 100% 5
pbtransferredphenotype 0.992 [0.988, 0.994] 100% 5
stageenrichedderived 0.701 [0.656, 0.746] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

26 · Find categories that split in two on another measurement (UMAP + HDBSCAN) -- untestable

Pattern 5: significant findings (q < 0.05) on a random half of the mapped genes; each is checked on the other half -- the minority group enriched again among the cluster's A-matching genes (hypergeometric p < 0.05), or the measurement bimodal again. Metric: share replicating. Null: 20 runs with B shuffled within the second half. Pass: above the null's 95th percentile by 0.2 with at least three findings.

Metric: share of findings that replicate. 15 runs, 3 settings, targets: (the strategy chooses its own).

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults minclustersize=40 nan [nan, nan] nan% [nan, nan] nan nan 0

Sensitivity (mean skill at each value, all other settings pooled):

27 · Find kinds of gene defined by two labels at once (UMAP + HDBSCAN) -- untestable

Pattern 5: conjunctions found on a random half; each is checked on the other half as the second label's enrichment in the cluster among genes carrying the first label (hypergeometric p < 0.05). Metric: share replicating. Null: 20 runs with the second label shuffled within the second half. Pass: above the null's 95th percentile by 0.2 with at least three findings.

Metric: share of findings that replicate. 15 runs, 3 settings, targets: (the strategy chooses its own).

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults minclustersize=15 nan [nan, nan] nan% [nan, nan] nan nan 0

Sensitivity (mean skill at each value, all other settings pooled):

28 · Find paralogs that changed jobs (profile correlation) -- weak

Paralog pairs with both genes labelled; the labels are withheld from the profiles. Metric: AUROC of divergence for pairs whose labels differ against pairs whose labels match. Null: 20 random reassignments of the divergence values to pairs. Pass: above the null's 95th percentile by 0.05.

Metric: AUROC of profile divergence for paralogs with different lopit_pf_location. 60 runs, 3 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults min_shared=10 0.112 [-0.137, 0.372] 60% [36, 80] 0.556 0.499 15
tuned min_shared=10 0.111 [-0.123, 0.362] 50% [19, 81] 0.556 0.501 6

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported 0.104 [0.099, 0.111] 80% 5
lopitpflocation 0.373 [0.364, 0.381] 100% 5
pbtransferredphenotype -0.142 [-0.157, -0.126] 0% 5
stageenrichedderived nan [nan, nan] nan% 0

Sensitivity (mean skill at each value, all other settings pooled):

29 · Carry what one parasite shows to the other (orthogroup mapping) -- reliable

Numeric target: 25% of the genes measured in both species hidden; the relation is learned on the rest. Metric: rank correlation of transferred and hidden values. Null: 20 runs with the ortholog values permuted among genes. Categorical target: Pattern 1 on genes with an ortholog value. Pass: above the null's 95th percentile by 0.1 (numeric) or 0.05 (categorical).

Metric: rank correlation of transferred and hidden values. 5 runs, 1 settings, targets: (the strategy chooses its own).

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults -- 0.313 [0.298, 0.328] 100% [57, 100] 0.316 0.004 5
tuned -- 0.322 [0.312, 0.333] 100% [34, 100] 0.329 0.01 2

30 · Test inference on the genes orthology cannot reach (kNN) -- weak

Pattern 1 restricted to the stratum: 25% of the label hidden by whole orthogroups; only hidden genes inside the stratum are scored. Null: 10 runs on shuffled labels. Pass: above the null's 95th percentile by 0.05.

Metric: correct calls per hidden gene. 80 runs, 4 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults stratum=lineage-specific -0.007 [-0.007, -0.007] 0% [0, 66] 0.659 0.661 2
tuned stratum=conserved 0.228 [0.037, 0.424] 50% [22, 78] 0.701 0.537 8

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported -0.525 [-1.084, 0.026] 0% 5
lopitpflocation 0.454 [0.440, 0.473] 100% 5
pbtransferredphenotype 0.374 [0.358, 0.390] 100% 5
stageenrichedderived 0.078 [-0.055, 0.246] 20% 5

Sensitivity (mean skill at each value, all other settings pooled):

31 · Call a gene only when independent strategies agree (kNN + logistic + network vote) -- works when tuned

Pattern 1 scored by precision: 25% of the label hidden; each method is trained on the visible genes and the agreed calls on hidden genes are scored. Metric: share of agreed calls that are correct. Null: 5 runs with visible labels shuffled. Pass: above the null's 95th percentile by 0.1. Single-method precisions are reported.

Metric: precision of calls on hidden genes. 60 runs, 3 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults min_agree=2 0.245 [-0.256, 0.547] 70% [48, 85] 0.768 0.554 20
tuned min_agree=3 0.663 [0.505, 0.810] 57% [25, 84] 0.866 0.546 7

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported 0.756 [0.632, 0.858] 0% 5
lopitpflocation 0.827 [0.787, 0.865] 100% 5
pbtransferredphenotype 0.568 [0.542, 0.592] 100% 5
stageenrichedderived 0.468 [0.468, 0.468] 0% 1

Sensitivity (mean skill at each value, all other settings pooled):

32 · Put the understudied genes first (kNN + logistic + network vote) -- works when tuned

Pattern 1 scored by precision and restricted to understudied genes: 25% of the label hidden; agreed calls are scored only on hidden understudied genes. Null: 5 runs with visible labels shuffled. Pass: above the null's 95th percentile by 0.1.

Metric: precision of calls on hidden genes. 60 runs, 3 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults min_agree=2 0.187 [-0.389, 0.545] 65% [43, 82] 0.768 0.567 20
tuned min_agree=3 0.708 [0.532, 0.883] 67% [30, 90] 0.877 0.524 6

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported 0.751 [0.610, 0.866] 0% 5
lopitpflocation 0.854 [0.827, 0.882] 100% 5
pbtransferredphenotype 0.567 [0.537, 0.601] 100% 5
stageenrichedderived nan [nan, nan] nan% 0

Sensitivity (mean skill at each value, all other settings pooled):

33 · Put every layer into one space and read a gene's neighbourhood (logistic edge model) -- reliable

A layer's edges are hidden by orthogroup -- whole groups at a time, so no hidden edge survives through a visible paralog -- and the layer is removed from its own features. Metric: AUROC of the hidden edges against DEGREE-MATCHED non-pairs, one per hidden edge, matched on connectivity at both ends. Null: 10 configuration-model rewirings of the hidden edges, which keep their degree sequence and destroy their topology, so a model reading fame cannot beat it. Pass: above the null's 95th percentile by 0.05. The AUROC against random non-pairs and the gap between the two are reported as numbers, not as the verdict.

Metric: AUROC of hidden coexpression edges against degree-matched non-pairs. 90 runs, 18 settings, targets: (the strategy chooses its own).

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults k=10; knn=15; layer=coexpression 0.392 [0.368, 0.414] 100% [57, 100] 0.707 0.52 5
tuned k=5; knn=15; layer=struct 0.744 [0.671, 0.817] 100% [34, 100] 0.889 0.559 2

Sensitivity (mean skill at each value, all other settings pooled):

34 · Train on the networks and rank the edges they are missing (logistic / spectral embedding) -- reliable

The chosen layer's edges are hidden by orthogroup and the layer is removed from its own features. Metric: AUROC of the hidden edges against DEGREE-MATCHED non-pairs. Null: 10 configuration-model rewirings of the hidden edges, which preserve their degree sequence, so fame alone cannot clear the bar. Pass: above the null's 95th percentile by 0.05. The random-null AUROC, the fame gap, precision@k, the Brier score and the reliability gap are all reported as numbers beside the verdict.

Metric: AUROC of hidden coexpression edges against degree-matched non-pairs. 60 runs, 12 settings, targets: (the strategy chooses its own).

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults fraction=0.25; layer=coexpression; model=logistic 0.392 [0.368, 0.416] 100% [57, 100] 0.707 0.52 5
tuned fraction=0.4; layer=struct; model=logistic 0.729 [0.708, 0.751] 100% [34, 100] 0.891 0.595 2

Sensitivity (mean skill at each value, all other settings pooled):

35 · Call genes with a stated error rate (split conformal prediction) -- reliable

25% of the label hidden by whole orthogroups; the model is trained and calibrated on the rest (itself split by orthogroup) and builds a set for every hidden gene. Metric: set efficiency, 1 - (mean set size - 1) / (classes - 1) -- how far the sets narrow the possibilities. Null: 10 runs on shuffled labels, where the model learns nothing and the sets must grow to keep the promise. Pass: above the null's 95th percentile by 0.05. Also reported: set coverage on the hidden genes beside the 1 - alpha promised, and the scorecard of the single-label calls.

Metric: set efficiency: 1 - (mean set size - 1) / (classes - 1). 120 runs, 6 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults alpha=0.1; model=logistic 0.633 [0.444, 0.775] 100% [84, 100] 0.7 0.199 20
tuned alpha=0.2; model=logistic 0.776 [0.555, 0.941] 100% [68, 100] 0.836 0.299 8

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported 0.993 [0.980, 1.000] 100% 5
lopitpflocation 0.846 [0.809, 0.883] 100% 5
pbtransferredphenotype 0.459 [0.435, 0.488] 100% 5
stageenrichedderived 0.849 [0.803, 0.885] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

36 · Smooth the measurements along the networks, then classify (graph convolution + logistic regression) -- reliable

Pattern 1: 25% of the label hidden by whole orthogroups; the smoothed features are computed from measurements only, and the model is trained on the visible genes. Metric: hidden genes called correctly. Null: the analytic chance level for the same predicted and true class mixes. Pass: above that level's 95% bound by 0.05.

Metric: correct calls per hidden gene. 120 runs, 6 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults C=0.5; hops=2 0.458 [0.367, 0.544] 100% [84, 100] 0.709 0.448 20
tuned C=0.1; hops=2 0.460 [0.375, 0.539] 100% [68, 100] 0.701 0.456 8

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported 0.443 [0.363, 0.522] 100% 5
lopitpflocation 0.519 [0.499, 0.537] 100% 5
pbtransferredphenotype 0.347 [0.329, 0.365] 100% 5
stageenrichedderived 0.461 [0.403, 0.518] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

37 · Let a random forest find what defines a label (random forest + permutation importance) -- reliable

Pattern 1: 25% of the label hidden by whole orthogroups; the forest is trained on the visible genes and calls the hidden ones. Metric: hidden genes called correctly. Null: the analytic chance level for the same predicted and true class mixes. Pass: above that level's 95% bound by 0.05.

Metric: correct calls per hidden gene. 120 runs, 6 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults min_leaf=2; trees=300 0.361 [0.186, 0.532] 75% [53, 89] 0.741 0.513 20
tuned min_leaf=5; trees=300 0.410 [0.265, 0.538] 75% [41, 93] 0.749 0.509 8

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported 0.269 [0.210, 0.338] 0% 5
lopitpflocation 0.575 [0.560, 0.590] 100% 5
pbtransferredphenotype 0.426 [0.410, 0.442] 100% 5
stageenrichedderived 0.411 [0.323, 0.502] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

38 · Learn how much to trust each kind of evidence (stacked logistic regression) -- reliable

Pattern 1: 25% of the label hidden by whole orthogroups; base predictions for the meta-model are out of fold within the visible genes only, so no hidden label reaches either level. Metric: hidden genes called correctly. Null: the analytic chance level for the same predicted and true class mixes. Pass: above that level's 95% bound by 0.05.

Metric: correct calls per hidden gene. 60 runs, 3 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults k=15 0.451 [0.384, 0.522] 95% [76, 99] 0.7 0.431 20
tuned k=15 0.460 [0.394, 0.522] 100% [68, 100] 0.704 0.439 8

Tuned setting, per held-out target:

target skill [95% CI] pass runs
is_exported 0.394 [0.350, 0.438] 80% 5
lopitpflocation 0.551 [0.533, 0.566] 100% 5
pbtransferredphenotype 0.381 [0.364, 0.406] 100% 5
stageenrichedderived 0.479 [0.435, 0.524] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):

39 · Predict a value with an interval that holds (gradient boosting / ridge + split conformal) -- reliable

Pattern 4: 20% of the measured values hidden; the model is trained and calibrated on the rest (split by orthogroup) and predicts the hidden ones. Metric: rank correlation of predicted and hidden values. Null: the chance distribution of a rank correlation, with refits on shuffled values reported. Pass: above its 95th percentile by 0.1. Also reported: the share of hidden values inside their intervals, beside the coverage promised.

Metric: rank correlation of predicted and hidden values. 60 runs, 4 settings, targets: exprschizont, meanplddt, piggybac_mis.

setting skill [95% CI] pass [95% CI] observed chance runs
at defaults alpha=0.1; model=boosted 0.500 [0.218, 0.837] 100% [80, 100] 0.5 0.0 15
tuned alpha=0.1; model=boosted 0.488 [0.185, 0.826] 100% [61, 100] 0.488 0.0 6

Tuned setting, per held-out target:

target skill [95% CI] pass runs
expr_schizont 0.206 [0.155, 0.268] 100% 5
mean_plddt 0.842 [0.826, 0.854] 100% 5
piggybac_mis 0.451 [0.440, 0.462] 100% 5

Sensitivity (mean skill at each value, all other settings pooled):