Strategy calibration
Generated by scripts/calibrate_strategies.py --publish on 2026-09-26 from 7,640 self-test runs (results/calibration_2026-09-26b).
Every strategy carries a self-test that hides information already known -- labels, set members, edges or values -- asks the strategy for it back, and compares the answer with the same procedure on shuffled data. Calibration runs that test over a grid of the strategy's settings, over several held-out labels and five seeds, so each number below is a mean with an interval rather than one draw.
- Skill = (observed - chance) / (1 - chance): 0 is the shuffled-data null, 1 is perfect. Negative means worse than shuffled.
- At defaults: the strategy as it opens in the application.
- Tuned: the setting with the best lower confidence bound on seeds 1-3, reported on seeds 4-5 only.
- 95% CI: two-stage bootstrap (held-out targets, then runs within a target).
- Pass: share of conclusive runs beating the null's 95th percentile by the strategy's minimum effect.
Grades: reliable (defaults beat chance with the interval above 0.05 and pass at least 60%), works when tuned (only the tuned setting does), weak (above chance on average but not reliably), no skill, untestable (fewer than five conclusive runs).
What a high skill here does and does not mean
Each strategy is scored against ITS OWN null, and a null can be too easy. The clearest case is edge prediction: strategy 18 scores its held-out pairs against non-pairs drawn at random, and a random pair of genes is usually a pair of obscure genes, so much of what such a test measures is that well-connected genes are well connected. Its skill here is therefore an upper bound (strategies 16, 33 and 34 use degree-matched non-pairs instead). docs/graphspace.md re-measures the same question against degree-matched and configuration-model nulls, where the honest figure is an AUROC near 0.67 rather than 0.97 -- and it reports the gap between the two nulls as its own quantity. Read a high number here as a reason to look at how the null was built, not as a result.
Toxoplasma gondii
01 · Hold out a category and search for a map that finds it (UMAP + HDBSCAN) -- weak
25% of the label is hidden (whole orthogroups together, so no gene is recovered through a visible paralog). The walk is built as set -- its genes per map (up to 4,000), feature sets and grids (up to three values each); the configuration, and the one cluster that best isolates each label, are both chosen using visible labels only. Metric: for each label, the F1 of its hidden genes against its chosen cluster, weighted by size -- does the structure found on known genes hold the unknown ones? Null: the same chosen clusters scored after permuting the hidden genes' labels, 100 times. (A second search on shuffled labels was the null once; it picks the largest cluster for every label, which scores F1 near 2p by size alone and made the null beat real labels.) Pass: above the null's 95th percentile by at least 0.05.
Metric: F1 of hidden genes in the cluster chosen for their label on known genes. 150 runs, 6 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | minclustersize=[20, 50]; n_neighbors=[15, 50]; selection=[eom, leaf] | 0.071 [0.034, 0.105] | 44% [27, 63] | 0.376 | 0.322 | 25 |
| tuned | minclustersize=[10, 25]; n_neighbors=[25, 100]; selection=[eom, leaf] | 0.070 [0.026, 0.115] | 40% [17, 69] | 0.394 | 0.342 | 10 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.074 [0.050, 0.103] | 60% | 5 |
| dtm_class | 0.078 [0.061, 0.096] | 0% | 5 |
| lopit_unified | 0.108 [0.095, 0.120] | 100% | 5 |
| screenanyphenotype | -0.005 [-0.023, 0.006] | 0% | 5 |
| stageenrichedderived | 0.159 [0.106, 0.208] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
min_cluster_size-- 10, 25: 0.06; 20, 50: 0.07n_neighbors-- 10, 30: 0.06; 15, 50: 0.06; 25, 100: 0.08selection-- eom, leaf: 0.07
02 · Find the map where your gene list is one cluster (UMAP + HDBSCAN) -- weak
30% of the set is hidden; the walk as set (genes per map up to 4,000, feature sets, up to three values of each grid) picks the cluster with the best F1 for the other 70%. Metric: F1 of the hidden members against that cluster's other genes -- precision is the share of the cluster's candidates that are hidden members, recall the share of hidden members among them. Null: 20 random sets of the same size through the same walk. Pass: above the null's 95th percentile by at least 0.05 (an F1 margin; random sets score about 0.04).
Metric: F1 of the hidden members against the best cluster's other genes. 100 runs, 4 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | minclustersize=[20, 50]; n_neighbors=[15, 50] | 0.025 [0.004, 0.042] | 16% [6, 35] | 0.058 | 0.033 | 25 |
| tuned | minclustersize=[20, 50]; n_neighbors=[10, 30] | 0.028 [0.002, 0.047] | 10% [2, 40] | 0.06 | 0.032 | 10 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.037 [0.029, 0.047] | 20% | 5 |
| dtm_class | 0.024 [0.008, 0.037] | 0% | 5 |
| lopit_unified | 0.040 [0.028, 0.054] | 20% | 5 |
| screenanyphenotype | 0.047 [0.032, 0.062] | 40% | 5 |
| stageenrichedderived | -0.009 [-0.019, 0.000] | 0% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
min_cluster_size-- 10, 25: 0.03; 20, 50: 0.03n_neighbors-- 10, 30: 0.03; 15, 50: 0.02
03 · Ask which categories the data can rediscover (UMAP + neighbour AUROC) -- reliable
30% of the label is hidden. The atlas is built from visible genes only (leave-one-out neighbour AUROC per category); hidden genes are then scored by their visible neighbours. Metric: mean hidden-gene AUROC over the categories the atlas ranks in its top half. Null: the same with the visible labels shuffled, 20 times. Pass: above the null's 95th percentile by 0.05. The rank agreement between atlas and hidden recovery is reported.
Metric: hidden-gene AUROC of the categories the atlas ranks in its top half. 75 runs, 3 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | k=15 | 0.499 [0.217, 0.679] | 80% [61, 91] | 0.75 | 0.501 | 25 |
| tuned | k=15 | 0.502 [0.222, 0.680] | 80% [49, 94] | 0.75 | 0.501 | 10 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.617 [0.544, 0.686] | 100% | 5 |
| dtm_class | 0.563 [0.539, 0.587] | 100% | 5 |
| lopit_unified | 0.699 [0.662, 0.734] | 100% | 5 |
| screenanyphenotype | -0.057 [-0.143, 0.024] | 0% | 5 |
| stageenrichedderived | 0.676 [0.632, 0.712] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
k-- 15: 0.50; 5: 0.45; 50: 0.52
04 · Keep only the modules that survive the whole walk (UMAP + HDBSCAN co-clustering) -- weak
Modules are built from a small walk on 1,500 genes without the label. 25% of the label is hidden; the module that best isolates each label is chosen on the visible genes. Metric: the size-weighted F1 of each label's hidden genes against its chosen module. Null: the same modules scored after permuting the hidden genes' labels, 100 times -- a large module scores the same either way and earns nothing. Pass: above the null's 95th percentile by 0.05.
Metric: F1 of hidden genes in the module chosen for their label on known genes. 75 runs, 3 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | threshold=0.5 | 0.073 [0.014, 0.127] | 56% [37, 73] | 0.386 | 0.336 | 25 |
| tuned | threshold=0.3 | 0.045 [-0.041, 0.104] | 50% [24, 76] | 0.37 | 0.336 | 10 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.063 [0.024, 0.096] | 60% | 5 |
| dtm_class | 0.087 [0.065, 0.109] | 40% | 5 |
| lopit_unified | 0.084 [0.039, 0.123] | 80% | 5 |
| screenanyphenotype | -0.033 [-0.119, 0.024] | 0% | 5 |
| stageenrichedderived | 0.149 [0.116, 0.186] | 80% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
threshold-- 0.3: 0.07; 0.5: 0.07; 0.8: 0.05
05 · Tune a map without labels, then read what it encodes (UMAP + HDBSCAN, chi-square / Kruskal-Wallis) -- reliable
Pattern 5. One label-free map on up to 2,000 genes. Held-out features significant at q < 0.05 on a random half of the genes are the findings; each is re-tested (p < 0.05) on the other half. Metric: share that replicate. Null: the same with the second half's cluster labels permuted, 10 times. Pass: above the null's 95th percentile by 0.2, with at least three findings.
Metric: share of first-half findings that replicate on the second half. 70 runs, 14 settings, targets: (the strategy chooses its own).
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | map_from=transcription | 0.949 [0.884, 0.997] | 100% [57, 100] | 0.952 | 0.059 | 5 |
| tuned | map_from=chemistry | 0.964 [0.947, 0.982] | 100% [34, 100] | 0.966 | 0.046 | 2 |
Sensitivity (mean skill at each value, all other settings pooled):
map_from-- PTM: 0.95; chemistry: 0.97; fitness: 0.94; host effect: 0.86; immunity: 0.64; localization: 0.96; metabolism: 0.84; phenotype: 0.80; protein abundance: 0.94; regulation: 0.91; relation: 0.92; sequence: 0.96; transcription: 0.95; translation: 0.97
06 · Find which kind of evidence carries a label (kNN ablation) -- reliable
25% of the label is hidden. Each kind of evidence is ranked by cross-validated accuracy on the visible labels; the top-ranked one then predicts the hidden labels. Metric: its hidden accuracy. Null: the hidden accuracy of every kind of evidence, i.e. choosing at random. Pass: above the null's 80th percentile by 0.02 (with ~15 kinds of evidence the 95th would demand the single best, which asks more than a ranking must deliver).
Metric: hidden accuracy of the evidence ranked first (transcription). 75 runs, 3 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | k=15 | 0.140 [0.060, 0.251] | 64% [45, 80] | 0.604 | 0.543 | 25 |
| tuned | k=5 | 0.159 [0.081, 0.247] | 70% [40, 89] | 0.59 | 0.525 | 10 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.094 [0.074, 0.114] | 100% | 5 |
| dtm_class | 0.078 [0.068, 0.087] | 80% | 5 |
| lopit_unified | 0.139 [0.127, 0.156] | 100% | 5 |
| screenanyphenotype | 0.165 [0.123, 0.227] | 0% | 5 |
| stageenrichedderived | 0.349 [0.316, 0.374] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
k-- 15: 0.14; 5: 0.16; 50: 0.10
07 · Call a gene by the genes that behave like it (kNN) -- reliable
Pattern 1: 25% of the label hidden by whole orthogroups; each hidden gene called by its k nearest visible genes with the same vote threshold. Metric: hidden genes called correctly. Null: 10 runs with visible labels shuffled. Pass: above the null's 95th percentile by 0.05.
Metric: correct calls per hidden gene. 225 runs, 9 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | k=15; min_share=0.3 | 0.225 [0.079, 0.358] | 60% [41, 77] | 0.619 | 0.476 | 25 |
| tuned | k=5; min_share=0.0 | 0.292 [0.163, 0.428] | 80% [49, 94] | 0.629 | 0.468 | 10 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.297 [0.278, 0.316] | 100% | 5 |
| dtm_class | 0.181 [0.173, 0.190] | 80% | 5 |
| lopit_unified | 0.373 [0.364, 0.383] | 100% | 5 |
| screenanyphenotype | 0.050 [-0.029, 0.111] | 0% | 5 |
| stageenrichedderived | 0.496 [0.454, 0.537] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
k-- 15: 0.20; 5: 0.27; 50: 0.12min_share-- 0.0: 0.22; 0.3: 0.23; 0.6: 0.15
08 · Call a gene by its neighbours on the map (UMAP + kNN) -- weak
Pattern 1 on the genes the map places (up to 2,500): 25% of the label hidden by whole orthogroups; hidden genes called by their k nearest visible genes in the map. Null: 20 runs on shuffled visible labels. Pass: above the null's 95th percentile by 0.05.
Metric: correct calls per hidden gene. 225 runs, 9 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | k=15; n_neighbors=25 | 0.125 [0.008, 0.236] | 60% [41, 77] | 0.565 | 0.474 | 25 |
| tuned | k=5; n_neighbors=60 | 0.114 [-0.037, 0.229] | 60% [31, 83] | 0.538 | 0.46 | 10 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.169 [0.148, 0.188] | 100% | 5 |
| dtm_class | 0.034 [0.004, 0.057] | 0% | 5 |
| lopit_unified | 0.231 [0.217, 0.243] | 100% | 5 |
| screenanyphenotype | -0.056 [-0.180, 0.028] | 0% | 5 |
| stageenrichedderived | 0.296 [0.260, 0.328] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
k-- 15: 0.13; 5: 0.15; 50: 0.13n_neighbors-- 10: 0.15; 25: 0.14; 60: 0.13
09 · Name a cluster by the label it is enriched for (UMAP + HDBSCAN, hypergeometric) -- reliable
Pattern 1 scored by precision, on the genes a blind map places (up to 2,500): 25% of the label hidden; the enrichment is recomputed from visible labels, and the calls it makes on hidden genes are scored -- the strategy abstains on noise and unenriched clusters by design, so what matters is how often a call is right. Null: 20 runs on shuffled visible labels. Pass: above the null's 95th percentile by 0.1.
Metric: precision of calls on hidden genes. 300 runs, 12 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | minclustersize=25; min_lift=1.5; selection=leaf | 0.332 [0.249, 0.423] | 100% [84, 100] | 0.333 | 0.002 | 20 |
| tuned | minclustersize=10; min_lift=3.0; selection=leaf | 0.409 [0.303, 0.501] | 100% [68, 100] | 0.411 | 0.004 | 8 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.367 [0.271, 0.460] | 100% | 5 |
| dtm_class | 0.276 [0.207, 0.356] | 100% | 5 |
| lopit_unified | 0.548 [0.493, 0.603] | 100% | 5 |
| screenanyphenotype | nan [nan, nan] | nan% | 0 |
| stageenrichedderived | 0.480 [0.402, 0.557] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
min_cluster_size-- 10: 0.35; 25: 0.31; 60: 0.22min_lift-- 1.5: 0.27; 3.0: 0.35selection-- eom: 0.21; leaf: 0.34
10 · Find genes whose label their neighbours contradict (kNN + network neighbours) -- reliable
5% of the labels (at least ten) are swapped to a wrong class, drawn in proportion to class size. Surprise is computed with the corrupted labels. Metric: AUROC of surprise for the swapped genes among all labelled genes. Null: 20 random sets of the same size. Pass: above the null's 95th percentile by 0.1.
Metric: AUROC of surprise for the swapped labels. 75 runs, 3 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | k=15 | 0.638 [0.490, 0.769] | 100% [87, 100] | 0.818 | 0.498 | 25 |
| tuned | k=15 | 0.650 [0.505, 0.776] | 100% [72, 100] | 0.823 | 0.496 | 10 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.608 [0.583, 0.631] | 100% | 5 |
| dtm_class | 0.796 [0.789, 0.804] | 100% | 5 |
| lopit_unified | 0.633 [0.610, 0.655] | 100% | 5 |
| screenanyphenotype | 0.342 [0.281, 0.396] | 100% | 5 |
| stageenrichedderived | 0.810 [0.797, 0.824] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
k-- 15: 0.64; 5: 0.63; 50: 0.63
11 · Diffuse a label across one measured network (random walk with restart) -- reliable
Pattern 1: 25% of the label hidden by whole orthogroups; the fields are seeded from visible genes only, so a hidden gene never seeds its own call. Metric: hidden genes called correctly (unreached genes count as misses). Null: 10 runs on shuffled labels. Pass: above the null's 95th percentile by 0.05.
Metric: correct calls per hidden gene. 900 runs, 36 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | layer=coexpression; restart=0.5 | 0.125 [0.082, 0.169] | 75% [53, 89] | 0.291 | 0.186 | 20 |
| tuned | layer=coexpression; restart=0.2 | 0.136 [0.095, 0.176] | 88% [53, 98] | 0.288 | 0.175 | 8 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.089 [0.082, 0.095] | 100% | 5 |
| dtm_class | 0.159 [0.150, 0.165] | 100% | 5 |
| lopit_unified | 0.188 [0.176, 0.199] | 100% | 5 |
| screenanyphenotype | 0.099 [0.088, 0.114] | 40% | 5 |
| stageenrichedderived | nan [nan, nan] | nan% | 0 |
Sensitivity (mean skill at each value, all other settings pooled):
layer-- coexpression: 0.13; cofitness: 0.07; comention: 0.01; comentionft: 0.06; cotranslation: 0.02; domain: 0.04; ipms: 0.00; orthogroup: 0.00; struct: 0.08; structuralhole: 0.01; unwritteninteraction: 0.07; xlms: 0.07restart-- 0.2: 0.05; 0.5: 0.05; 0.8: 0.04
12 · Let every network vote, weighted by what it has earned (chance-weighted ensemble vote) -- weak
Pattern 1: 25% of the label hidden; the weights are learned on an inner holdout of the visible labels only, then hidden genes are called. Null: 5 runs with visible labels shuffled (weights relearned each time). Pass: above the null's 95th percentile by 0.05.
Metric: correct calls per hidden gene. 75 runs, 3 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | k=15 | 0.300 [-0.034, 0.536] | 80% [61, 91] | 0.577 | 0.383 | 25 |
| tuned | k=15 | 0.246 [-0.208, 0.542] | 80% [49, 94] | 0.559 | 0.392 | 10 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.321 [0.313, 0.330] | 100% | 5 |
| dtm_class | 0.495 [0.406, 0.570] | 100% | 5 |
| lopit_unified | 0.394 [0.362, 0.432] | 100% | 5 |
| screenanyphenotype | -0.334 [-0.780, 0.088] | 0% | 5 |
| stageenrichedderived | 0.625 [0.529, 0.723] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
k-- 15: 0.30; 5: 0.30; 50: 0.33
13 · Place a protein by the proteins it physically touches (weighted partner vote) -- reliable
Pattern 1 restricted to genes with at least one physical partner: 25% of the label hidden by whole orthogroups; hidden genes called by their visible partners' weighted vote. Null: 20 runs on shuffled labels. Pass: above the null's 95th percentile by 0.05.
Metric: correct calls per hidden gene. 25 runs, 1 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | -- | 0.249 [0.076, 0.428] | 60% [41, 77] | 0.385 | 0.194 | 25 |
| tuned | -- | 0.257 [0.075, 0.426] | 60% [31, 83] | 0.39 | 0.189 | 10 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.428 [0.414, 0.443] | 100% | 5 |
| dtm_class | 0.302 [0.280, 0.335] | 100% | 5 |
| lopit_unified | 0.488 [0.462, 0.510] | 100% | 5 |
| screenanyphenotype | 0.008 [-0.013, 0.021] | 0% | 5 |
| stageenrichedderived | 0.021 [0.008, 0.035] | 0% | 5 |
14 · Annotate function through shared fold (TM-score-weighted vote) -- reliable
Pattern 1 restricted to proteins with a structural neighbour: 25% of the annotation (at the chosen level) hidden by whole orthogroups; hidden proteins called by their visible structural neighbours. Null: 20 runs on shuffled annotations. Pass: above the null's 95th percentile by 0.05.
Metric: correct calls per hidden gene. 15 runs, 3 settings, targets: (the strategy chooses its own).
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | level=1 | 0.637 [0.602, 0.664] | 100% [57, 100] | 0.733 | 0.266 | 5 |
| tuned | level=3 | 0.713 [0.706, 0.719] | 100% [34, 100] | 0.736 | 0.083 | 2 |
Sensitivity (mean skill at each value, all other settings pooled):
level-- 1: 0.64; 2: 0.70; 3: 0.71
15 · Find the communities several networks agree on (modularity + Louvain consensus) -- weak
On genes placed in a community: 25% of the label hidden; the community that best isolates each label is chosen on the visible genes (the communities themselves never see labels). Metric: the size-weighted F1 of each label's hidden genes against its chosen community. Null: the same communities scored after permuting the hidden genes' labels, 100 times. Pass: above the null's 95th percentile by 0.05.
Metric: F1 of hidden genes in the community chosen for their label on known genes. 150 runs, 6 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | agreement=0.5; resolution=1.0 | 0.054 [0.017, 0.085] | 40% [23, 59] | 0.345 | 0.304 | 25 |
| tuned | agreement=0.5; resolution=1.0 | 0.048 [-0.001, 0.087] | 40% [17, 69] | 0.338 | 0.299 | 10 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.041 [0.037, 0.045] | 0% | 5 |
| dtm_class | 0.071 [0.062, 0.080] | 60% | 5 |
| lopit_unified | 0.098 [0.091, 0.104] | 100% | 5 |
| screenanyphenotype | -0.015 [-0.038, 0.007] | 0% | 5 |
| stageenrichedderived | 0.076 [0.053, 0.101] | 40% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
agreement-- 0.3: 0.04; 0.5: 0.04resolution-- 0.5: 0.01; 1.0: 0.05; 2.0: 0.05
16 · Predict the contacts an interactome missed (logistic regression) -- reliable
Pattern 3: 20% of the layer's edges hidden; the model is trained on the rest against DEGREE-MATCHED non-edges -- each with a gene of similar degree at both ends -- and scores the hidden edges against fresh degree-matched non-edges. Against random non-edges this test read AUROC 0.99 on every correlation layer, because a random pair is usually two obscure genes and degree separates them. The layer's source measurements leave the similarity feature with it, and derived, annotation and literature layers are refused as targets. Metric: AUROC. Null: 5 models trained with each gene's evidence read from a random other gene (identities permuted). Pass: above the null's 95th percentile by 0.05.
Metric: AUROC of hidden pairs against degree-matched non-pairs. 10 runs, 2 settings, targets: (the strategy chooses its own).
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | layer=xlms | 0.606 [0.590, 0.623] | 100% [57, 100] | 0.805 | 0.504 | 5 |
| tuned | layer=struct | 0.926 [0.925, 0.927] | 100% [34, 100] | 0.963 | 0.506 | 2 |
Sensitivity (mean skill at each value, all other settings pooled):
layer-- struct: 0.93; xlms: 0.61
17 · Read the literature for biology, not fame (publication-count residual) -- reliable
The label is never used to build the literature layer. Among co-mentioned pairs with both genes labelled, the top k by corrected residual are taken (k = the chosen number, capped at a fifth of the pool). Metric: the share of those pairs sharing a label. Null: 20 random sets of k co-mentioned pairs. Pass: above the null's 95th percentile by 0.05.
Metric: share of the top 50 corrected pairs sharing a compartment label. 75 runs, 3 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | top=200 | 0.376 [0.209, 0.545] | 80% [61, 91] | 0.707 | 0.523 | 25 |
| tuned | top=200 | 0.379 [0.210, 0.547] | 80% [49, 94] | 0.707 | 0.521 | 10 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.587 [0.585, 0.589] | 100% | 5 |
| dtm_class | 0.155 [0.141, 0.169] | 100% | 5 |
| lopit_unified | 0.609 [0.602, 0.615] | 100% | 5 |
| screenanyphenotype | 0.183 [0.169, 0.197] | 0% | 5 |
| stageenrichedderived | 0.344 [0.322, 0.358] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
top-- 1000: 0.26; 200: 0.38; 50: 0.41
18 · List what the data says and the literature has not written (multi-layer support count) -- reliable
Pattern 3 with the literature as truth: co-mentioned pairs among genes in the measurement layers against 5 times as many random pairs. Score: number of measurement layers linking the pair. Metric: AUROC. Null: 10 runs with the measurement layers' gene identities permuted. Pass: above the null's 95th percentile by 0.02.
Metric: AUROC of hidden pairs against random non-pairs. 15 runs, 3 settings, targets: (the strategy chooses its own).
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | min_layers=2 | 0.112 [0.111, 0.112] | 100% [57, 100] | 0.556 | 0.5 | 5 |
| tuned | min_layers=1 | 0.111 [0.111, 0.111] | 100% [34, 100] | 0.556 | 0.5 | 2 |
Sensitivity (mean skill at each value, all other settings pooled):
min_layers-- 1: 0.11; 2: 0.11; 3: 0.11
19 · Train a classifier on the known genes and call the rest (logistic regression) -- reliable
Pattern 1: 25% of the label hidden by whole orthogroups; the model is trained on the visible genes and calls the hidden ones. Metric: hidden genes called correctly. Null: the analytic chance level for the same predicted and true class mixes (re-running a multinomial fit on shuffled labels ten times would take longer than it tells). Pass: above that level's 95% bound by 0.05.
Metric: correct calls per hidden gene. 100 runs, 4 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | C=0.01 | 0.318 [0.180, 0.465] | 80% [61, 91] | 0.542 | 0.326 | 25 |
| tuned | C=10.0 | 0.369 [0.203, 0.498] | 80% [49, 94] | 0.612 | 0.362 | 10 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.407 [0.394, 0.417] | 100% | 5 |
| dtm_class | 0.339 [0.322, 0.360] | 100% | 5 |
| lopit_unified | 0.496 [0.488, 0.505] | 100% | 5 |
| screenanyphenotype | 0.043 [-0.018, 0.111] | 0% | 5 |
| stageenrichedderived | 0.573 [0.542, 0.603] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
C-- 0.01: 0.32; 0.1: 0.37; 1.0: 0.38; 10.0: 0.37
20 · Learn what makes your list special, from positives alone (PU bagging, logistic regression) -- reliable
Pattern 2: 30% of the set hidden; the other 70% are the positives. Metric: AUROC of the hidden members against every other non-query gene. Null: 10 random sets of the same size. Pass: above the null's 95th percentile by 0.1.
Metric: AUROC of hidden members against every other gene. 75 runs, 3 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | bags=15 | 0.787 [0.722, 0.849] | 100% [87, 100] | 0.895 | 0.504 | 25 |
| tuned | bags=5 | 0.800 [0.741, 0.853] | 100% [72, 100] | 0.9 | 0.5 | 10 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.880 [0.869, 0.892] | 100% | 5 |
| dtm_class | 0.663 [0.635, 0.691] | 100% | 5 |
| lopit_unified | 0.840 [0.827, 0.853] | 100% | 5 |
| screenanyphenotype | 0.787 [0.768, 0.805] | 100% | 5 |
| stageenrichedderived | 0.757 [0.745, 0.777] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
bags-- 15: 0.79; 40: 0.79; 5: 0.79
21 · Predict a measurement, and find the genes that defy the prediction (gradient boosting / ridge) -- reliable
Pattern 4: 20% of the measured values hidden; the model is trained on the rest. Metric: rank correlation between predicted and hidden values. Null: 3 models trained on the visible values shuffled among the measured genes. Pass: above the null's 95th percentile by 0.1.
Metric: rank correlation of predicted and hidden values. 60 runs, 4 settings, targets: exprtachy, fitinvitrohff, fitinvivo_PE.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | model=boosted; own_kind=leave out | 0.560 [0.051, 0.909] | 67% [42, 85] | 0.56 | 0.0 | 15 |
| tuned | model=boosted; own_kind=include | 0.666 [0.148, 0.968] | 100% [61, 100] | 0.666 | 0.0 | 6 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| expr_tachy | 0.969 [0.965, 0.972] | 100% | 5 |
| fitinvitrohff | 0.892 [0.890, 0.893] | 100% | 5 |
| fitinvivoPE | 0.128 [0.114, 0.148] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
model-- boosted: 0.61; ridge: 0.57own_kind-- include: 0.64; leave out: 0.55
22 · Fill in what was never measured, and say where that is honest (soft-impute, low-rank SVD) -- reliable
Pattern 4 across the whole table: 10% of every column's measured entries hidden, the table completed at the chosen rank. Metric: median over columns of the rank correlation on hidden entries. Null: 3 completions of a table whose columns were each shuffled independently. Pass: above the null's 95th percentile by 0.1.
Metric: median per-column rank correlation on hidden entries. 15 runs, 3 settings, targets: (the strategy chooses its own).
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | rank=20 | 0.868 [0.864, 0.873] | 100% [57, 100] | 0.868 | 0.001 | 5 |
| tuned | rank=60 | 0.903 [0.902, 0.904] | 100% [34, 100] | 0.903 | -0.001 | 2 |
Sensitivity (mean skill at each value, all other settings pooled):
rank-- 20: 0.87; 5: 0.74; 60: 0.90
23 · Find what matters more in one condition, and why (residual + gradient boosting / ridge) -- weak
Pattern 4 on the shift: 20% of genes with both measurements hidden; a model trained on the rest predicts their shift. Metric: rank correlation on hidden genes. Null: 3 models trained on shuffled shifts. Pass: above the null's 95th percentile by 0.1.
Metric: rank correlation of predicted and hidden values. 20 runs, 4 settings, targets: (the strategy chooses its own).
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | model=boosted; own_kind=leave out | 0.073 [0.060, 0.085] | 0% [0, 43] | 0.073 | 0.0 | 5 |
| tuned | model=ridge; own_kind=include | 0.068 [0.055, 0.082] | 0% [0, 66] | 0.068 | 0.0 | 2 |
Sensitivity (mean skill at each value, all other settings pooled):
model-- boosted: 0.10; ridge: 0.08own_kind-- include: 0.11; leave out: 0.07
24 · Describe what your gene list has in common (hypergeometric + rank-sum) -- reliable
Pattern 2: 40% of the set hidden; the profile is built from the other 60% and scores every gene. Metric: AUROC of the hidden members against every other non-query gene. Null: 10 random sets of the same size. Pass: above the null's 95th percentile by 0.1.
Metric: AUROC of hidden members against every other gene. 25 runs, 1 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | -- | 0.615 [0.510, 0.719] | 100% [87, 100] | 0.807 | 0.499 | 25 |
| tuned | -- | 0.644 [0.518, 0.749] | 100% [72, 100] | 0.822 | 0.5 | 10 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.791 [0.766, 0.814] | 100% | 5 |
| dtm_class | 0.565 [0.476, 0.658] | 100% | 5 |
| lopit_unified | 0.678 [0.643, 0.706] | 100% | 5 |
| screenanyphenotype | 0.598 [0.528, 0.672] | 100% | 5 |
| stageenrichedderived | 0.442 [0.389, 0.498] | 100% | 5 |
25 · Grow your gene list along the networks (random walk with restart) -- reliable
Pattern 2: 30% of the set hidden; the walk is seeded from the other 70%. Metric: AUROC of the hidden members against every other non-seed gene. Null: 10 random seed sets of the same size. Pass: above the null's 95th percentile by 0.1.
Metric: AUROC of hidden members against every other gene. 150 runs, 6 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | mode=networks + measurements; restart=0.3 | 0.594 [0.488, 0.683] | 100% [87, 100] | 0.799 | 0.506 | 25 |
| tuned | mode=networks + measurements; restart=0.6 | 0.608 [0.533, 0.669] | 100% [72, 100] | 0.806 | 0.504 | 10 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.696 [0.655, 0.737] | 100% | 5 |
| dtm_class | 0.604 [0.601, 0.609] | 100% | 5 |
| lopit_unified | 0.698 [0.661, 0.729] | 100% | 5 |
| screenanyphenotype | 0.590 [0.565, 0.624] | 100% | 5 |
| stageenrichedderived | 0.391 [0.339, 0.445] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
mode-- networks: 0.44; networks + measurements: 0.59restart-- 0.1: 0.52; 0.3: 0.52; 0.6: 0.52
26 · Find categories that split in two on another measurement (UMAP + HDBSCAN) -- untestable
Pattern 5: significant findings (q < 0.05) on a random half of the mapped genes; each is checked on the other half -- the minority group enriched again among the cluster's A-matching genes (hypergeometric p < 0.05), or the measurement bimodal again. Metric: share replicating. Null: 20 runs with B shuffled within the second half. Pass: above the null's 95th percentile by 0.2 with at least three findings.
Metric: share of findings that replicate. 15 runs, 3 settings, targets: (the strategy chooses its own).
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | minclustersize=40 | nan [nan, nan] | nan% [nan, nan] | nan | nan | 0 |
Sensitivity (mean skill at each value, all other settings pooled):
min_cluster_size--
27 · Find kinds of gene defined by two labels at once (UMAP + HDBSCAN) -- reliable
Pattern 5: conjunctions found on a random half; each is checked on the other half as the second label's enrichment in the cluster among genes carrying the first label (hypergeometric p < 0.05). Metric: share replicating. Null: 20 runs with the second label shuffled within the second half. Pass: above the null's 95th percentile by 0.2 with at least three findings.
Metric: share of first-half findings that replicate on the second half. 15 runs, 3 settings, targets: (the strategy chooses its own).
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | minclustersize=15 | 0.490 [0.376, 0.623] | 100% [57, 100] | 0.502 | 0.027 | 5 |
| tuned | minclustersize=15 | 0.531 [0.322, 0.740] | 100% [34, 100] | 0.542 | 0.027 | 2 |
Sensitivity (mean skill at each value, all other settings pooled):
min_cluster_size-- 15: 0.49; 30: 0.66; 8: 0.48
28 · Find paralogs that changed jobs (profile correlation) -- reliable
Paralog pairs with both genes labelled; the labels are withheld from the profiles. Metric: AUROC of divergence for pairs whose labels differ against pairs whose labels match. Null: 20 random reassignments of the divergence values to pairs. Pass: above the null's 95th percentile by 0.05.
Metric: AUROC of profile divergence for paralogs with different compartment. 75 runs, 3 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | min_shared=10 | 0.164 [0.057, 0.312] | 60% [39, 78] | 0.582 | 0.5 | 20 |
| tuned | min_shared=10 | 0.164 [0.056, 0.312] | 62% [31, 86] | 0.582 | 0.5 | 8 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.094 [0.085, 0.103] | 40% | 5 |
| dtm_class | 0.026 [0.022, 0.030] | 0% | 5 |
| lopit_unified | 0.151 [0.141, 0.160] | 100% | 5 |
| screenanyphenotype | nan [nan, nan] | nan% | 0 |
| stageenrichedderived | 0.384 [0.376, 0.392] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
min_shared-- 10: 0.16; 30: 0.16; 5: 0.16
29 · Carry what one parasite shows to the other (orthogroup mapping) -- reliable
Numeric target: 25% of the genes measured in both species hidden; the relation is learned on the rest. Metric: rank correlation of transferred and hidden values. Null: 20 runs with the ortholog values permuted among genes. Categorical target: Pattern 1 on genes with an ortholog value. Pass: above the null's 95th percentile by 0.1 (numeric) or 0.05 (categorical).
Metric: rank correlation of transferred and hidden values. 5 runs, 1 settings, targets: (the strategy chooses its own).
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | -- | 0.314 [0.298, 0.333] | 100% [57, 100] | 0.316 | 0.003 | 5 |
| tuned | -- | 0.321 [0.293, 0.349] | 100% [34, 100] | 0.323 | 0.003 | 2 |
30 · Test inference on the genes orthology cannot reach (kNN) -- reliable
Pattern 1 restricted to the stratum: 25% of the label hidden by whole orthogroups; only hidden genes inside the stratum are scored. Null: 10 runs on shuffled labels. Pass: above the null's 95th percentile by 0.05.
Metric: correct calls per hidden gene. 100 runs, 4 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | stratum=lineage-specific | 0.264 [0.146, 0.345] | 75% [53, 89] | 0.586 | 0.416 | 20 |
| tuned | stratum=lineage-specific | 0.281 [0.166, 0.358] | 75% [41, 93] | 0.591 | 0.412 | 8 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.291 [0.251, 0.321] | 100% | 5 |
| dtm_class | 0.086 [0.066, 0.107] | 0% | 5 |
| lopit_unified | 0.349 [0.339, 0.361] | 100% | 5 |
| screenanyphenotype | nan [nan, nan] | nan% | 0 |
| stageenrichedderived | 0.332 [0.290, 0.371] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
stratum-- conserved: 0.25; hypothetical protein: 0.21; lineage-specific: 0.26; understudied: 0.21
31 · Call a gene only when independent strategies agree (kNN + logistic + network vote) -- reliable
Pattern 1 scored by precision: 25% of the label hidden; each method is trained on the visible genes and the agreed calls on hidden genes are scored. Metric: share of agreed calls that are correct. Null: 5 runs with visible labels shuffled. Pass: above the null's 95th percentile by 0.1. Single-method precisions are reported.
Metric: precision of calls on hidden genes. 75 runs, 3 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | min_agree=2 | 0.348 [0.174, 0.485] | 60% [41, 77] | 0.705 | 0.513 | 25 |
| tuned | min_agree=3 | 0.608 [0.331, 0.774] | 80% [49, 94] | 0.835 | 0.523 | 10 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.723 [0.676, 0.769] | 100% | 5 |
| dtm_class | 0.790 [0.761, 0.819] | 100% | 5 |
| lopit_unified | 0.730 [0.711, 0.747] | 100% | 5 |
| screenanyphenotype | -0.034 [-0.152, 0.070] | 0% | 5 |
| stageenrichedderived | 0.781 [0.752, 0.818] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
min_agree-- 1: 0.28; 2: 0.35; 3: 0.60
32 · Put the understudied genes first (kNN + logistic + network vote) -- reliable
Pattern 1 scored by precision and restricted to understudied genes: 25% of the label hidden; agreed calls are scored only on hidden understudied genes. Null: 5 runs with visible labels shuffled. Pass: above the null's 95th percentile by 0.1.
Metric: precision of calls on hidden genes. 75 runs, 3 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | min_agree=2 | 0.320 [0.134, 0.456] | 60% [41, 77] | 0.702 | 0.53 | 25 |
| tuned | min_agree=3 | 0.585 [0.318, 0.754] | 80% [49, 94] | 0.831 | 0.545 | 10 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.689 [0.615, 0.760] | 100% | 5 |
| dtm_class | 0.773 [0.741, 0.805] | 100% | 5 |
| lopit_unified | 0.702 [0.680, 0.724] | 100% | 5 |
| screenanyphenotype | -0.060 [-0.208, 0.038] | 0% | 5 |
| stageenrichedderived | 0.754 [0.728, 0.790] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
min_agree-- 1: 0.26; 2: 0.32; 3: 0.57
33 · Put every layer into one space and read a gene's neighbourhood (logistic edge model) -- reliable
A layer's edges are hidden by orthogroup -- whole groups at a time, so no hidden edge survives through a visible paralog -- and the layer is removed from its own features. Metric: AUROC of the hidden edges against DEGREE-MATCHED non-pairs, one per hidden edge, matched on connectivity at both ends. Null: 10 configuration-model rewirings of the hidden edges, which keep their degree sequence and destroy their topology, so a model reading fame cannot beat it. Pass: above the null's 95th percentile by 0.05. The AUROC against random non-pairs and the gap between the two are reported as numbers, not as the verdict.
Metric: AUROC of hidden coexpression edges against degree-matched non-pairs. 150 runs, 30 settings, targets: (the strategy chooses its own).
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | k=10; knn=15; layer=coexpression | 0.587 [0.582, 0.592] | 100% [57, 100] | 0.79 | 0.491 | 5 |
| tuned | k=10; knn=15; layer=cotranslation | 0.901 [0.864, 0.939] | 100% [34, 100] | 0.949 | 0.481 | 2 |
Sensitivity (mean skill at each value, all other settings pooled):
k-- 10: 0.47; 25: 0.47; 5: 0.47knn-- 15: 0.47; 5: 0.47layer-- coexpression: 0.59; cofitness: 0.14; cotranslation: 0.89; struct: 0.49; xlms: 0.26
34 · Train on the networks and rank the edges they are missing (logistic / spectral embedding) -- reliable
The chosen layer's edges are hidden by orthogroup and the layer is removed from its own features. Metric: AUROC of the hidden edges against DEGREE-MATCHED non-pairs. Null: 10 configuration-model rewirings of the hidden edges, which preserve their degree sequence, so fame alone cannot clear the bar. Pass: above the null's 95th percentile by 0.05. The random-null AUROC, the fame gap, precision@k, the Brier score and the reliability gap are all reported as numbers beside the verdict.
Metric: AUROC of hidden coexpression edges against degree-matched non-pairs. 100 runs, 20 settings, targets: (the strategy chooses its own).
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | fraction=0.25; layer=coexpression; model=logistic | 0.587 [0.582, 0.592] | 100% [57, 100] | 0.79 | 0.491 | 5 |
| tuned | fraction=0.4; layer=cotranslation; model=logistic | 0.880 [0.877, 0.884] | 100% [34, 100] | 0.939 | 0.492 | 2 |
Sensitivity (mean skill at each value, all other settings pooled):
fraction-- 0.25: 0.50; 0.4: 0.48layer-- coexpression: 0.59; cofitness: 0.14; cotranslation: 0.89; struct: 0.50; xlms: 0.34model-- embedding: 0.51; logistic: 0.47
35 · Call genes with a stated error rate (split conformal prediction) -- reliable
25% of the label hidden by whole orthogroups; the model is trained and calibrated on the rest (itself split by orthogroup) and builds a set for every hidden gene. Metric: set efficiency, 1 - (mean set size - 1) / (classes - 1) -- how far the sets narrow the possibilities. Null: 10 runs on shuffled labels, where the model learns nothing and the sets must grow to keep the promise. Pass: above the null's 95th percentile by 0.05. Also reported: set coverage on the hidden genes beside the 1 - alpha promised, and the scorecard of the single-label calls.
Metric: set efficiency: 1 - (mean set size - 1) / (classes - 1). 150 runs, 6 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | alpha=0.1; model=logistic | 0.500 [0.211, 0.729] | 80% [61, 91] | 0.588 | 0.176 | 25 |
| tuned | alpha=0.2; model=kNN | 0.496 [0.203, 0.765] | 80% [49, 94] | 0.609 | 0.186 | 10 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.449 [0.405, 0.495] | 100% | 5 |
| dtm_class | 0.399 [0.378, 0.418] | 100% | 5 |
| lopit_unified | 0.652 [0.574, 0.725] | 100% | 5 |
| screenanyphenotype | 0.071 [-0.040, 0.156] | 0% | 5 |
| stageenrichedderived | 0.884 [0.868, 0.903] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
alpha-- 0.05: 0.23; 0.1: 0.38; 0.2: 0.55model-- kNN: 0.28; logistic: 0.50
36 · Smooth the measurements along the networks, then classify (graph convolution + logistic regression) -- reliable
Pattern 1: 25% of the label hidden by whole orthogroups; the smoothed features are computed from measurements only, and the model is trained on the visible genes. Metric: hidden genes called correctly. Null: the analytic chance level for the same predicted and true class mixes. Pass: above that level's 95% bound by 0.05.
Metric: correct calls per hidden gene. 150 runs, 6 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | C=0.5; hops=2 | 0.399 [0.164, 0.569] | 80% [61, 91] | 0.629 | 0.359 | 25 |
| tuned | C=0.5; hops=1 | 0.401 [0.185, 0.560] | 80% [49, 94] | 0.629 | 0.358 | 10 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.441 [0.431, 0.451] | 100% | 5 |
| dtm_class | 0.441 [0.425, 0.454] | 100% | 5 |
| lopit_unified | 0.522 [0.510, 0.534] | 100% | 5 |
| screenanyphenotype | -0.045 [-0.088, -0.001] | 0% | 5 |
| stageenrichedderived | 0.646 [0.621, 0.673] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
C-- 0.1: 0.39; 0.5: 0.40hops-- 1: 0.39; 2: 0.39; 3: 0.39
37 · Let a random forest find what defines a label (random forest + permutation importance) -- reliable
Pattern 1: 25% of the label hidden by whole orthogroups; the forest is trained on the visible genes and calls the hidden ones. Metric: hidden genes called correctly. Null: the analytic chance level for the same predicted and true class mixes. Pass: above that level's 95% bound by 0.05.
Metric: correct calls per hidden gene. 150 runs, 6 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | min_leaf=2; trees=300 | 0.454 [0.216, 0.647] | 80% [61, 91] | 0.701 | 0.444 | 25 |
| tuned | min_leaf=5; trees=300 | 0.451 [0.210, 0.659] | 80% [49, 94] | 0.696 | 0.43 | 10 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.426 [0.406, 0.446] | 100% | 5 |
| dtm_class | 0.770 [0.756, 0.788] | 100% | 5 |
| lopit_unified | 0.502 [0.491, 0.511] | 100% | 5 |
| screenanyphenotype | -0.009 [-0.052, 0.030] | 0% | 5 |
| stageenrichedderived | 0.651 [0.628, 0.673] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
min_leaf-- 1: 0.44; 2: 0.45; 5: 0.46trees-- 100: 0.45; 300: 0.45
38 · Learn how much to trust each kind of evidence (stacked logistic regression) -- reliable
Pattern 1: 25% of the label hidden by whole orthogroups; base predictions for the meta-model are out of fold within the visible genes only, so no hidden label reaches either level. Metric: hidden genes called correctly. Null: the analytic chance level for the same predicted and true class mixes. Pass: above that level's 95% bound by 0.05.
Metric: correct calls per hidden gene. 75 runs, 3 settings, targets: compartment, dtmclass, lopitunified, screenanyphenotype, stageenrichedderived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | k=15 | 0.408 [0.183, 0.573] | 80% [61, 91] | 0.619 | 0.342 | 25 |
| tuned | k=10 | 0.406 [0.198, 0.559] | 80% [49, 94] | 0.613 | 0.337 | 10 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| compartment | 0.435 [0.419, 0.450] | 100% | 5 |
| dtm_class | 0.456 [0.442, 0.470] | 100% | 5 |
| lopit_unified | 0.541 [0.533, 0.550] | 100% | 5 |
| screenanyphenotype | -0.021 [-0.049, 0.007] | 0% | 5 |
| stageenrichedderived | 0.636 [0.613, 0.658] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
k-- 10: 0.41; 15: 0.41; 30: 0.41
39 · Predict a value with an interval that holds (gradient boosting / ridge + split conformal) -- weak
Pattern 4: 20% of the measured values hidden; the model is trained and calibrated on the rest (split by orthogroup) and predicts the hidden ones. Metric: rank correlation of predicted and hidden values. Null: the chance distribution of a rank correlation, with refits on shuffled values reported. Pass: above its 95th percentile by 0.1. Also reported: the share of hidden values inside their intervals, beside the coverage promised.
Metric: rank correlation of predicted and hidden values. 60 runs, 4 settings, targets: exprtachy, fitinvitrohff, fitinvivo_PE.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | alpha=0.1; model=boosted | 0.552 [0.036, 0.906] | 67% [42, 85] | 0.552 | 0.0 | 15 |
| tuned | alpha=0.2; model=ridge | 0.532 [0.029, 0.858] | 67% [30, 90] | 0.532 | 0.0 | 6 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| expr_tachy | 0.855 [0.841, 0.868] | 100% | 5 |
| fitinvitrohff | 0.701 [0.695, 0.708] | 100% | 5 |
| fitinvivoPE | 0.046 [0.029, 0.063] | 0% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
alpha-- 0.1: 0.54; 0.2: 0.54model-- boosted: 0.55; ridge: 0.53
Plasmodium falciparum
01 · Hold out a category and search for a map that finds it (UMAP + HDBSCAN) -- weak
25% of the label is hidden (whole orthogroups together, so no gene is recovered through a visible paralog). The walk is built as set -- its genes per map (up to 4,000), feature sets and grids (up to three values each); the configuration, and the one cluster that best isolates each label, are both chosen using visible labels only. Metric: for each label, the F1 of its hidden genes against its chosen cluster, weighted by size -- does the structure found on known genes hold the unknown ones? Null: the same chosen clusters scored after permuting the hidden genes' labels, 100 times. (A second search on shuffled labels was the null once; it picks the largest cluster for every label, which scores F1 near 2p by size alone and made the null beat real labels.) Pass: above the null's 95th percentile by at least 0.05.
Metric: F1 of hidden genes in the cluster chosen for their label on known genes. 120 runs, 6 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | minclustersize=[20, 50]; n_neighbors=[15, 50]; selection=[eom, leaf] | 0.087 [0.009, 0.195] | 25% [11, 47] | 0.536 | 0.503 | 20 |
| tuned | minclustersize=[20, 50]; n_neighbors=[10, 30]; selection=[eom, leaf] | 0.114 [0.019, 0.287] | 25% [7, 59] | 0.566 | 0.535 | 8 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | 0.345 [0.214, 0.475] | 0% | 5 |
| lopitpflocation | 0.094 [0.078, 0.109] | 100% | 5 |
| pbtransferredphenotype | -0.000 [-0.000, 0.000] | 0% | 5 |
| stageenrichedderived | 0.036 [0.014, 0.062] | 0% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
min_cluster_size-- 10, 25: 0.06; 20, 50: 0.11n_neighbors-- 10, 30: 0.09; 15, 50: 0.07; 25, 100: 0.09selection-- eom, leaf: 0.09
02 · Find the map where your gene list is one cluster (UMAP + HDBSCAN) -- reliable
30% of the set is hidden; the walk as set (genes per map up to 4,000, feature sets, up to three values of each grid) picks the cluster with the best F1 for the other 70%. Metric: F1 of the hidden members against that cluster's other genes -- precision is the share of the cluster's candidates that are hidden members, recall the share of hidden members among them. Null: 20 random sets of the same size through the same walk. Pass: above the null's 95th percentile by at least 0.05 (an F1 margin; random sets score about 0.04).
Metric: F1 of the hidden members against the best cluster's other genes. 80 runs, 4 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | minclustersize=[20, 50]; n_neighbors=[15, 50] | 0.214 [0.093, 0.352] | 90% [70, 97] | 0.23 | 0.022 | 20 |
| tuned | minclustersize=[20, 50]; n_neighbors=[15, 50] | 0.217 [0.085, 0.353] | 75% [41, 93] | 0.234 | 0.022 | 8 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | 0.242 [0.211, 0.274] | 100% | 5 |
| lopitpflocation | 0.151 [0.119, 0.200] | 100% | 5 |
| pbtransferredphenotype | 0.411 [0.388, 0.430] | 100% | 5 |
| stageenrichedderived | 0.051 [0.035, 0.067] | 60% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
min_cluster_size-- 10, 25: 0.21; 20, 50: 0.21n_neighbors-- 10, 30: 0.21; 15, 50: 0.21
03 · Ask which categories the data can rediscover (UMAP + neighbour AUROC) -- reliable
30% of the label is hidden. The atlas is built from visible genes only (leave-one-out neighbour AUROC per category); hidden genes are then scored by their visible neighbours. Metric: mean hidden-gene AUROC over the categories the atlas ranks in its top half. Null: the same with the visible labels shuffled, 20 times. Pass: above the null's 95th percentile by 0.05. The rank agreement between atlas and hidden recovery is reported.
Metric: hidden-gene AUROC of the categories the atlas ranks in its top half. 60 runs, 3 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | k=15 | 0.686 [0.565, 0.793] | 100% [84, 100] | 0.843 | 0.503 | 20 |
| tuned | k=15 | 0.652 [0.555, 0.717] | 100% [68, 100] | 0.828 | 0.506 | 8 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | 0.771 [0.705, 0.856] | 100% | 5 |
| lopitpflocation | 0.668 [0.641, 0.696] | 100% | 5 |
| pbtransferredphenotype | 0.506 [0.452, 0.562] | 100% | 5 |
| stageenrichedderived | 0.798 [0.721, 0.875] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
k-- 15: 0.69; 5: 0.61; 50: 0.67
04 · Keep only the modules that survive the whole walk (UMAP + HDBSCAN co-clustering) -- weak
Modules are built from a small walk on 1,500 genes without the label. 25% of the label is hidden; the module that best isolates each label is chosen on the visible genes. Metric: the size-weighted F1 of each label's hidden genes against its chosen module. Null: the same modules scored after permuting the hidden genes' labels, 100 times -- a large module scores the same either way and earns nothing. Pass: above the null's 95th percentile by 0.05.
Metric: F1 of hidden genes in the module chosen for their label on known genes. 60 runs, 3 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | threshold=0.5 | 0.145 [0.034, 0.284] | 45% [26, 66] | 0.503 | 0.446 | 20 |
| tuned | threshold=0.5 | 0.186 [0.047, 0.387] | 50% [22, 78] | 0.519 | 0.453 | 8 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | 0.358 [0.222, 0.471] | 0% | 5 |
| lopitpflocation | 0.140 [0.114, 0.166] | 100% | 5 |
| pbtransferredphenotype | 0.076 [0.032, 0.108] | 80% | 5 |
| stageenrichedderived | 0.007 [-0.000, 0.020] | 0% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
threshold-- 0.3: 0.13; 0.5: 0.15; 0.8: 0.05
05 · Tune a map without labels, then read what it encodes (UMAP + HDBSCAN, chi-square / Kruskal-Wallis) -- reliable
Pattern 5. One label-free map on up to 2,000 genes. Held-out features significant at q < 0.05 on a random half of the genes are the findings; each is re-tested (p < 0.05) on the other half. Metric: share that replicate. Null: the same with the second half's cluster labels permuted, 10 times. Pass: above the null's 95th percentile by 0.2, with at least three findings.
Metric: share of first-half findings that replicate on the second half. 250 runs, 50 settings, targets: (the strategy chooses its own).
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | map_from=transcription | 0.876 [0.817, 0.934] | 100% [57, 100] | 0.882 | 0.054 | 5 |
| tuned | map_from=isoelectric | 0.953 [0.906, 1.000] | 100% [34, 100] | 0.956 | 0.061 | 2 |
Sensitivity (mean skill at each value, all other settings pooled):
map_from-- antisense: 0.96; chromprox: 0.94; codon: 0.94; committed: 0.92; complex: 0.90; dnds: 0.95; engaged: 0.96; ev: 0.57; exon: 0.80; export: 0.90; expr: 0.95; febrile: 0.97; fertility: 0.81; field: 0.87; gametocyte: 0.92; has: 0.94; hsp90: 0.93; idc: 0.86; is: 0.93; isoelectric: 0.98; latency: 0.87; length: 0.93; lopit: 0.96; m6a: 0.97; mean: 0.98; melting: 0.93; molecular: 0.95; mrna: 0.92; n: 0.96; novel: 0.92; ortholog: 0.92; paralog: 0.91; pb: 0.95; pf6: 0.97; piggybac: 0.85; plddt: 0.92; polysomal: 0.97; protein: 0.97; proteome: 0.95; resistance: 0.77; riboseq: 0.95; rna: 0.96; sir2: 0.83; sir2a: 0.85; sir2b: 0.87; snp: 0.95; stage: 0.86; steady: 0.96; transcript: 0.93; transcription: 0.88
06 · Find which kind of evidence carries a label (kNN ablation) -- reliable
25% of the label is hidden. Each kind of evidence is ranked by cross-validated accuracy on the visible labels; the top-ranked one then predicts the hidden labels. Metric: its hidden accuracy. Null: the hidden accuracy of every kind of evidence, i.e. choosing at random. Pass: above the null's 80th percentile by 0.02 (with ~15 kinds of evidence the 95th would demand the single best, which asks more than a ranking must deliver).
Metric: hidden accuracy of the evidence ranked first (expr). 60 runs, 3 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | k=15 | 0.133 [0.066, 0.191] | 65% [43, 82] | 0.643 | 0.569 | 20 |
| tuned | k=5 | 0.145 [0.077, 0.203] | 75% [41, 93] | 0.63 | 0.553 | 8 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | 0.106 [0.053, 0.159] | 0% | 5 |
| lopitpflocation | 0.180 [0.167, 0.189] | 100% | 5 |
| pbtransferredphenotype | 0.150 [0.130, 0.173] | 100% | 5 |
| stageenrichedderived | 0.164 [0.094, 0.234] | 80% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
k-- 15: 0.13; 5: 0.15; 50: 0.14
07 · Call a gene by the genes that behave like it (kNN) -- weak
Pattern 1: 25% of the label hidden by whole orthogroups; each hidden gene called by its k nearest visible genes with the same vote threshold. Metric: hidden genes called correctly. Null: 10 runs with visible labels shuffled. Pass: above the null's 95th percentile by 0.05.
Metric: correct calls per hidden gene. 180 runs, 9 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | k=15; min_share=0.3 | 0.076 [-0.369, 0.417] | 50% [30, 70] | 0.684 | 0.532 | 20 |
| tuned | k=15; min_share=0.3 | 0.219 [0.020, 0.423] | 50% [22, 78] | 0.699 | 0.539 | 8 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | -0.523 [-1.101, 0.041] | 0% | 5 |
| lopitpflocation | 0.453 [0.438, 0.471] | 100% | 5 |
| pbtransferredphenotype | 0.375 [0.359, 0.391] | 100% | 5 |
| stageenrichedderived | -0.000 [-0.055, 0.055] | 0% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
k-- 15: 0.05; 5: 0.14; 50: -0.02min_share-- 0.0: 0.08; 0.3: 0.08; 0.6: 0.00
08 · Call a gene by its neighbours on the map (UMAP + kNN) -- weak
Pattern 1 on the genes the map places (up to 2,500): 25% of the label hidden by whole orthogroups; hidden genes called by their k nearest visible genes in the map. Null: 20 runs on shuffled visible labels. Pass: above the null's 95th percentile by 0.05.
Metric: correct calls per hidden gene. 180 runs, 9 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | k=15; n_neighbors=25 | -0.003 [-0.531, 0.286] | 60% [39, 78] | 0.635 | 0.525 | 20 |
| tuned | k=50; n_neighbors=25 | 0.148 [0.025, 0.270] | 62% [31, 86] | 0.643 | 0.537 | 8 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | -0.581 [-1.235, -0.007] | 0% | 5 |
| lopitpflocation | 0.262 [0.250, 0.273] | 100% | 5 |
| pbtransferredphenotype | 0.273 [0.255, 0.285] | 100% | 5 |
| stageenrichedderived | 0.084 [0.019, 0.146] | 20% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
k-- 15: -0.01; 5: 0.03; 50: -0.00n_neighbors-- 10: 0.01; 25: 0.00; 60: 0.00
09 · Name a cluster by the label it is enriched for (UMAP + HDBSCAN, hypergeometric) -- reliable
Pattern 1 scored by precision, on the genes a blind map places (up to 2,500): 25% of the label hidden; the enrichment is recomputed from visible labels, and the calls it makes on hidden genes are scored -- the strategy abstains on noise and unenriched clusters by design, so what matters is how often a call is right. Null: 20 runs on shuffled visible labels. Pass: above the null's 95th percentile by 0.1.
Metric: precision of calls on hidden genes. 240 runs, 12 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | minclustersize=25; min_lift=3.0; selection=leaf | 0.418 [0.326, 0.603] | 100% [76, 100] | 0.422 | 0.008 | 12 |
| tuned | minclustersize=25; min_lift=1.5; selection=leaf | 0.560 [0.447, 0.703] | 100% [68, 100] | 0.564 | 0.011 | 8 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | 0.358 [0.212, 0.496] | 100% | 5 |
| lopitpflocation | 0.456 [0.408, 0.516] | 100% | 5 |
| pbtransferredphenotype | 0.760 [0.696, 0.814] | 100% | 5 |
| stageenrichedderived | 0.565 [0.417, 0.700] | 100% | 3 |
Sensitivity (mean skill at each value, all other settings pooled):
min_cluster_size-- 10: 0.47; 25: 0.47; 60: 0.38min_lift-- 1.5: 0.48; 3.0: 0.38selection-- eom: 0.36; leaf: 0.47
10 · Find genes whose label their neighbours contradict (kNN + network neighbours) -- reliable
5% of the labels (at least ten) are swapped to a wrong class, drawn in proportion to class size. Surprise is computed with the corrupted labels. Metric: AUROC of surprise for the swapped genes among all labelled genes. Null: 20 random sets of the same size. Pass: above the null's 95th percentile by 0.1.
Metric: AUROC of surprise for the swapped labels. 60 runs, 3 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | k=15 | 0.740 [0.573, 0.918] | 100% [84, 100] | 0.871 | 0.501 | 20 |
| tuned | k=15 | 0.765 [0.612, 0.924] | 100% [68, 100] | 0.882 | 0.497 | 8 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | 0.992 [0.991, 0.993] | 100% | 5 |
| lopitpflocation | 0.720 [0.700, 0.742] | 100% | 5 |
| pbtransferredphenotype | 0.521 [0.492, 0.547] | 100% | 5 |
| stageenrichedderived | 0.727 [0.659, 0.818] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
k-- 15: 0.74; 5: 0.74; 50: 0.70
11 · Diffuse a label across one measured network (random walk with restart) -- reliable
Pattern 1: 25% of the label hidden by whole orthogroups; the fields are seeded from visible genes only, so a hidden gene never seeds its own call. Metric: hidden genes called correctly (unreached genes count as misses). Null: 10 runs on shuffled labels. Pass: above the null's 95th percentile by 0.05.
Metric: correct calls per hidden gene. 480 runs, 24 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | layer=coexpression; restart=0.5 | 0.217 [0.137, 0.331] | 93% [70, 99] | 0.446 | 0.317 | 15 |
| tuned | layer=coexpression; restart=0.2 | 0.244 [0.135, 0.442] | 100% [61, 100] | 0.45 | 0.308 | 6 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | 0.376 [0.265, 0.469] | 100% | 5 |
| lopitpflocation | 0.142 [0.135, 0.152] | 100% | 5 |
| pbtransferredphenotype | 0.177 [0.152, 0.201] | 100% | 5 |
| stageenrichedderived | nan [nan, nan] | nan% | 0 |
Sensitivity (mean skill at each value, all other settings pooled):
layer-- coexpression: 0.22; comention: 0.01; cotranslation: 0.05; domain: 0.11; ip_ms: -0.00; orthogroup: 0.00; struct: 0.03; xlms: 0.01restart-- 0.2: 0.05; 0.5: 0.05; 0.8: 0.05
12 · Let every network vote, weighted by what it has earned (chance-weighted ensemble vote) -- reliable
Pattern 1: 25% of the label hidden; the weights are learned on an inner holdout of the visible labels only, then hidden genes are called. Null: 5 runs with visible labels shuffled (weights relearned each time). Pass: above the null's 95th percentile by 0.05.
Metric: correct calls per hidden gene. 60 runs, 3 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | k=15 | 0.512 [0.415, 0.650] | 65% [43, 82] | 0.646 | 0.323 | 20 |
| tuned | k=15 | 0.496 [0.395, 0.612] | 62% [31, 86] | 0.655 | 0.334 | 8 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | 0.705 [0.478, 0.847] | 40% | 5 |
| lopitpflocation | 0.468 [0.445, 0.490] | 100% | 5 |
| pbtransferredphenotype | 0.445 [0.423, 0.476] | 100% | 5 |
| stageenrichedderived | 0.429 [0.284, 0.531] | 20% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
k-- 15: 0.51; 5: 0.55; 50: 0.33
13 · Place a protein by the proteins it physically touches (weighted partner vote) -- weak
Pattern 1 restricted to genes with at least one physical partner: 25% of the label hidden by whole orthogroups; hidden genes called by their visible partners' weighted vote. Null: 20 runs on shuffled labels. Pass: above the null's 95th percentile by 0.05.
Metric: correct calls per hidden gene. 20 runs, 1 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | -- | 0.216 [0.015, 0.468] | 33% [15, 58] | 0.539 | 0.371 | 15 |
| tuned | -- | 0.210 [0.026, 0.524] | 33% [10, 70] | 0.542 | 0.355 | 6 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | 0.149 [0.037, 0.341] | 0% | 5 |
| lopitpflocation | 0.481 [0.401, 0.524] | 100% | 5 |
| pbtransferredphenotype | 0.018 [-0.107, 0.109] | 0% | 5 |
| stageenrichedderived | nan [nan, nan] | nan% | 0 |
14 · Annotate function through shared fold (TM-score-weighted vote) -- reliable
Pattern 1 restricted to proteins with a structural neighbour: 25% of the annotation (at the chosen level) hidden by whole orthogroups; hidden proteins called by their visible structural neighbours. Null: 20 runs on shuffled annotations. Pass: above the null's 95th percentile by 0.05.
Metric: correct calls per hidden gene. 15 runs, 3 settings, targets: (the strategy chooses its own).
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | level=1 | 0.587 [0.544, 0.613] | 100% [57, 100] | 0.693 | 0.258 | 5 |
| tuned | level=2 | 0.588 [0.561, 0.615] | 100% [34, 100] | 0.644 | 0.137 | 2 |
Sensitivity (mean skill at each value, all other settings pooled):
level-- 1: 0.59; 2: 0.66; 3: 0.75
15 · Find the communities several networks agree on (modularity + Louvain consensus) -- weak
On genes placed in a community: 25% of the label hidden; the community that best isolates each label is chosen on the visible genes (the communities themselves never see labels). Metric: the size-weighted F1 of each label's hidden genes against its chosen community. Null: the same communities scored after permuting the hidden genes' labels, 100 times. Pass: above the null's 95th percentile by 0.05.
Metric: F1 of hidden genes in the community chosen for their label on known genes. 120 runs, 6 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | agreement=0.5; resolution=1.0 | 0.029 [-0.032, 0.088] | 50% [30, 70] | 0.249 | 0.225 | 20 |
| tuned | agreement=0.3; resolution=2.0 | 0.055 [0.002, 0.100] | 50% [22, 78] | 0.194 | 0.148 | 8 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | 0.025 [0.020, 0.030] | 0% | 5 |
| lopitpflocation | 0.095 [0.085, 0.108] | 100% | 5 |
| pbtransferredphenotype | 0.083 [0.074, 0.094] | 100% | 5 |
| stageenrichedderived | 0.021 [-0.058, 0.100] | 20% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
agreement-- 0.3: 0.03; 0.5: 0.04resolution-- 0.5: 0.03; 1.0: 0.03; 2.0: 0.06
16 · Predict the contacts an interactome missed (logistic regression) -- reliable
Pattern 3: 20% of the layer's edges hidden; the model is trained on the rest against DEGREE-MATCHED non-edges -- each with a gene of similar degree at both ends -- and scores the hidden edges against fresh degree-matched non-edges. Against random non-edges this test read AUROC 0.99 on every correlation layer, because a random pair is usually two obscure genes and degree separates them. The layer's source measurements leave the similarity feature with it, and derived, annotation and literature layers are refused as targets. Metric: AUROC. Null: 5 models trained with each gene's evidence read from a random other gene (identities permuted). Pass: above the null's 95th percentile by 0.05.
Metric: AUROC of hidden pairs against degree-matched non-pairs. 5 runs, 1 settings, targets: (the strategy chooses its own).
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | layer=struct | 0.958 [0.957, 0.961] | 100% [57, 100] | 0.98 | 0.508 | 5 |
| tuned | layer=struct | 0.961 [0.958, 0.963] | 100% [34, 100] | 0.98 | 0.504 | 2 |
Sensitivity (mean skill at each value, all other settings pooled):
layer-- struct: 0.96
17 · Read the literature for biology, not fame (publication-count residual) -- weak
The label is never used to build the literature layer. Among co-mentioned pairs with both genes labelled, the top k by corrected residual are taken (k = the chosen number, capped at a fifth of the pool). Metric: the share of those pairs sharing a label. Null: 20 random sets of k co-mentioned pairs. Pass: above the null's 95th percentile by 0.05.
Metric: share of the top 48 corrected pairs sharing a lopit_pf_location label. 60 runs, 3 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | top=200 | 0.263 [0.229, 0.287] | 40% [20, 64] | 0.727 | 0.624 | 15 |
| tuned | top=1000 | 0.258 [0.213, 0.297] | 33% [10, 70] | 0.727 | 0.628 | 6 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | 0.237 [0.187, 0.289] | 0% | 5 |
| lopitpflocation | 0.267 [0.262, 0.271] | 100% | 5 |
| pbtransferredphenotype | 0.285 [0.258, 0.309] | 20% | 5 |
| stageenrichedderived | nan [nan, nan] | nan% | 0 |
Sensitivity (mean skill at each value, all other settings pooled):
top-- 1000: 0.26; 200: 0.26; 50: 0.13
18 · List what the data says and the literature has not written (multi-layer support count) -- reliable
Pattern 3 with the literature as truth: co-mentioned pairs among genes in the measurement layers against 5 times as many random pairs. Score: number of measurement layers linking the pair. Metric: AUROC. Null: 10 runs with the measurement layers' gene identities permuted. Pass: above the null's 95th percentile by 0.02.
Metric: AUROC of hidden pairs against random non-pairs. 15 runs, 3 settings, targets: (the strategy chooses its own).
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | min_layers=2 | 0.108 [0.106, 0.109] | 100% [57, 100] | 0.554 | 0.5 | 5 |
| tuned | min_layers=1 | 0.110 [0.110, 0.110] | 100% [34, 100] | 0.554 | 0.499 | 2 |
Sensitivity (mean skill at each value, all other settings pooled):
min_layers-- 1: 0.11; 2: 0.11; 3: 0.11
19 · Train a classifier on the known genes and call the rest (logistic regression) -- reliable
Pattern 1: 25% of the label hidden by whole orthogroups; the model is trained on the visible genes and calls the hidden ones. Metric: hidden genes called correctly. Null: the analytic chance level for the same predicted and true class mixes (re-running a multinomial fit on shuffled labels ten times would take longer than it tells). Pass: above that level's 95% bound by 0.05.
Metric: correct calls per hidden gene. 80 runs, 4 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | C=0.01 | 0.388 [0.351, 0.435] | 100% [84, 100] | 0.63 | 0.409 | 20 |
| tuned | C=0.1 | 0.440 [0.360, 0.509] | 100% [68, 100] | 0.687 | 0.445 | 8 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | 0.418 [0.353, 0.483] | 100% | 5 |
| lopitpflocation | 0.510 [0.498, 0.522] | 100% | 5 |
| pbtransferredphenotype | 0.346 [0.326, 0.367] | 100% | 5 |
| stageenrichedderived | 0.456 [0.411, 0.506] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
C-- 0.01: 0.39; 0.1: 0.43; 1.0: 0.43; 10.0: 0.41
20 · Learn what makes your list special, from positives alone (PU bagging, logistic regression) -- reliable
Pattern 2: 30% of the set hidden; the other 70% are the positives. Metric: AUROC of the hidden members against every other non-query gene. Null: 10 random sets of the same size. Pass: above the null's 95th percentile by 0.1.
Metric: AUROC of hidden members against every other gene. 60 runs, 3 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | bags=15 | 0.903 [0.801, 0.976] | 100% [84, 100] | 0.951 | 0.498 | 20 |
| tuned | bags=40 | 0.906 [0.809, 0.978] | 100% [68, 100] | 0.953 | 0.506 | 8 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | 0.917 [0.904, 0.933] | 100% | 5 |
| lopitpflocation | 0.944 [0.937, 0.952] | 100% | 5 |
| pbtransferredphenotype | 0.995 [0.994, 0.996] | 100% | 5 |
| stageenrichedderived | 0.757 [0.740, 0.777] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
bags-- 15: 0.90; 40: 0.90; 5: 0.90
21 · Predict a measurement, and find the genes that defy the prediction (gradient boosting / ridge) -- reliable
Pattern 4: 20% of the measured values hidden; the model is trained on the rest. Metric: rank correlation between predicted and hidden values. Null: 3 models trained on the visible values shuffled among the measured genes. Pass: above the null's 95th percentile by 0.1.
Metric: rank correlation of predicted and hidden values. 60 runs, 4 settings, targets: exprschizont, meanplddt, piggybac_mis.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | model=boosted; own_kind=leave out | 0.518 [0.247, 0.849] | 100% [80, 100] | 0.518 | 0.0 | 15 |
| tuned | model=boosted; own_kind=include | 0.523 [0.246, 0.847] | 100% [61, 100] | 0.523 | 0.0 | 6 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| expr_schizont | 0.241 [0.195, 0.284] | 100% | 5 |
| mean_plddt | 0.852 [0.843, 0.861] | 100% | 5 |
| piggybac_mis | 0.463 [0.449, 0.476] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
model-- boosted: 0.52; ridge: 0.46own_kind-- include: 0.49; leave out: 0.49
22 · Fill in what was never measured, and say where that is honest (soft-impute, low-rank SVD) -- reliable
Pattern 4 across the whole table: 10% of every column's measured entries hidden, the table completed at the chosen rank. Metric: median over columns of the rank correlation on hidden entries. Null: 3 completions of a table whose columns were each shuffled independently. Pass: above the null's 95th percentile by 0.1.
Metric: median per-column rank correlation on hidden entries. 15 runs, 3 settings, targets: (the strategy chooses its own).
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | rank=20 | 0.678 [0.672, 0.685] | 100% [57, 100] | 0.679 | 0.001 | 5 |
| tuned | rank=20 | 0.679 [0.678, 0.681] | 100% [34, 100] | 0.68 | 0.002 | 2 |
Sensitivity (mean skill at each value, all other settings pooled):
rank-- 20: 0.68; 5: 0.50; 60: 0.65
23 · Find what matters more in one condition, and why (residual + gradient boosting / ridge) -- reliable
Pattern 4 on the shift: 20% of genes with both measurements hidden; a model trained on the rest predicts their shift. Metric: rank correlation on hidden genes. Null: 3 models trained on shuffled shifts. Pass: above the null's 95th percentile by 0.1.
Metric: rank correlation of predicted and hidden values. 20 runs, 4 settings, targets: (the strategy chooses its own).
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | model=boosted; own_kind=leave out | 0.642 [0.637, 0.647] | 100% [57, 100] | 0.642 | 0.0 | 5 |
| tuned | model=boosted; own_kind=include | 0.642 [0.641, 0.644] | 100% [34, 100] | 0.642 | 0.0 | 2 |
Sensitivity (mean skill at each value, all other settings pooled):
model-- boosted: 0.64; ridge: 0.59own_kind-- include: 0.62; leave out: 0.62
24 · Describe what your gene list has in common (hypergeometric + rank-sum) -- reliable
Pattern 2: 40% of the set hidden; the profile is built from the other 60% and scores every gene. Metric: AUROC of the hidden members against every other non-query gene. Null: 10 random sets of the same size. Pass: above the null's 95th percentile by 0.1.
Metric: AUROC of hidden members against every other gene. 20 runs, 1 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | -- | 0.852 [0.755, 0.937] | 100% [84, 100] | 0.926 | 0.501 | 20 |
| tuned | -- | 0.854 [0.756, 0.937] | 100% [68, 100] | 0.927 | 0.501 | 8 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | 0.923 [0.917, 0.930] | 100% | 5 |
| lopitpflocation | 0.837 [0.823, 0.852] | 100% | 5 |
| pbtransferredphenotype | 0.949 [0.946, 0.952] | 100% | 5 |
| stageenrichedderived | 0.700 [0.664, 0.730] | 100% | 5 |
25 · Grow your gene list along the networks (random walk with restart) -- reliable
Pattern 2: 30% of the set hidden; the walk is seeded from the other 70%. Metric: AUROC of the hidden members against every other non-seed gene. Null: 10 random seed sets of the same size. Pass: above the null's 95th percentile by 0.1.
Metric: AUROC of hidden members against every other gene. 120 runs, 6 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | mode=networks + measurements; restart=0.3 | 0.859 [0.744, 0.959] | 100% [84, 100] | 0.93 | 0.501 | 20 |
| tuned | mode=networks + measurements; restart=0.6 | 0.854 [0.734, 0.961] | 100% [68, 100] | 0.93 | 0.508 | 8 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | 0.907 [0.884, 0.937] | 100% | 5 |
| lopitpflocation | 0.846 [0.823, 0.870] | 100% | 5 |
| pbtransferredphenotype | 0.992 [0.988, 0.994] | 100% | 5 |
| stageenrichedderived | 0.701 [0.656, 0.746] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
mode-- networks: 0.45; networks + measurements: 0.86restart-- 0.1: 0.65; 0.3: 0.65; 0.6: 0.66
26 · Find categories that split in two on another measurement (UMAP + HDBSCAN) -- untestable
Pattern 5: significant findings (q < 0.05) on a random half of the mapped genes; each is checked on the other half -- the minority group enriched again among the cluster's A-matching genes (hypergeometric p < 0.05), or the measurement bimodal again. Metric: share replicating. Null: 20 runs with B shuffled within the second half. Pass: above the null's 95th percentile by 0.2 with at least three findings.
Metric: share of findings that replicate. 15 runs, 3 settings, targets: (the strategy chooses its own).
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | minclustersize=40 | nan [nan, nan] | nan% [nan, nan] | nan | nan | 0 |
Sensitivity (mean skill at each value, all other settings pooled):
min_cluster_size--
27 · Find kinds of gene defined by two labels at once (UMAP + HDBSCAN) -- untestable
Pattern 5: conjunctions found on a random half; each is checked on the other half as the second label's enrichment in the cluster among genes carrying the first label (hypergeometric p < 0.05). Metric: share replicating. Null: 20 runs with the second label shuffled within the second half. Pass: above the null's 95th percentile by 0.2 with at least three findings.
Metric: share of findings that replicate. 15 runs, 3 settings, targets: (the strategy chooses its own).
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | minclustersize=15 | nan [nan, nan] | nan% [nan, nan] | nan | nan | 0 |
Sensitivity (mean skill at each value, all other settings pooled):
min_cluster_size--
28 · Find paralogs that changed jobs (profile correlation) -- weak
Paralog pairs with both genes labelled; the labels are withheld from the profiles. Metric: AUROC of divergence for pairs whose labels differ against pairs whose labels match. Null: 20 random reassignments of the divergence values to pairs. Pass: above the null's 95th percentile by 0.05.
Metric: AUROC of profile divergence for paralogs with different lopit_pf_location. 60 runs, 3 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | min_shared=10 | 0.112 [-0.137, 0.372] | 60% [36, 80] | 0.556 | 0.499 | 15 |
| tuned | min_shared=10 | 0.111 [-0.123, 0.362] | 50% [19, 81] | 0.556 | 0.501 | 6 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | 0.104 [0.099, 0.111] | 80% | 5 |
| lopitpflocation | 0.373 [0.364, 0.381] | 100% | 5 |
| pbtransferredphenotype | -0.142 [-0.157, -0.126] | 0% | 5 |
| stageenrichedderived | nan [nan, nan] | nan% | 0 |
Sensitivity (mean skill at each value, all other settings pooled):
min_shared-- 10: 0.11; 30: 0.11; 5: 0.11
29 · Carry what one parasite shows to the other (orthogroup mapping) -- reliable
Numeric target: 25% of the genes measured in both species hidden; the relation is learned on the rest. Metric: rank correlation of transferred and hidden values. Null: 20 runs with the ortholog values permuted among genes. Categorical target: Pattern 1 on genes with an ortholog value. Pass: above the null's 95th percentile by 0.1 (numeric) or 0.05 (categorical).
Metric: rank correlation of transferred and hidden values. 5 runs, 1 settings, targets: (the strategy chooses its own).
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | -- | 0.313 [0.298, 0.328] | 100% [57, 100] | 0.316 | 0.004 | 5 |
| tuned | -- | 0.322 [0.312, 0.333] | 100% [34, 100] | 0.329 | 0.01 | 2 |
30 · Test inference on the genes orthology cannot reach (kNN) -- weak
Pattern 1 restricted to the stratum: 25% of the label hidden by whole orthogroups; only hidden genes inside the stratum are scored. Null: 10 runs on shuffled labels. Pass: above the null's 95th percentile by 0.05.
Metric: correct calls per hidden gene. 80 runs, 4 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | stratum=lineage-specific | -0.007 [-0.007, -0.007] | 0% [0, 66] | 0.659 | 0.661 | 2 |
| tuned | stratum=conserved | 0.228 [0.037, 0.424] | 50% [22, 78] | 0.701 | 0.537 | 8 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | -0.525 [-1.084, 0.026] | 0% | 5 |
| lopitpflocation | 0.454 [0.440, 0.473] | 100% | 5 |
| pbtransferredphenotype | 0.374 [0.358, 0.390] | 100% | 5 |
| stageenrichedderived | 0.078 [-0.055, 0.246] | 20% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
stratum-- conserved: 0.10; lineage-specific: -0.01; understudied: 0.04
31 · Call a gene only when independent strategies agree (kNN + logistic + network vote) -- works when tuned
Pattern 1 scored by precision: 25% of the label hidden; each method is trained on the visible genes and the agreed calls on hidden genes are scored. Metric: share of agreed calls that are correct. Null: 5 runs with visible labels shuffled. Pass: above the null's 95th percentile by 0.1. Single-method precisions are reported.
Metric: precision of calls on hidden genes. 60 runs, 3 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | min_agree=2 | 0.245 [-0.256, 0.547] | 70% [48, 85] | 0.768 | 0.554 | 20 |
| tuned | min_agree=3 | 0.663 [0.505, 0.810] | 57% [25, 84] | 0.866 | 0.546 | 7 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | 0.756 [0.632, 0.858] | 0% | 5 |
| lopitpflocation | 0.827 [0.787, 0.865] | 100% | 5 |
| pbtransferredphenotype | 0.568 [0.542, 0.592] | 100% | 5 |
| stageenrichedderived | 0.468 [0.468, 0.468] | 0% | 1 |
Sensitivity (mean skill at each value, all other settings pooled):
min_agree-- 1: 0.13; 2: 0.25; 3: 0.70
32 · Put the understudied genes first (kNN + logistic + network vote) -- works when tuned
Pattern 1 scored by precision and restricted to understudied genes: 25% of the label hidden; agreed calls are scored only on hidden understudied genes. Null: 5 runs with visible labels shuffled. Pass: above the null's 95th percentile by 0.1.
Metric: precision of calls on hidden genes. 60 runs, 3 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | min_agree=2 | 0.187 [-0.389, 0.545] | 65% [43, 82] | 0.768 | 0.567 | 20 |
| tuned | min_agree=3 | 0.708 [0.532, 0.883] | 67% [30, 90] | 0.877 | 0.524 | 6 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | 0.751 [0.610, 0.866] | 0% | 5 |
| lopitpflocation | 0.854 [0.827, 0.882] | 100% | 5 |
| pbtransferredphenotype | 0.567 [0.537, 0.601] | 100% | 5 |
| stageenrichedderived | nan [nan, nan] | nan% | 0 |
Sensitivity (mean skill at each value, all other settings pooled):
min_agree-- 1: 0.07; 2: 0.19; 3: 0.72
33 · Put every layer into one space and read a gene's neighbourhood (logistic edge model) -- reliable
A layer's edges are hidden by orthogroup -- whole groups at a time, so no hidden edge survives through a visible paralog -- and the layer is removed from its own features. Metric: AUROC of the hidden edges against DEGREE-MATCHED non-pairs, one per hidden edge, matched on connectivity at both ends. Null: 10 configuration-model rewirings of the hidden edges, which keep their degree sequence and destroy their topology, so a model reading fame cannot beat it. Pass: above the null's 95th percentile by 0.05. The AUROC against random non-pairs and the gap between the two are reported as numbers, not as the verdict.
Metric: AUROC of hidden coexpression edges against degree-matched non-pairs. 90 runs, 18 settings, targets: (the strategy chooses its own).
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | k=10; knn=15; layer=coexpression | 0.392 [0.368, 0.414] | 100% [57, 100] | 0.707 | 0.52 | 5 |
| tuned | k=5; knn=15; layer=struct | 0.744 [0.671, 0.817] | 100% [34, 100] | 0.889 | 0.559 | 2 |
Sensitivity (mean skill at each value, all other settings pooled):
k-- 10: 0.55; 25: 0.55; 5: 0.55knn-- 15: 0.55; 5: 0.55layer-- coexpression: 0.39; cotranslation: 0.47; struct: 0.79
34 · Train on the networks and rank the edges they are missing (logistic / spectral embedding) -- reliable
The chosen layer's edges are hidden by orthogroup and the layer is removed from its own features. Metric: AUROC of the hidden edges against DEGREE-MATCHED non-pairs. Null: 10 configuration-model rewirings of the hidden edges, which preserve their degree sequence, so fame alone cannot clear the bar. Pass: above the null's 95th percentile by 0.05. The random-null AUROC, the fame gap, precision@k, the Brier score and the reliability gap are all reported as numbers beside the verdict.
Metric: AUROC of hidden coexpression edges against degree-matched non-pairs. 60 runs, 12 settings, targets: (the strategy chooses its own).
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | fraction=0.25; layer=coexpression; model=logistic | 0.392 [0.368, 0.416] | 100% [57, 100] | 0.707 | 0.52 | 5 |
| tuned | fraction=0.4; layer=struct; model=logistic | 0.729 [0.708, 0.751] | 100% [34, 100] | 0.891 | 0.595 | 2 |
Sensitivity (mean skill at each value, all other settings pooled):
fraction-- 0.25: 0.55; 0.4: 0.54layer-- coexpression: 0.38; cotranslation: 0.46; struct: 0.79model-- embedding: 0.54; logistic: 0.55
35 · Call genes with a stated error rate (split conformal prediction) -- reliable
25% of the label hidden by whole orthogroups; the model is trained and calibrated on the rest (itself split by orthogroup) and builds a set for every hidden gene. Metric: set efficiency, 1 - (mean set size - 1) / (classes - 1) -- how far the sets narrow the possibilities. Null: 10 runs on shuffled labels, where the model learns nothing and the sets must grow to keep the promise. Pass: above the null's 95th percentile by 0.05. Also reported: set coverage on the hidden genes beside the 1 - alpha promised, and the scorecard of the single-label calls.
Metric: set efficiency: 1 - (mean set size - 1) / (classes - 1). 120 runs, 6 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | alpha=0.1; model=logistic | 0.633 [0.444, 0.775] | 100% [84, 100] | 0.7 | 0.199 | 20 |
| tuned | alpha=0.2; model=logistic | 0.776 [0.555, 0.941] | 100% [68, 100] | 0.836 | 0.299 | 8 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | 0.993 [0.980, 1.000] | 100% | 5 |
| lopitpflocation | 0.846 [0.809, 0.883] | 100% | 5 |
| pbtransferredphenotype | 0.459 [0.435, 0.488] | 100% | 5 |
| stageenrichedderived | 0.849 [0.803, 0.885] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
alpha-- 0.05: 0.46; 0.1: 0.51; 0.2: 0.68model-- kNN: 0.43; logistic: 0.67
36 · Smooth the measurements along the networks, then classify (graph convolution + logistic regression) -- reliable
Pattern 1: 25% of the label hidden by whole orthogroups; the smoothed features are computed from measurements only, and the model is trained on the visible genes. Metric: hidden genes called correctly. Null: the analytic chance level for the same predicted and true class mixes. Pass: above that level's 95% bound by 0.05.
Metric: correct calls per hidden gene. 120 runs, 6 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | C=0.5; hops=2 | 0.458 [0.367, 0.544] | 100% [84, 100] | 0.709 | 0.448 | 20 |
| tuned | C=0.1; hops=2 | 0.460 [0.375, 0.539] | 100% [68, 100] | 0.701 | 0.456 | 8 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | 0.443 [0.363, 0.522] | 100% | 5 |
| lopitpflocation | 0.519 [0.499, 0.537] | 100% | 5 |
| pbtransferredphenotype | 0.347 [0.329, 0.365] | 100% | 5 |
| stageenrichedderived | 0.461 [0.403, 0.518] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
C-- 0.1: 0.44; 0.5: 0.46hops-- 1: 0.45; 2: 0.45; 3: 0.45
37 · Let a random forest find what defines a label (random forest + permutation importance) -- reliable
Pattern 1: 25% of the label hidden by whole orthogroups; the forest is trained on the visible genes and calls the hidden ones. Metric: hidden genes called correctly. Null: the analytic chance level for the same predicted and true class mixes. Pass: above that level's 95% bound by 0.05.
Metric: correct calls per hidden gene. 120 runs, 6 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | min_leaf=2; trees=300 | 0.361 [0.186, 0.532] | 75% [53, 89] | 0.741 | 0.513 | 20 |
| tuned | min_leaf=5; trees=300 | 0.410 [0.265, 0.538] | 75% [41, 93] | 0.749 | 0.509 | 8 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | 0.269 [0.210, 0.338] | 0% | 5 |
| lopitpflocation | 0.575 [0.560, 0.590] | 100% | 5 |
| pbtransferredphenotype | 0.426 [0.410, 0.442] | 100% | 5 |
| stageenrichedderived | 0.411 [0.323, 0.502] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
min_leaf-- 1: 0.33; 2: 0.36; 5: 0.42trees-- 100: 0.37; 300: 0.37
38 · Learn how much to trust each kind of evidence (stacked logistic regression) -- reliable
Pattern 1: 25% of the label hidden by whole orthogroups; base predictions for the meta-model are out of fold within the visible genes only, so no hidden label reaches either level. Metric: hidden genes called correctly. Null: the analytic chance level for the same predicted and true class mixes. Pass: above that level's 95% bound by 0.05.
Metric: correct calls per hidden gene. 60 runs, 3 settings, targets: isexported, lopitpflocation, pbtransferredphenotype, stageenriched_derived.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | k=15 | 0.451 [0.384, 0.522] | 95% [76, 99] | 0.7 | 0.431 | 20 |
| tuned | k=15 | 0.460 [0.394, 0.522] | 100% [68, 100] | 0.704 | 0.439 | 8 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| is_exported | 0.394 [0.350, 0.438] | 80% | 5 |
| lopitpflocation | 0.551 [0.533, 0.566] | 100% | 5 |
| pbtransferredphenotype | 0.381 [0.364, 0.406] | 100% | 5 |
| stageenrichedderived | 0.479 [0.435, 0.524] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
k-- 10: 0.45; 15: 0.45; 30: 0.45
39 · Predict a value with an interval that holds (gradient boosting / ridge + split conformal) -- reliable
Pattern 4: 20% of the measured values hidden; the model is trained and calibrated on the rest (split by orthogroup) and predicts the hidden ones. Metric: rank correlation of predicted and hidden values. Null: the chance distribution of a rank correlation, with refits on shuffled values reported. Pass: above its 95th percentile by 0.1. Also reported: the share of hidden values inside their intervals, beside the coverage promised.
Metric: rank correlation of predicted and hidden values. 60 runs, 4 settings, targets: exprschizont, meanplddt, piggybac_mis.
| setting | skill [95% CI] | pass [95% CI] | observed | chance | runs | |
|---|---|---|---|---|---|---|
| at defaults | alpha=0.1; model=boosted | 0.500 [0.218, 0.837] | 100% [80, 100] | 0.5 | 0.0 | 15 |
| tuned | alpha=0.1; model=boosted | 0.488 [0.185, 0.826] | 100% [61, 100] | 0.488 | 0.0 | 6 |
Tuned setting, per held-out target:
| target | skill [95% CI] | pass | runs |
|---|---|---|---|
| expr_schizont | 0.206 [0.155, 0.268] | 100% | 5 |
| mean_plddt | 0.842 [0.826, 0.854] | 100% | 5 |
| piggybac_mis | 0.451 [0.440, 0.462] | 100% | 5 |
Sensitivity (mean skill at each value, all other settings pooled):
alpha-- 0.1: 0.48; 0.2: 0.48model-- boosted: 0.50; ridge: 0.45