Strategy scorecards
Generated by scripts/calibrate_strategies.py --publish on 2026-09-26 from 7,640 self-test runs.
A self-test hides information that is already known, asks the strategy for it back, and judges the answer against the same procedure on shuffled data. The verdict rests on one metric. The scorecard reports, on the same hidden genes, every standard metric for the kind of task the strategy performs -- so strategies that do the same thing can be compared number by number, and a strategy's strengths and blind spots are visible: precise but narrow, good at the top of a list but not throughout, right about common classes and blind to rare ones.
Every value below is a mean with a 95% interval: a two-stage bootstrap over held-out labels, then over runs (seeds) within a label, at the strategy's default settings unless the table says tuned. From Python: strategies.metrics(), strategies.calibration(key)['default']['scorecard'], and TestResult.card() for a test you run yourself.
The verdict block (every test)
| Field | Meaning |
|---|---|
| Verdict | PASS: above the 95th percentile of the null AND by the stated margin. FAIL otherwise. INCONCLUSIVE when too little could be hidden to score. |
| Judged on | The one metric the verdict rests on, chosen per strategy as the honest test of its claim. |
| Observed | That metric on the hidden genes. |
| Chance | The same metric for the same procedure on shuffled labels, random sets or permuted identities -- measured, not assumed. |
| Bar | The 95th percentile of the null runs: what luck reaches one time in twenty. |
| p | Share of null runs at least as good as the observed, (1 + k) / (1 + runs). |
| Skill | (observed - chance) / (1 - chance): 0 is chance, 1 is perfect, negative is worse than chance. Puts every metric on one scale. |
| Hidden | How many genes, pairs or findings the test scored. |
Label calls
Call a label for genes that lack it. What is hidden: a share of a label's genes, whole orthogroups at a time. Strategies: 06, 07, 08, 09, 11, 12, 13, 14, 19, 30, 31, 32, 35, 36, 37, 38.
| Metric | Definition | Range | Chance | How to read it |
|---|---|---|---|---|
| Accuracy | Hidden genes called with their true label, divided by ALL hidden genes. A gene the strategy declined to call counts as wrong. | 0 to 1 | about the sum of squared class shares (Cohen's chance term) | The headline for label calls. Counting abstentions as errors stops a strategy looking accurate by calling only the easy genes; coverage and precision of calls split it apart. |
| Coverage | Share of hidden genes that received any call. | 0 to 1 | not applicable (a property of the strategy, not of luck) | How far the strategy reaches. A network strategy cannot call a gene with no edges; a low coverage with a high precision of calls is a precise but narrow tool. |
| Precision of calls | Correct calls divided by calls made (also called selective accuracy). | 0 to 1 | as accuracy, among the genes called | How far to trust one call. Equals accuracy when coverage is 1. |
| Macro precision | For each true class: of the genes called that class, the share that truly are; averaged over classes with equal weight. A class never called scores 0. | 0 to 1 | about the average class share | Whether calls of the RARE classes can be trusted too; a strategy that only ever calls the commonest class scores low. |
| Macro recall (balanced accuracy) | For each true class: the share of its hidden genes called correctly; averaged over classes with equal weight. Identical to balanced accuracy. | 0 to 1 | 1 / number of classes, for a caller that ignores the data | Whether every class is found, not just the large ones. Compare with accuracy: a large gap means the strategy lives on the big classes. |
| Macro F1 | Per class, the harmonic mean of precision and recall; averaged over classes with equal weight. | 0 to 1 | low; roughly the average class share | One number balancing finding each class and being right when calling it, with rare classes counted as much as common ones. |
| Weighted F1 | Per-class F1 averaged with each class weighted by its number of hidden genes. | 0 to 1 | about the sum of squared class shares | Macro F1's counterpart that follows the class sizes; close to accuracy when coverage is high. |
| Cohen's kappa | Agreement between calls and truth corrected for the agreement their class frequencies alone would give: (observed - expected) / (1 - expected). No call is its own category. | -1 to 1 | 0 | Accuracy with chance removed: 0 is no better than matching class frequencies, 1 is perfect. Comparable across labels with different numbers and sizes of classes. |
| Matthews correlation (MCC) | The multiclass Matthews correlation coefficient (Gorodkin's R_K) between calls and truth, with no call as its own category. | -1 to 1 | 0 | A correlation between the call and the truth that stays honest under strong class imbalance; often the single most informative number for an unbalanced label. |
| Macro AUROC | For each class, the probability that a random hidden gene of that class gets a higher score for the class than a random hidden gene of another class; averaged over classes. Needs per-class scores. | 0 to 1 | 0.5 | How well the strategy's scores separate each class from the rest, before any threshold is chosen. Missing for strategies that call labels without scoring every class. |
| Macro AUPRC | For each class, the area under the precision-recall curve of its one-vs-rest scores (average precision); averaged over classes. Needs per-class scores. | 0 to 1 | the class's share among hidden genes, averaged | Like macro AUROC but dominated by the top of each ranking, which is where calls are made; far below AUROC means good separation overall but a noisy top. |
Ranking
Rank candidates so the true ones come first. What is hidden: set members, edges or corrupted labels, ranked among negatives. Strategies: 03, 10, 16, 17, 18, 20, 24, 25, 28, 33, 34.
| Metric | Definition | Range | Chance | How to read it |
|---|---|---|---|---|
| AUROC | Probability that a random hidden positive is ranked above a random negative (area under the ROC curve); ties count half. | 0 to 1 | 0.5 | Separation over the whole ranking. Insensitive to how rare positives are, so a high AUROC can coexist with a poor top of the list -- read it with AUPRC. |
| AUPRC | Average precision: the mean, over the hidden positives, of the precision of the ranking down to that positive (area under the precision-recall curve). | 0 to 1 | the prevalence (share of positives among everything ranked) | How clean the top of the ranking is. Its chance level is the prevalence, so compare it with that (AUPRC lift), never with 0.5. |
| AUPRC lift | AUPRC divided by the prevalence, its value for a random ordering. | 0 to 1/prevalence | 1 | How many times better than a random ordering the ranking is, where it matters. Comparable across tests with different prevalences, unlike AUPRC itself. |
| Prevalence | Hidden positives divided by everything ranked. | 0 to 1 | not applicable (a property of the test) | The chance level of AUPRC and of precision at any depth; context for every other number. |
| Partial AUROC (FPR <= 10%) | AUROC restricted to false-positive rates up to 10%, rescaled (McClish) so 0.5 is chance and 1 perfect. | 0.5 to 1 (after rescaling) | 0.5 | Separation in the part of the ranking anyone would act on; what matters for screening. |
| R-precision | Precision among the top R items, where R is the number of hidden positives. Equal to recall at that depth. | 0 to 1 | the prevalence | If you took as many candidates as there are true positives, the share that would be right. |
| Precision @ top 1% | Share of positives among the top 1% of the ranking (at least one item). | 0 to 1 | the prevalence | How good the very first candidates are -- the list anyone would test first. |
| Enrichment @ top 1% | Precision at the top 1% divided by the prevalence. | 0 to 1/prevalence | 1 | How many times more positives the first 1% holds than a random 1% would. |
| Recall @ top 10% | Share of all hidden positives found in the top 10% of the ranking. | 0 to 1 | 0.1 | How much of the answer a short list contains. |
| Best F1 | The highest F1 (harmonic mean of precision and recall) over every cut-off of the ranking. | 0 to 1 | about 2 x prevalence / (1 + prevalence) (taking everything) | The best single-threshold trade-off the ranking allows; optimistic, since the cut-off is chosen on the answer. |
| nDCG | Normalised discounted cumulative gain: positives count 1 / log2(rank + 1), divided by the value of a perfect ranking. | 0 to 1 | depends on prevalence; low when positives are rare | A whole-ranking score that weights the top most, smoothly rather than at one cut-off. |
Set retrieval
Return a set of genes that belong with a query. What is hidden: part of a gene set, to be returned among all other genes. Strategies: 02.
| Metric | Definition | Range | Chance | How to read it |
|---|---|---|---|---|
| Precision | Share of the returned genes that are hidden members. | 0 to 1 | the members' share of the candidate genes | How many of the candidates are real. |
| Recall | Share of the hidden members that were returned. | 0 to 1 | the returned share of the candidate genes | How much of the set was found. |
| F1 | Harmonic mean of precision and recall. | 0 to 1 | low; about the members' share for a random return of the same size | One number that is only high when both are. |
| Jaccard index | Overlap of returned and hidden sets divided by their union. | 0 to 1 | near 0 | How close the returned set is to being exactly the hidden set; stricter than F1. |
| Matthews correlation (MCC) | Correlation between 'returned' and 'is a hidden member' over all candidate genes. | -1 to 1 | 0 | Accounts for the genes correctly NOT returned as well; honest when the set is small. |
| Fold enrichment | Precision divided by the members' share of the candidate genes. | 0 to 1/share | 1 | How many times more members the returned set holds than a random set of its size. |
| Genes returned | How many genes the strategy returned as the set. | 0 to all | not applicable | Context: a high recall from returning half the genome is not a finding. |
Cluster recovery
Find clusters that correspond to a label nobody showed them. What is hidden: a share of a label, scored against clusters chosen on the rest. Strategies: 01, 04, 15.
| Metric | Definition | Range | Chance | How to read it |
|---|---|---|---|---|
| Weighted F1 | For each label, F1 of its hidden genes against the ONE cluster chosen for it on visible genes; averaged with labels weighted by their hidden genes. A label with no cluster scores 0. | 0 to 1 | set by permuting hidden labels over the same clusters (the verdict's null) | Whether a label that nobody showed the clustering falls out as a cluster, scored on genes the choice never saw. |
| Weighted precision | For each label, the share of hidden genes in its chosen cluster that carry it; label-size weighted. | 0 to 1 | about the label's share | How pure the chosen clusters are. |
| Weighted recall | For each label, the share of its hidden genes that landed in its chosen cluster; label-size weighted. | 0 to 1 | about the cluster's share of the genes | How completely each label is gathered into one cluster. |
| Adjusted Rand index (ARI) | Agreement between the clustering and the hidden labels over all pairs of hidden genes, corrected so random partitions of the same sizes score 0. | -0.5 to 1 | 0 | Whole-partition agreement, independent of which cluster was chosen for which label. |
| Normalised mutual information (NMI) | Mutual information between clusters and hidden labels, divided by the mean of their entropies. | 0 to 1 | small but above 0 for many small clusters | How much knowing a gene's cluster tells you about its label. |
| Homogeneity | 1 minus the uncertainty about the label left once the cluster is known, relative to the label's own entropy. | 0 to 1 | near 0 | Whether each cluster holds one label. |
| Completeness | 1 minus the uncertainty about the cluster left once the label is known, relative to the clustering's entropy. | 0 to 1 | near 0 | Whether each label sits in one cluster. Its harmonic mean with homogeneity is the V-measure, which equals NMI. |
| Unclustered share | Share of hidden genes HDBSCAN left as noise (in no cluster). | 0 to 1 | not applicable | Context: noise genes cannot be recovered by any cluster. |
Values
Predict a measured value for genes without one. What is hidden: a share of a measurement's values. Strategies: 21, 22, 23, 29, 39.
| Metric | Definition | Range | Chance | How to read it |
|---|---|---|---|---|
| Spearman rho | Rank correlation between predicted and hidden measured values. | -1 to 1 | 0 (standard deviation 1/sqrt(n - 1)) | Whether the ORDER of the genes is predicted, whatever the scale. The headline for values, robust to outliers and to a prediction on a different scale. |
| Pearson r | Linear correlation between predicted and hidden values. | -1 to 1 | 0 | Like Spearman but on the values themselves, so a few extreme genes can dominate it. |
| Kendall tau-b | Share of gene pairs ordered the same way by prediction and truth, minus the share ordered oppositely, corrected for ties. | -1 to 1 | 0 | The most directly interpretable rank agreement: (1 + tau) / 2 is the chance a random pair is ordered correctly. |
| R-squared (out of sample) | 1 minus the squared prediction error over the variance of the hidden values around their own mean. | below 0 to 1 | 0 or below (predicting the mean scores 0) | The share of the hidden values' variance the prediction explains. Negative when the prediction is further off than the mean -- common for a good ranking on the wrong scale. |
| Normalised RMSE | Root-mean-square error divided by the standard deviation of the hidden values. | 0 upward | 1 (predicting the mean) | Typical error in units of the measurement's own spread: below 1 beats the mean. |
| Mean absolute error | Mean absolute difference between predicted and hidden values, in the measurement's units. | 0 upward | the mean absolute deviation of the hidden values | Typical size of an error, in the same units as the data. |
| Top-decile recall | Share of the genes in the true top 10% of hidden values that are also in the predicted top 10%. | 0 to 1 | 0.1 | Whether the extreme genes -- usually the interesting ones -- are predicted extreme. |
| Bottom-decile recall | The same for the bottom 10% -- for fitness scores, the most essential genes. | 0 to 1 | 0.1 | Whether the genes at the other extreme are found. |
| Coverage | Share of hidden values that received a prediction. | 0 to 1 | not applicable | Every other value metric is computed on these genes. |
Replication
Make findings that hold beyond the genes they were made on. What is hidden: half of the genes, on which first-half findings are checked. Strategies: 05, 26, 27.
| Metric | Definition | Range | Chance | How to read it |
|---|---|---|---|---|
| Replication rate | Share of the findings made on one half of the genes that hold on the other half. | 0 to 1 | the rate with the second half's evidence scrambled | Whether the findings are properties of the genes or of the sample they were found in. |
| Findings made | Number of findings made on the first half. | 0 upward | not applicable | Context: how much there was to replicate. |
| Findings replicated | Number of those findings that held on the second half. | 0 to findings | findings x the null rate | The findings worth reading first; each is a claim that held twice. |
| Replication by chance | The replication rate with the second half's evidence scrambled, averaged over runs. | 0 to 1 | is itself the chance level | What replication looks like for findings with nothing behind them. |
| Replication lift | Replication rate divided by the replication rate by chance. | 0 upward | 1 | How many times more often real findings replicate than chance ones. |
Toxoplasma gondii, at the default settings
label calls -- Call a label for genes that lack it. Hidden: a share of a label's genes, whole orthogroups at a time.
| # | Strategy | Grade | Accuracy | Coverage | Precision of calls | Macro precision | Macro recall (balanced accuracy) | Macro F1 | Weighted F1 | Cohen's kappa | Matthews correlation (MCC) | Macro AUROC | Macro AUPRC |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 06 | Find which kind of evidence carries a label (kNN ablation) | reliable | 0.60 [0.44, 0.77] | 1.00 [1.00, 1.00] | 0.60 [0.44, 0.77] | 0.48 [0.34, 0.67] | 0.41 [0.30, 0.52] | 0.39 [0.28, 0.54] | 0.54 [0.39, 0.70] | 0.23 [0.09, 0.38] | 0.26 [0.12, 0.42] | 0.71 [0.59, 0.82] | 0.46 [0.32, 0.64] |
| 07 | Call a gene by the genes that behave like it (kNN) | reliable | 0.62 [0.47, 0.76] | 0.93 [0.84, 1.00] | 0.65 [0.54, 0.76] | 0.60 [0.49, 0.74] | 0.44 [0.34, 0.54] | 0.44 [0.34, 0.57] | 0.59 [0.46, 0.72] | 0.27 [0.13, 0.42] | 0.31 [0.16, 0.46] | 0.77 [0.64, 0.87] | 0.54 [0.44, 0.68] |
| 08 | Call a gene by its neighbours on the map (UMAP + kNN) | weak | 0.56 [0.38, 0.74] | 0.91 [0.81, 1.00] | 0.60 [0.46, 0.74] | 0.42 [0.31, 0.53] | 0.38 [0.27, 0.48] | 0.37 [0.26, 0.48] | 0.53 [0.37, 0.69] | 0.19 [0.05, 0.33] | 0.20 [0.06, 0.34] | 0.70 [0.57, 0.79] | 0.44 [0.33, 0.54] |
| 09 | Name a cluster by the label it is enriched for (UMAP + HDBSCAN, hypergeometric) | reliable | 0.06 [0.03, 0.08] | 0.18 [0.09, 0.27] | 0.33 [0.25, 0.42] | 0.20 [0.08, 0.32] | 0.10 [0.07, 0.14] | 0.10 [0.06, 0.15] | 0.07 [0.03, 0.11] | 0.04 [0.02, 0.06] | 0.08 [0.06, 0.10] | 0.76 [0.70, 0.82] | 0.25 [0.14, 0.37] |
| 11 | Diffuse a label across one measured network (random walk with restart) | reliable | 0.29 [0.18, 0.40] | 0.80 [0.78, 0.82] | 0.37 [0.22, 0.51] | 0.32 [0.19, 0.46] | 0.32 [0.24, 0.40] | 0.28 [0.17, 0.41] | 0.33 [0.19, 0.47] | 0.10 [0.06, 0.15] | 0.11 [0.07, 0.16] | 0.69 [0.61, 0.77] | 0.33 [0.19, 0.49] |
| 12 | Let every network vote, weighted by what it has earned (chance-weighted ensemble vote) | weak | 0.58 [0.40, 0.75] | 0.90 [0.69, 1.00] | 0.64 [0.48, 0.78] | 0.58 [0.43, 0.74] | 0.42 [0.32, 0.53] | 0.44 [0.33, 0.57] | 0.56 [0.39, 0.72] | 0.33 [0.19, 0.44] | 0.36 [0.21, 0.48] | 0.71 [0.63, 0.76] | 0.48 [0.41, 0.56] |
| 13 | Place a protein by the proteins it physically touches (weighted partner vote) | reliable | 0.39 [0.17, 0.61] | 0.57 [0.28, 0.84] | 0.65 [0.56, 0.75] | 0.47 [0.34, 0.59] | 0.32 [0.13, 0.51] | 0.34 [0.17, 0.52] | 0.43 [0.22, 0.65] | 0.24 [0.06, 0.41] | 0.25 [0.07, 0.43] | 0.72 [0.61, 0.81] | 0.42 [0.36, 0.49] |
| 14 | Annotate function through shared fold (TM-score-weighted vote) | reliable | 0.73 [0.70, 0.76] | 0.78 [0.75, 0.80] | 0.94 [0.94, 0.95] | 0.90 [0.87, 0.93] | 0.67 [0.65, 0.68] | 0.75 [0.75, 0.76] | 0.82 [0.80, 0.83] | 0.66 [0.62, 0.68] | 0.68 [0.66, 0.71] | 0.96 [0.95, 0.97] | 0.70 [0.68, 0.72] |
| 19 | Train a classifier on the known genes and call the rest (logistic regression) | reliable | 0.54 [0.41, 0.68] | 1.00 [1.00, 1.00] | 0.54 [0.41, 0.68] | 0.47 [0.36, 0.59] | 0.57 [0.48, 0.69] | 0.48 [0.36, 0.61] | 0.56 [0.43, 0.70] | 0.32 [0.18, 0.46] | 0.33 [0.19, 0.48] | 0.80 [0.66, 0.90] | 0.55 [0.45, 0.70] |
| 30 | Test inference on the genes orthology cannot reach (kNN) | reliable | 0.59 [0.39, 0.78] | 0.89 [0.76, 1.00] | 0.64 [0.50, 0.78] | 0.59 [0.43, 0.76] | 0.35 [0.24, 0.49] | 0.38 [0.27, 0.54] | 0.56 [0.38, 0.72] | 0.30 [0.22, 0.40] | 0.35 [0.26, 0.46] | 0.80 [0.76, 0.86] | 0.49 [0.36, 0.68] |
| 31 | Call a gene only when independent strategies agree (kNN + logistic + network vote) | reliable | 0.63 [0.48, 0.76] | 0.87 [0.76, 0.96] | 0.71 [0.61, 0.80] | 0.66 [0.55, 0.78] | 0.47 [0.40, 0.55] | 0.50 [0.43, 0.61] | 0.62 [0.50, 0.75] | 0.33 [0.17, 0.46] | 0.35 [0.18, 0.48] | 0.83 [0.68, 0.92] | 0.63 [0.54, 0.74] |
| 32 | Put the understudied genes first (kNN + logistic + network vote) | reliable | 0.62 [0.48, 0.77] | 0.87 [0.76, 0.97] | 0.70 [0.60, 0.80] | 0.63 [0.53, 0.76] | 0.46 [0.38, 0.54] | 0.48 [0.41, 0.59] | 0.62 [0.49, 0.75] | 0.31 [0.16, 0.43] | 0.33 [0.17, 0.45] | 0.82 [0.67, 0.91] | 0.62 [0.53, 0.73] |
| 35 | Call genes with a stated error rate (split conformal prediction) | reliable | 0.17 [0.05, 0.38] | 0.21 [0.06, 0.44] | 0.81 [0.65, 0.92] | 0.57 [0.37, 0.75] | 0.19 [0.05, 0.42] | 0.25 [0.09, 0.48] | 0.23 [0.08, 0.47] | 0.10 [0.01, 0.23] | 0.16 [0.05, 0.32] | 0.81 [0.65, 0.92] | 0.59 [0.49, 0.73] |
| 36 | Smooth the measurements along the networks, then classify (graph convolution + logistic regression) | reliable | 0.63 [0.53, 0.74] | 1.00 [1.00, 1.00] | 0.63 [0.53, 0.74] | 0.54 [0.46, 0.64] | 0.61 [0.52, 0.72] | 0.56 [0.48, 0.67] | 0.64 [0.54, 0.76] | 0.40 [0.16, 0.57] | 0.40 [0.16, 0.58] | 0.82 [0.65, 0.92] | 0.61 [0.51, 0.74] |
| 37 | Let a random forest find what defines a label (random forest + permutation importance) | reliable | 0.70 [0.57, 0.83] | 1.00 [1.00, 1.00] | 0.70 [0.57, 0.83] | 0.67 [0.56, 0.80] | 0.55 [0.47, 0.64] | 0.56 [0.46, 0.68] | 0.67 [0.53, 0.81] | 0.45 [0.21, 0.65] | 0.47 [0.23, 0.66] | 0.85 [0.68, 0.94] | 0.65 [0.55, 0.77] |
| 38 | Learn how much to trust each kind of evidence (stacked logistic regression) | reliable | 0.62 [0.51, 0.74] | 1.00 [1.00, 1.00] | 0.62 [0.51, 0.74] | 0.54 [0.47, 0.64] | 0.62 [0.53, 0.73] | 0.56 [0.47, 0.67] | 0.63 [0.52, 0.76] | 0.41 [0.18, 0.57] | 0.41 [0.18, 0.58] | 0.83 [0.66, 0.93] | 0.63 [0.53, 0.76] |
ranking -- Rank candidates so the true ones come first. Hidden: set members, edges or corrupted labels, ranked among negatives.
| # | Strategy | Grade | AUROC | AUPRC | AUPRC lift | Prevalence | Partial AUROC (FPR <= 10%) | R-precision | Precision @ top 1% | Enrichment @ top 1% | Recall @ top 10% | Best F1 | nDCG |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 03 | Ask which categories the data can rediscover (UMAP + neighbour AUROC) | reliable | 0.75 [0.61, 0.84] | 0.59 [0.44, 0.74] | 9.56 [1.61, 18.93] | 0.37 [0.14, 0.61] | 0.65 [0.57, 0.73] | 0.59 [0.46, 0.72] | 0.63 [0.42, 0.82] | 10.2 [1.8, 18.9] | 0.35 [0.17, 0.54] | 0.67 [0.54, 0.79] | 0.83 [0.73, 0.91] |
| 10 | Find genes whose label their neighbours contradict (kNN + network neighbours) | reliable | 0.82 [0.74, 0.88] | 0.25 [0.17, 0.33] | 5.04 [3.51, 6.68] | 0.05 [0.05, 0.05] | 0.66 [0.59, 0.72] | 0.30 [0.20, 0.39] | 0.38 [0.25, 0.51] | 7.77 [5.19, 10.36] | 0.48 [0.35, 0.61] | 0.35 [0.27, 0.44] | 0.71 [0.61, 0.79] |
| 16 | Predict the contacts an interactome missed (logistic regression) | reliable | 0.80 [0.80, 0.81] | 0.61 [0.59, 0.62] | 3.62 [3.54, 3.69] | 0.17 [0.17, 0.17] | 0.72 [0.71, 0.73] | 0.57 [0.56, 0.59] | 0.95 [0.92, 0.98] | 5.64 [5.47, 5.82] | 0.44 [0.43, 0.45] | 0.59 [0.58, 0.60] | 0.92 [0.92, 0.92] |
| 17 | Read the literature for biology, not fame (publication-count residual) | reliable | 0.62 [0.60, 0.64] | 0.64 [0.58, 0.68] | 1.23 [1.14, 1.31] | 0.52 [0.46, 0.57] | 0.54 [0.52, 0.56] | 0.59 [0.54, 0.65] | 0.76 [0.61, 0.92] | 1.50 [1.09, 1.87] | 0.14 [0.12, 0.16] | 0.70 [0.66, 0.74] | 0.93 [0.91, 0.94] |
| 18 | List what the data says and the literature has not written (multi-layer support count) | reliable | 0.56 [0.56, 0.56] | 0.24 [0.24, 0.24] | 1.46 [1.46, 1.47] | 0.17 [0.17, 0.17] | 0.55 [0.55, 0.55] | 0.97 [0.97, 0.97] | 1.00 [0.99, 1.00] | 5.98 [5.97, 5.99] | 0.57 [0.57, 0.57] | 0.99 [0.99, 0.99] | 0.99 [0.99, 0.99] |
| 20 | Learn what makes your list special, from positives alone (PU bagging, logistic regression) | reliable | 0.89 [0.86, 0.92] | 0.10 [0.06, 0.14] | 15.5 [9.8, 21.6] | 0.01 [0.01, 0.01] | 0.71 [0.66, 0.77] | 0.14 [0.08, 0.21] | 0.12 [0.07, 0.16] | 17.8 [11.2, 24.9] | 0.66 [0.55, 0.77] | 0.17 [0.11, 0.24] | 0.55 [0.48, 0.60] |
| 24 | Describe what your gene list has in common (hypergeometric + rank-sum) | reliable | 0.81 [0.75, 0.86] | 0.07 [0.04, 0.10] | 8.10 [3.93, 12.65] | 0.01 [0.01, 0.01] | 0.63 [0.58, 0.69] | 0.11 [0.05, 0.16] | 0.10 [0.05, 0.15] | 11.5 [5.1, 18.2] | 0.47 [0.32, 0.61] | 0.14 [0.09, 0.19] | 0.51 [0.45, 0.57] |
| 25 | Grow your gene list along the networks (random walk with restart) | reliable | 0.80 [0.74, 0.84] | 0.05 [0.03, 0.07] | 7.73 [4.42, 11.50] | 0.01 [0.01, 0.01] | 0.64 [0.60, 0.68] | 0.09 [0.05, 0.12] | 0.08 [0.05, 0.12] | 12.6 [7.3, 19.8] | 0.47 [0.38, 0.56] | 0.13 [0.08, 0.18] | 0.46 [0.41, 0.51] |
| 28 | Find paralogs that changed jobs (profile correlation) | reliable | 0.58 [0.53, 0.66] | 0.38 [0.26, 0.58] | 1.20 [1.04, 1.38] | 0.30 [0.23, 0.41] | 0.53 [0.50, 0.57] | 0.35 [0.24, 0.52] | 0.54 [0.34, 0.71] | 1.80 [1.26, 2.35] | 0.13 [0.10, 0.17] | 0.48 [0.40, 0.60] | 0.81 [0.74, 0.88] |
| 33 | Put every layer into one space and read a gene's neighbourhood (logistic edge model) | reliable | 0.79 [0.79, 0.79] | 0.80 [0.80, 0.81] | 1.60 [1.59, 1.61] | 0.50 [0.50, 0.50] | 0.65 [0.64, 0.66] | 0.72 [0.71, 0.72] | 1.00 [1.00, 1.00] | 2.00 [2.00, 2.00] | 0.19 [0.18, 0.19] | 0.73 [0.73, 0.74] | 0.97 [0.97, 0.97] |
| 34 | Train on the networks and rank the edges they are missing (logistic / spectral embedding) | reliable | 0.79 [0.79, 0.79] | 0.80 [0.80, 0.81] | 1.60 [1.59, 1.61] | 0.50 [0.50, 0.50] | 0.65 [0.64, 0.66] | 0.72 [0.71, 0.72] | 1.00 [1.00, 1.00] | 2.00 [2.00, 2.00] | 0.19 [0.18, 0.19] | 0.73 [0.73, 0.74] | 0.97 [0.97, 0.97] |
set retrieval -- Return a set of genes that belong with a query. Hidden: part of a gene set, to be returned among all other genes.
| # | Strategy | Grade | Precision | Recall | F1 | Jaccard index | Matthews correlation (MCC) | Fold enrichment | Genes returned |
|---|---|---|---|---|---|---|---|---|---|
| 02 | Find the map where your gene list is one cluster (UMAP + HDBSCAN) | weak | 0.04 [0.03, 0.05] | 0.22 [0.12, 0.34] | 0.06 [0.04, 0.07] | 0.03 [0.02, 0.04] | 0.04 [0.01, 0.06] | 2.45 [1.49, 3.36] | 433.0 [165.8, 802.2] |
cluster recovery -- Find clusters that correspond to a label nobody showed them. Hidden: a share of a label, scored against clusters chosen on the rest.
| # | Strategy | Grade | Weighted F1 | Weighted precision | Weighted recall | Adjusted Rand index (ARI) | Normalised mutual information (NMI) | Homogeneity | Completeness | Unclustered share |
|---|---|---|---|---|---|---|---|---|---|---|
| 01 | Hold out a category and search for a map that finds it (UMAP + HDBSCAN) | weak | 0.38 [0.19, 0.57] | 0.51 [0.35, 0.66] | 0.40 [0.14, 0.71] | 0.03 [0.01, 0.06] | 0.27 [0.12, 0.43] | 0.53 [0.24, 0.81] | 0.20 [0.11, 0.31] | 0.45 [0.19, 0.70] |
| 04 | Keep only the modules that survive the whole walk (UMAP + HDBSCAN co-clustering) | weak | 0.39 [0.27, 0.49] | 0.46 [0.27, 0.65] | 0.46 [0.40, 0.54] | 0.04 [0.02, 0.07] | 0.17 [0.10, 0.24] | 0.21 [0.14, 0.29] | 0.18 [0.08, 0.29] | 0.11 [0.07, 0.16] |
| 15 | Find the communities several networks agree on (modularity + Louvain consensus) | weak | 0.35 [0.24, 0.46] | 0.45 [0.26, 0.63] | 0.37 [0.30, 0.47] | 0.02 [0.01, 0.03] | 0.05 [0.03, 0.06] | 0.05 [0.04, 0.06] | 0.05 [0.02, 0.07] | 0.00 [0.00, 0.00] |
values -- Predict a measured value for genes without one. Hidden: a share of a measurement's values.
| # | Strategy | Grade | Spearman rho | Pearson r | Kendall tau-b | R-squared (out of sample) | Normalised RMSE | Mean absolute error | Top-decile recall | Bottom-decile recall | Coverage |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 21 | Predict a measurement, and find the genes that defy the prediction (gradient boosting / ridge) | reliable | 0.56 [0.05, 0.91] | 0.57 [0.06, 0.93] | 0.44 [0.04, 0.76] | 0.43 [-0.09, 0.87] | 0.70 [0.36, 1.04] | 5.21 [0.19, 13.97] | 0.45 [0.14, 0.76] | 0.40 [0.15, 0.69] | 1.00 [1.00, 1.00] |
| 22 | Fill in what was never measured, and say where that is honest (soft-impute, low-rank SVD) | reliable | 0.87 [0.86, 0.87] | 0.87 [0.87, 0.87] | 0.69 [0.69, 0.70] | 0.75 [0.75, 0.76] | 0.50 [0.49, 0.50] | -- | 0.68 [0.67, 0.68] | 0.56 [0.56, 0.57] | 1.00 [1.00, 1.00] |
| 23 | Find what matters more in one condition, and why (residual + gradient boosting / ridge) | weak | 0.07 [0.06, 0.09] | 0.07 [0.04, 0.09] | 0.05 [0.04, 0.06] | -0.04 [-0.06, -0.03] | 1.02 [1.02, 1.03] | 0.25 [0.25, 0.25] | 0.11 [0.10, 0.12] | 0.18 [0.16, 0.20] | 1.00 [1.00, 1.00] |
| 29 | Carry what one parasite shows to the other (orthogroup mapping) | reliable | 0.32 [0.30, 0.33] | 0.32 [0.30, 0.33] | 0.21 [0.20, 0.23] | -1.60 [-1.70, -1.50] | 1.61 [1.58, 1.64] | 3.00 [2.95, 3.05] | 0.14 [0.12, 0.15] | 0.13 [0.11, 0.15] | 1.00 [1.00, 1.00] |
| 39 | Predict a value with an interval that holds (gradient boosting / ridge + split conformal) | weak | 0.55 [0.04, 0.91] | 0.56 [0.04, 0.93] | 0.43 [0.02, 0.75] | 0.41 [-0.13, 0.87] | 0.71 [0.36, 1.06] | 5.42 [0.19, 14.65] | 0.45 [0.14, 0.76] | 0.39 [0.15, 0.68] | 1.00 [1.00, 1.00] |
replication -- Make findings that hold beyond the genes they were made on. Hidden: half of the genes, on which first-half findings are checked.
| # | Strategy | Grade | Replication rate | Findings made | Findings replicated | Replication by chance | Replication lift |
|---|---|---|---|---|---|---|---|
| 05 | Tune a map without labels, then read what it encodes (UMAP + HDBSCAN, chi-square / Kruskal-Wallis) | reliable | 0.95 [0.89, 1.00] | 53.2 [38.8, 63.0] | 51.4 [36.4, 62.0] | 0.06 [0.04, 0.08] | 20.5 [12.7, 28.3] |
| 26 | Find categories that split in two on another measurement (UMAP + HDBSCAN) | untestable | -- | -- | -- | -- | -- |
| 27 | Find kinds of gene defined by two labels at once (UMAP + HDBSCAN) | reliable | 0.50 [0.39, 0.64] | 5.20 [3.60, 7.00] | 2.60 [1.80, 3.40] | 0.03 [0.02, 0.04] | 21.7 [15.0, 31.0] |
Toxoplasma gondii, at the tuned setting (chosen on seeds 1-3, reported on seeds 4-5)
label calls -- Call a label for genes that lack it. Hidden: a share of a label's genes, whole orthogroups at a time.
| # | Strategy | Grade | Accuracy | Coverage | Precision of calls | Macro precision | Macro recall (balanced accuracy) | Macro F1 | Weighted F1 | Cohen's kappa | Matthews correlation (MCC) | Macro AUROC | Macro AUPRC |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 06 | Find which kind of evidence carries a label (kNN ablation) | reliable | 0.59 [0.40, 0.77] | 1.00 [1.00, 1.00] | 0.59 [0.40, 0.77] | 0.42 [0.29, 0.58] | 0.40 [0.28, 0.52] | 0.38 [0.26, 0.52] | 0.54 [0.36, 0.71] | 0.21 [0.08, 0.36] | 0.23 [0.09, 0.38] | 0.66 [0.56, 0.77] | 0.40 [0.26, 0.58] |
| 07 | Call a gene by the genes that behave like it (kNN) | reliable | 0.63 [0.48, 0.77] | 1.00 [1.00, 1.00] | 0.63 [0.48, 0.77] | 0.54 [0.42, 0.69] | 0.47 [0.38, 0.57] | 0.48 [0.38, 0.60] | 0.60 [0.46, 0.74] | 0.31 [0.17, 0.45] | 0.33 [0.18, 0.48] | 0.74 [0.65, 0.81] | 0.51 [0.40, 0.64] |
| 08 | Call a gene by its neighbours on the map (UMAP + kNN) | weak | 0.54 [0.37, 0.70] | 0.97 [0.94, 1.00] | 0.55 [0.39, 0.70] | 0.39 [0.29, 0.49] | 0.36 [0.27, 0.46] | 0.37 [0.27, 0.47] | 0.52 [0.36, 0.68] | 0.16 [0.00, 0.28] | 0.16 [0.00, 0.29] | 0.63 [0.49, 0.73] | 0.38 [0.29, 0.48] |
| 09 | Name a cluster by the label it is enriched for (UMAP + HDBSCAN, hypergeometric) | reliable | 0.04 [0.01, 0.07] | 0.09 [0.03, 0.15] | 0.41 [0.30, 0.50] | 0.26 [0.14, 0.37] | 0.08 [0.03, 0.13] | 0.10 [0.05, 0.15] | 0.06 [0.02, 0.11] | 0.04 [0.01, 0.06] | 0.09 [0.04, 0.13] | 0.74 [0.70, 0.79] | 0.26 [0.17, 0.35] |
| 11 | Diffuse a label across one measured network (random walk with restart) | reliable | 0.29 [0.17, 0.41] | 0.79 [0.77, 0.81] | 0.37 [0.21, 0.54] | 0.33 [0.19, 0.48] | 0.33 [0.24, 0.41] | 0.28 [0.17, 0.42] | 0.33 [0.18, 0.48] | 0.10 [0.07, 0.15] | 0.12 [0.08, 0.16] | 0.70 [0.62, 0.77] | 0.34 [0.19, 0.51] |
| 12 | Let every network vote, weighted by what it has earned (chance-weighted ensemble vote) | weak | 0.56 [0.36, 0.75] | 0.88 [0.60, 1.00] | 0.64 [0.48, 0.78] | 0.54 [0.35, 0.72] | 0.41 [0.27, 0.52] | 0.43 [0.28, 0.56] | 0.54 [0.35, 0.72] | 0.32 [0.15, 0.43] | 0.35 [0.17, 0.48] | 0.72 [0.65, 0.76] | 0.48 [0.41, 0.55] |
| 13 | Place a protein by the proteins it physically touches (weighted partner vote) | reliable | 0.39 [0.17, 0.59] | 0.56 [0.27, 0.84] | 0.68 [0.59, 0.77] | 0.49 [0.34, 0.62] | 0.33 [0.13, 0.50] | 0.36 [0.17, 0.52] | 0.44 [0.22, 0.64] | 0.25 [0.06, 0.41] | 0.27 [0.08, 0.43] | 0.72 [0.60, 0.82] | 0.44 [0.37, 0.51] |
| 14 | Annotate function through shared fold (TM-score-weighted vote) | reliable | 0.74 [0.73, 0.74] | 0.78 [0.78, 0.79] | 0.94 [0.94, 0.95] | 0.83 [0.80, 0.86] | 0.66 [0.65, 0.67] | 0.71 [0.71, 0.72] | 0.80 [0.79, 0.81] | 0.72 [0.72, 0.72] | 0.74 [0.74, 0.74] | 0.97 [0.96, 0.97] | 0.65 [0.64, 0.67] |
| 19 | Train a classifier on the known genes and call the rest (logistic regression) | reliable | 0.61 [0.53, 0.70] | 1.00 [1.00, 1.00] | 0.61 [0.53, 0.70] | 0.53 [0.46, 0.61] | 0.57 [0.50, 0.66] | 0.54 [0.47, 0.62] | 0.63 [0.54, 0.72] | 0.37 [0.20, 0.50] | 0.38 [0.21, 0.51] | 0.81 [0.66, 0.90] | 0.58 [0.49, 0.69] |
| 30 | Test inference on the genes orthology cannot reach (kNN) | reliable | 0.59 [0.40, 0.77] | 0.90 [0.76, 1.00] | 0.64 [0.51, 0.77] | 0.61 [0.46, 0.76] | 0.36 [0.25, 0.49] | 0.39 [0.29, 0.55] | 0.56 [0.40, 0.72] | 0.32 [0.24, 0.41] | 0.36 [0.29, 0.47] | 0.80 [0.76, 0.86] | 0.50 [0.37, 0.69] |
| 31 | Call a gene only when independent strategies agree (kNN + logistic + network vote) | reliable | 0.35 [0.22, 0.49] | 0.41 [0.28, 0.54] | 0.83 [0.76, 0.91] | 0.65 [0.48, 0.83] | 0.22 [0.17, 0.27] | 0.28 [0.22, 0.34] | 0.44 [0.30, 0.57] | 0.14 [0.08, 0.20] | 0.21 [0.11, 0.27] | 0.83 [0.71, 0.91] | 0.62 [0.54, 0.72] |
| 32 | Put the understudied genes first (kNN + logistic + network vote) | reliable | 0.35 [0.20, 0.50] | 0.41 [0.26, 0.55] | 0.83 [0.75, 0.91] | 0.63 [0.46, 0.82] | 0.22 [0.16, 0.26] | 0.27 [0.21, 0.33] | 0.43 [0.28, 0.58] | 0.13 [0.07, 0.19] | 0.19 [0.10, 0.25] | 0.83 [0.70, 0.91] | 0.61 [0.53, 0.71] |
| 35 | Call genes with a stated error rate (split conformal prediction) | reliable | 0.18 [0.01, 0.42] | 0.25 [0.01, 0.54] | 0.45 [0.12, 0.75] | 0.29 [0.06, 0.52] | 0.17 [0.01, 0.40] | 0.19 [0.01, 0.42] | 0.22 [0.01, 0.48] | 0.09 [-0.00, 0.25] | 0.10 [-0.00, 0.28] | 0.76 [0.64, 0.86] | 0.53 [0.43, 0.66] |
| 36 | Smooth the measurements along the networks, then classify (graph convolution + logistic regression) | reliable | 0.63 [0.53, 0.73] | 1.00 [1.00, 1.00] | 0.63 [0.53, 0.73] | 0.54 [0.46, 0.63] | 0.61 [0.53, 0.71] | 0.56 [0.48, 0.66] | 0.64 [0.54, 0.75] | 0.40 [0.18, 0.55] | 0.41 [0.19, 0.56] | 0.83 [0.67, 0.92] | 0.61 [0.52, 0.72] |
| 37 | Let a random forest find what defines a label (random forest + permutation importance) | reliable | 0.70 [0.56, 0.83] | 1.00 [1.00, 1.00] | 0.70 [0.56, 0.83] | 0.61 [0.50, 0.73] | 0.58 [0.49, 0.68] | 0.58 [0.47, 0.69] | 0.68 [0.54, 0.82] | 0.45 [0.21, 0.66] | 0.45 [0.21, 0.66] | 0.85 [0.69, 0.94] | 0.65 [0.54, 0.76] |
| 38 | Learn how much to trust each kind of evidence (stacked logistic regression) | reliable | 0.61 [0.50, 0.73] | 1.00 [1.00, 1.00] | 0.61 [0.50, 0.73] | 0.55 [0.47, 0.63] | 0.62 [0.54, 0.71] | 0.55 [0.46, 0.65] | 0.63 [0.52, 0.75] | 0.41 [0.20, 0.55] | 0.41 [0.20, 0.56] | 0.82 [0.66, 0.92] | 0.63 [0.54, 0.74] |
ranking -- Rank candidates so the true ones come first. Hidden: set members, edges or corrupted labels, ranked among negatives.
| # | Strategy | Grade | AUROC | AUPRC | AUPRC lift | Prevalence | Partial AUROC (FPR <= 10%) | R-precision | Precision @ top 1% | Enrichment @ top 1% | Recall @ top 10% | Best F1 | nDCG |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 03 | Ask which categories the data can rediscover (UMAP + neighbour AUROC) | reliable | 0.75 [0.61, 0.84] | 0.62 [0.44, 0.79] | 8.81 [1.50, 18.16] | 0.41 [0.14, 0.67] | 0.65 [0.57, 0.73] | 0.61 [0.45, 0.76] | 0.66 [0.44, 0.86] | 9.31 [1.61, 18.06] | 0.34 [0.15, 0.54] | 0.68 [0.53, 0.82] | 0.84 [0.73, 0.93] |
| 10 | Find genes whose label their neighbours contradict (kNN + network neighbours) | reliable | 0.82 [0.75, 0.89] | 0.26 [0.19, 0.34] | 5.27 [3.86, 6.77] | 0.05 [0.05, 0.05] | 0.66 [0.60, 0.72] | 0.31 [0.21, 0.39] | 0.40 [0.28, 0.52] | 8.03 [5.67, 10.43] | 0.49 [0.35, 0.62] | 0.36 [0.28, 0.44] | 0.72 [0.62, 0.80] |
| 16 | Predict the contacts an interactome missed (logistic regression) | reliable | 0.96 [0.96, 0.96] | 0.82 [0.81, 0.83] | 4.83 [4.79, 4.86] | 0.17 [0.17, 0.17] | 0.83 [0.83, 0.84] | 0.77 [0.76, 0.77] | 0.98 [0.98, 0.99] | 5.77 [5.76, 5.79] | 0.48 [0.48, 0.49] | 0.77 [0.77, 0.78] | 0.97 [0.97, 0.98] |
| 17 | Read the literature for biology, not fame (publication-count residual) | reliable | 0.62 [0.60, 0.64] | 0.64 [0.58, 0.68] | 1.23 [1.14, 1.31] | 0.52 [0.46, 0.57] | 0.54 [0.52, 0.56] | 0.59 [0.54, 0.65] | 0.76 [0.61, 0.92] | 1.50 [1.11, 1.87] | 0.14 [0.12, 0.16] | 0.70 [0.65, 0.74] | 0.93 [0.91, 0.94] |
| 18 | List what the data says and the literature has not written (multi-layer support count) | reliable | 0.56 [0.56, 0.56] | 0.24 [0.24, 0.24] | 1.46 [1.46, 1.46] | 0.17 [0.17, 0.17] | 0.55 [0.55, 0.55] | 0.97 [0.97, 0.97] | 1.00 [1.00, 1.00] | 5.98 [5.97, 5.99] | 0.57 [0.57, 0.57] | 0.99 [0.99, 0.99] | 0.99 [0.99, 0.99] |
| 20 | Learn what makes your list special, from positives alone (PU bagging, logistic regression) | reliable | 0.90 [0.87, 0.93] | 0.11 [0.07, 0.14] | 17.1 [11.6, 23.6] | 0.01 [0.01, 0.01] | 0.72 [0.66, 0.78] | 0.16 [0.09, 0.22] | 0.14 [0.08, 0.18] | 20.0 [12.4, 28.0] | 0.66 [0.55, 0.77] | 0.18 [0.12, 0.24] | 0.56 [0.49, 0.61] |
| 24 | Describe what your gene list has in common (hypergeometric + rank-sum) | reliable | 0.82 [0.76, 0.87] | 0.07 [0.04, 0.11] | 9.08 [4.13, 13.91] | 0.01 [0.01, 0.01] | 0.65 [0.59, 0.70] | 0.11 [0.05, 0.17] | 0.10 [0.04, 0.16] | 12.3 [4.7, 20.3] | 0.50 [0.33, 0.65] | 0.15 [0.09, 0.21] | 0.52 [0.45, 0.59] |
| 25 | Grow your gene list along the networks (random walk with restart) | reliable | 0.81 [0.77, 0.84] | 0.05 [0.03, 0.08] | 8.54 [5.21, 12.99] | 0.01 [0.01, 0.01] | 0.65 [0.61, 0.69] | 0.10 [0.06, 0.14] | 0.08 [0.04, 0.13] | 13.0 [6.9, 21.0] | 0.50 [0.43, 0.57] | 0.14 [0.09, 0.19] | 0.47 [0.42, 0.52] |
| 28 | Find paralogs that changed jobs (profile correlation) | reliable | 0.58 [0.53, 0.66] | 0.38 [0.26, 0.58] | 1.20 [1.04, 1.38] | 0.30 [0.23, 0.41] | 0.53 [0.50, 0.57] | 0.35 [0.24, 0.52] | 0.54 [0.34, 0.71] | 1.80 [1.25, 2.35] | 0.13 [0.10, 0.17] | 0.48 [0.40, 0.60] | 0.81 [0.74, 0.88] |
| 33 | Put every layer into one space and read a gene's neighbourhood (logistic edge model) | reliable | 0.95 [0.93, 0.97] | 0.95 [0.94, 0.97] | 1.90 [1.87, 1.94] | 0.50 [0.50, 0.50] | 0.87 [0.82, 0.91] | 0.89 [0.86, 0.92] | 1.00 [1.00, 1.00] | 2.00 [2.00, 2.00] | 0.20 [0.20, 0.20] | 0.89 [0.87, 0.92] | 0.99 [0.99, 1.00] |
| 34 | Train on the networks and rank the edges they are missing (logistic / spectral embedding) | reliable | 0.94 [0.94, 0.94] | 0.94 [0.94, 0.95] | 1.89 [1.88, 1.89] | 0.50 [0.50, 0.50] | 0.83 [0.83, 0.84] | 0.88 [0.87, 0.88] | 1.00 [1.00, 1.00] | 2.00 [2.00, 2.00] | 0.20 [0.20, 0.20] | 0.88 [0.87, 0.89] | 0.99 [0.99, 0.99] |
set retrieval -- Return a set of genes that belong with a query. Hidden: part of a gene set, to be returned among all other genes.
| # | Strategy | Grade | Precision | Recall | F1 | Jaccard index | Matthews correlation (MCC) | Fold enrichment | Genes returned |
|---|---|---|---|---|---|---|---|---|---|
| 02 | Find the map where your gene list is one cluster (UMAP + HDBSCAN) | weak | 0.04 [0.03, 0.06] | 0.14 [0.07, 0.26] | 0.06 [0.04, 0.08] | 0.03 [0.02, 0.04] | 0.04 [0.01, 0.06] | 2.63 [1.59, 3.59] | 212.2 [93.9, 363.7] |
cluster recovery -- Find clusters that correspond to a label nobody showed them. Hidden: a share of a label, scored against clusters chosen on the rest.
| # | Strategy | Grade | Weighted F1 | Weighted precision | Weighted recall | Adjusted Rand index (ARI) | Normalised mutual information (NMI) | Homogeneity | Completeness | Unclustered share |
|---|---|---|---|---|---|---|---|---|---|---|
| 01 | Hold out a category and search for a map that finds it (UMAP + HDBSCAN) | weak | 0.39 [0.17, 0.60] | 0.51 [0.34, 0.67] | 0.43 [0.12, 0.74] | 0.03 [-0.01, 0.06] | 0.26 [0.10, 0.43] | 0.49 [0.20, 0.81] | 0.20 [0.10, 0.31] | 0.41 [0.13, 0.69] |
| 04 | Keep only the modules that survive the whole walk (UMAP + HDBSCAN co-clustering) | weak | 0.37 [0.27, 0.46] | 0.44 [0.25, 0.63] | 0.45 [0.39, 0.52] | 0.03 [0.01, 0.05] | 0.13 [0.08, 0.18] | 0.15 [0.09, 0.26] | 0.14 [0.06, 0.23] | 0.08 [0.05, 0.12] |
| 15 | Find the communities several networks agree on (modularity + Louvain consensus) | weak | 0.34 [0.24, 0.44] | 0.45 [0.25, 0.64] | 0.36 [0.30, 0.44] | 0.01 [0.00, 0.03] | 0.05 [0.03, 0.06] | 0.05 [0.04, 0.06] | 0.05 [0.03, 0.08] | 0.00 [0.00, 0.00] |
values -- Predict a measured value for genes without one. Hidden: a share of a measurement's values.
| # | Strategy | Grade | Spearman rho | Pearson r | Kendall tau-b | R-squared (out of sample) | Normalised RMSE | Mean absolute error | Top-decile recall | Bottom-decile recall | Coverage |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 21 | Predict a measurement, and find the genes that defy the prediction (gradient boosting / ridge) | reliable | 0.67 [0.16, 0.97] | 0.69 [0.20, 0.97] | 0.55 [0.11, 0.86] | 0.58 [0.01, 0.95] | 0.55 [0.22, 1.00] | 4.73 [0.11, 13.16] | 0.54 [0.22, 0.87] | 0.52 [0.20, 0.80] | 1.00 [1.00, 1.00] |
| 22 | Fill in what was never measured, and say where that is honest (soft-impute, low-rank SVD) | reliable | 0.90 [0.90, 0.90] | 0.90 [0.90, 0.90] | 0.74 [0.73, 0.74] | 0.81 [0.81, 0.81] | 0.43 [0.43, 0.43] | -- | 0.73 [0.72, 0.73] | 0.60 [0.60, 0.60] | 1.00 [1.00, 1.00] |
| 23 | Find what matters more in one condition, and why (residual + gradient boosting / ridge) | weak | 0.07 [0.05, 0.08] | 0.08 [0.07, 0.09] | 0.05 [0.04, 0.05] | -0.04 [-0.05, -0.03] | 1.02 [1.02, 1.02] | 0.25 [0.25, 0.26] | 0.12 [0.11, 0.14] | 0.21 [0.19, 0.23] | 1.00 [1.00, 1.00] |
| 29 | Carry what one parasite shows to the other (orthogroup mapping) | reliable | 0.32 [0.30, 0.35] | 0.32 [0.30, 0.34] | 0.22 [0.20, 0.24] | -1.59 [-1.71, -1.47] | 1.61 [1.57, 1.64] | 3.01 [2.96, 3.06] | 0.12 [0.12, 0.13] | 0.13 [0.12, 0.15] | 1.00 [1.00, 1.00] |
| 39 | Predict a value with an interval that holds (gradient boosting / ridge + split conformal) | weak | 0.53 [0.02, 0.86] | 0.54 [0.04, 0.89] | 0.40 [0.02, 0.68] | 0.40 [-0.07, 0.80] | 0.74 [0.45, 1.04] | 5.43 [0.27, 14.73] | 0.38 [0.11, 0.62] | 0.35 [0.13, 0.59] | 1.00 [1.00, 1.00] |
replication -- Make findings that hold beyond the genes they were made on. Hidden: half of the genes, on which first-half findings are checked.
| # | Strategy | Grade | Replication rate | Findings made | Findings replicated | Replication by chance | Replication lift |
|---|---|---|---|---|---|---|---|
| 05 | Tune a map without labels, then read what it encodes (UMAP + HDBSCAN, chi-square / Kruskal-Wallis) | reliable | 0.97 [0.95, 0.98] | 58.5 [57.0, 60.0] | 56.5 [56.0, 57.0] | 0.05 [0.04, 0.05] | 21.6 [17.8, 25.5] |
| 26 | Find categories that split in two on another measurement (UMAP + HDBSCAN) | untestable | -- | -- | -- | -- | -- |
| 27 | Find kinds of gene defined by two labels at once (UMAP + HDBSCAN) | reliable | 0.54 [0.33, 0.75] | 3.50 [3.00, 4.00] | 2.00 [1.00, 3.00] | 0.03 [0.02, 0.04] | 20.0 [20.0, 20.0] |
Plasmodium falciparum, at the default settings
label calls -- Call a label for genes that lack it. Hidden: a share of a label's genes, whole orthogroups at a time.
| # | Strategy | Grade | Accuracy | Coverage | Precision of calls | Macro precision | Macro recall (balanced accuracy) | Macro F1 | Weighted F1 | Cohen's kappa | Matthews correlation (MCC) | Macro AUROC | Macro AUPRC |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 06 | Find which kind of evidence carries a label (kNN ablation) | reliable | 0.64 [0.44, 0.85] | 1.00 [1.00, 1.00] | 0.64 [0.44, 0.85] | 0.46 [0.35, 0.55] | 0.43 [0.32, 0.51] | 0.42 [0.31, 0.49] | 0.60 [0.39, 0.82] | 0.19 [0.07, 0.28] | 0.20 [0.08, 0.29] | 0.68 [0.60, 0.75] | 0.47 [0.35, 0.56] |
| 07 | Call a gene by the genes that behave like it (kNN) | weak | 0.68 [0.53, 0.85] | 0.96 [0.87, 1.00] | 0.71 [0.61, 0.86] | 0.66 [0.57, 0.76] | 0.51 [0.45, 0.58] | 0.51 [0.47, 0.55] | 0.66 [0.54, 0.84] | 0.30 [0.19, 0.41] | 0.33 [0.25, 0.43] | 0.84 [0.77, 0.90] | 0.63 [0.58, 0.68] |
| 08 | Call a gene by its neighbours on the map (UMAP + kNN) | weak | 0.64 [0.42, 0.84] | 0.94 [0.83, 1.00] | 0.66 [0.49, 0.84] | 0.54 [0.45, 0.65] | 0.48 [0.35, 0.58] | 0.47 [0.36, 0.57] | 0.62 [0.42, 0.83] | 0.24 [0.17, 0.31] | 0.26 [0.19, 0.32] | 0.79 [0.71, 0.85] | 0.56 [0.42, 0.69] |
| 09 | Name a cluster by the label it is enriched for (UMAP + HDBSCAN, hypergeometric) | reliable | 0.06 [0.02, 0.09] | 0.14 [0.07, 0.22] | 0.42 [0.33, 0.60] | 0.21 [0.17, 0.34] | 0.19 [0.09, 0.27] | 0.17 [0.11, 0.22] | 0.07 [0.02, 0.12] | 0.05 [0.02, 0.07] | 0.16 [0.11, 0.21] | 0.80 [0.70, 0.87] | 0.38 [0.17, 0.59] |
| 11 | Diffuse a label across one measured network (random walk with restart) | reliable | 0.45 [0.17, 0.72] | 0.95 [0.95, 0.96] | 0.47 [0.18, 0.75] | 0.39 [0.18, 0.56] | 0.47 [0.23, 0.76] | 0.38 [0.17, 0.53] | 0.48 [0.19, 0.80] | 0.14 [0.12, 0.16] | 0.18 [0.13, 0.24] | 0.73 [0.62, 0.85] | 0.42 [0.18, 0.64] |
| 12 | Let every network vote, weighted by what it has earned (chance-weighted ensemble vote) | reliable | 0.65 [0.53, 0.78] | 0.94 [0.81, 1.00] | 0.70 [0.57, 0.86] | 0.63 [0.55, 0.72] | 0.50 [0.43, 0.57] | 0.50 [0.44, 0.56] | 0.64 [0.52, 0.80] | 0.30 [0.18, 0.43] | 0.33 [0.22, 0.44] | 0.71 [0.67, 0.77] | 0.49 [0.45, 0.54] |
| 13 | Place a protein by the proteins it physically touches (weighted partner vote) | weak | 0.54 [0.33, 0.75] | 0.71 [0.61, 0.80] | 0.75 [0.48, 0.97] | 0.53 [0.46, 0.62] | 0.41 [0.30, 0.52] | 0.45 [0.35, 0.56] | 0.61 [0.43, 0.83] | 0.17 [0.00, 0.44] | 0.20 [0.00, 0.49] | 0.65 [0.49, 0.86] | 0.47 [0.43, 0.50] |
| 14 | Annotate function through shared fold (TM-score-weighted vote) | reliable | 0.69 [0.66, 0.72] | 0.75 [0.71, 0.77] | 0.93 [0.92, 0.94] | 0.78 [0.73, 0.84] | 0.53 [0.47, 0.59] | 0.60 [0.54, 0.67] | 0.77 [0.74, 0.80] | 0.59 [0.56, 0.62] | 0.63 [0.60, 0.65] | 0.90 [0.87, 0.93] | 0.54 [0.49, 0.59] |
| 19 | Train a classifier on the known genes and call the rest (logistic regression) | reliable | 0.63 [0.46, 0.82] | 1.00 [1.00, 1.00] | 0.63 [0.46, 0.82] | 0.55 [0.46, 0.63] | 0.68 [0.52, 0.83] | 0.55 [0.43, 0.67] | 0.64 [0.45, 0.84] | 0.39 [0.35, 0.43] | 0.42 [0.36, 0.49] | 0.85 [0.77, 0.92] | 0.61 [0.55, 0.68] |
| 30 | Test inference on the genes orthology cannot reach (kNN) | weak | 0.66 [0.32, 1.00] | 1.00 [1.00, 1.00] | 0.66 [0.32, 1.00] | 0.55 [0.11, 1.00] | 0.67 [0.33, 1.00] | 0.58 [0.16, 1.00] | 0.58 [0.15, 1.00] | 0.00 [0.00, 0.00] | 0.00 [0.00, 0.00] | 0.84 [0.84, 0.84] | 0.86 [0.86, 0.86] |
| 31 | Call a gene only when independent strategies agree (kNN + logistic + network vote) | works when tuned | 0.67 [0.54, 0.84] | 0.86 [0.77, 0.96] | 0.77 [0.67, 0.87] | 0.70 [0.63, 0.78] | 0.53 [0.48, 0.59] | 0.55 [0.51, 0.58] | 0.69 [0.57, 0.84] | 0.34 [0.23, 0.43] | 0.37 [0.29, 0.45] | 0.87 [0.79, 0.94] | 0.68 [0.61, 0.74] |
| 32 | Put the understudied genes first (kNN + logistic + network vote) | works when tuned | 0.67 [0.56, 0.84] | 0.87 [0.78, 0.96] | 0.77 [0.68, 0.87] | 0.65 [0.55, 0.76] | 0.50 [0.44, 0.58] | 0.51 [0.45, 0.56] | 0.69 [0.58, 0.84] | 0.32 [0.20, 0.43] | 0.36 [0.25, 0.45] | 0.87 [0.80, 0.94] | 0.68 [0.61, 0.74] |
| 35 | Call genes with a stated error rate (split conformal prediction) | reliable | 0.32 [0.10, 0.56] | 0.39 [0.12, 0.68] | 0.85 [0.79, 0.94] | 0.59 [0.47, 0.67] | 0.36 [0.11, 0.62] | 0.36 [0.17, 0.55] | 0.40 [0.15, 0.66] | 0.14 [0.06, 0.22] | 0.24 [0.17, 0.31] | 0.87 [0.77, 0.94] | 0.65 [0.59, 0.72] |
| 36 | Smooth the measurements along the networks, then classify (graph convolution + logistic regression) | reliable | 0.71 [0.59, 0.84] | 1.00 [1.00, 1.00] | 0.71 [0.59, 0.84] | 0.60 [0.55, 0.66] | 0.68 [0.56, 0.80] | 0.62 [0.55, 0.69] | 0.72 [0.60, 0.85] | 0.46 [0.37, 0.54] | 0.47 [0.38, 0.55] | 0.87 [0.78, 0.94] | 0.67 [0.60, 0.73] |
| 37 | Let a random forest find what defines a label (random forest + permutation importance) | reliable | 0.74 [0.66, 0.86] | 1.00 [1.00, 1.00] | 0.74 [0.66, 0.86] | 0.65 [0.58, 0.73] | 0.57 [0.54, 0.60] | 0.55 [0.51, 0.59] | 0.71 [0.62, 0.85] | 0.36 [0.19, 0.53] | 0.39 [0.23, 0.54] | 0.88 [0.80, 0.94] | 0.68 [0.62, 0.72] |
| 38 | Learn how much to trust each kind of evidence (stacked logistic regression) | reliable | 0.70 [0.60, 0.82] | 1.00 [1.00, 1.00] | 0.70 [0.60, 0.82] | 0.60 [0.57, 0.64] | 0.70 [0.59, 0.81] | 0.62 [0.57, 0.67] | 0.72 [0.61, 0.84] | 0.45 [0.38, 0.52] | 0.48 [0.41, 0.54] | 0.87 [0.79, 0.94] | 0.67 [0.61, 0.71] |
ranking -- Rank candidates so the true ones come first. Hidden: set members, edges or corrupted labels, ranked among negatives.
| # | Strategy | Grade | AUROC | AUPRC | AUPRC lift | Prevalence | Partial AUROC (FPR <= 10%) | R-precision | Precision @ top 1% | Enrichment @ top 1% | Recall @ top 10% | Best F1 | nDCG |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 03 | Ask which categories the data can rediscover (UMAP + neighbour AUROC) | reliable | 0.84 [0.78, 0.90] | 0.66 [0.43, 0.90] | 5.18 [1.47, 9.52] | 0.37 [0.07, 0.75] | 0.72 [0.66, 0.76] | 0.62 [0.45, 0.86] | 0.84 [0.58, 0.99] | 7.38 [1.88, 13.50] | 0.39 [0.18, 0.60] | 0.69 [0.54, 0.89] | 0.87 [0.74, 0.97] |
| 10 | Find genes whose label their neighbours contradict (kNN + network neighbours) | reliable | 0.87 [0.79, 0.94] | 0.44 [0.18, 0.74] | 8.82 [3.66, 14.74] | 0.05 [0.05, 0.05] | 0.74 [0.60, 0.88] | 0.46 [0.21, 0.71] | 0.56 [0.28, 0.82] | 11.4 [5.6, 16.7] | 0.61 [0.37, 0.82] | 0.51 [0.27, 0.73] | 0.76 [0.65, 0.90] |
| 16 | Predict the contacts an interactome missed (logistic regression) | reliable | 0.98 [0.98, 0.98] | 0.90 [0.89, 0.90] | 5.35 [5.31, 5.38] | 0.17 [0.17, 0.17] | 0.92 [0.91, 0.92] | 0.84 [0.84, 0.85] | 0.99 [0.99, 1.00] | 5.90 [5.85, 5.94] | 0.55 [0.55, 0.55] | 0.87 [0.87, 0.87] | 0.98 [0.98, 0.99] |
| 17 | Read the literature for biology, not fame (publication-count residual) | weak | 0.56 [0.46, 0.67] | 0.71 [0.56, 0.98] | 1.18 [1.00, 1.41] | 0.63 [0.41, 0.98] | 0.53 [0.51, 0.56] | 0.66 [0.48, 0.98] | 0.83 [0.50, 1.00] | 1.49 [1.02, 2.42] | 0.14 [0.10, 0.18] | 0.77 [0.66, 0.99] | 0.91 [0.83, 1.00] |
| 18 | List what the data says and the literature has not written (multi-layer support count) | reliable | 0.55 [0.55, 0.55] | 0.24 [0.24, 0.25] | 1.47 [1.44, 1.49] | 0.17 [0.17, 0.17] | 0.55 [0.55, 0.55] | 0.98 [0.97, 0.99] | 0.99 [0.98, 1.00] | 5.96 [5.89, 6.00] | 0.58 [0.58, 0.59] | 0.99 [0.99, 0.99] | 0.99 [0.99, 0.99] |
| 20 | Learn what makes your list special, from positives alone (PU bagging, logistic regression) | reliable | 0.95 [0.90, 0.99] | 0.34 [0.09, 0.70] | 42.5 [11.6, 83.0] | 0.01 [0.00, 0.01] | 0.82 [0.66, 0.95] | 0.31 [0.08, 0.61] | 0.28 [0.08, 0.54] | 35.2 [12.5, 63.0] | 0.84 [0.58, 0.99] | 0.38 [0.13, 0.70] | 0.67 [0.43, 0.89] |
| 24 | Describe what your gene list has in common (hypergeometric + rank-sum) | reliable | 0.93 [0.88, 0.97] | 0.20 [0.09, 0.32] | 19.1 [12.7, 25.8] | 0.01 [0.01, 0.01] | 0.78 [0.69, 0.87] | 0.25 [0.15, 0.37] | 0.26 [0.13, 0.40] | 24.9 [19.3, 31.1] | 0.76 [0.58, 0.92] | 0.30 [0.20, 0.42] | 0.62 [0.50, 0.74] |
| 25 | Grow your gene list along the networks (random walk with restart) | reliable | 0.93 [0.88, 0.98] | 0.31 [0.07, 0.60] | 37.7 [12.9, 71.9] | 0.01 [0.00, 0.01] | 0.83 [0.72, 0.93] | 0.32 [0.11, 0.58] | 0.29 [0.08, 0.51] | 35.2 [15.3, 62.3] | 0.82 [0.67, 0.96] | 0.37 [0.15, 0.62] | 0.65 [0.45, 0.86] |
| 28 | Find paralogs that changed jobs (profile correlation) | weak | 0.56 [0.44, 0.68] | 0.46 [0.33, 0.54] | 1.20 [1.07, 1.45] | 0.38 [0.31, 0.46] | 0.52 [0.50, 0.54] | 0.41 [0.32, 0.50] | 0.80 [0.39, 1.00] | 2.02 [1.26, 2.67] | 0.13 [0.10, 0.15] | 0.59 [0.49, 0.64] | 0.82 [0.81, 0.84] |
| 33 | Put every layer into one space and read a gene's neighbourhood (logistic edge model) | reliable | 0.71 [0.68, 0.73] | 0.72 [0.68, 0.75] | 1.43 [1.36, 1.50] | 0.50 [0.50, 0.50] | 0.59 [0.56, 0.63] | 0.65 [0.63, 0.66] | 0.95 [0.90, 0.99] | 1.89 [1.80, 1.97] | 0.17 [0.16, 0.19] | 0.69 [0.68, 0.70] | 0.96 [0.95, 0.97] |
| 34 | Train on the networks and rank the edges they are missing (logistic / spectral embedding) | reliable | 0.71 [0.68, 0.73] | 0.72 [0.68, 0.75] | 1.43 [1.36, 1.50] | 0.50 [0.50, 0.50] | 0.59 [0.56, 0.63] | 0.65 [0.63, 0.66] | 0.95 [0.91, 0.99] | 1.89 [1.81, 1.97] | 0.17 [0.16, 0.19] | 0.69 [0.68, 0.70] | 0.96 [0.95, 0.97] |
set retrieval -- Return a set of genes that belong with a query. Hidden: part of a gene set, to be returned among all other genes.
| # | Strategy | Grade | Precision | Recall | F1 | Jaccard index | Matthews correlation (MCC) | Fold enrichment | Genes returned |
|---|---|---|---|---|---|---|---|---|---|
| 02 | Find the map where your gene list is one cluster (UMAP + HDBSCAN) | reliable | 0.16 [0.08, 0.25] | 0.42 [0.17, 0.69] | 0.23 [0.11, 0.37] | 0.14 [0.06, 0.23] | 0.24 [0.11, 0.40] | 10.9 [7.5, 15.9] | 102.0 [58.3, 132.6] |
cluster recovery -- Find clusters that correspond to a label nobody showed them. Hidden: a share of a label, scored against clusters chosen on the rest.
| # | Strategy | Grade | Weighted F1 | Weighted precision | Weighted recall | Adjusted Rand index (ARI) | Normalised mutual information (NMI) | Homogeneity | Completeness | Unclustered share |
|---|---|---|---|---|---|---|---|---|---|---|
| 01 | Hold out a category and search for a map that finds it (UMAP + HDBSCAN) | weak | 0.54 [0.23, 0.81] | 0.55 [0.36, 0.80] | 0.70 [0.30, 0.98] | 0.08 [0.01, 0.18] | 0.20 [0.03, 0.45] | 0.33 [0.05, 0.71] | 0.58 [0.32, 0.86] | 0.22 [0.01, 0.54] |
| 04 | Keep only the modules that survive the whole walk (UMAP + HDBSCAN co-clustering) | weak | 0.50 [0.23, 0.78] | 0.57 [0.35, 0.84] | 0.56 [0.21, 0.91] | 0.11 [0.01, 0.25] | 0.23 [0.07, 0.42] | 0.35 [0.13, 0.57] | 0.41 [0.16, 0.76] | 0.21 [0.00, 0.42] |
| 15 | Find the communities several networks agree on (modularity + Louvain consensus) | weak | 0.25 [0.14, 0.35] | 0.47 [0.24, 0.74] | 0.24 [0.12, 0.34] | 0.02 [-0.00, 0.04] | 0.12 [0.03, 0.23] | 0.23 [0.07, 0.48] | 0.10 [0.02, 0.18] | 0.00 [0.00, 0.00] |
values -- Predict a measured value for genes without one. Hidden: a share of a measurement's values.
| # | Strategy | Grade | Spearman rho | Pearson r | Kendall tau-b | R-squared (out of sample) | Normalised RMSE | Mean absolute error | Top-decile recall | Bottom-decile recall | Coverage |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 21 | Predict a measurement, and find the genes that defy the prediction (gradient boosting / ridge) | reliable | 0.52 [0.25, 0.85] | 0.58 [0.40, 0.85] | 0.40 [0.21, 0.66] | -3.29 [-14.04, 0.72] | 1.26 [0.52, 2.67] | 126.5 [0.3, 360.2] | 0.35 [0.23, 0.51] | 0.37 [0.20, 0.59] | 1.00 [1.00, 1.00] |
| 22 | Fill in what was never measured, and say where that is honest (soft-impute, low-rank SVD) | reliable | 0.68 [0.67, 0.69] | 0.69 [0.69, 0.70] | 0.51 [0.51, 0.52] | 0.45 [0.44, 0.46] | 0.74 [0.73, 0.75] | -- | 0.48 [0.47, 0.50] | 0.40 [0.38, 0.42] | 1.00 [1.00, 1.00] |
| 23 | Find what matters more in one condition, and why (residual + gradient boosting / ridge) | reliable | 0.64 [0.64, 0.65] | 0.63 [0.63, 0.64] | 0.46 [0.46, 0.47] | 0.40 [0.40, 0.41] | 0.77 [0.77, 0.78] | 0.14 [0.13, 0.14] | 0.39 [0.37, 0.42] | 0.41 [0.39, 0.42] | 1.00 [1.00, 1.00] |
| 29 | Carry what one parasite shows to the other (orthogroup mapping) | reliable | 0.32 [0.30, 0.33] | 0.31 [0.30, 0.33] | 0.21 [0.20, 0.22] | -85.8 [-88.3, -83.7] | 9.31 [9.20, 9.45] | 2.97 [2.92, 3.01] | 0.17 [0.13, 0.20] | 0.13 [0.11, 0.15] | 1.00 [1.00, 1.00] |
| 39 | Predict a value with an interval that holds (gradient boosting / ridge + split conformal) | reliable | 0.50 [0.22, 0.84] | 0.55 [0.35, 0.84] | 0.38 [0.17, 0.65] | -3.49 [-14.82, 0.71] | 1.28 [0.54, 2.73] | 139.9 [0.3, 401.4] | 0.35 [0.23, 0.50] | 0.37 [0.20, 0.58] | 1.00 [1.00, 1.00] |
replication -- Make findings that hold beyond the genes they were made on. Hidden: half of the genes, on which first-half findings are checked.
| # | Strategy | Grade | Replication rate | Findings made | Findings replicated | Replication by chance | Replication lift |
|---|---|---|---|---|---|---|---|
| 05 | Tune a map without labels, then read what it encodes (UMAP + HDBSCAN, chi-square / Kruskal-Wallis) | reliable | 0.88 [0.83, 0.94] | 54.0 [51.8, 55.6] | 47.6 [44.6, 50.6] | 0.05 [0.04, 0.07] | 18.3 [13.6, 23.7] |
| 26 | Find categories that split in two on another measurement (UMAP + HDBSCAN) | untestable | -- | -- | -- | -- | -- |
| 27 | Find kinds of gene defined by two labels at once (UMAP + HDBSCAN) | untestable | -- | -- | -- | -- | -- |
Plasmodium falciparum, at the tuned setting (chosen on seeds 1-3, reported on seeds 4-5)
label calls -- Call a label for genes that lack it. Hidden: a share of a label's genes, whole orthogroups at a time.
| # | Strategy | Grade | Accuracy | Coverage | Precision of calls | Macro precision | Macro recall (balanced accuracy) | Macro F1 | Weighted F1 | Cohen's kappa | Matthews correlation (MCC) | Macro AUROC | Macro AUPRC |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 06 | Find which kind of evidence carries a label (kNN ablation) | reliable | 0.63 [0.42, 0.84] | 1.00 [1.00, 1.00] | 0.63 [0.42, 0.84] | 0.45 [0.36, 0.54] | 0.44 [0.34, 0.51] | 0.43 [0.33, 0.52] | 0.61 [0.41, 0.82] | 0.19 [0.06, 0.28] | 0.19 [0.06, 0.29] | 0.63 [0.55, 0.71] | 0.44 [0.32, 0.52] |
| 07 | Call a gene by the genes that behave like it (kNN) | weak | 0.70 [0.55, 0.88] | 0.96 [0.87, 1.00] | 0.73 [0.62, 0.88] | 0.68 [0.55, 0.83] | 0.49 [0.45, 0.53] | 0.50 [0.47, 0.54] | 0.68 [0.55, 0.84] | 0.28 [0.15, 0.41] | 0.32 [0.21, 0.43] | 0.83 [0.77, 0.90] | 0.64 [0.58, 0.71] |
| 08 | Call a gene by its neighbours on the map (UMAP + kNN) | weak | 0.64 [0.40, 0.86] | 0.89 [0.66, 1.00] | 0.70 [0.55, 0.86] | 0.50 [0.40, 0.63] | 0.42 [0.28, 0.51] | 0.43 [0.29, 0.53] | 0.62 [0.40, 0.83] | 0.20 [0.07, 0.29] | 0.22 [0.08, 0.32] | 0.79 [0.72, 0.86] | 0.56 [0.43, 0.69] |
| 09 | Name a cluster by the label it is enriched for (UMAP + HDBSCAN, hypergeometric) | reliable | 0.08 [0.04, 0.12] | 0.15 [0.06, 0.24] | 0.56 [0.45, 0.69] | 0.34 [0.23, 0.46] | 0.15 [0.09, 0.23] | 0.18 [0.13, 0.24] | 0.12 [0.06, 0.18] | 0.05 [0.03, 0.09] | 0.16 [0.12, 0.22] | 0.79 [0.70, 0.87] | 0.39 [0.21, 0.55] |
| 11 | Diffuse a label across one measured network (random walk with restart) | reliable | 0.45 [0.17, 0.76] | 0.95 [0.94, 0.96] | 0.47 [0.18, 0.79] | 0.39 [0.19, 0.57] | 0.47 [0.24, 0.75] | 0.38 [0.17, 0.55] | 0.48 [0.18, 0.82] | 0.15 [0.12, 0.17] | 0.18 [0.13, 0.24] | 0.73 [0.61, 0.86] | 0.45 [0.20, 0.70] |
| 12 | Let every network vote, weighted by what it has earned (chance-weighted ensemble vote) | reliable | 0.65 [0.56, 0.77] | 0.94 [0.77, 1.00] | 0.71 [0.57, 0.87] | 0.67 [0.54, 0.84] | 0.48 [0.45, 0.51] | 0.50 [0.47, 0.54] | 0.65 [0.55, 0.76] | 0.28 [0.12, 0.43] | 0.31 [0.18, 0.44] | 0.71 [0.63, 0.78] | 0.50 [0.44, 0.57] |
| 13 | Place a protein by the proteins it physically touches (weighted partner vote) | weak | 0.54 [0.31, 0.76] | 0.71 [0.63, 0.78] | 0.75 [0.48, 0.97] | 0.49 [0.48, 0.51] | 0.38 [0.30, 0.44] | 0.42 [0.36, 0.46] | 0.61 [0.41, 0.84] | 0.19 [-0.02, 0.49] | 0.20 [-0.04, 0.53] | 0.63 [0.49, 0.79] | 0.46 [0.43, 0.50] |
| 14 | Annotate function through shared fold (TM-score-weighted vote) | reliable | 0.64 [0.62, 0.67] | 0.72 [0.66, 0.78] | 0.90 [0.86, 0.94] | 0.65 [0.64, 0.67] | 0.53 [0.52, 0.55] | 0.56 [0.56, 0.56] | 0.72 [0.69, 0.75] | 0.60 [0.57, 0.62] | 0.63 [0.62, 0.64] | 0.94 [0.89, 0.98] | 0.55 [0.54, 0.56] |
| 19 | Train a classifier on the known genes and call the rest (logistic regression) | reliable | 0.69 [0.56, 0.84] | 1.00 [1.00, 1.00] | 0.69 [0.56, 0.84] | 0.58 [0.52, 0.65] | 0.68 [0.56, 0.82] | 0.61 [0.53, 0.70] | 0.70 [0.57, 0.85] | 0.44 [0.36, 0.51] | 0.46 [0.37, 0.54] | 0.86 [0.77, 0.94] | 0.66 [0.58, 0.74] |
| 30 | Test inference on the genes orthology cannot reach (kNN) | weak | 0.70 [0.55, 0.88] | 0.95 [0.86, 1.00] | 0.73 [0.62, 0.88] | 0.67 [0.55, 0.81] | 0.50 [0.45, 0.53] | 0.51 [0.48, 0.55] | 0.68 [0.55, 0.85] | 0.30 [0.15, 0.41] | 0.33 [0.21, 0.43] | 0.83 [0.75, 0.89] | 0.63 [0.58, 0.70] |
| 31 | Call a gene only when independent strategies agree (kNN + logistic + network vote) | works when tuned | 0.43 [0.18, 0.75] | 0.48 [0.21, 0.78] | 0.87 [0.78, 0.95] | 0.50 [0.37, 0.58] | 0.27 [0.14, 0.41] | 0.32 [0.20, 0.42] | 0.50 [0.27, 0.79] | 0.17 [0.09, 0.21] | 0.23 [0.14, 0.27] | 0.87 [0.78, 0.94] | 0.70 [0.61, 0.82] |
| 32 | Put the understudied genes first (kNN + logistic + network vote) | works when tuned | 0.49 [0.24, 0.88] | 0.54 [0.26, 0.90] | 0.88 [0.77, 0.98] | 0.51 [0.49, 0.52] | 0.31 [0.17, 0.47] | 0.36 [0.24, 0.48] | 0.56 [0.32, 0.90] | 0.20 [0.17, 0.23] | 0.27 [0.24, 0.30] | 0.87 [0.73, 0.95] | 0.67 [0.58, 0.79] |
| 35 | Call genes with a stated error rate (split conformal prediction) | reliable | 0.41 [0.14, 0.67] | 0.50 [0.20, 0.80] | 0.80 [0.71, 0.87] | 0.62 [0.55, 0.69] | 0.42 [0.15, 0.69] | 0.42 [0.23, 0.62] | 0.49 [0.21, 0.75] | 0.18 [0.09, 0.28] | 0.26 [0.18, 0.36] | 0.86 [0.77, 0.93] | 0.65 [0.58, 0.73] |
| 36 | Smooth the measurements along the networks, then classify (graph convolution + logistic regression) | reliable | 0.70 [0.57, 0.85] | 1.00 [1.00, 1.00] | 0.70 [0.57, 0.85] | 0.59 [0.53, 0.67] | 0.68 [0.56, 0.82] | 0.61 [0.54, 0.71] | 0.71 [0.58, 0.86] | 0.46 [0.38, 0.54] | 0.47 [0.38, 0.56] | 0.87 [0.78, 0.93] | 0.67 [0.60, 0.75] |
| 37 | Let a random forest find what defines a label (random forest + permutation importance) | reliable | 0.75 [0.64, 0.88] | 1.00 [1.00, 1.00] | 0.75 [0.64, 0.88] | 0.65 [0.59, 0.73] | 0.59 [0.56, 0.63] | 0.58 [0.54, 0.61] | 0.73 [0.62, 0.86] | 0.41 [0.26, 0.54] | 0.43 [0.31, 0.54] | 0.89 [0.81, 0.95] | 0.68 [0.62, 0.74] |
| 38 | Learn how much to trust each kind of evidence (stacked logistic regression) | reliable | 0.70 [0.59, 0.84] | 1.00 [1.00, 1.00] | 0.70 [0.59, 0.84] | 0.60 [0.55, 0.65] | 0.70 [0.59, 0.82] | 0.62 [0.56, 0.69] | 0.72 [0.60, 0.85] | 0.46 [0.40, 0.52] | 0.48 [0.41, 0.54] | 0.87 [0.79, 0.94] | 0.69 [0.62, 0.75] |
ranking -- Rank candidates so the true ones come first. Hidden: set members, edges or corrupted labels, ranked among negatives.
| # | Strategy | Grade | AUROC | AUPRC | AUPRC lift | Prevalence | Partial AUROC (FPR <= 10%) | R-precision | Precision @ top 1% | Enrichment @ top 1% | Recall @ top 10% | Best F1 | nDCG |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 03 | Ask which categories the data can rediscover (UMAP + neighbour AUROC) | reliable | 0.83 [0.78, 0.86] | 0.65 [0.46, 0.88] | 4.76 [1.47, 9.12] | 0.37 [0.09, 0.74] | 0.68 [0.61, 0.74] | 0.63 [0.48, 0.85] | 0.81 [0.62, 0.97] | 6.80 [1.69, 12.54] | 0.38 [0.18, 0.58] | 0.69 [0.56, 0.89] | 0.87 [0.76, 0.96] |
| 10 | Find genes whose label their neighbours contradict (kNN + network neighbours) | reliable | 0.88 [0.81, 0.96] | 0.46 [0.20, 0.77] | 9.41 [4.05, 15.39] | 0.05 [0.05, 0.05] | 0.76 [0.62, 0.90] | 0.48 [0.25, 0.73] | 0.56 [0.28, 0.83] | 11.3 [5.6, 16.9] | 0.63 [0.41, 0.87] | 0.53 [0.31, 0.76] | 0.77 [0.66, 0.90] |
| 16 | Predict the contacts an interactome missed (logistic regression) | reliable | 0.98 [0.98, 0.98] | 0.91 [0.90, 0.91] | 5.38 [5.37, 5.39] | 0.17 [0.17, 0.17] | 0.92 [0.92, 0.92] | 0.85 [0.84, 0.85] | 0.99 [0.98, 1.00] | 5.89 [5.84, 5.94] | 0.55 [0.55, 0.55] | 0.87 [0.87, 0.88] | 0.99 [0.99, 0.99] |
| 17 | Read the literature for biology, not fame (publication-count residual) | weak | 0.56 [0.46, 0.67] | 0.71 [0.56, 0.98] | 1.18 [1.00, 1.41] | 0.63 [0.41, 0.98] | 0.53 [0.51, 0.56] | 0.66 [0.48, 0.98] | 0.83 [0.50, 1.00] | 1.49 [1.02, 2.42] | 0.14 [0.10, 0.18] | 0.77 [0.66, 0.99] | 0.91 [0.83, 1.00] |
| 18 | List what the data says and the literature has not written (multi-layer support count) | reliable | 0.55 [0.55, 0.55] | 0.25 [0.25, 0.25] | 1.49 [1.49, 1.49] | 0.17 [0.17, 0.17] | 0.55 [0.55, 0.55] | 0.99 [0.99, 0.99] | 1.00 [1.00, 1.00] | 6.00 [6.00, 6.00] | 0.59 [0.59, 0.59] | 0.99 [0.99, 0.99] | 0.99 [0.99, 0.99] |
| 20 | Learn what makes your list special, from positives alone (PU bagging, logistic regression) | reliable | 0.95 [0.91, 0.99] | 0.34 [0.08, 0.72] | 43.0 [12.5, 84.5] | 0.01 [0.00, 0.01] | 0.83 [0.67, 0.95] | 0.32 [0.07, 0.63] | 0.28 [0.08, 0.54] | 34.7 [12.4, 63.3] | 0.84 [0.59, 0.98] | 0.38 [0.12, 0.73] | 0.67 [0.43, 0.89] |
| 24 | Describe what your gene list has in common (hypergeometric + rank-sum) | reliable | 0.93 [0.88, 0.97] | 0.19 [0.09, 0.31] | 18.3 [12.2, 24.5] | 0.01 [0.01, 0.01] | 0.78 [0.70, 0.87] | 0.25 [0.16, 0.35] | 0.24 [0.12, 0.38] | 23.6 [18.7, 29.8] | 0.78 [0.62, 0.93] | 0.30 [0.19, 0.41] | 0.61 [0.48, 0.74] |
| 25 | Grow your gene list along the networks (random walk with restart) | reliable | 0.93 [0.88, 0.98] | 0.30 [0.07, 0.59] | 37.2 [14.5, 72.1] | 0.01 [0.00, 0.01] | 0.83 [0.72, 0.93] | 0.31 [0.10, 0.58] | 0.29 [0.09, 0.52] | 36.3 [17.5, 62.9] | 0.80 [0.61, 0.95] | 0.37 [0.16, 0.62] | 0.65 [0.46, 0.84] |
| 28 | Find paralogs that changed jobs (profile correlation) | weak | 0.56 [0.44, 0.68] | 0.46 [0.33, 0.54] | 1.20 [1.07, 1.45] | 0.38 [0.31, 0.46] | 0.52 [0.50, 0.54] | 0.41 [0.32, 0.50] | 0.80 [0.39, 1.00] | 2.02 [1.26, 2.67] | 0.13 [0.10, 0.15] | 0.59 [0.49, 0.64] | 0.82 [0.81, 0.84] |
| 33 | Put every layer into one space and read a gene's neighbourhood (logistic edge model) | reliable | 0.89 [0.86, 0.91] | 0.92 [0.90, 0.94] | 1.84 [1.80, 1.88] | 0.50 [0.50, 0.50] | 0.86 [0.82, 0.90] | 0.81 [0.77, 0.84] | 1.00 [1.00, 1.00] | 2.00 [2.00, 2.00] | 0.20 [0.20, 0.20] | 0.84 [0.79, 0.88] | 0.99 [0.98, 0.99] |
| 34 | Train on the networks and rank the edges they are missing (logistic / spectral embedding) | reliable | 0.89 [0.88, 0.90] | 0.92 [0.91, 0.93] | 1.84 [1.83, 1.85] | 0.50 [0.50, 0.50] | 0.85 [0.84, 0.86] | 0.80 [0.79, 0.81] | 1.00 [1.00, 1.00] | 2.00 [2.00, 2.00] | 0.20 [0.20, 0.20] | 0.82 [0.81, 0.83] | 0.99 [0.99, 0.99] |
set retrieval -- Return a set of genes that belong with a query. Hidden: part of a gene set, to be returned among all other genes.
| # | Strategy | Grade | Precision | Recall | F1 | Jaccard index | Matthews correlation (MCC) | Fold enrichment | Genes returned |
|---|---|---|---|---|---|---|---|---|---|
| 02 | Find the map where your gene list is one cluster (UMAP + HDBSCAN) | reliable | 0.16 [0.07, 0.26] | 0.43 [0.18, 0.68] | 0.23 [0.10, 0.37] | 0.14 [0.06, 0.23] | 0.25 [0.10, 0.40] | 10.9 [6.8, 16.0] | 102.5 [53.1, 139.6] |
cluster recovery -- Find clusters that correspond to a label nobody showed them. Hidden: a share of a label, scored against clusters chosen on the rest.
| # | Strategy | Grade | Weighted F1 | Weighted precision | Weighted recall | Adjusted Rand index (ARI) | Normalised mutual information (NMI) | Homogeneity | Completeness | Unclustered share |
|---|---|---|---|---|---|---|---|---|---|---|
| 01 | Hold out a category and search for a map that finds it (UMAP + HDBSCAN) | weak | 0.57 [0.25, 0.84] | 0.56 [0.36, 0.79] | 0.75 [0.31, 0.99] | 0.15 [0.01, 0.33] | 0.24 [0.05, 0.47] | 0.30 [0.04, 0.69] | 0.52 [0.29, 0.85] | 0.20 [0.00, 0.55] |
| 04 | Keep only the modules that survive the whole walk (UMAP + HDBSCAN co-clustering) | weak | 0.52 [0.23, 0.81] | 0.61 [0.38, 0.85] | 0.57 [0.20, 0.95] | 0.15 [0.02, 0.36] | 0.25 [0.11, 0.41] | 0.34 [0.14, 0.54] | 0.38 [0.19, 0.63] | 0.19 [0.00, 0.39] |
| 15 | Find the communities several networks agree on (modularity + Louvain consensus) | weak | 0.19 [0.09, 0.27] | 0.50 [0.23, 0.79] | 0.16 [0.07, 0.24] | 0.02 [0.00, 0.05] | 0.13 [0.04, 0.23] | 0.28 [0.10, 0.50] | 0.11 [0.03, 0.19] | 0.00 [0.00, 0.00] |
values -- Predict a measured value for genes without one. Hidden: a share of a measurement's values.
| # | Strategy | Grade | Spearman rho | Pearson r | Kendall tau-b | R-squared (out of sample) | Normalised RMSE | Mean absolute error | Top-decile recall | Bottom-decile recall | Coverage |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 21 | Predict a measurement, and find the genes that defy the prediction (gradient boosting / ridge) | reliable | 0.52 [0.24, 0.84] | 0.54 [0.29, 0.84] | 0.40 [0.20, 0.65] | 0.02 [-0.77, 0.71] | 0.93 [0.54, 1.33] | 117.4 [0.3, 338.2] | 0.37 [0.22, 0.54] | 0.38 [0.20, 0.57] | 1.00 [1.00, 1.00] |
| 22 | Fill in what was never measured, and say where that is honest (soft-impute, low-rank SVD) | reliable | 0.68 [0.68, 0.68] | 0.70 [0.69, 0.70] | 0.51 [0.51, 0.52] | 0.46 [0.44, 0.47] | 0.74 [0.73, 0.75] | -- | 0.48 [0.46, 0.50] | 0.40 [0.37, 0.43] | 1.00 [1.00, 1.00] |
| 23 | Find what matters more in one condition, and why (residual + gradient boosting / ridge) | reliable | 0.64 [0.64, 0.64] | 0.64 [0.63, 0.64] | 0.46 [0.46, 0.46] | 0.40 [0.40, 0.41] | 0.77 [0.77, 0.78] | 0.13 [0.13, 0.14] | 0.40 [0.36, 0.44] | 0.41 [0.40, 0.42] | 1.00 [1.00, 1.00] |
| 29 | Carry what one parasite shows to the other (orthogroup mapping) | reliable | 0.33 [0.32, 0.34] | 0.32 [0.31, 0.34] | 0.22 [0.22, 0.23] | -87.4 [-90.0, -84.7] | 9.40 [9.26, 9.54] | 3.01 [2.99, 3.03] | 0.14 [0.11, 0.17] | 0.13 [0.09, 0.17] | 1.00 [1.00, 1.00] |
| 39 | Predict a value with an interval that holds (gradient boosting / ridge + split conformal) | reliable | 0.49 [0.19, 0.83] | 0.51 [0.25, 0.84] | 0.37 [0.15, 0.64] | 0.04 [-0.71, 0.70] | 0.93 [0.55, 1.31] | 131.3 [0.3, 368.4] | 0.36 [0.22, 0.53] | 0.36 [0.20, 0.56] | 1.00 [1.00, 1.00] |
replication -- Make findings that hold beyond the genes they were made on. Hidden: half of the genes, on which first-half findings are checked.
| # | Strategy | Grade | Replication rate | Findings made | Findings replicated | Replication by chance | Replication lift |
|---|---|---|---|---|---|---|---|
| 05 | Tune a map without labels, then read what it encodes (UMAP + HDBSCAN, chi-square / Kruskal-Wallis) | reliable | 0.96 [0.91, 1.00] | 55.5 [54.0, 57.0] | 53.0 [52.0, 54.0] | 0.06 [0.06, 0.06] | 15.7 [14.4, 16.9] |
| 26 | Find categories that split in two on another measurement (UMAP + HDBSCAN) | untestable | -- | -- | -- | -- | -- |
| 27 | Find kinds of gene defined by two labels at once (UMAP + HDBSCAN) | untestable | -- | -- | -- | -- | -- |
Techniques
What each strategy's method is built from. The Strategies tab shows the same explanations beside each strategy; strategies.techniques() returns this table.
| Technique | Kind | What it does | Why a strategy uses it | Used by |
|---|---|---|---|---|
| UMAP | embedding | Uniform Manifold Approximation and Projection: places every gene in three dimensions so that genes with similar measurements stay close, preserving local neighbourhoods over global distances. nneighbors sets how local, mindist how tightly points may pack. | Makes the combined measurements into a map whose structure can be clustered and looked at; the map is only ever built from columns the held-out label's closure permits. | 01, 02, 03, 04, 05, 08, 09, 26, 27 |
| HDBSCAN | clustering | Hierarchical density-based clustering: finds dense groups of any shape and leaves sparse genes unclustered (noise) instead of forcing them into a group. minclustersize is the smallest group it reports; 'eom' keeps the most persistent clusters, 'leaf' the finest. | Clusters a map without being told how many clusters there are, and admits when a gene belongs nowhere. | 01, 02, 04, 05, 09, 26, 27 |
| Settings walk | clustering | Builds and clusters the map under every combination of settings on a grid (and optionally each family of measurements alone), and keeps the combination that scores best on labels visible to it. | No single map setting is right for every question; the walk searches them, and the self-test then scores the chosen one on labels the choice never saw. | 01, 02, 04 |
| Co-association consensus | clustering | Counts, for every pair of genes, the share of clusterings that put them together, then joins genes by average linkage on one minus that share. | Keeps only structure that survives the arbitrary choices a map requires. | 04 |
| Label-free map quality | clustering | Scores a clustering by how much of the proteome it places and how evenly, without looking at any label. | Lets a map be tuned without the labels it will later be tested against. | 05 |
| k-nearest neighbours (kNN) | neighbours | Finds the k genes most similar to a gene (Euclidean distance on standardised measurements, or on map coordinates) and lets their labels vote, each weighted by 1 / distance. | The plainest guilt by association: no model, no clustering, and every call can be traced to the genes that made it. | 03, 06, 07, 08, 10, 12, 30, 31, 32, 35, 38 |
| kNN graph | neighbours | Links every gene to its k nearest neighbours in measurement space, symmetrised and degree-normalised, so a table can be walked like a network. | Lets measurement similarity and measured networks be combined in one graph. | 25, 33 |
| Network neighbour vote | networks | Calls a gene by the labels of its direct partners in one or more edge layers, each partner weighted by the edge's weight (contact counts, correlation strength). Abstains where a gene has no labelled partner. | Uses measured relationships -- crosslinks, pulldowns, co-expression -- directly as evidence. | 10, 12, 13, 14, 31, 32, 38 |
| Random walk with restart | networks | Diffuses from seed genes along the network: at each step a walker moves to a neighbour or, with the restart probability, jumps back to a seed. The stationary visiting frequency scores every gene. | Reaches genes several steps from the seeds while still favouring close ones; restart sets how far the influence spreads. | 11, 25 |
| Greedy modularity communities | networks | Merges groups of genes greedily while the network's modularity -- edges inside groups beyond what degree alone predicts -- increases. The resolution parameter sets the community size. | Finds communities in each layer separately, before any layer is allowed to dominate. | 15 |
| Louvain consensus | networks | Joins genes that a sufficient share of the layers put together, and splits the joined graph into communities by Louvain modularity optimisation. | Keeps groupings several independent kinds of evidence agree on without chaining everything into one component. | 15 |
| Structural similarity (TM-score) | networks | Foldseek TM-score between predicted structures: 0 to 1, above about 0.5 usually the same fold, regardless of sequence identity. | Fold is conserved long after sequence is not, so it can annotate proteins no sequence search reaches. | 14 |
| Chance-weighted voting | networks | Scores each evidence source on an inner holdout of the known labels and weights its vote by its accuracy minus its own chance level; a source that only guesses the commonest class gets no vote. | Lets the evidence that has earned trust for this label count most, and says which that is. | 12 |
| Agreement of independent callers | networks | Calls a gene only where at least a set number of callers built on different evidence give the same label. | Trades reach for precision: independent errors rarely agree. | 31, 32 |
| Logistic regression | models | A linear model of the log-odds of a class (or of an edge) from standardised features, with classes balanced and L2 regularisation (C: smaller is stronger). | Interpretable -- each feature has a weight -- and hard to overfit; the baseline any fancier model must beat. | 16, 19, 20, 31, 32, 33, 34, 35, 36, 38 |
| Gradient boosting | models | An ensemble of shallow decision trees, each fitted to the errors of the ones before (scikit-learn's histogram gradient boosting). Handles missing values and non-linear effects. | Captures interactions and thresholds a linear model cannot, at the cost of interpretability. | 21, 23, 39 |
| Ridge regression | models | Linear regression with an L2 penalty that shrinks every coefficient towards zero. | A stable linear alternative to boosting, with missing values imputed first. | 21, 23, 39 |
| Positive-unlabelled bagging | models | Trains many classifiers, each on the positives against a small random draw of other genes, and scores every gene only by the models that did not train on it (out of bag). | When all you have are positives, treating every other gene as a negative teaches the model to reject the hidden positives you want to find; small random draws rarely contain them. | 20 |
| Soft-impute (low-rank SVD) | models | Completes a matrix with missing entries by repeatedly filling them from a truncated singular value decomposition of the current estimate, at a chosen rank. | Borrows strength across all measurements at once; columns are checked on hidden entries before any imputed value is trusted. | 22 |
| Residualisation on a baseline | models | Regresses the condition's measurement on its baseline (both rank-scaled) and keeps the residual: the part the baseline does not explain. | Separates what matters only in one condition from what matters everywhere. | 23 |
| Spectral embedding | models | Node vectors from a truncated SVD of a layer's adjacency matrix; a pair is described by the elementwise product of its two vectors. | Captures network position beyond direct evidence; kept only where it beats the interpretable baseline on a degree-matched null. | 34 |
| Shared partners (triadic closure) | networks | Counts the partners two genes share in a layer; two genes linked to the same partners are often linked themselves. | The classic link-prediction feature, used alongside support from other layers. | 16 |
| Degree-matched non-pairs | statistics | Negative pairs drawn so each end has a degree similar to the positives' ends. | Against random non-pairs, a model scores highly just by knowing which genes are well connected; matched non-pairs leave only what the evidence says about the pair. | 16, 33, 34 |
| Attention-corrected co-mention | statistics | Replaces each pair's co-mention count by its residual over what the two genes' own publication counts predict (log2 observed / expected). | Stops famous genes being linked just because they are famous. | 17 |
| Multi-layer support count | networks | Counts, for each gene pair, how many independent measurement layers link it. | Independent kinds of evidence agreeing on a pair is stronger than any one of them. | 18 |
| Profile correlation | statistics | One minus the correlation of two genes' profiles across every permitted measurement both have. | Measures how differently two paralogs behave, whatever they are called. | 28 |
| Orthogroup mapping | statistics | Aggregates a measurement or label over each orthogroup in the other species (median, or the commonest label) and learns its relation to this species' values on genes measured in both. | Carries evidence across species without merging the two tables. | 29 |
| Hypergeometric enrichment | statistics | The probability of drawing at least the observed number of a label's genes into a group of that size by chance (one-sided Fisher's exact test). | The standard test that a cluster or gene list holds more of a label than chance. | 09, 24, 26, 27 |
| Rank-sum test | statistics | Mann-Whitney U: whether a measurement's values are higher (or lower) in one group of genes than in the rest, using ranks only. | Tests a measurement against a gene list without assuming a distribution. | 24 |
| Chi-square test (Cramer's V) | statistics | Tests whether a categorical label is distributed differently across clusters; Cramer's V is the effect size, 0 to 1. | Asks what a map built from one kind of measurement encodes about another kind. | 05 |
| Kruskal-Wallis test (eta-squared) | statistics | Rank-based test that a measurement differs between clusters; eta-squared is the effect size. | The numeric counterpart of the chi-square test above. | 05 |
| Benjamini-Hochberg FDR | statistics | Adjusts many p-values together so the expected share of false discoveries among the significant ones stays below the chosen rate (q). | Every strategy that tests many clusters, labels or measurements at once corrects for it. | 05, 09, 24, 26, 27 |
| Orthogroup-grouped cross-validation | statistics | Splits genes into folds by whole orthogroups, so a gene is never predicted by a model that saw it or its paralog. | Paralogs share measurements; without grouping, a model can look accurate by recognising family members. | 06, 19, 21, 23, 38, 39 |
| Evidence ablation | statistics | Scores each kind of measurement alone and the combination without it, out of fold. | Says which experiments carry which biology and which could be dropped without loss. | 06 |
| Bimodal split | statistics | Divides a cluster's values on a second measurement into two modes and requires them to be at least a set distance apart. | Finds categories whose members fall into two kinds on another measurement. | 26 |
| Split conformal prediction | models | Calibrates any model on genes it did not train on: the errors (or the scores of the true labels) on those genes set a threshold so that the returned sets or intervals contain the truth for at least 1 - alpha of new genes. | Turns scores into a stated error rate that holds whatever the model, and says which genes the data cannot decide. | 35, 39 |
| Simplified graph convolution | networks | Multiplies the measurement matrix by the degree-normalised network (with self-loops) once per step, so each gene gains features that average its neighbours' measurements. | Lets a linear model learn from a gene's network neighbourhood -- the core of a graph neural network -- while staying transparent and leak-free. | 36 |
| Random forest | models | Hundreds of decision trees, each grown on a bootstrap sample of genes with a random subset of measurements at every split, voting on the class; classes are re-weighted to balance. | Finds thresholds and interactions a linear model cannot, and is robust to scale and outliers. | 37 |
| Permutation importance | statistics | Shuffles one measurement among held-out genes and records how much balanced accuracy drops; repeated to give a mean and a spread. | Measures what a model actually relies on for new genes, unlike impurity importance, which rewards measurements with many distinct values. | 37 |
| Stacking | models | Trains a meta-model on the out-of-fold predictions of several base models, so the combination is learned on genes no base model trained on. | Learns per label and per class how far to trust each kind of evidence, without rewarding a base model for memorising its training genes. | 38 |