Every model in the family measured on the same held-out
molecules, as the training set grew from 1,000,000 to 2,114,521.
Why these numbers are comparable. Both training
scripts reserve chunks 0 to 19 of the screening collection, 9,975 molecules, and
train on everything after. Every model in the table was therefore scored on the
identical held-out set. The two scripts name the metric differently,
test_median_mcc and
val_median_mcc, but it is one measurement on one set.
2.1x
Training set growth across the family
0.893
Agreement at 2,114,521 molecules
58,039
Peptide ensembles in the newest model
6
Models in the family
Agreement against training set size
Brown points are collection-only models, green points include
peptides. Note the vertical scale.
Model
Training molecules
Collection
Peptides
Agreement
PharmCast-S 1.0M
1,000,000
1,000,000
none
0.888
PharmCast-S 1.2M
1,192,657
1,192,657
none
0.889
PharmCast-S 1.36M
1,357,547
1,357,547
none
0.891
PharmCast-SP v1
1,368,086
1,357,547
10,539
0.890
PharmCast-SP v2
1,724,833
1,688,794
36,039
0.891
PharmCast-SP v3
2,114,521
2,056,482
58,039
0.893
What the held-out set does and does not cover.
Chunks 0 to 19 are screening collection molecules only, so this metric measures
catalogue chemistry. Peptide performance is not in it and is measured separately
in the PharmCast-SP report. The screening collection is still being
fingerprinted and the peptide corpus is still being extracted, so the training
set continues to grow.