Enamine screening collection · June 2026
A reference sheet for the model builds and assessments that use this library, including the pharmacophoric similarity surrogate, which is trained and tested on it. Two questions are asked here: what the compounds look like, and how much of the collection is genuinely distinct chemistry rather than variations on the same theme.
Where this collection came from
Enamine is a commercial supplier of screening compounds and building blocks. It markets this as the world's largest screening collection, and the point it presses hardest is physical availability: these are compounds synthesised in its own labs and held in stock, supplied as neat samples of roughly 50 to 150 milligrams or as 10 millimolar DMSO solutions, not a virtual enumeration. Around 3.12 million are held in US stock with two to three day delivery inside the USA, and a further 1.08 million in Ukraine stock. Roughly 225,000 new compounds are added each year, and material is quality controlled by LCMS or proton NMR to at least 90 percent purity.
That distinction matters for how this page should be read. Everything counted here is real, orderable material. Enamine separately publishes much larger virtual spaces, most notably REAL Space, enumerated by applying validated reactions to in-stock building blocks. None of that is included here. The source is the June 2026 SDF release, which currently states 4,774,670 compounds.
What the compounds look like
| Property | 5th pct | Median | 95th pct |
|---|---|---|---|
| Molecular weight | 255 | 344 | 465 |
| Bonds | 19 | 26 | 36 |
| Rings | 2 | 3 | 5 |
| Aromatic rings | 1 | 2 | 4 |
| Rotatable bonds | 2 | 5 | 8 |
| H-bond donors | 0 | 1 | 3 |
| H-bond acceptors | 2 | 4 | 7 |
| Polar surface area | 38 | 71 | 114 |
| cLogP | 0.7 | 2.7 | 4.7 |
| Fraction sp3 carbon | 0.09 | 0.35 | 0.71 |
Properties from a random sample of 200,000 compounds. Heavy atom counts are from all 4,612,044. Bond counts run from 8 to 50.
This is a lead-like library rather than a fragment or a biologics-adjacent one. The median compound weighs 344, carries three rings of which two are aromatic, five rotatable bonds, and sits at cLogP 2.7. Only 8 percent carry defined stereochemistry, so the collection is predominantly flat, achiral scaffolding decorated with substituents.
Scaffold diversity
Reducing each molecule to its Bemis-Murcko scaffold, the ring systems and the linkers between them with all side chains stripped, gives a direct read on how many distinct structural frameworks the library is built from.
The framework diversity is high. Roughly seven in ten molecules bring a scaffold no other molecule in the sample shares, and no single framework dominates: the commonest is plain benzene at 2.09 percent, and the top ten together account for only 4.0 percent.
| Commonest scaffolds | Share, percent |
|---|---|
c1ccccc1 | 2.09 |
O=C(Nc1ccccc1)c1ccccc1 | 0.39 |
O=S(=O)(Nc1ccccc1)c1ccccc1 | 0.37 |
c1ccncc1 | 0.27 |
O=C(NCc1ccccc1)c1ccccc1 | 0.19 |
O=C(c1ccccc1)N1CCCCC1 | 0.14 |
How much is distinct chemistry
Scaffold counts describe frameworks. They say nothing about how close any two whole molecules are. For that, every compound was compared against every other compound in the collection, with no sampling, keeping one number per compound: the highest Tanimoto similarity it reaches against any other, itself excluded. Compounds are represented by 2,048-bit Morgan fingerprints at radius 2. That is 10.6 trillion pairwise comparisons, computed exhaustively over 66.8 hours on six cores. Nothing here is sampled or estimated.
| Threshold | Compounds | Percent |
|---|---|---|
| Has a neighbour at 0.95 or closer | 115,471 | 2.5 |
| Has a neighbour at 0.90 or closer | 199,333 | 4.3 |
| Has a neighbour at 0.80 or closer | 985,227 | 21.4 |
| Has a neighbour at 0.70 or closer | 2,532,866 | 54.9 |
| Has a neighbour at 0.60 or closer | 3,894,561 | 84.4 |
| Has a neighbour at 0.50 or closer | 4,497,684 | 97.5 |
The spread is narrow. The tenth percentile sits at 0.569 and the ninetieth at 0.844, so almost the whole library lives in a band between roughly 0.57 and 0.84. The single most isolated compound in the collection still reaches 0.12 against its closest relative, which means there is no truly orphaned chemistry in here at all.
A median of 0.714 is high. Picking a compound and then picking its nearest neighbour gets you two molecules that share most of their substructural features. The sharper number is at the top end: 64,473 compounds, 1.4 percent of the library, have at least one other compound with an identical 2D fingerprint. Those are not near neighbours; at this resolution they are indistinguishable. A further 4.3 percent sit at 0.90 or above. Deduplicating or down-weighting that fraction costs almost no chemical coverage.
Reading the two results together
These two measurements pull in opposite directions, and both are true. At the framework level the library is broad: 679 distinct scaffolds per thousand molecules, most of them appearing once. At the whole-molecule level it is dense: the median compound has a neighbour at 0.714 and 55 percent sit within 0.70 of something else.
The resolution is that the collection is wide in frameworks and deep in analogs around them. Many scaffolds are present, and each is decorated many times over with closely related substituent patterns. For model building this matters directly. A random split will place near-identical analogs on both sides and report an optimistic number, so scaffold-aware splitting is the honest choice. For screening, uniform sampling spends much of its budget re-asking the same structural question.
Nearest-neighbour similarity computed exhaustively over all 4,612,044
compounds. Properties and scaffolds from a random sample of 200,000; heavy atom counts
from the full set. Figures on this page are read at build time from
redundancy_report.json and collection_profile.json.