Screening Collection Profile

What is in a 4.6 million compound library, and how much of it is genuinely distinct chemistry

Enamine screening collection · June 2026

A reference sheet for the model builds and assessments that use this library, including the pharmacophoric similarity surrogate, which is trained and tested on it. Two questions are asked here: what the compounds look like, and how much of the collection is genuinely distinct chemistry rather than variations on the same theme.

Where this collection came from

Enamine is a commercial supplier of screening compounds and building blocks. It markets this as the world's largest screening collection, and the point it presses hardest is physical availability: these are compounds synthesised in its own labs and held in stock, supplied as neat samples of roughly 50 to 150 milligrams or as 10 millimolar DMSO solutions, not a virtual enumeration. Around 3.12 million are held in US stock with two to three day delivery inside the USA, and a further 1.08 million in Ukraine stock. Roughly 225,000 new compounds are added each year, and material is quality controlled by LCMS or proton NMR to at least 90 percent purity.

That distinction matters for how this page should be read. Everything counted here is real, orderable material. Enamine separately publishes much larger virtual spaces, most notably REAL Space, enumerated by applying validated reactions to in-stock building blocks. None of that is included here. The source is the June 2026 SDF release, which currently states 4,774,670 compounds.

This is a filtered subset. Property filters were applied on ingest, keeping compounds up to 600 molecular weight, at most 10 rotatable bonds, and no more than one Lipinski violation. 4,612,044 of the 4,774,670 survived, so about 3 percent of the catalogue was set aside. Every number on this page describes that filtered set.

What the compounds look like

4,612,044compounds after filtering
344median molecular weight, range 116 to 599
24median heavy atoms, range 5 to 45
8%carry defined stereochemistry
Property5th pctMedian95th pct
Molecular weight255344465
Bonds192636
Rings235
Aromatic rings124
Rotatable bonds258
H-bond donors013
H-bond acceptors247
Polar surface area3871114
cLogP0.72.74.7
Fraction sp3 carbon0.090.350.71

Properties from a random sample of 200,000 compounds. Heavy atom counts are from all 4,612,044. Bond counts run from 8 to 50.

This is a lead-like library rather than a fragment or a biologics-adjacent one. The median compound weighs 344, carries three rings of which two are aromatic, five rotatable bonds, and sits at cLogP 2.7. Only 8 percent carry defined stereochemistry, so the collection is predominantly flat, achiral scaffolding decorated with substituents.

Scaffold diversity

Reducing each molecule to its Bemis-Murcko scaffold, the ring systems and the linkers between them with all side chains stripped, gives a direct read on how many distinct structural frameworks the library is built from.

135,768distinct scaffolds in 200,000 molecules
679distinct scaffolds per 1,000 molecules
88%scaffolds appearing exactly once
4.0%share held by the ten commonest scaffolds

The framework diversity is high. Roughly seven in ten molecules bring a scaffold no other molecule in the sample shares, and no single framework dominates: the commonest is plain benzene at 2.09 percent, and the top ten together account for only 4.0 percent.

Commonest scaffoldsShare, percent
c1ccccc12.09
O=C(Nc1ccccc1)c1ccccc10.39
O=S(=O)(Nc1ccccc1)c1ccccc10.37
c1ccncc10.27
O=C(NCc1ccccc1)c1ccccc10.19
O=C(c1ccccc1)N1CCCCC10.14

How much is distinct chemistry

Scaffold counts describe frameworks. They say nothing about how close any two whole molecules are. For that, every compound was compared against every other compound in the collection, with no sampling, keeping one number per compound: the highest Tanimoto similarity it reaches against any other, itself excluded. Compounds are represented by 2,048-bit Morgan fingerprints at radius 2. That is 10.6 trillion pairwise comparisons, computed exhaustively over 66.8 hours on six cores. Nothing here is sampled or estimated.

0.714median similarity to the closest other compound
55%have a neighbour at 0.70 or closer
64,473have a 2D-identical twin
0.569 to 0.844where the middle 80 percent sits
0166k332k median 0.714 0.00.20.40.60.81.0 Highest Tanimoto similarity to any other compound Compounds
One value per compound: the highest Tanimoto it reaches against any other compound in the collection, itself excluded. Pale bars are compounds below 0.70, the ones carrying genuinely distinct chemistry. The darkest bars at the right are compounds at 0.90 or above, near copies of something already present.
ThresholdCompoundsPercent
Has a neighbour at 0.95 or closer115,4712.5
Has a neighbour at 0.90 or closer199,3334.3
Has a neighbour at 0.80 or closer985,22721.4
Has a neighbour at 0.70 or closer2,532,86654.9
Has a neighbour at 0.60 or closer3,894,56184.4
Has a neighbour at 0.50 or closer4,497,68497.5

The spread is narrow. The tenth percentile sits at 0.569 and the ninetieth at 0.844, so almost the whole library lives in a band between roughly 0.57 and 0.84. The single most isolated compound in the collection still reaches 0.12 against its closest relative, which means there is no truly orphaned chemistry in here at all.

A median of 0.714 is high. Picking a compound and then picking its nearest neighbour gets you two molecules that share most of their substructural features. The sharper number is at the top end: 64,473 compounds, 1.4 percent of the library, have at least one other compound with an identical 2D fingerprint. Those are not near neighbours; at this resolution they are indistinguishable. A further 4.3 percent sit at 0.90 or above. Deduplicating or down-weighting that fraction costs almost no chemical coverage.

Reading the two results together

These two measurements pull in opposite directions, and both are true. At the framework level the library is broad: 679 distinct scaffolds per thousand molecules, most of them appearing once. At the whole-molecule level it is dense: the median compound has a neighbour at 0.714 and 55 percent sit within 0.70 of something else.

The resolution is that the collection is wide in frameworks and deep in analogs around them. Many scaffolds are present, and each is decorated many times over with closely related substituent patterns. For model building this matters directly. A random split will place near-identical analogs on both sides and report an optimistic number, so scaffold-aware splitting is the honest choice. For screening, uniform sampling spends much of its budget re-asking the same structural question.

Everything here measures 2D structural similarity. Two compounds that look alike on paper can present different three-dimensional pharmacophores, and two that look unrelated can present similar ones. This describes the catalogue, not what the molecules do. It is also why the surrogate's independence from 2D similarity matters: a library that is dense in two dimensions is not necessarily dense in pharmacophore space.

Nearest-neighbour similarity computed exhaustively over all 4,612,044 compounds. Properties and scaffolds from a random sample of 200,000; heavy atom counts from the full set. Figures on this page are read at build time from redundancy_report.json and collection_profile.json.