AI Pioneer | Drug Discovery Expert | Creative Innovator | Serial Inventor
CEO of Eidogen-Sertanty with over 40 years at the intersection of AI, chemistry, biology, and drug discovery. Host of the Renaissance Circle podcast. Builder of Study with Hannah, AI-Steve, Food Health Scan, and AI-Dad.
Several active projects across drug discovery, AI study tools, personal productivity, health technology, and creative tools. Each one solves a real problem in my own life or work.
Drag to rotate
Forward screening asks which of many ligands bind one target. Reverse Screen asks the opposite, which is the only question available for a molecule a generative method has just proposed.
Docking one compound into every characterized binding site answers it directly and takes 40.5 hours. Instead every co-crystal ligand in the Protein Data Bank is indexed by the three-dimensional pharmacophore fingerprint it presents, predicted from flat structure by PharmCast, together with the UniProt accessions it was solved against. A query is fingerprinted in 4 milliseconds and compared against all 27,797 indexed ligands, covering 28,579 target sites, in 40 milliseconds. The proteins its nearest neighbors were crystallized against come back as the candidate pool, and AutoDock Vina docks only into those.
Across 3,000 held-out molecules, pooling the five most similar indexed ligands gives 21.1 candidate proteins and contains the molecule's own known target 48.8 percent of the time; pooling twenty-five gives 104.5 proteins and 60.8 percent.
Loading the three dimensional view
acceptor ×2 donor ×3 hydrophobic ×4
Nirmatrelvir, the active ingredient in Paxlovid, is a designed peptidomimetic built against the peptide substrate of the SARS-CoV-2 main protease. Search a library of real protein loops by the features they present, and a five residue loop from an unrelated crystal structure comes back presenting the same pharmacophore at 0.913, while sharing only 0.167 of the drug's two dimensional structure.
The drug. Nirmatrelvir, the Paxlovid active ingredient, a designed peptidomimetic covalent inhibitor built against the enzyme's peptide substrate. Ligand ZGW in PDB entry 8GFU.
The loop. Residues 65 to 69 of chain A in PDB entry 2JHX, a RhoGDI mutant crystallized for an entirely different purpose. Sequence AMVPN, relative solvent accessibility 0.679, resolution 1.60 Angstrom. Heavy dark sticks are the drug; light sticks are the loop in the conformation it adopts in 2JHX, not a generated conformer.
What the numbers say. Against a generated low energy conformer of the same sequence the fit is 87.3 with 10 shared features. The real loop gives 79.6 with 9, a change of -7.7, so the conformation nature holds the loop in costs a little of the match rather than making it.
The point of the exercise is direction. A peptidomimetic is normally designed from a peptide toward a drug, and this asks the question backward: given a drug, which real protein loops already present its pharmacophore? Loops come from deposited crystal structures, so each one is a conformation some protein was observed to hold, and the comparison runs on the same PharmPrint three dimensional pharmacophore fingerprints used elsewhere on this site.
That backward question is worth asking because the answer works both ways. Where a peptide reproduces what a drug presents, it is a starting point for rules of peptide mimicry in either direction: what a small molecule has to carry to stand in for a peptide, and equally what a short peptide has to carry to stand in for a small molecule. This case is one result from the wider reverse peptide mimetics study.
What a kinase ligand is actually doing in its pocket, measured from the crystal structure: every hydrogen bond, halogen bond, salt bridge and stacking contact, drawn with its distance and angle, beside the compound's measured potency in the Kinase Knowledgebase. A new structure on every visit.
There are 3,961 figures across 3,893 PDB entries, and the page draws a different one each time it loads. Bound is not necessarily potent, a companion figure, shows one pocket's ligands spanning a millionfold in potency.

Every co-crystal ligand of a protein superposed into one frame, with the measured potency beside each. Within a single pocket the numbers span up to six orders of magnitude, so a solved structure tells you a compound binds and very little about how well.
A three-dimensional pharmacophore fingerprint records the binding features a molecule can present. It is a description of a hand in search of a glove. The descriptor has stayed a niche tool for thirty years because its cost is dominated by conformer generation, so we removed the conformational stage: PharmCast predicts all 10,549 bits of the ensemble fingerprint directly from a SMILES string.
The two routes to a fingerprint, worked through on saquinavir and indinavir, two HIV-1 protease inhibitors of unrelated scaffold. The conventional route generates a conformer ensemble and runs the reference calculation over it; PharmCast predicts the same ensemble record from the two-dimensional structure and skips the ensemble entirely. The reference calculation puts the pair at a pharmacophore Tanimoto of 0.841, against a two-dimensional Morgan Tanimoto of 0.303. The predicted value is shown in panel E.
The descriptor is the one described in the PharmPrint work: each of 10,549 bits is one three-point pharmacophore, three typed features and the three binned distances between them, and a molecule sets a bit when any accessible conformation presents that triangle. Because the bit records that a molecule can present that triangle in some accessible conformation, without asserting that it presents it in any particular pose, it is closer to a property of molecular constitution than of any single geometry, and constitution is what a two-dimensional structure encodes. It is a weaker claim than predicting a conformer, which is why it is tractable where conformer prediction remains hard. The network is deliberately simple: a 2,048 bit Morgan fingerprint of radius 2 alongside 11 whole molecule descriptors, 2,059 inputs in all, through two hidden layers of 1,024 and 512 units to all 10,549 bits, for 8,045,877 parameters trained by minimizing binary cross entropy.
Timed on catalog compounds, the conventional route requires 2.86 s per molecule, of which 2.82 s is conformer generation and 0.039 s is the fingerprint calculation itself. Since the bit calculation is only 1.7% of the total, accelerating it changes nothing; anything that materially moves the economics has to remove the conformational stage. One PharmCast fingerprint costs 0.288 ms in a library-scale batch, and starting from two SMILES strings we obtain a similarity in 0.584 ms against 5.71 s. Screening the complete 4.65 million compound Enamine collection against a reference takes 29.3 minutes, measured end to end at 2,644 compounds per second on six worker processes, against an extrapolated 3,697 core hours by the conventional route.
All timings were measured on one machine, an Apple M3 Ultra with 28 CPU cores, a 60 core GPU and 256 GB of unified memory, running macOS 26.5.1 with PyTorch 2.9.1 on CPU. The GPU was not used, and both routes were timed on the same hardware.
Version 10 was trained on 5,887,229 molecules drawn from a screening collection, activity-backed ChEMBL compounds from 142 to 1000 Da, and peptide loops excised from crystal structures, with a molecular-weight-stratified one percent held back for early stopping. It is evaluated on three chemically distinct populations, all unseen during training: 155,648 real, purchasable Enamine screening collection compounds that the training-set ingest filter excluded, checked by canonical SMILES against all 4,617,292 screening collection structures; 139,700 activity-backed ChEMBL molecules not present in the training set; and 13,500 peptide loops reserved for testing, 9.9% of the full peptide set. Median fingerprint error, Pearson r and pairwise ranking accuracy are 0.008, 0.980 and 0.936 for screening collection chemistry; 0.016, 0.984 and 0.952 for loop peptides; and 0.027, 0.936 and 0.889 for activity-backed ChEMBL compounds. The reference calculation reproduces itself at an error of 0.006 and r of 0.995. Taken one molecule at a time, the median Matthews correlation coefficient is 0.881 on the screening collection, 0.914 on loop peptides and 0.860 on ChEMBL.
Where it is weakest. ChEMBL is the weakest population, with a median error of 0.027 and roughly one pair in four falling outside 0.05, and agreement declines at the upper end of the mass range where the source sets are thinnest. The model predicts the ORed ensemble fingerprint, so where an application requires the matching geometry of a particular conformer, that must come from the reference calculation. Predicted similarity inflates under optimization pressure, since a search that maximizes it selects the molecules the model overestimates; we therefore recommend the surrogate for ranking, the reference calculation for final decisions, and a broad survivor set for rescoring.
PharmCast is joint work with Malcolm J. McGregor, co-author of the original pharmacophore fingerprinting papers the descriptor comes from, and is described in a paper posted to bioRxiv on 7 September 2026, doi 10.64898/2026.09.02.748999. The code, model utilities, documentation and trained weights are public at github.com/smuskal/pharmcast under Apache-2.0; the full method, benchmarks and applicability domain are on pharmcast.ai, and the weights and their SHA-256 digests on pharmcast.ai/models.
Take a marketed drug. Design a molecule that presents the same three-dimensional arrangement of binding features on a scaffold that looks nothing like it, reachable in a few steps from catalog building blocks. The campaign below delivers two of them.
See the CampaignChIP searches reaction routes rather than molecules. A gene is a synthesis route: it names the transforms and the catalog blocks fed into them, is enumerated into the products it would make, and those products are built in 3D and fingerprinted against the reference with PharmPrint-style 3D pharmacophore fingerprints. Because every candidate is by construction the product of real reactions on orderable material, the output is synthesizable by design. The two easy alternatives both fail here: screening a catalog by 2D similarity returns analogs, which is the thing being avoided, and generative models return molecules that may not be makeable.
A campaign runs in three phases. Phase 1 searches reaction sequences, asking which sequence of transforms can reach the reference pharmacophore at all, and stops when the population collapses onto one. Phase 2 freezes that sequence and searches the catalog blocks that fill its slots, so a gene becomes one purchasable block per slot and enumerates to exactly one molecule; this is the phase that delivers the purchasable design. Phase 3 keeps route and slots fixed and edits the parts of each block the reaction does not touch, asking how much better the design would be if the catalog were slightly richer. Phase 3 products are virtual satellites of the Phase 2 design rather than orderable compounds.
The reference is Orforglipron, Lilly's oral GLP-1 receptor agonist approved in April 2026. Fitness was pharmacophore similarity and nothing else; 2D Morgan similarity was recorded for every product but never entered selection. Two campaigns were run against it and each converged on a different reaction sequence, so the report delivers two independent designs. Design 1 was scored by the reference conformer calculation. Design 2 came from a separate campaign scored by PharmCast, the surrogate that predicts the same fingerprint from 2D structure, which is what made a second search of that size affordable.
The two designs, side by side
| Design 1 | Design 2 | |
|---|---|---|
| Pharmacophore similarity to Orforglipron | 0.838 | 0.810 |
| 2D Morgan similarity | 0.138 | 0.145 |
| Shape Tanimoto in the overlay | 0.438 | 0.495 |
| Best single conformer pair | 0.659 | 0.633 |
| Formula and weight | C35H32N2O9, 624.6 | C39H35BrF2N6O4, 769.6 |
| Route | 2 steps, 3 blocks | 3 steps, 4 blocks |
| Scored during the search by | reference calculation | the PharmCast surrogate |
Both designs do the thing the search was after: pharmacophore similarity high, 2D similarity at the floor, and that 2D value stayed low on its own without ever being penalized. Across design 1's 115 retained designs, pharmacophore similarity ranges from 0.824 to 0.838 while Morgan similarity stays between 0.081 and 0.174. These are not Orforglipron analogs; they are structurally distinct scaffolds presenting a similar arrangement of features in space.
Limitations. This is a computational simulation. No potency model and no target structure were used anywhere in either search, and nothing here predicts activity at the GLP-1 receptor or any other target. No molecule described has been synthesized and no assay has been run. A high pharmacophore Tanimoto means the candidate can present a similar arrangement of features across its accessible conformers; it does not mean the candidate adopts that conformation when bound, or that it binds at all. The routes were enumerated by reaction transforms and have not been reviewed by a synthetic chemist: no protecting group strategy or order of addition was considered, regiochemistry and stereochemistry were not checked, and no route has been run. Design 1 leaves two stereocenters undefined, so material made to that specification would be a mixture. Conformer ensembles inside the search loop are kept to 100 per molecule for cost, so ranking near the top of the list is not sharp and survivors should be rescored at full ensemble settings. Building block availability came from catalog identifiers and was not confirmed against live supplier stock, pricing or purity. Read the whole thing as a search result, not a biological claim. It is a different demonstration from the Kinase Foundation Model, where the objective was predicted potency.
Searchable medical school study cards with USMLE prep notes, medical images, and an AI study tutor.
Request AccessTens of thousands of medical-school flashcards and lecture notes across anatomy, physiology, pathology, pharmacology, and immunology, with keyword plus semantic image search. Hannah keeps the archive current throughout her studies; the public landing page introduces the system, while the study content is access-controlled.
Counts update from the Study with Hannah public archive feed.
My daily companion for research recall, journaling, and family knowledge capture. AI-Steve pulls in emails and attachments, calendar events, to-do items, iMessages, Mac Photos plus curated imports, chat sessions, Q&A, and Socratic pairs extracted from mail, embedding them into PostgreSQL so Claude can respond with grounded context. Search Content, sentiment, and health intelligence fuel daily reflections (including 1, 3, 5, 7, and 10-year lookbacks) plus end-of-week recaps that project the week ahead. Built entirely through natural-language coding (speaking into Wispr Flow driving agentic CLIs like Droid, Claude Code, Codex, Gemini CLI) and able to handle small coding or automation projects on demand, similar to how I built Toast apps and AI-Dad.
Features at a glance
An innovative application of RAG (Retrieval-Augmented Generation) technology to create an interactive AI assistant embodying 60+ years of intellectual property legal expertise and family wisdom. Built using natural language programming and Claude Code, this deeply personal project preserves my father's extensive knowledge in IP law alongside decades of family history and personal interactions, making his guidance on both legal and life matters accessible for future generations.
Legal Expertise
Family Wisdom
AI-Dad: Always here for you - Combining decades of legal expertise with heartfelt family wisdom
A CLIP (Contrastive Language-Image Pre-training) + pgvector-powered explorer for my photo archives. Nightly clustering keeps similar sets together so I can browse clusters, select many images at once, and annotate entire groups without touching each file. I can also annotate straight from similarity search, including sub-images and video frames, accelerating how visuals become structured context for AI-Steve’s RAG. Built by speaking English into Wispr Flow to orchestrate Droid, Claude Code, Codex, Gemini CLI, similar to Toast apps, AI-Dad, and AI-Steve.
Added into my AI-Steve infrastructure is the ability to auto-code projects in a domain-specific way using simple natural language project descriptions on top of Droid, Claude, and/or Codex. What brought it to the next level is using AI-Steve’s RAG system to wrap a project direction in my voice - imagine enabling all your coders to code given their own past projects and insights. It is akin to saying: build this new OS in the voice of Linus Torvalds.
Domain-specific guidance + RAG voice overlay drive code generation, review loops, and polished reporting.
Key elements: domain prompts, RAG context, peer review, self-healing retries, and packaged reports.
Designed to make project requests feel like they were built by the same voice that created the original system.
Food Health Scan turns meals into structured, reviewable nutrition data with photo capture, note capture, AI analysis, and fast edit flows. It is the practical product expression of the March 8, 2026 Renaissance Circle piece, Food Is Medicine. But First It Has to Become Data.
The existing walkthrough shows the current meal-to-data capture flow on the live app.
The newer V4 video adds another quick look at the updated Food Health Scan experience.
The diabetes-aware companion to Food Health Scan. AI carb and macro estimation paired with continuous glucose monitor (CGM) data, so every meal links directly to the glucose response it produced. Built for people with Type 1, Type 2, and prediabetes who want to see what their food is actually doing - and for the wider 88 to 93 percent of American adults living with at least one metabolic risk factor.
What you see on the right
Companion essay (May 8, 2026): The Metabolic Crisis Is Not a News Story on Substack or Medium.
This preview loads the live Food Showcase page, so updates on `foodhealthscan.com` show up here automatically. The showcase already includes TxD plates.
Daily Apple Health exports power a dedicated analytics pipeline inside AI-Steve. Correlation engines, lag analysis, and a sleep-concentration model surface the behaviors most tied to deep sleep and REM recovery. The resulting plots are stored as visual artifacts so they’re searchable and reviewable alongside photos and other memory assets.
Face → Health treats the face as a sensor, not a narrative. Daily images support both retrospective (last night) and predictive (tonight) sleep modeling with strict temporal alignment. Health-specific vision analysis, dual embeddings, and a materialized ML view make this a defensible, longitudinal wellness signal.
A real-world build story: I received a USB drive full of DICOM CT slices after a root canal, and instead of using a PC-only viewer, I described the problem to an AI coding assistant and walked away. Minutes later, I had a working browser-based viewer with 3D rotation, cross-sectional MPR views, and contrast controls.
Ask which, not how much. Two models trained directly on the comparison, so they rank compounds and kinases instead of predicting a potency.
Explore the ModelScreening is a prioritization problem: a project needs to know which compound to make next and which kinase a series is likely to hit, and an ordering does not require predicting a potency first. Version 2 puts the entire question into one input row and returns the probability that one side wins. A protein sequence is encoded by ESM2 into 480 numbers, a compound becomes a 1,024-bit Morgan count fingerprint plus 14 descriptors, and neither half needs a structure, a docked pose, or a binding-site definition. Everything is trained on the Kinase Knowledgebase and tested on ChEMBL comparisons the models never saw.
Version 2, the current release: two models, each trained directly on ordered comparisons.
Why the comparison lives inside the row: predicting a potency for each side and subtracting only works if both predictions share a scale, and per-target scores do not, so a difference between them reports assay scale as if it were selectivity. Training on the ordered pair removes the intermediate quantity altogether. The models rank; they do not estimate a potency, and there is no predicted IC50 to put in a table. Predictions are for research use and are not a substitute for measurement.
Both version 2 models also run locally and offline with nothing transmitted, which is the practical requirement when the compounds are unpublished. Version 1, the original pointwise pIC50 scorer this grew out of, remains published in full on the site. For the companion demonstration that optimizes pharmacophore similarity instead of predicted potency, see ChIP.
Developing machine learning models for predicting drug-target interactions and optimizing lead compounds to accelerate the path from discovery to clinical trials.
View Publications →Multi-conformer 3-point pharmacophore fingerprinting technology for AI-driven virtual screening, food-based drug discovery, and cross-species conservation analysis. Originally published in J. Chem. Inf. Comput. Sci. (1999, 2000) and reborn for AI-era natural-product matching.
Triplet-based pharmacophore matching across drug compounds, natural products, and target binding sites.
The same fingerprint is the objective function inside ChIP, which searches reaction routes for purchasable molecules that match a reference drug's pharmacophore on an unrelated scaffold. Almost everything useful you do with a pharmacophore fingerprint is comparative, and PharmCast now answers those comparisons from 2D structure alone. This is not a new way of mapping pharmacophores: the features, the triplets and the 10,549-bit encoding are unchanged, and what changes is the cost of asking how similar two molecules are, and a complete comparison of two molecules from their SMILES strings takes 0.584 ms against 5.71 s for the reference route, because it never builds a three-dimensional conformer at all. Building the peptide corpus that trains the composite model also turned up a result worth its own look: for short protein loops, sequence does not determine structure. Conformer generation is 98% of the real cost and the surrogate skips it, so ranking the complete 4.65 million compound Enamine collection against a reference takes 29.3 minutes, measured on one Apple M3 Ultra with PyTorch on CPU. That is what makes design search, catalog triage, 3D-aware clustering and scaffold hopping tractable at collection scale, with the per-conformer calculation still the arbiter that says which conformation carries the overlap. The same speed made it practical to profile that collection properly: all 4,612,044 in-stock Enamine compounds compared against every other, 10.6 trillion pairwise comparisons with no sampling. The library turns out to be wide in frameworks and deep in analogs around them, 679 distinct Bemis-Murcko scaffolds per thousand molecules yet a median nearest-neighbor similarity of 0.714, with 64,473 compounds carrying a 2D-identical twin. That is structural redundancy, not pharmacophoric.
Remarkable 3D molecular alignment between Pravastatin (cholesterol drug) and Ganoderic acid from Reishi mushrooms. This innovative "reverse screening" approach explores how natural compounds in food mirror pharmaceutical drugs. Visit DrugToTable.com →
While the world focuses on self-promotion, we're building a place to honor the people around you. Share stories, create tributes, and celebrate the lives that touch yours.
Each app focuses on different communities and ways to celebrate friendships through meaningful connections.
The newest Renaissance Circle post, Toast Our Friend, turns the idea behind Toast Our Friend into a larger call to move tributes upstream. Instead of saving the best stories for memorials, it asks people to gather video toasts, photos, gratitude, and memories while friends, mentors, and family can still hear them. The essay connects the first toast circles for Scott Miller, Mike Szwajkowski, Steve Cooper, Tom Shearer, and Sung-Hou Kim to the broader AI-Steve thread: use technology to preserve relationship, amplify appreciation, and make it easier to say the important thing now.