Blog
How to Choose a Protein Target and Ligands for a Molecular Docking Study (Step-by-Step)
- July 20, 2026
- Posted by: Stem Skills Lab
- Category: Molecular Modeling

Choose a target whose structure you can defend and a ligand set whose members you can justify one by one. In practice that means a high-resolution experimental structure with a complete binding site, a co-crystallised ligand you can redock, and a ligand library built from known actives plus your test compounds, all assembled before you launch a single docking run.
Every docking tutorial starts at the same place: you already have a protein and you already have ligands. The tutorial shows you how to prepare them, set a grid box and read the scores. The step nobody teaches is the one that actually decides whether your project holds up, which is how those two files were chosen in the first place. This guide from the StemSkills Lab team (10+ years in sequence and structural bioinformatics, drug discovery and design, and multiscale molecular modeling) walks through the selection logic examiners probe. It sits inside our pillar guide on learning molecular docking.
How do you choose a protein target for molecular docking?
Start from the biological question, not from the structure database. A target earns its place in your study when three things are true: it has a documented role in the disease or process you are studying, it has a known or clearly identifiable binding site, and there is a structure of it good enough to dock against. If any one of those fails, you have a weak project no matter how clean your docking output looks.
Working in that order matters. Students often reverse it, searching the Protein Data Bank for a structure that looks nice and then writing a rationale backwards from it. An examiner spots that immediately, because the introduction will not explain why this protein and not the twenty others in the same pathway. Write the one-sentence justification first: I am docking against protein X because inhibiting it is an established strategy in condition Y, and here is the reference. If you cannot write that sentence with a citation attached, change targets.
The supply of structures is not the constraint. As of July 2026 the RCSB Protein Data Bank holds 256,448 experimentally determined structures, alongside more than 990,000 computed models mirrored from the AlphaFold Database. The constraint is picking the right one out of that pile, and often a single target has dozens of entries that differ in resolution, mutations, bound ligands and which loops were actually resolved.
What makes a PDB structure good enough to dock against?
Four properties decide it: resolution, completeness of the binding site, the identity of the construct, and what is bound in the pocket. Check all four before you download anything.
Resolution. For X-ray structures, lower numbers mean more experimental detail behind each atomic position. Structures near 1.5 Angstrom show side-chain positions and ordered waters clearly. Around 2.5 to 3.0 Angstrom, side-chain placement is increasingly modelled rather than directly observed, which matters because docking scores depend heavily on exactly where pocket side chains point. You can filter for this directly: the RCSB advanced search exposes Refinement Resolution as a searchable attribute under Methods, so you can restrict results with a less-than operator instead of opening entries one at a time.
| Resolution range | What you can trust | Suitability for docking |
|---|---|---|
| Better than 2.0 A | Side-chain orientations, many ordered waters, ligand geometry | Preferred. Use this if available. |
| 2.0 to 2.5 A | Backbone and most side chains, ligand position reliable | Usually fine. State the resolution in your methods. |
| 2.5 to 3.0 A | Backbone and fold, side chains partly modelled | Workable if nothing better exists. Justify the choice. |
| Worse than 3.0 A | Overall fold and domain arrangement | Weak for docking. Look for an alternative entry. |
| NMR ensembles | Solution conformations, multiple models | Possible, but you must state which model you used and why. |
Resolution is a global number, so it does not tell you whether your specific pocket is well ordered. For that, look at B-factors (temperature factors) for the residues lining the site. High values mean the atom was poorly localised in the experiment. The RCSB Guide to Understanding PDB Data puts the practical threshold at roughly 50 and above, at which point an atom is barely visible in the density, a situation common for long surface side chains that stay mobile in solution. If the residues forming your binding pocket sit in that range, any interaction you report with them is on thin ice.
Completeness. Crystallographers deposit what they can see. Flexible loops frequently go unmodelled, leaving gaps in the coordinate file. A gap somewhere on the far side of the protein is harmless. A gap in a loop that forms one wall of your binding pocket is fatal, because docking software will happily place a ligand into the artificial hole where that loop should be and hand you a confident score for a pose that cannot exist. Compare the SEQRES record against the observed residues, or simply open the structure and look for chain breaks near the site. Our walkthrough on how to validate a protein structure covers the checks in detail.
Construct identity. Read the entry title and the sequence details rather than trusting the protein name. Structures are routinely solved from truncated constructs, single domains, fusion partners, or point mutants introduced to aid crystallisation. Confirm you have the domain that contains the site you care about, on a chain that is complete through that region.
What is bound. An entry with a co-crystallised inhibitor is worth far more to you than an empty apo structure, for two reasons. It tells you exactly where the pocket is and which residues line it, and it gives you a ready-made validation experiment. Strip the ligand out, dock it back in, and measure the RMSD between your top pose and the crystallographic pose. That is the standard way to show your protocol works before you trust it on unknowns, and we walk through it in our guide to validating a docking protocol by redocking.
Should you use the wild-type structure or a mutant?
Use the wild type unless the mutant is the point of your study. Many deposited structures carry substitutions introduced for crystallisation or stability, and those changes can sit inside or adjacent to the binding site. If you dock against such an entry without noticing, your results describe a protein that does not exist in the patient.
There is one clear exception. If your research question is about resistance mutations or a disease-associated variant, the mutant structure is exactly what you want, and the strongest design is to dock the same ligand set against both the wild type and the variant and compare. That comparison gives you a real result rather than a list of scores. Either way, name the construct and any mutations explicitly in your methods section, which we template in our guide to writing the methods section of a docking study.
What if there is no experimental structure for your target?
You have two fallbacks, and both need to be declared honestly. Homology modelling builds your protein from a related structure with sufficient sequence identity. Deep-learning prediction, most commonly AlphaFold, gives you a model with per-residue confidence scores attached.
The trap with predicted models is that the confidence is not uniform. A model can be excellent across the folded core and unreliable in the loops, and loops are frequently what shape a binding pocket. Before docking against any predicted structure, inspect the per-residue confidence values specifically for the pocket-lining residues, not the global average. Predicted models also arrive with no ligand bound, so you lose your redocking validation and must identify the site another way. Our comparison of AlphaFold, homology modelling and experimental structures lays out when each is defensible, and our guide to finding the binding site covers pocket detection when no co-crystallised ligand exists.
Want the guided, hands-on version?
Our live Molecular Modeling & MD Simulations cohort bootcamp takes you from zero to running real docking and MD workflows, with a portfolio project for your grad-school applications.
How do you build a defensible ligand set?
A defensible ligand set has three parts: positive controls, your test compounds, and enough chemical consistency that comparing scores means something. Missing the controls is the single most common weakness in student docking projects.
Positive controls are compounds already known to bind your target: the co-crystallised ligand, approved drugs against that protein, or published inhibitors with measured activity. They anchor your score scale. Without them you have a ranked list with no reference point, and a score of -8.5 kcal/mol means nothing on its own. With them you can say your best hit scores in the same range as a known nanomolar inhibitor, which is a claim a reader can evaluate.
Test compounds are whatever your question is actually about: a natural-product series from a plant of interest, an analogue series around a scaffold, a curated subset of a screening library. Keep the set internally coherent. Docking fifty random molecules from different chemical classes produces a ranking dominated by molecular size rather than by fit, because most scoring functions reward buried surface area.
Here is where to get each part, with the trade-offs that matter for a student project.
| Source | Best used for | What to watch |
|---|---|---|
| RCSB PDB co-crystallised ligands | The primary positive control and redocking validation | Extract the correct chain and residue; check occupancy |
| ChEMBL | Known actives with measured bioactivity values against your target | Filter by assay type and confidence score, never by target name alone |
| PubChem | Broad structure retrieval, natural products, analogue searching | Little activity curation; 2D records need 3D conformer generation |
| DrugBank | Approved drugs for repurposing studies and controls | Free tier has limits on bulk download |
| Published papers on your target | Series-specific analogues tied to real SAR | You must draw structures yourself and verify each one |
Whichever source you use, every ligand needs correct protonation at physiological pH, sensible tautomers, defined stereochemistry, and a generated 3D conformer before it reaches the docking program. Skipping that step silently corrupts your entire ranking. Our guide on preparing protein and ligand files for docking covers the preparation pipeline end to end.
How many ligands should an MSc docking study include?
Enough to support a comparison, few enough that you can inspect every top pose by eye. For a typical thesis project that usually means somewhere between twenty and a few hundred compounds, including at least two or three positive controls.
The reason to cap the number is not compute time. AutoDock Vina is fast; Trott and Olson reported “approximately two orders of magnitude speed-up” over the earlier AutoDock 4 in their original paper (Journal of Computational Chemistry, 2010, 31(2), 455 to 461), and the Vina 1.2.0 release described by Eberhardt and colleagues (Journal of Chemical Information and Modeling, 2021, 61(8), 3891 to 3898) added macrocycle handling and Python bindings that make batch work straightforward. The reason to cap it is that a docking score is a rough estimate, and the scientific content of your project comes from examining the binding modes of your best hits, not from the length of your table. A study of thirty compounds with careful interaction analysis beats a screen of five thousand with none.
What is the step-by-step target and ligand selection checklist?
- Write the biological rationale for the target in one sentence, with a citation.
- Search the PDB for that protein and filter by Refinement Resolution, keeping entries better than 2.5 Angstrom where possible.
- Prefer entries with a co-crystallised ligand in the site of interest.
- Check the construct: correct organism, correct domain, note any mutations.
- Inspect the binding site for chain breaks, missing side chains and high B-factors.
- Record the exact PDB ID, chain and resolution for your methods section.
- Extract the co-crystallised ligand as your redocking control and confirm the protocol reproduces its pose.
- Collect positive controls from ChEMBL or the literature, with their reported activity values.
- Assemble your test compounds as a chemically coherent series.
- Prepare every ligand: protonation, tautomers, stereochemistry, 3D conformer.
- Document the source and identifier of each ligand in a table before you dock anything.
Steps 6 and 11 are the ones students skip and then regret. Six months later, writing up, nobody remembers whether the structure was chain A or chain B, or which of three PubChem entries a compound came from. Record it while you are doing it. If you are still choosing a project shape, our list of docking project ideas for an MSc thesis gives you target and ligand pairings that already satisfy this checklist, and the computational biology skills roadmap shows where docking fits among the other skills to build.
What selection mistakes fail a thesis viva?
Examiners ask predictable questions, and nearly all of them trace back to selection rather than to the docking run itself. The recurring failures are: choosing a target with no stated rationale, using a low-resolution or incomplete structure without acknowledging it, docking against a mutant construct unknowingly, running no positive control so the scores have no reference, comparing compounds of wildly different sizes, and reporting hydrogen bonds to residues that were never resolved in the crystal. Each of these is cheap to prevent at the start and impossible to fix at the end.
FAQ
Can I dock against an apo structure with no ligand bound?
Yes, but you take on two extra obligations. You need an independent way to locate and justify the binding site, and you lose redocking as a validation. Apo pockets can also be partly collapsed, since side chains often rearrange on ligand binding. If a holo structure of the same protein exists, prefer it.
Which chain should I use when a PDB entry has several?
Usually chain A, but check first. Many entries contain multiple copies of the same protein in the asymmetric unit, and they can differ in how much is resolved. Pick the chain with the most complete binding site and the co-crystallised ligand present, then state that choice in your methods.
Should I remove crystallographic water molecules before docking?
Most standard protocols remove all waters, and that is the usual default. The exception is a water that is conserved across multiple structures of the target and mediates a known interaction bridging ligand and protein. If you keep such a water, say which one and why; if you remove everything, say that too.
Is AlphaFold good enough for a docking project?
It can be, for targets with no experimental structure, provided the pocket-lining residues carry high per-residue confidence. It is a weaker choice when a good crystal structure already exists, because you give up both experimental evidence and your redocking control for no gain.
How do I get 3D structures for compounds I only have as names or SMILES?
Retrieve the canonical SMILES from PubChem or ChEMBL, then generate 3D conformers with a chemistry toolkit such as RDKit or Open Babel before docking. Do not feed a flat 2D structure to a docking program and assume it will sort the geometry out for you.
Want the guided, hands-on version?
Our live Molecular Modeling & MD Simulations cohort bootcamp takes you from zero to running real docking and MD workflows, with a portfolio project for your grad-school applications.