Blog
How to Dock a Multi-Epitope Vaccine Construct with TLR4: ClusPro, HDOCK and HADDOCK Compared
- August 22, 2026
- Posted by: Ragini Mishra
- Category: Bioinformatics

To dock a multi-epitope vaccine construct with TLR4, model and validate the construct’s 3D structure, download a TLR4 ectodomain structure such as PDB entry 3FXI, strip the non-protein records, then run ClusPro for cluster-ranked poses, HDOCK for a fast global scan, and HADDOCK when you can name the contact residues in advance.
This picks up exactly where the screening step ends. You have a construct sequence that survived antigenicity, allergenicity and toxicity screening, built from epitopes and linkers as described in the guide to designing a multi-epitope vaccine construct. The next question a reviewer asks is whether the thing you designed can physically engage the innate immune receptor you claim it engages.
Docking answers a narrower question than most student manuscripts admit. It produces a plausible geometry and a relative ranking. It does not produce a binding constant, and it does not demonstrate immunogenicity. This guide covers how to get a defensible pose, how to read the numbers each server gives you, and where the failure modes are.
Which receptor should a multi-epitope construct be docked against?
The receptor follows from the adjuvant you attached to the construct, not from what other papers happened to use. Toll-like receptors recognize distinct molecular patterns, so the choice is a claim about mechanism.
If your construct carries a flagellin or flagellin-derived adjuvant, TLR5 is the mechanistically correct partner. If it carries a lipopolysaccharide-mimicking or beta-defensin type adjuvant frequently paired with TLR4 in the literature, TLR4 is the defensible choice. Docking against a receptor your adjuvant has no relationship with is the single most common criticism these manuscripts attract.
Worth stating plainly, because it affects how you cite prior work: the StemSkills Lab team’s own published constructs were docked against TLR5, not TLR4. Both the dengue construct in Frontiers in Immunology and the monkeypox construct in Viruses used TLR5, because both carried flagellin-based adjuvants. Those papers are on the team’s research page and they are a good template for the workflow shape, but they are not TLR4 precedent. Cite them for method, and name the receptor they actually used.
Which TLR4 structure do you download from the PDB?
Search the RCSB Protein Data Bank rather than accepting whatever PDB code a previous paper printed. Three entries account for most TLR4 docking work, and they are not interchangeable.
| PDB entry | Contents | Method | Resolution | Use it when |
|---|---|---|---|---|
| 3FXI | Human TLR4, human MD-2, E. coli LPS Ra complex | X-ray diffraction | 3.1 A | Default choice for a human TLR4 study |
| 4G8A | Human TLR4 polymorphic variant D299G and T399I with MD-2 and LPS | X-ray diffraction | 2.4 A | You are specifically studying the D299G/T399I variant |
| 3VQ2 | Mouse TLR4, MD-2, LPS complex | X-ray diffraction | 2.48 A | Your downstream model is a mouse immunization study |
Two facts about 3FXI change how you prepare it, and both are readable from the RCSB entry page. First, it holds two polymer entities: TLR4 as chains A and B, with a crystallized construct of 605 residues per chain, and MD-2 (annotated as lymphocyte antigen 96) as chains C and D at 142 residues each. It is a 2:2 assembly, not a single receptor copy. Second, full-length human TLR4 is 839 amino acids according to UniProt entry O00206, so the crystallized portion is the ectodomain. The transmembrane and TIR signaling domains are absent. That is the correct region to dock against, but say “TLR4 ectodomain” in your methods rather than implying you docked the whole receptor.
The structural work behind 3FXI is Park and colleagues, published in Nature in 2009 (PubMed 19252480). Cite the structure paper alongside the four-character code.
Do you keep MD-2 in the receptor file?
Decide deliberately and then say what you did. MD-2 is the co-receptor that actually binds lipid A, and the TLR4 dimer interface in 3FXI forms through MD-2. Keeping chains A and C together models the physiological unit. Deleting MD-2 gives you a cleaner single-chain receptor but removes a real part of the binding surface. Most published multi-epitope work keeps one TLR chain plus its MD-2 partner and discards the second copy. Whichever you choose, record the chains in your methods so the run is reproducible.
How do you prepare the two PDB files before docking?
Your construct needs a 3D structure first. The sequence alone is not dockable. Model it, then validate the model before you dock anything, because docking a bad model produces a confident-looking answer to the wrong question.
The monkeypox study from the team used AlphaFold v2.0 for the tertiary structure of both partners, then validated with ProCheck for the Ramachandran plot and ProSA-web for the Z-score, alongside the AlphaFold pLDDT scores. That combination is a reasonable default: one predictor, two independent validators. If a region of your construct comes back with low pLDDT, that region is a linker or a disordered terminus in most cases, and you should say so rather than pretending the whole model is equally reliable.
Once both structures exist, the preparation is mechanical:
- Remove non-protein records from the receptor. The 3FXI file contains six non-polymer entities, including the LPS components and sugars. Docking servers expect protein. Leaving them in is the direct cause of the most common ClusPro submission error.
- Keep only the chains you decided to keep. For 3FXI that usually means chains A and C, deleting B and D.
- Check for HETATM records sitting in ATOM lines. The ClusPro help page names this exactly: the error that a file has unknown residues happens because “some record in your pdb file is marked as ATOM, but is not one of the 20 standard amino acids or an RNA base.” Some programs write HETATMs into ATOM records. Edit them out or convert them back to HETATM.
- Confirm there are no chain-break gaps across your intended interface. Missing loops in a crystal structure are normal, but a gap at the contact surface makes any interface analysis you run afterwards misleading.
- Name the larger molecule as the receptor. The HDOCK documentation recommends this directly for docking efficiency when one molecule is much larger than the other, which is always true here, because a TLR4 ectodomain is far larger than a multi-epitope construct.
Want the guided, hands-on version?
Our live Molecular Modeling & MD Simulations cohort bootcamp takes you from zero to running real docking and MD workflows, with a portfolio project for your grad-school applications.
ClusPro, HDOCK or HADDOCK: which server should you use?
These three are not competitors doing the same job. They differ in what they assume you already know about the interface.
| ClusPro | HDOCK | HADDOCK 2.4 | |
|---|---|---|---|
| Approach | FFT rigid-body docking with PIPER, then clustering | Hybrid template-based and template-free, FFT global search | Information-driven docking with restraints |
| Input accepted | Two PDB files | PDB file, PDB ID with chain (for example 1CGI:E), or FASTA sequence | PDB files plus a restraint definition |
| Interface knowledge needed | None | Optional | Required for a meaningful run |
| Primary ranking metric | Cluster size (member count) | Docking score, plus a derived confidence score | HADDOCK score of the top cluster |
| Typical runtime | Under 4 hours per the developers | Normally within 30 minutes | Hours, varies with restraint count |
| Account required | Yes, free | No | Yes, free for non-profit users |
| Best used for | Blind docking where you claim nothing about the site | A fast first look, or when you only have a sequence | Reproducing or testing a specific binding hypothesis |
The practical answer for most student projects: run ClusPro and HDOCK as independent blind searches, and add HADDOCK only if you can justify the residues you feed it. Agreement between two independent methods is a far stronger argument than a good score from one.
How do you run ClusPro?
Create a free account at cluspro.bu.edu, upload the receptor and ligand PDB files, and leave the balanced coefficient set selected unless you have a reason not to. The ClusPro FAQ is direct about this: if you have no prior knowledge of which forces dominate in your complex, use the balanced coefficients. The antibody mode exists for antibody-antigen pairs and is not what you want here.
Understanding the output requires knowing what the server actually did. From the official help page, the procedure is: rotate the ligand through 70,000 rotations, translate on a grid for each rotation, keep the best-scoring translation per rotation, take the 1,000 lowest-scoring rotation and translation combinations, then greedily cluster those 1,000 positions using a 9 A C-alpha RMSD radius. The developers point out that the initial sampling covers roughly 109 positions, so the retained 1,000 are in the top millionth.
That last detail is why ClusPro tells you not to rank by score. The help page states it without hedging: “the best way to rank models is by cluster size, which is how the models are ranked coming out of Cluspro”, and separately, “we strongly encourage you to not judge models based on these scores because that is not what the scoring was designed for”. Students routinely quote the lowest ClusPro energy in a results table. That is the wrong column. Report the member count of the top cluster.
Kozakov and colleagues describe the server in Nature Protocols in 2017 (PubMed 28079879), and note that six energy functions are available and that “docking with each energy parameter set results in ten models defined by centers of highly populated clusters of low-energy docked structures”.
How do you run HDOCK?
HDOCK needs no login. Note that the server answers on http://hdock.phys.hust.edu.cn/ and refuses connections on port 443, so use the http address rather than assuming the https version exists.
Upload the receptor and ligand, or paste sequences. If you supply only a sequence, the server builds a model from a homologous template using HH-suite, Clustal W2 and MODELLER, and the documentation notes that this pipeline is designed for single-chain proteins, so upload your own file for anything multi-chain. Binding site residues, if you have them, go in as 195:A, 203-206:A, 108:B. Distance restraints take the form 195:A 236:B 8, meaning those residues should end up within 8 A.
Yan and colleagues published the server protocol in Nature Protocols in 2020 (PubMed 32269383), reporting that HDOCK had processed more than 30,000 docking jobs since its 2017 release.
How do you run HADDOCK, and where do the active residues come from?
HADDOCK 2.4 at wenmr.science.uu.nl is different in kind. It converts your knowledge of the interface into ambiguous interaction restraints, so it needs you to nominate active residues: the ones you believe are in direct contact. Passive residues are their surface neighbours. A HADDOCK run with arbitrary active residues is not a blind docking with extra steps; it is a docking of your assumption.
The team’s monkeypox study shows where legitimate active residues come from. The methods state that the potential binding regions identified for flagellin and human TLR5 were LQRVRELAVQ and EILDISRNQL, and that “during the docking experiments, these sequences were defined as the part of ‘Active Residues’ while running the HADDOCK program”. The residues came from published binding-region analysis of that specific receptor-ligand pair, not from inspection of the model. For a TLR4 run, the equivalent source is the lipid A and dimerization interface described in the 3FXI structure paper.
How do you read the output without over-claiming?
Each server reports something different, and the numbers are not comparable across servers.
ClusPro. Report the number of members in the top cluster and the coefficient set used, and report the second cluster’s size next to it. The comparison is what carries the information: a top cluster well clear of the rest means the search converged on one site, while a top cluster barely larger than the next two or three means it did not, and that is worth stating rather than hiding behind the first row of the table. There is no published cutoff that turns a member count into a verdict, so do not invent one.
HDOCK. The docking score comes from the ITScorePP knowledge-based scoring function. The documentation is explicit that “a more negative docking score means a more possible binding model, but the score should not be treated as the true binding affinity of two molecules because it has not been calibrated to the experimental data.” The confidence score is a sigmoid transform of that score, defined on the help page as:
Confidence_score = 1.0 / [1.0 + e^(0.02 * (Docking_Score + 150))]
The published interpretation bands are above 0.7 for very likely to bind, 0.5 to 0.7 for possible, and below 0.5 for unlikely, with the server’s own caveat that the score should be used carefully given its empirical nature. HDOCK also reports interface residues as all residue pairs within 5.0 A, and warns that ligand RMSD “is not necessarily a metric of the accuracy for the corresponding model”.
HADDOCK. The score is a weighted sum, and the weights are published, which means you can state exactly what your number contains. From the HADDOCK 2.4 documentation, the water-refinement stage uses:
HADDOCKscore-water = 1.0 Evdw + 0.2 Eelec + 1.0 Edesol + 0.1 Eair
where Evdw is the van der Waals intermolecular energy, Eelec the electrostatic term, Edesol the desolvation energy and Eair the restraint energy. The documentation adds that “the structure with the smallest weighted sum will be ranked first”. Report the top cluster’s HADDOCK score with its standard deviation, plus the cluster size, never a single pose in isolation.
What should you do after you pick a pose?
Characterize the interface with tools that do not know about docking scores at all, which is what makes them useful as an independent check.
- PDBsum generates a schematic of interface residues, hydrogen bonds and salt bridges. This is the figure a reviewer actually looks at.
- PDBePISA reports buried surface area and an assessment of whether an interface looks biologically meaningful rather than a crystal contact.
- PRODIGY estimates binding affinity from the 3D structure using a model based on intermolecular contacts and non-interface surface properties (Xue and colleagues, Bioinformatics 2016, PubMed 27503228). It gives you a number in kcal/mol, which is more defensible to quote than a raw docking score, though it is still a prediction.
For a sense of what a reported interface analysis looks like in practice, the monkeypox paper computed distance range maps with COCOMAPS using a 5 A cut-off between atoms, and reported 46 contacts between hydrophilic residues, 50 between hydrophilic and hydrophobic residues, and 11 between two hydrophobic residues. That level of specificity is what turns a picture of two proteins touching into a result.
What goes wrong, and how do you fix it?
| Error or symptom | Cause | Fix |
|---|---|---|
| ClusPro reports unknown residues in your file | Non-standard residues written as ATOM records, often LPS or sugars from 3FXI | Delete them or convert them back to HETATM records before resubmitting |
| Only the receptor appears when you open a result model | Your viewer does not support multiple PDB entries in one file | Open it in PyMOL, or strip the separators: grep -Ev '^(HEADER)|(END)' model.000.00.pdb > model.000.00.stripped.pdb |
| Top ClusPro cluster has very few members | The search found no dominant binding site | Report it honestly, try a second coefficient set, and check the construct model is not disordered at the presumed interface |
| HDOCK https link fails to connect | The server does not serve port 443 | Use the http address |
| HDOCK confidence score sits near 0.5 | Borderline docking score around minus 150 | Treat it as inconclusive rather than positive; look for agreement with ClusPro instead |
| HADDOCK returns clusters that all look alike | Restraints were too tight or covered too few residues | Widen the passive residue selection, or re-derive active residues from published interface data |
| The construct binds the wrong face of TLR4 | Blind docking has no reason to prefer the physiological site | Check whether the pose contacts the known dimerization or MD-2 associated surface; a pose on the convex back of the horseshoe is not a signaling-competent complex |
| Your model has no interface hydrogen bonds | Steric-only contact, common with rigid-body poses | Refine the complex before analysis, and be cautious about claiming a stable interaction |
What comes next after you have a ranked complex?
A docked pose is a static snapshot from a rigid or semi-rigid search. The standard next step is molecular dynamics, which tests whether the interface survives when both partners are allowed to move in solvent. The monkeypox study ran three independent 100 ns simulations from different initial velocities and found the complex stabilized after 30 ns, with an RMSD of 1.08 plus or minus 0.1 nm afterwards. Running one short simulation and declaring stability is not the same thing.
The mechanics of setting that up belong to a different skill set. If you have not run GROMACS before, start with the molecular dynamics pillar rather than learning it inside a vaccine project. Similarly, if the docking concepts here are new, the molecular docking pillar covers scoring functions and search algorithms in general terms. The rest of this workflow sits on the immunoinformatics pillar, and if you are still assembling the underlying skills, the computational biology skills roadmap sequences them.
Frequently asked questions
Can I dock a vaccine construct using only its sequence?
HDOCK accepts a FASTA sequence and will build a model from a homologous template using HH-suite, Clustal W2 and MODELLER. For a multi-epitope construct this is usually a poor idea, because the construct is an artificial sequence with no natural homolog, so the template search has little to work with. Model it yourself with a structure predictor and validate the model first.
Is a lower docking score always a better result?
No, and this is the most common misreading. ClusPro explicitly advises against ranking by its score and ranks by cluster size instead. HDOCK’s docking score is not calibrated to experimental affinity. HADDOCK’s score is a weighted sum whose lowest value ranks first, but only within one run. Scores from different servers cannot be compared with each other at all.
Do I need both TLR4 and MD-2 in my receptor file?
MD-2 is part of the physiological binding unit in PDB entry 3FXI, so keeping one TLR4 chain with its MD-2 partner is the more realistic model. Removing MD-2 is defensible if your claim is about the TLR4 surface specifically. What is not defensible is failing to say which chains you used.
How many docking servers should I report?
At least two independent ones. A single server’s top pose is a single method’s opinion. Two blind methods converging on the same interface is evidence; two disagreeing tells you the result is not settled, which is also worth reporting.
Does a good docking result mean my vaccine candidate will work?
No. Docking predicts a geometry and a relative ranking. It says nothing about expression, folding in vivo, immunogenicity, safety or protection. The team’s dengue construct went through docking, simulation and immune simulation, and still required in vivo rabbit validation before any claim about a cross-reactive response could be made. Write your conclusions at the level the method supports.
Want the guided, hands-on version?
Our live Molecular Modeling & MD Simulations cohort bootcamp takes you from zero to running real docking and MD workflows, with a portfolio project for your grad-school applications.
Written by the StemSkills Lab team, which has more than 10 years of combined work in sequence and structural bioinformatics, drug discovery and design, and multiscale molecular modeling, including peer-reviewed multi-epitope vaccine constructs validated in vivo.