Multi-Epitope Vaccine Structure Prediction and Validation
Skip to content

Multi-Epitope Vaccine Structure: Model, Refine, Validate

Multi-Epitope Vaccine Structure: Model, Refine, Validate

A multi-epitope vaccine construct has no natural template, so predict it with AlphaFold2 or another de novo method rather than homology modelling, refine the result with GalaxyRefine, then validate it with ProSA-web and the SAVES suite. Expect confident epitope and adjuvant regions and low confidence at the linkers. That pattern is normal, not a failed model.

By this point in a design project you have a single FASTA sequence: adjuvant, epitopes, linkers, sometimes a His tag, all stitched together. No organism has ever made that protein. When you push it through a structure predictor and the report comes back with bright blue epitope blocks and orange segments between them, the usual reaction is that the model failed and the construct needs redesigning. Almost always it did not, and it does not. This guide is about reading that output correctly and getting a structure you can defend in a thesis or a paper.

We are only covering what is different about a chimera here. If you have never run a structure predictor before, our team has separate walkthroughs on predicting protein structure with AlphaFold and on validating a protein structure. Read those for the mechanics. Read this for the judgement calls.

Why is homology modelling the wrong tool for a vaccine construct?

Because there is nothing to model it on. Homology modelling works by threading your sequence onto an experimentally solved relative, and the whole method rests on that relative existing. A construct you assembled yesterday out of eight epitopes and three linkers has no relative in the PDB. A template search will still return hits, because the adjuvant is usually a real protein and individual epitopes come from real proteins, but those hits cover fragments of your sequence and say nothing about how the fragments sit relative to each other.

That is exactly the question a vaccine construct model has to answer. You already know the epitopes fold; you took them out of folded proteins. What you do not know is whether the assembled chimera packs into something compact or hangs open, and whether an epitope you need exposed ends up buried against the adjuvant.

So use homology modelling on the natural proteins in your project, such as a receptor with a solved relative, and use a de novo predictor on the construct itself. Our MODELLER tutorial and SWISS-MODEL tutorial cover the template-based route when you do have a template, and AlphaFold vs homology modelling vs experimental structures compares the three families of method in general.

Which predictor should you use for a de novo construct?

Four are in common use in published immunoinformatics work and all four are free to students. The differences that matter for a chimera are whether the method needs a template, and whether it hands back per-residue confidence you can inspect at the linkers.

ToolNeeds a templatePer-residue confidencePractical notes for a construct
AlphaFold ServerNopLDDT and PAEBrowser based, no install, daily job limit. PAE tells you how confident the model is about the relative position of two domains, which is the chimera question.
ColabFold AlphaFold2 notebookNopLDDT and PAERuns in Google Colab on a free GPU, gives you the PDB files and plots to keep. The route most student projects end up using.
trRosettaNoPredicted inter-residue geometry and estimated accuracyDeep learning on predicted distances and orientations. A useful second opinion when one predictor gives an implausible arrangement.
I-TASSERUses threading templatesC-score and per-residue error estimateStill widely cited in vaccine papers, but the threading step is working against you on a sequence with no relative. Queue times are long.

Our own published work used AlphaFold. In the monkeypox multi-epitope vaccine study from our team, the methods state that “the Alphafoldv2.0 program was employed to predict the 3-D structures of both the MPXV-MEV construct and immunogenic TLR5”, and that “Alphafoldv2.0, ProCheck and ProSA web server were then used to validate the tertiary structures” (Vaccines, 2022, PMC9693848). The method itself is described in Jumper et al., Highly accurate protein structure prediction with AlphaFold, Nature 596:583-589, 2021, and the Colab implementation in Mirdita et al., ColabFold: making protein folding accessible to all, Nature Methods 19:679-682, 2022.

Want the guided, hands-on version?

Our live Molecular Modeling & MD Simulations cohort bootcamp takes you from zero to running real docking and MD workflows, with a portfolio project for your grad-school applications.

Join the waitlist (free) →

Why is pLDDT low at the linkers, and is the model broken?

The model is not broken. Low confidence at a flexible linker is the physically correct answer, and it is what a predictor should return for a segment that genuinely has no single conformation.

First, fix what the numbers mean. The AlphaFold Protein Structure Database paper states the bands plainly: “Residues with pLDDT >= 90 have very high model confidence, while residues with 90 > pLDDT >= 70 are classified as confident. Residues with 70 > pLDDT >= 50 have low confidence, and residues with pLDDT < 50 correspond to very low confidence” (Varadi et al., Nucleic Acids Research, PMC8728224). The same paper adds the sentence that settles the linker argument: “very low confidence pLDDT scores correlate with high propensities for intrinsic disorder”. A GS or EAAAK linker is built to be flexible. A predictor that reported it as rigid with pLDDT 95 would be telling you something false.

Second, look at what our team’s own construct did. In the monkeypox study, AlphaFold gave “high pLDDT values with confidence scores > 90% for the N- and C-terminal regions of the vaccine construct. However, a low confidence score with pLDDT values of <50% were predicted for the regions where adjuvants/linkers and epitope sequences were present (amino acids from 142-303)”. That model went on to be refined, validated, docked and published. If your construct shows the same pattern, you are looking at a normal result for this class of molecule, and the right move is to say so in your write-up rather than quietly cropping the figure.

The check worth doing, that most students skip: read the per-residue confidence for your linker residues specifically instead of eyeballing the colour. AlphaFold DB “stores these values in the B-factor fields of the mmCIF and PDB files”, so the number for every residue is already sitting in your downloaded model. Open the PDB in a text editor or in ChimeraX, pull the B-factor column for the residue range of each linker, and write those ranges into your methods. That turns a vague claim about flexible linkers into a stated observation with residue numbers attached.

PAE is the other half of the picture and it answers a different question: how sure is the model about where domain A sits relative to domain B. Those values “are measured in Angstroms and capped at 31.75 A”. For a chimera with three or four epitope blocks, high PAE between blocks means the model is not committing to a single relative arrangement, which again is honest rather than wrong.

How do you refine the predicted model?

Refinement is a separate step, and it targets local problems that prediction leaves behind: side chains in poor rotamers, small clashes, backbone strain. The free route used in most published vaccine work is GalaxyRefine from the Seok lab, described in Heo, Park and Seok, GalaxyRefine: Protein structure refinement driven by side-chain repacking, Nucleic Acids Research 41:W384-W388, 2013. Upload the predicted PDB, and the server returns several refined models with their own quality scores.

Two rules keep this step useful. Pick the refined model on the validation scores, not on the order the server lists them. And keep the input model, because refinement can make a structure score worse. If every refined model validates below your starting model, report the unrefined one and say why.

How do you validate a chimera’s structure?

Run more than one validator, because each answers a narrow question. None of them can tell you whether a flexible linker’s modelled conformation is the one it adopts in solution, and no validation score should be presented as if they can.

ToolWhat it actually answers
ProSA-webWhether the model’s overall energy z-score falls inside the range occupied by experimentally determined proteins of similar size. A single global sanity check.
PROCHECK, via SAVES v6.1Backbone dihedral angles against the Ramachandran plot: how many residues sit in favoured, allowed and disallowed regions.
ERRAT, via SAVESPatterns of non-bonded atomic interactions compared with reliable high-resolution structures.
VERIFY3D, via SAVESWhether each residue’s 3D environment is compatible with its amino acid type.
MolProbityAll-atom contacts, steric clashes and rotamer outliers, including hydrogens.
PDBsumA readable structural summary with Ramachandran and secondary structure figures for the write-up.

ProSA-web is documented in Wiederstein and Sippl, Nucleic Acids Research 35:W407-W410, 2007. PROCHECK (Laskowski et al., Journal of Applied Crystallography 26:283-291, 1993) and ERRAT (Colovos and Yeates, “Verification of protein structures: patterns of nonbonded atomic interactions”) are both run and credited from the SAVES server.

A worked set of numbers. For the published MPXV-MEV construct, the Ramachandran analysis gave 96.7% of residues in favoured regions, 2.9% in allowed and generously allowed regions, and 0.4% disallowed. For the receptor in the same study, TLR5, the figures were 98.9% in core acceptable regions and 1.1% in allowed and generously allowed regions. Report your own numbers in that form, for the construct and the receptor separately. Do not carry over a threshold from another paper as though it were a standard: state what you got, and where the outliers sit. In a chimera the outliers usually sit in the linkers, which is consistent with everything above.

What do you hand to the docking step?

One refined, validated PDB of the construct, plus a receptor structure prepared the same way. In the monkeypox study the receptor was TLR5, taken from AlphaFold because no experimental structure exists for it: UniProt D1CS82, and only the ectodomain, residues 1 to 639, with residues 640 to 836 covering the transmembrane and TIR regions excluded. That trimming decision is worth copying. Docking the membrane-spanning part of a receptor that sits in a bilayer produces a complex that means nothing.

Note that the published study docked TLR5, not TLR4, using information-driven docking in HADDOCK 2.4. TLR4 is the more common choice in the literature and it is the one our TLR4 docking walkthrough uses. Pick the receptor your antigen’s immunology justifies and say why in the methods.

Carry the confidence information forward rather than dropping it at the file boundary. If your linker region modelled below 50 pLDDT, then a docking pose whose interface is formed mainly by linker residues is weakly supported, whatever the docking score says. Interfaces formed by high confidence epitope and adjuvant surfaces are the ones to trust. The same caution applies to the following step, our GROMACS MD simulation guide, where a flexible linker is also where most of the RMSD you observe will come from.

What do you do when the model goes wrong?

Six failures that actually happen, and what each one means.

  • The model is a floating string of helices with no packed core. Usually a construct with long linkers and few interacting surfaces. Check PAE between blocks: if it is high everywhere, the predictor is telling you the arrangement is undetermined, not that this extended shape is the structure. Consider shortening the linkers in the design and re-running, and see our construct assembly guide for linker choices.
  • pLDDT is below 50 across the entire construct, not just the linkers. That is a different problem. Confirm you submitted the right sequence, that the epitope blocks are in frame and not frameshifted during assembly, and that the adjuvant sequence is complete rather than truncated.
  • Ramachandran outliers all sit in the linker. Expected, and reportable. Run refinement, then report the outlier residue numbers and note that they fall in the flexible region.
  • The ProSA z-score sits outside the range shown for native proteins of that size. Treat it as a signal to re-run refinement and to check for missing residues or broken chains in the file rather than as a verdict on the design. Our guide on fixing missing residues and loops in a PDB structure covers the repair.
  • The refined model scores worse than the input. Keep the input. Refinement is an optimisation with its own assumptions and it does not always help.
  • The docking server rejects your PDB. Almost always a file problem: multiple models in one file, missing chain identifiers, alternate locations, or non-standard residues. Keep a single model, one chain identifier, and standard residues only.

One last check before you move on. Run the sequence through ProtParam if you have not already, so the physicochemical profile and the structural model are reported together.

Frequently asked questions

Does a low pLDDT linker mean I should redesign my construct?

No. Linkers are designed to be flexible, and the AlphaFold DB paper notes that very low confidence scores correlate with high propensities for intrinsic disorder. Redesign when the epitope or adjuvant regions themselves model poorly, or when an epitope you need exposed is buried.

Can I use AlphaFold3 or the AlphaFold Server instead of AlphaFold2?

Yes, and the reading of the output does not change. Whichever version you use, state it in the methods with the date you ran it, because these models are updated and a reviewer will ask.

Do I need to refine the model at all?

Refinement is standard in published vaccine design workflows and it is cheap to run. It is not compulsory. What is compulsory is validating whichever model you take forward, and reporting which one you chose.

Which single validation number should I report?

There is no single one. Report the Ramachandran percentages, the ProSA z-score and at least one all-atom check such as ERRAT or MolProbity, for both the construct and the receptor.

Can I use the AlphaFold Protein Structure Database instead of predicting?

For natural proteins such as a TLR receptor, yes, and that is what our monkeypox study did for TLR5. For your construct, no. It is a sequence that has never existed, so no database holds it.

Who wrote this

This guide comes from the StemSkills Lab team, which has more than ten years of combined work in sequence and structural bioinformatics, drug discovery and design, and multiscale molecular modeling. The monkeypox multi-epitope vaccine study quoted throughout is our own published work, and every number taken from it is quoted from the paper rather than paraphrased. The study page has the full reference. Tool links were checked on 17 September 2026.

This is step six of nine in our immunoinformatics roadmap. With a validated model in hand, the next step is docking the construct to a TLR receptor, followed by molecular dynamics on the complex. If structural modelling is new to you as a whole, our computational biology skills roadmap sets out the order to learn it in, and the free assessments are a quick way to check where you stand.

Want the guided, hands-on version?

Our live Molecular Modeling & MD Simulations cohort bootcamp takes you from zero to running real docking and MD workflows, with a portfolio project for your grad-school applications.

Join the waitlist (free) →

Think you know Biomolecular Modeling & Simulations?
Take the free StemSkills assessment and earn a verifiable certificate you can download and add to your LinkedIn profile.
Start the free assessment

Keep going

How to Select Target Antigens for Reverse Vaccinology Start a vaccine design project right. Five filters, in order, that cut a whole proteome down to the… ProtParam for a Vaccine Construct: Reading Every Number Run ProtParam on your multi-epitope construct and learn what each number means, which cutoff is real, and what… How to Run an IEDB Population Coverage Analysis for Your Epitope Set Run the IEDB Population Coverage tool step by step, format the input correctly, read PC90 properly, and fix…
See live workshops