Blog
How to Design a Multi-Epitope Vaccine Construct: Linkers, Adjuvant and Order
- August 8, 2026
- Posted by: Ragini Mishra
- Category: Bioinformatics

A multi-epitope vaccine construct is assembled in a fixed order: adjuvant at the N-terminus, an EAAAK linker, then the epitope blocks joined by class-specific spacers (AAY for cytotoxic T-cell epitopes, GPGPG for helper epitopes, KK for B-cell epitopes), and an optional histidine tag. Validate the finished sequence in ProtParam before doing anything else.
Most immunoinformatics tutorials stop at the shortlist. You have run BepiPred and ABCpred for your B-cell epitopes and NetMHCpan and IEDB for your T-cell epitopes, you have a table of peptides that passed the percentile-rank cut-offs, and then the guide ends. The next step is the one that decides whether the project works: physically joining those peptides into one designed protein sequence.
This is a design task with real constraints, not a copy-paste step. The sequence you place between two epitopes changes whether either of them is presented at all. This guide covers the parts list, the order, the linker set, the adjuvant, and the validation you run before the construct goes anywhere near a docking program.
What actually goes into a multi-epitope vaccine construct?
A construct has five kinds of component, and only one of them is your epitopes:
- An adjuvant at the N-terminus, usually a protein or peptide that engages a Toll-like receptor, included because a string of short peptides is poorly immunogenic on its own.
- A rigid linker separating the adjuvant from the epitope block so the two do not fold into each other.
- The epitope blocks, grouped by class: cytotoxic T-lymphocyte (CTL) epitopes, helper T-lymphocyte (HTL) epitopes, and linear B-cell epitopes.
- Class-specific spacers between the epitopes inside each block.
- An optional purification tag, typically 6 histidines at the C-terminus.
The scaffold is a larger fraction of the final protein than students expect. Take a common layout: six 9-mer CTL epitopes joined by AAY, four 15-mer HTL epitopes joined by GPGPG, five 16-mer B-cell epitopes joined by KK, two GPGPG spacers joining the three blocks, an A(EAAAK)3A linker after the adjuvant, and a 6xHis tag. The epitopes contribute 194 residues. The linkers and tag contribute 71. Before the adjuvant is even added, roughly 27% of your construct is scaffold. Linker choice is not cosmetic.
What order should the blocks go in, and does it matter?
There is no order mandated by a standards body, and any guide that tells you otherwise is overstating the evidence. What is defensible are three constraints that follow from how the components work.
The adjuvant goes first. It has to be accessible to a receptor on the outside of a cell, so burying it in the middle of a 300-residue chain works against its only purpose. Placing it at the N-terminus and separating it with a rigid linker keeps it exposed.
Epitopes stay grouped by class. Each class takes a different spacer, so mixing a B-cell epitope into the middle of the CTL block forces you either to use the wrong spacer or to switch spacers mid-block for no reason. Grouping keeps the design explainable in your methods section.
The tag goes last. A histidine tag at the C-terminus stays away from the adjuvant and away from the first epitopes, and it is trivial to remove from the sequence later if a reviewer objects to it.
Beyond that, order is a choice you justify rather than a rule you follow. Say what you chose and why in your methods, then test the consequences structurally.
Why does the sequence between two epitopes change whether they work?
This is the part that separates a designed construct from a concatenated one, and there is direct experimental evidence for it.
Bergmann and colleagues linked H-2d-restricted HIV-1 and mouse hepatitis virus CTL epitopes using a range of spacer residues and expressed the resulting 20 to 31 amino acid peptides in recombinant vaccinia viruses. Their finding, published in The Journal of Immunology in 1996, is the empirical basis for the whole spacer convention: “Flanking amino acids with aromatic (tyrosine), basic (lysine), and small aliphatic side chains (alanine) supported efficient CTL recognition of both epitopes. By contrast, acidic and helix breaking residues (glycine, proline) specifically inhibited recognition of the adjacent amino-terminal epitope.”
Read that again with the standard linker set in mind. Alanine, alanine, tyrosine. Lysine, lysine. Glycine and proline used only where class I presentation is not what you are protecting. The convention is not arbitrary tradition, it is a direct reading of which flanking residues supported recognition and which suppressed it.
The same paper reports how large the effect is: the ratios of peptide-specific CTL precursors primed by the tandem epitopes varied up to 50-fold depending on molecular context. A 50-fold swing caused by the three residues you typed between two peptides is why this step deserves a section of its own.
Which linker goes where?
Chen, Zaro and Shen’s review in Advanced Drug Delivery Reviews (2013) groups empirically designed linkers into three structural classes: flexible, rigid, and in vivo cleavable. Every linker below belongs to one of those three, which is the useful way to remember them.
| Linker | Type | Where it belongs | Why that one |
|---|---|---|---|
| A(EAAAK)nA | Rigid, alpha-helical | Between the adjuvant and the first epitope block | Holds two domains apart instead of letting them interact |
| AAY | Short, proteasome-friendly | Between CTL (MHC class I) epitopes | Alanine and tyrosine flanks supported efficient CTL recognition in Bergmann et al. |
| GPGPG | Flexible, helix-breaking | Between HTL (MHC class II) epitopes, and to join blocks | Designed so that glycine and proline are not primary anchor residues, which disrupts junctional epitopes |
| KK | Cleavable (lysosomal) | Between linear B-cell epitopes | Basic lysine flanks supported recognition, and the site is a reported cathepsin B target |
| (GGGGS)n | Flexible | General-purpose, when domains must stay mobile | Standard flexible linker where separation, not rigidity, is the goal |
Why EAAAK and not a glycine-serine linker after the adjuvant?
Because you want distance you can rely on. Arai and colleagues inserted helix-forming A(EAAAK)nA linkers with n = 2 to 5 between two green fluorescent protein variants and measured what happened. Fluorescence resonance energy transfer from EBFP to EGFP decreased as the length of the linkers increased, and circular dichroism showed the linkers forming an alpha-helix whose helical content rose with length (Protein Engineering, 2001). A flexible glycine-serine linker lets the two ends find each other. A helix does not.
Why GPGPG between helper epitopes specifically?
The GPGPG spacer was introduced by Livingston and colleagues in The Journal of Immunology (2002) as a universal spacer for HTL epitopes. The reasoning is that neither glycine nor proline is used as a primary anchor residue by known class II molecules, so a GPGPG junction is unlikely to accidentally create a new epitope that competes with the ones you designed in. Note the tension with Bergmann: glycine and proline are exactly the residues that suppressed class I recognition. That is not a contradiction, it is the reason the two spacers are not interchangeable. Use GPGPG where you are protecting class II presentation, not where you are protecting class I.
Want the guided, hands-on version?
Our live Molecular Modeling & MD Simulations cohort bootcamp takes you from zero to running real docking and MD workflows, with a portfolio project for your grad-school applications.
What adjuvant should you use, and where does it go?
The adjuvant is an immunostimulant fused into the same open reading frame as your epitopes, chosen to engage a specific Toll-like receptor. Commonly used choices in published in-silico constructs include beta-defensins, bacterial flagellin, the 50S ribosomal protein L7/L12, and short TLR-agonist peptides.
The honest guidance is that the choice is a claim you have to defend, not a default. Whichever adjuvant you pick, you should be able to name the receptor it is supposed to engage, and you should test that interaction structurally rather than assert it. That is what the published work does. In the StemSkills Lab team’s own cross-reactive dengue construct (Frontiers in Immunology, 2022), ten conserved interferon-inducing epitopes were assembled and the finished construct was docked and simulated against TLR5, and one epitope was synthesised and tested in a rabbit model. The team’s monkeypox construct (Viruses, 2022) used seven epitopes screened as antigenic, non-allergenic, non-toxic and interferon-activating, again combined with adjuvants and checked for stable binding to human TLR5. In both cases the receptor claim was tested, not assumed. You can see the full set on the team’s research and publications page.
Do you also need a universal helper epitope?
Many constructs include PADRE, a pan-DR epitope engineered by Alexander and colleagues to provide T-cell help across diverse HLA-DR types. The original Immunity paper reports that the pan-DR binding peptides “bound 10 of 10 DR molecules tested, with affinities, in most cases, in the nanomolar range”, and that in one example of their capacity to elicit T help they were approximately 1000 times more powerful than natural T cell epitopes. If your HTL epitope set has narrow allele coverage, adding a universal helper epitope is a reasonable design decision you can cite. If your coverage is already broad, it is extra sequence you have to justify.
How do you validate the construct sequence before doing anything else?
Paste the full assembled sequence into ExPASy ProtParam and read five numbers. These are the checks a reviewer will look for, and every threshold below comes from the ProtParam documentation, not from convention.
- Length and molecular weight. Sanity-check the length against your own arithmetic. If ProtParam reports a different residue count than your epitopes plus linkers, you have a copy-paste error, usually a duplicated or dropped linker.
- Theoretical pI. Relevant for downstream purification planning, and worth reporting.
- Instability index. The documented rule is explicit: “A protein whose instability index is smaller than 40 is predicted as stable, a value above 40 predicts that the protein may be unstable.” The index is computed from dipeptide composition following Guruprasad, Reddy and Pandit (Protein Engineering, 1990).
- Aliphatic index. The relative volume occupied by aliphatic side chains, calculated as X(Ala) + 2.9 x X(Val) + 3.9 x (X(Ile) + X(Leu)), from Ikai (Journal of Biochemistry, 1980). Higher values are associated with thermostability in globular proteins.
- GRAVY. The grand average of hydropathicity, calculated as the sum of hydropathy values of all the amino acids divided by the number of residues, using the Kyte and Doolittle scale (Journal of Molecular Biology, 1982). A negative GRAVY indicates a hydrophilic protein, which is what you generally want for a soluble construct.
What about antigenicity, allergenicity and toxicity?
Three more screens belong here, run on the whole construct and not only on the individual epitopes, because assembly can change the answer. Predict overall antigenicity with VaxiJen, screen for allergenicity with AllerTOP, and check toxicity with ToxinPred. Report the model and threshold you used for each, since these servers offer more than one. A construct that passes epitope-level screening and then fails at the construct level is a normal result, not a mistake, and the fix is usually a substitution or removal at the offending position rather than starting over.
Troubleshooting: what to do when the construct fails a check
These are the failures students actually hit, in rough order of frequency.
The instability index comes back above 40
The index is derived from dipeptide composition, so the junctions you created are a plausible contributor. Check whether swapping a spacer at one or two junctions, or reordering two epitopes within a block so a different dipeptide pair forms, brings the value down. Do not delete a validated epitope to chase the number. Report the value honestly if it stays high and say what you tried.
GRAVY is positive
A positive GRAVY suggests an overall hydrophobic protein and predicts solubility problems on expression. The usual cause is one strongly hydrophobic epitope, not the construct as a whole. Compute GRAVY for each epitope separately, find the outlier, and decide whether it earns its place.
The whole construct is predicted allergenic even though every epitope passed
This is the classic assembly artefact. The screen is being run on a sequence that did not exist when you screened the parts, including all your junctions. Try a different spacer at the block boundaries first, since those are the newest sequence in the design, before touching the epitopes.
The server rejects your sequence
Almost always a formatting problem rather than a biology problem. Strip any non-standard characters, remove line numbers pasted along with the sequence, make sure there are no B, J, O, U, X or Z residues carried over from a database entry, and confirm you pasted plain sequence rather than FASTA with a stray header where the tool did not expect one.
You suspect a junction created a new epitope
Test it rather than guessing. Take a window spanning each junction, roughly nine residues either side, and run those windows back through the same class I and class II predictors you used originally. If a junction window scores as a strong binder, you have manufactured an epitope that was not in your design, and the spacer at that position needs to change.
What happens after the sequence passes?
A validated sequence is a starting point, not a result. The construct now has to be shown to fold plausibly and to engage the receptor its adjuvant targets. That means predicting a structure, refining it, and then running the two workflows that live outside immunoinformatics: docking the construct against its target Toll-like receptor, and running a molecular dynamics simulation to check the complex holds together over time.
Those are separate skills with their own tooling, and re-teaching them here would do you no favours. Work through the molecular docking pillar for the docking half and the GROMACS molecular dynamics pillar for the simulation half. The construct you just designed is the input to both. If you are still mapping out which of these skills to learn in what order, the computational biology skills roadmap sequences them, and the wider immunoinformatics pillar covers the steps on either side of this one.
Frequently asked questions
Is there a required order for CTL, HTL and B-cell epitopes in a construct?
No. No standard specifies one. What is defensible is keeping each class grouped so it can take its own spacer, putting the adjuvant at the N-terminus where it stays accessible, and putting any purification tag at the C-terminus. Beyond that, state your order and your reason in the methods.
Can I use the same linker everywhere to keep it simple?
You can, but you should expect to lose recognition at some junctions. Bergmann and colleagues showed that flanking residues that suppressed recognition of one epitope class were the same residues that worked for another, and that the effect on primed CTL precursor ratios reached 50-fold. A single linker throughout means accepting that trade at every junction you did not tune.
How long should a multi-epitope construct be?
There is no published cut-off, and any specific number you see quoted as a rule is usually someone’s example rather than a standard. The practical limits are the ones the downstream tools impose: prediction servers, structure prediction and expression systems all have their own size constraints, so check those before committing to a design. Report your actual length and let it be justified by the epitopes you needed.
Do I need the histidine tag?
Only if the construct is intended for expression and purification. It is a downstream convenience, not part of the immunological design. Many purely in-silico studies include it anyway to show the construct is expressible in principle, which is fine as long as you say that is why it is there.
Should I screen antigenicity on each epitope or on the whole construct?
Both, and in that order. Epitope-level screening is how you build the shortlist. Construct-level screening is how you catch what assembly created. A result that changes between the two is informative, since it points at your junctions.
Where does codon optimisation fit in?
After the protein sequence is final. Codon optimisation and in-silico cloning translate a finished, validated construct into a nucleotide sequence for a chosen expression host. Doing it before the sequence is settled means redoing it.
Who wrote this
This guide was written by the StemSkills Lab team, whose members have more than a decade of combined work in sequence and structural bioinformatics, drug discovery and design, and multiscale molecular modeling. The team has published multi-epitope vaccine designs against dengue, monkeypox, canine circovirus, feline infectious peritonitis virus, rotavirus and Candida dubliniensis, two of them with in-vivo validation. The full list is on the research and publications page.
Want the guided, hands-on version?
Our live Molecular Modeling & MD Simulations cohort bootcamp takes you from zero to running real docking and MD workflows, with a portfolio project for your grad-school applications.