ProtParam for a Vaccine Construct: Reading Every Number

ProtParam takes your assembled multi-epitope sequence and returns nine sequence-derived properties in a single pass. Only one of them carries a published cutoff: an instability index below 40 predicts a stable protein, and a value above 40 predicts the protein may be unstable. The rest are descriptive. Read them as a panel, then fix the construct rather than the number.
Almost every immunoinformatics paper prints a ProtParam table. Very few say what any of the numbers mean, and almost none say what you should do when one of them looks bad. That gap is what this guide fills. If you have already assembled a construct by following our walkthrough on designing a multi-epitope vaccine construct, this is the next step, and it sits directly before codon optimization for expression in E. coli. It is part of the wider workflow mapped out on our immunoinformatics pillar guide.
What do you need before you run ProtParam?
One thing: the final amino acid sequence of the assembled construct, as a single continuous string. That means the epitopes, every linker between them, the adjuvant, and any tag you intend to keep, in the exact order you plan to express them.
Three points from the official ProtParam documentation decide how much the output is worth:
- ProtParam sees sequence only. The documentation states plainly that “no additional information is required about the protein under consideration”. That is a feature, and it is also the limit. Every number is derived from composition or from the N-terminal residue.
- It knows nothing about post-translational modification. The doc says it is “not possible to specify post-translational modification for your protein”. If your construct will be glycosylated or N-terminally modified, the reported values describe the unmodified chain.
- It does not know your protein is a multimer. The doc gives an explicit workaround: “If you do know that your protein forms a dimer, you may just duplicate your sequence”. That works because every calculation rests on composition or the N-terminus.
Run the antigenicity, allergenicity and toxicity screens first, because a construct that fails those is not worth characterising. Our guide on epitope antigenicity, allergenicity and toxicity screening covers that stage.
How do you run ProtParam on a multi-epitope construct?
Go to web.expasy.org/protparam, paste the raw sequence into the box, and submit. The tool also accepts a UniProtKB accession or ID, which is useful when you are characterising a native antigen rather than a designed construct, and in that case it offers an intermediate page for selecting a chain, a domain or a residue range.
Two mechanical details save time. The documentation says that in a raw sequence “space and numbers are ignored”, so a sequence copied out of a numbered alignment will work. Anything that is not an amino acid letter is a different matter, so strip the FASTA header line, any trailing asterisk from a translated stop codon, and any X placeholder before pasting.
Record the output in full, along with the date you ran it. Server-side tools change, and a methods section that names the tool and the date it was used is reproducible in a way that “analysed with ProtParam” is not.
What does each ProtParam number actually mean?
ProtParam returns molecular weight, theoretical pI, amino acid composition, atomic composition, extinction coefficient, estimated half-life, instability index, aliphatic index and the grand average of hydropathicity (GRAVY). The table below is the interpretation layer. Read the third column carefully: only one row contains a published numeric rule, and any page that gives you a target value for the others is inventing it.
| Parameter | What it measures | Published interpretation rule | What to do if the value looks wrong |
|---|---|---|---|
| Molecular weight | Mass of the chain, computed as in Compute pI/Mw | None. It is a fact, not a score | Nothing. Use it to pick a ladder and to sanity-check the construct length |
| Theoretical pI | pH at which the chain carries no net charge | None | Plan purification buffers away from this pH. A very basic pI often traces to a lysine-rich or arginine-rich adjuvant |
| Amino acid composition, plus counts of negatively and positively charged residues | Residue frequencies across the whole construct | None | Read it as the explanation for pI and for solubility predictions, not as a result in itself |
| Extinction coefficient | Absorbance at 280 nm, from Tyr, Trp and cystine content | None, but the doc warns of more than 10% error for proteins without Trp | If the construct has no tryptophan, report the caveat or measure concentration another way |
| Estimated half-life | N-end rule lookup on the N-terminal residue, for three model systems | None. It is a table lookup, not a calculation on your design | Change it only by changing the N-terminal residue. Do not present it as evidence your design is stable |
| Instability index | Weighted sum over consecutive dipeptides | Below 40 predicted stable, above 40 may be unstable. The only real cutoff in the panel | Change the linkers and the epitope order. Both alter the dipeptides at every junction |
| Aliphatic index | Relative volume occupied by Ala, Val, Ile and Leu side chains | None. Higher values are associated with thermostability, with no threshold defined | Nothing directly. It is a description of composition, not a pass or fail |
| GRAVY | Sum of Kyte-Doolittle hydropathy values divided by residue count | None. Negative means hydrophilic on average, positive means hydrophobic on average | A positive GRAVY is a prompt to check solubility with a dedicated tool, not a failure |
The instability index is the only rule you can quote
The ProtParam documentation states it without hedging: “A protein whose instability index is smaller than 40 is predicted as stable, a value above 40 predicts that the protein may be unstable.” The method comes from Guruprasad, Reddy and Pandit, published in Protein Engineering 4:155-161 in 1990 (PMID 2075190). Its derivation matters for how much weight you put on it: the documentation describes a “Statistical analysis of 12 unstable and 32 stable proteins”, from which the authors assigned an instability weight to each of the 400 possible dipeptides.
Forty-four proteins is a small training set, and the index is a sum over consecutive residue pairs. For a multi-epitope construct that second fact is the useful one. Your epitopes were chosen on binding and coverage grounds and should not be discarded lightly, but every junction between them is a dipeptide that the index counts. Swapping a flexible linker for a rigid one, or reordering two epitopes, changes several dipeptides at once and can move the index without touching a single epitope. Re-run the whole panel after any such change.
Why the aliphatic index and GRAVY have no pass mark
The aliphatic index is defined in the documentation as “the relative volume occupied by aliphatic side chains (alanine, valine, isoleucine, and leucine)”, and it “may be regarded as a positive factor for the increase of thermostability of globular proteins”. The formula is X(Ala) + a * X(Val) + b * (X(Ile) + X(Leu)), where the mole percentages are weighted by the relative side-chain volumes a = 2.9 for valine and b = 3.9 for leucine and isoleucine, taken from Ikai, Journal of Biochemistry 88:1895-1898, 1980 (PMID 7462208). Notice what is absent: any cutoff. “May be regarded as a positive factor” is not a threshold.
GRAVY is even simpler. The documentation defines it as the sum of the hydropathy values of all residues divided by the number of residues, using the scale from Kyte and Doolittle, Journal of Molecular Biology 157:105-132, 1982 (PMID 7108955). A negative value means the average residue is hydrophilic. There is no published good value for a vaccine construct, and a paper that reports one has borrowed it from another paper that also invented it.
Why your half-life matches every other paper
ProtParam predicts half-life with the N-end rule, which relates stability to the identity of the N-terminal residue alone, and reports it for mammalian reticulocytes in vitro, yeast in vivo and E. coli in vivo. The documentation’s own table gives methionine as 30 hours, more than 20 hours and more than 10 hours across those three systems. Glycine carries the same row. Since most designed constructs begin with a methionine or a glycine-containing linker, that triple appears in paper after paper. It is a lookup on one residue, and the documentation adds that the estimate “is not applicable for N-terminally modified proteins”.
A worked example from our own published construct
The StemSkills Lab team ran exactly this step on a dengue vaccine construct, published in Frontiers in Immunology 13:865180, 2022 (PMID 35799781, open access at PMC9254734). The final construct had 309 amino acid residues, a molecular weight of 32.6 kDa, a theoretical pI of 9.99, 22 negatively and 32 positively charged residues, and an aliphatic index of 88.19, which the paper reports as indicating a protein stable across a broad temperature range. Half-life came back as the familiar 30 hours, more than 20 hours and more than 10 hours. Antigenicity was 0.463 by VaxiJen 2.0, and predicted solubility on overexpression in E. coli was 0.6909.
Two honest observations about that table. First, these are one construct’s values, not targets to aim at; a pI of 9.99 is a consequence of that particular adjuvant and epitope set, not a design goal. Second, the paper does not report an instability index or a GRAVY value for the final construct, which is common and is worth doing better in your own work. Report the full panel. You can see the team’s other work on the research page.
Want the guided, hands-on version?
Our live Molecular Modeling & MD Simulations cohort bootcamp takes you from zero to running real docking and MD workflows, with a portfolio project for your grad-school applications.
How do you predict whether the construct will be soluble in E. coli?
Not with ProtParam. Solubility is not among the nine parameters it computes, and GRAVY is a compositional average rather than a solubility prediction. You need a dedicated tool.
The practical choice today is Protein-Sol, a free web suite built by the Warwicker group at the University of Manchester. Its method paper is Hebditch, Carballo-Amador, Charonis, Curtis and Warwicker, Bioinformatics 33(19):3098-3100 (PMID 28575391), and the abstract tells you what the prediction is anchored to: “Using available data for Escherichia coli protein solubility in a cell-free expression system, 35 sequence-based properties are calculated.” Read that carefully before you quote the output. The reference dataset is a cell-free E. coli system, so the number speaks to that context, not to every expression host. The server returns a predicted solubility along with an indication of which features deviate most from average values, and that second output is the actionable half, because it names the composition problem instead of just scoring it.
A note on the tool most published methods sections still cite. SCRATCH, whose SOLpro module produced the 0.6909 figure in the dengue paper above, is hosted at scratch.proteomics.ics.uci.edu, and that host did not resolve on either HTTP or HTTPS when we checked on 10 September 2026. It has now failed repeated checks. Cite the method if you are describing prior work, but send readers to a server that answers.
If solubility comes back poor, the fixes are construct-level rather than sequence-cosmetic: reconsider the adjuvant, revisit linker choice, or plan a solubility-enhancing fusion partner. Codon choice is a separate problem that affects translation rather than intrinsic solubility, and it is covered in our guide to codon optimizing a vaccine construct for E. coli.
Should you engineer disulfide bonds into the construct?
Sometimes, and only after you have a three-dimensional model, because disulfide engineering is a geometry question rather than a sequence question. Building and checking that model belongs to a different part of the workflow: see our guides on predicting protein structure with AlphaFold and on validating a protein structure before you get here.
The standard tool is Disulfide by Design 2, version 2.13, released 20 August 2020, at cptweb.cpt.wayne.edu/DbD2. Note the plain HTTP address; there is no working HTTPS endpoint, so expect a browser warning. The method paper is Craig and Dombkowski, BMC Bioinformatics 14:346 (PMID 24289175), which describes the added B-factor analysis as a feature that “facilitates the identification of potential disulfides that are not only likely to form but are also expected to provide improved thermal stability to the protein”.
The defaults, read off the live interface and its user guide, are the parameters you should report:
- Chain scope. Intra-chain and inter-chain are both checked by default.
- Chi-3 torsion angle. Default +97 degrees plus or minus 30, and -87 degrees plus or minus 30. The guide explains the bimodal distribution behind those two peaks and suggests narrowing the tolerance when the default returns too many candidates.
- C-alpha to C-beta to S-gamma angle. Default 114.6 degrees plus or minus 10, against an observed range of roughly 105 to 125 degrees.
- Build C-beta for Gly. Checked by default on the live form. Glycine has no C-beta atom, so it cannot be assessed at all unless the tool builds one from the backbone coordinates. Uncheck it if you would rather leave glycine positions out of the candidate list.
- Upload limit. 2 MB for a local PDB file, or fetch by PDB identifier.
Results list each candidate pair with its chi-3 angle, a bond energy in kcal/mol, and the summed B-factor of the two residues. For calibration, the user guide reports that across a reference set of 331 proteins containing 1418 known disulfide bonds, the mean energy computed by the same equations is 0.89 kcal/mol and the maximum is 8.35 kcal/mol. Those two numbers are documented; the energy cutoffs that circulate in the literature usually are not, so quote the distribution and say which candidates you picked and why.
One consequence students forget: introducing cysteines changes the construct, so the ProtParam panel has to be re-run. The extinction coefficient in particular is reported as two values, one assuming all cysteines are paired as cystines and one assuming all are reduced, and adding cysteine pairs widens the gap between them.
What goes wrong, and how do you fix it?
Your instability index is above 40. This is not a verdict on the epitopes. The index sums weights over consecutive dipeptides, so junction chemistry contributes throughout. Change the linker set or the epitope order and re-run. Do not delete an epitope that earned its place on binding and population-coverage grounds; if you are unsure which those are, our guide to IEDB population coverage analysis is the place to check before cutting anything.
ProtParam rejects your sequence or returns a length you do not expect. The form wants a raw sequence. Spaces and digits are ignored, per the documentation, but a FASTA header, a stop-codon asterisk or an X placeholder are not amino acids. Strip them and resubmit, then confirm the residue count matches your intended construct.
Your half-life is identical to every paper you have read. It should be, if your construct begins with methionine or glycine. It is an N-end rule lookup on one residue. Report it, and do not build an argument on it.
The extinction coefficient looks unreliable. If the construct contains no tryptophan, it probably is. The documentation states that computation is reliable for proteins containing Trp, but “there may be more than 10% error for proteins without Trp residues”. Report the caveat alongside the value.
Disulfide by Design will not load your PDB. Its user guide names the cause under known issues: some files fail to load, “most often due to non-standard residue numbering”, because the program expects numbering to start at 1. Renumber the model from 1 and reload. This bites predicted models more often than deposited structures.
A reviewer asks for the SCRATCH solubility number. Say the host no longer resolves, give the date you checked, and supply a Protein-Sol value instead, naming the tool and the date. A dead link in a methods section is a reproducibility problem, and an honest substitution is the correct fix.
Frequently asked questions
Is there a good GRAVY value for a multi-epitope vaccine construct?
No. The ProtParam documentation defines GRAVY as a sum of Kyte-Doolittle hydropathy values divided by residue count and gives no threshold. A negative value tells you the average residue is hydrophilic, which is a description, not a pass mark.
Does a high aliphatic index mean my vaccine will work?
No. The documentation says the aliphatic index “may be regarded as a positive factor for the increase of thermostability of globular proteins”. Thermostability is not immunogenicity, and no cutoff is defined for either.
Can ProtParam tell me whether my construct is antigenic?
No. Antigenicity, allergenicity and toxicity are separate predictions from separate tools, covered in our guide to epitope screening. ProtParam computes physicochemical properties only.
Do I need a three-dimensional structure to run ProtParam?
No. ProtParam works from sequence alone. Disulfide by Design does need a structure, which is why disulfide engineering comes later in the workflow, after modelling and structure validation.
What is the difference between ProtParam, Compute pI/Mw and ProtScale?
ProtParam returns one value per property for the whole sequence, and its documentation states that molecular weight and theoretical pI are calculated as in Compute pI/Mw. ProtScale is the windowed version: it plots an amino acid scale, such as hydropathy, along the sequence, which is how you find the hydrophobic stretch behind a positive GRAVY.
Which parameters should I put in my thesis or paper?
All nine, with the tool name and the date you ran it. Then interpret only what can be interpreted: the instability index against its published cutoff, and everything else as description. Reporting the whole panel and interpreting it honestly reads as more competent than quoting an invented threshold.
Where this step sits in the workflow
Physicochemical characterisation is the checkpoint between design and expression. Before it comes epitope selection, screening and assembly. After it comes structure prediction, then docking the construct against a Toll-like receptor, then molecular dynamics on the complex, and finally codon optimization for the expression host. The full sequence of steps is laid out on our immunoinformatics pillar, and if you are still deciding which computational skills to build first, the computational biology skills roadmap puts this work in context.
Who wrote this
This guide was written by the StemSkills Lab team, which has more than ten years of combined work in sequence and structural bioinformatics, drug discovery and design, and multiscale molecular modeling. The dengue construct used as the worked example above is the team’s own published work, and every parameter definition, quotation and cutoff in this article is taken from the ProtParam documentation or from the primary papers it cites, all of which are linked above.
Want the guided, hands-on version?
Our live Molecular Modeling & MD Simulations cohort bootcamp takes you from zero to running real docking and MD workflows, with a portfolio project for your grad-school applications.
