How to Run an IEDB Population Coverage Analysis for Your Epitope Set

The IEDB Population Coverage tool estimates the fraction of a population that carries at least one HLA allele able to present your epitopes. Enter each epitope with its restricted alleles, pick the populations, and choose class I, class II or combined. It returns projected coverage, average epitope hits, and PC90, a count and not a percentage.
This is the step that reviewers of a reverse vaccinology manuscript ask for and that most student projects skip. You have a shortlist of predicted binders. What you do not yet know is whether those binders are presentable by the HLA types that people in your target region actually carry. This guide picks up exactly where our post on T-cell epitope and MHC binding prediction with NetMHCpan and IEDB ends: a filtered set of strong binders, with allele names already in hand.
What do you need before you start?
Three things, and nothing else. No structures, no docking output, no sequences beyond the peptides themselves.
- A final epitope shortlist. Peptide sequences, or names you can map back to sequences.
- The restricting alleles for each epitope. These come straight out of your class I or class II binding prediction run. Every allele must be written at two-field resolution, for example HLA-A*02:01, which matters more than students expect and is covered in the troubleshooting section below.
- A decision about which populations you care about. Region, country, ethnicity, or a custom population of your own.
Run coverage before you commit to a construct, not after. It is much cheaper to add two more epitopes now than to rebuild a linker-joined construct later. If your shortlist has not yet been through antigenicity, allergenicity and toxicity filters, do that screening step first, because epitopes you are about to drop should not be inflating your coverage number.
What exactly is population coverage, and what is PC90?
T cells recognise a complex of a peptide and a specific MHC molecule, so an epitope only produces a response in people who carry an MHC molecule that binds it. Human MHC is extremely polymorphic. The IEDB Population Coverage tutorial states that over a thousand different human HLA alleles are known and that different HLA types are expressed at dramatically different frequencies in different ethnicities. Its warning is the reason this step exists: without careful consideration, “a vaccine or diagnostic with ethnically biased population coverage could result”.
The tool answers that by combining your epitope-to-allele map with HLA genotypic frequencies. For every population you select, it returns exactly three numbers.
- Projected population coverage. The percentage of individuals predicted to respond to at least one epitope in your set.
- Average number of epitope hits / HLA combinations recognised by the population. The mean depth of the response, not its breadth.
- PC90. Defined by the help page as the “minimum number of epitope hits / HLA combinations recognized by 90% of the population”.
PC90 is a count, not a percentage. This is the single most common error on competing tutorials and on student posters. A PC90 of 4.5 does not mean 4.5% of anything. It means that 90% of that population is predicted to recognise at least 4.5 epitope-allele combinations from your set. Higher is better, and a set with high coverage but a PC90 near 1 is fragile, because most responders are hanging on a single epitope.
The frequency data behind all of this comes from the Allele Frequency Net Database, credited on the tool page to Derek Middleton. The IEDB help page states that the final merged set contains frequencies of “3,245 alleles (including class I and class II) for the world, 16 geographical areas, 21 ethnicities, 115 countries and ethnicities by country”. Frequencies are estimated as AF = a/2n, where a is the number of allele copies and n the number of subjects.
The canonical method paper is Bui HH, Sidney J, Dinh K, Southwood S, Newman MJ and Sette A, Predicting population coverage of T-cell epitope-based diagnostics and vaccines, BMC Bioinformatics 7:153 (2006), PMID 16545123. Cite that, not the tool URL alone. The IEDB reference page prints the volume as 17:153; the article is in volume 7.
How do you format the input file?
There are two file formats and students routinely confuse them.
The epitope and allele file
Plain text, two tab-delimited columns, and no header line. Column one is the epitope name or sequence. Column two is a comma-delimited list of the MHC alleles that epitope is restricted to. One epitope-allele combination per line. The examples printed on the IEDB help page look like this:
FMKAVCVEV HLA-A*02:01,HLA-A*02:02,HLA-A*02:03,HLA-A*02:06,HLA-A*68:02
FLIFFDLFLV HLA-A*02:01,HLA-A*02:02,HLA-A*02:03,HLA-A*02:06,HLA-A*68:02
GLIMVLSFL HLA-A*02:01,HLA-A*02:02,HLA-A*02:03,HLA-A*02:06,HLA-A*23:01The separator between the two columns is a tab. The separator inside column two is a comma with no space. Spreadsheets love to turn tabs into spaces on export, which is the number one reason an upload silently fails.
The user population file
You only need this if your target population is not one of the built-in ones. Here a header line is required. The first three column headings must be “MHC class”, “MHC locus” and “MHC allele”, and every column after that is a population name. Column one holds the class (I or II), column two the locus (HLA-A, HLA-B and so on), column three the two-field allele, and the remaining columns hold genotypic frequencies from 0 to 1. The help page prints this example:
| MHC Class | MHC Locus | MHC Allele | Asian | Black | European-Caucasian | North-America-Caucasian |
|---|---|---|---|---|---|---|
| I | HLA-A | HLA-A*01:01 | 0.007594 | 0.035415 | 0.167566 | 0.159952 |
| I | HLA-A | HLA-A*02:01 | 0.082624 | 0.103242 | 0.258922 | 0.175362 |
| I | HLA-A | HLA-A*02:02 | 0.000944 | 0.044661 | 0.007503 | 0.018633 |
Read those numbers slowly, because they are the whole argument for running this analysis. HLA-A*01:01 sits at 0.167566 in the European-Caucasian column and 0.007594 in the Asian column, a difference of more than twentyfold. HLA-A*02:01, the allele that dominates most published epitope shortlists, is roughly three times more frequent in the European-Caucasian column than in the Asian one. An epitope set tuned on HLA-A*02:01 will look strong globally and thin in South Asia.
Order matters. The help page is explicit that if you want to add user populations, you must do it before entering the epitope set. Adding them afterwards means starting the form again.
Which calculation option should you pick?
The tool offers exactly three calculation options, because class I and class II restricted epitopes elicit responses from two different T-cell populations, CTL and HTL respectively. Pick based on what your construct contains and what your reader will ask for.
| Calculation option | What it answers | T-cell population | What the input must contain | What a reviewer expects |
|---|---|---|---|---|
| Class I separate | What fraction of the population can present at least one of my CD8 epitopes? | CTL | Only class I alleles (HLA-A, HLA-B, HLA-C) | Reported for any construct claiming cytotoxic responses |
| Class II separate | What fraction can present at least one of my CD4 epitopes? | HTL | Only class II alleles (HLA-DRB1, HLA-DQ, HLA-DP) | Reported for any construct claiming helper responses |
| Class I and II combined | What fraction can mount both arms of the response from my set? | CTL and HTL together | Both class I and class II epitopes, in one file | The headline number for a multi-epitope construct |
Run all three when your construct carries both epitope types. The combined figure is the one that belongs in your abstract, but the two separate figures are what tell you which arm is weak. Multiple populations can be calculated at once, and the tool also generates an average across them.
Want the guided, hands-on version?
Our live Molecular Modeling & MD Simulations cohort bootcamp takes you from zero to running real docking and MD workflows, with a portfolio project for your grad-school applications.
How do you select populations, and where does India sit?
Populations can be queried by area, by country, or by ethnicity. The hierarchy is geographical area, then country, then ethnicity. India sits under the area South Asia, with the country entry India and the ethnicity entry Asian. Pakistan and Sri Lanka sit in the same area, with Pakistan carrying both an Asian and a Mixed ethnicity entry.
For a project aimed at an Indian cohort, run at least three selections side by side: South Asia as the area, India as the country, and one comparison population such as Europe or North America. Then read the gap rather than the single best number. If your India figure trails your Europe figure by a wide margin, your shortlist is carrying European-frequent alleles and needs work, and no amount of reporting the world average will hide that from a reviewer.
How do you read the output?
Take the three returned numbers in order and ask a different question of each.
- Projected coverage. Breadth. Is anyone left out? Compare it across populations, never in isolation.
- Average epitope hits. Depth. A high average with low coverage means your epitopes are piled onto a few common alleles.
- PC90. Margin for the least well served responders. If PC90 is close to 1, the bottom decile of your covered population responds through a single epitope-allele combination, and losing one epitope to a later toxicity or allergenicity filter would take them out entirely.
For a sense of what published numbers look like, our own monkeypox virus construct reported a world population coverage of 85.97% for its two CD4 epitopes and 87.03% for its four CD8 epitopes, using this tool (Akhtar et al., Viruses, 2022). Those are that construct’s results on that epitope set, not a target you should aim at, and your own numbers will differ with your antigen, your allele panel and your populations.
How do you improve poor coverage?
There is one wrong answer and three right ones.
The wrong answer: loosening the %Rank or IC50 threshold in your binding prediction until coverage improves. That does not add coverage, it adds weak binders that will not be presented, and it invalidates the shortlist you already screened. Never fix a coverage problem by relaxing a binding threshold.
The three that work:
- Add epitopes restricted to supertypes you have not represented. Look at which loci your set leans on. A shortlist that is all HLA-A will underperform in any population where HLA-B diversity carries the response.
- Go back to prediction with a wider allele panel. Re-run your class I and class II predictions including the alleles that are frequent in your weakest population, then take the strong binders from that run. This is the same workflow as before, just with a different allele list, so nothing about the method changes.
- Accept a longer construct. Coverage is bought with epitopes. If two more validated epitopes fix South Asia, add them and revisit linker design when you assemble the construct.
Once coverage passes, the chain continues: assemble the construct, then take it into receptor docking. The full sequence of steps, and where each one sits, is laid out on our immunoinformatics pillar guide.
Troubleshooting: real errors and their fixes
| What you see | Why it happens | Fix |
|---|---|---|
| An allele contributes nothing to coverage although it is in your file | You wrote it at one-field resolution, for example HLA-A*07. The merged frequency set excludes alleles with less than two fields | Rewrite at two-field resolution, for example HLA-A*02:01 |
| A four-field allele appears to be ignored | Alleles at higher resolution such as HLA-A*02:01:01:02 are collapsed to two fields before the frequencies are computed | Truncate to two fields yourself so your input matches what the tool scores |
| Coverage is far lower than expected for one country | The allele is genuinely rare or absent in the frequency set for that population, which is a real result, not a bug | Report it and add epitopes restricted to alleles that are frequent there |
| The combined option returns nothing useful | Your file has only class I alleles, or only class II, so there is no second arm to combine | Run the relevant separate option, or add the missing epitope class to the file |
| Your user populations are not offered on the form | They were added after the epitope set was entered | Restart and add user populations first, as the help page instructs |
| The upload is rejected or parsed into one column | Tabs were converted to spaces, usually by a spreadsheet export | Save as tab-delimited text and check with a plain text editor before uploading |
| An allele name is rejected outright | Legacy or non-standard nomenclature, for example HLA-A2 instead of HLA-A*02:01 | Check the current name against hla.alleles.org or the IPD-IMGT/HLA database |
Frequently asked questions
Is PC90 a percentage?
No. PC90 is a count. The IEDB help page defines it as the minimum number of epitope hits or HLA combinations recognised by 90% of the population. A PC90 of 3 means the least well served 10% boundary of your covered population still sees at least three epitope-allele combinations.
What coverage percentage is good enough?
There is no published threshold, and any page that gives you one without a citation is guessing. What reviewers look for is that you reported coverage for the populations your vaccine targets, that you reported class I and class II separately as well as combined, and that you did not hide a weak region behind a world average.
Can I run population coverage without predicting epitopes first?
No. The tool takes an epitope-to-allele map as input and does not predict anything itself. Run class I and class II binding prediction first, filter to strong binders, and bring the surviving allele lists here.
Where do the HLA frequencies come from?
From the Allele Frequency Net Database. Individual population studies from around the world are merged into a hierarchy of geographical area, country and ethnicity, and frequencies are estimated at two-field resolution using AF = a/2n.
Do I need to install anything?
No. The analysis runs in the browser at tools.iedb.org/population, part of the Immune Epitope Database and Analysis Resource. A downloadable version is offered from the same site if you need to script it, and the reference page lists the citations to include in your manuscript.
Should I run coverage before or after allergenicity and toxicity screening?
After. Screening removes epitopes, and an epitope you are about to discard should not be inflating your coverage figure. Screen first, then compute coverage on what survives.
Who wrote this
This guide was written by the StemSkills Lab team, whose members have more than a decade of combined work in sequence and structural bioinformatics, drug discovery and design, and multiscale molecular modeling. The team has published multi-epitope vaccine designs against dengue, monkeypox, canine circovirus, feline infectious peritonitis virus, rotavirus and Candida dubliniensis, two of them with in-vivo validation. The full list is on the research and publications page. If you are working out where this step sits in a wider skill set, the computational biology skills roadmap maps the whole path.
Want the guided, hands-on version?
Our live Molecular Modeling & MD Simulations cohort bootcamp takes you from zero to running real docking and MD workflows, with a portfolio project for your grad-school applications.
