Predict IFN-gamma, IL-4 and IL-10 Inducing Epitopes
Skip to content

How to Predict IFN-gamma, IL-4 and IL-10 Inducing Epitopes

How to Predict IFN-gamma, IL-4 and IL-10 Inducing Epitopes

A cytokine induction screen asks which way an epitope pushes the immune response, not whether it binds. Run IFNepitope or IFNepitope2 for interferon gamma, IL4pred for IL-4 and IL10pred for IL-10, and run them after your antigenicity and toxicity filters. Treat every score as a ranking signal, never as a measurement.

What does a cytokine induction screen actually tell you?

Every other filter in a reverse vaccinology pipeline asks the same question in a different accent: is this epitope any good? MHC binding prediction asks whether it will be presented. Antigenicity, allergenicity and toxicity screening asks whether it is safe and visible to the immune system. Conservancy analysis asks whether it will still be there in the next strain.

The cytokine screen asks something different: which direction does this peptide push a CD4+ response. A peptide that drives interferon gamma release is pushing towards Th1, which is what you want against an intracellular pathogen. A peptide that drives IL-4 is pushing towards Th2. A peptide that drives IL-10 is pushing towards immune suppression.

That difference has a practical consequence, and it is the single thing most tutorials get wrong. This is the only screen in the chain where a high score can be a rejection. An IL-10 inducing peptide is a reason to drop a candidate from a prophylactic construct, not a reason to keep it. The IL10pred authors were explicit about this when they built the tool: their paper is titled “Computer-aided designing of immunosuppressive peptides based on IL-10 inducing potential”, and immunosuppression is a feature for an autoimmune or allergy application and a defect for a vaccine meant to protect.

Which cytokine should you screen for?

The answer follows from what your construct is supposed to do, not from convention. Most published multi-epitope papers screen for interferon gamma because most of them target intracellular pathogens, and copying that choice without checking your own goal is how a student ends up filtering out the epitopes they needed.

Your construct’s goalSignal you want highSignal you want lowWhy
Prophylactic vaccine against an intracellular pathogen (bacteria inside macrophages, most viruses)IFN-gammaIL-10Interferon gamma release is the arm of the Th1 response that controls intracellular pathogens
Prophylactic vaccine where antibody is the protective mechanism (many helminths, some toxins)IL-4IL-10IL-4 is a Th2 signal, and Th2 help is what drives the antibody response you are after
Therapeutic cancer constructIFN-gammaIL-10Interferon gamma carries anti-tumour activity; IL-10 in the tumour setting works against you
Tolerising peptide for autoimmunity or allergyIL-10IFN-gammaThis is the case IL10pred was actually built for: designing immunosuppressive peptides

Note what the table does not say. It does not give you a score cutoff. The servers let you set your own threshold and the papers do not publish a universal one, so any page that tells you “keep everything above 0.5” is inventing a standard that does not exist. Pick a threshold, apply it to every candidate identically, and state it in your methods.

Which tool predicts which cytokine?

All four of the standard servers come from Professor G. P. S. Raghava’s group at IIIT-Delhi, which is why they share an interface and an output style. They do not share a training set, a model or a reported metric, so read the comparison carefully before you quote a number in your dissertation.

ServerCytokineHost specific?Best model reportedReported performance (and which metric)Primary citation
IFNepitopeIFN-gammaNoHybrid (motif plus SVM)Accuracy 82.10% with MCC 0.62 on the main datasetDhanda et al. 2013, Biology Direct
IFNepitope2IFN-gammaYes, human and mouseHybrid, extra-trees basedAUROC 0.90 (human) and 0.85 (mouse). AUROC, not accuracyDhall et al. 2024, Scientific Reports
IL4predIL-4NoHybrid (amino acid pairs plus motifs)Accuracy 75.76% with MCC 0.51Dhanda et al. 2013, Clin Dev Immunol
IL10predIL-10NoRandom Forest on dipeptide compositionAccuracy 81.24% with MCC 0.59Nagpal et al. 2017, Scientific Reports

The group’s server list also carries TNFepitope, IL6Pred and HLAncPred. Those are outside the three-cytokine screen described here and were not checked while this guide was written, so look them up yourself before you build one into a protocol.

How does IFNepitope decide whether a peptide is an inducer?

Knowing the mechanism is what lets you explain a strange result to an examiner. The algorithm page sets out the training data and the three models plainly.

The training set was built from MHC class II binders in IEDB. Binders whose interferon gamma potential had never been tested were removed, which left, in the authors’ words, “3705 IFN-gamma inducing and 6728 non-inducing MHC class II binders”. That imbalance is worth carrying in your head: the negative class is nearly twice the positive one.

Three models sit on top of that data. The motif model uses the MERCI pattern discovery software to find motifs present in inducing peptides and absent from non-inducing ones, and vice versa. The SVM model, built with SVM_Light, uses amino acid and dipeptide composition. The hybrid model checks motifs first and falls back to the SVM only when no known motif is found.

That fallback order is the mechanism behind the result students most often query: two peptides with almost identical composition can receive different verdicts, because one of them carries a motif and short-circuits the SVM entirely. If a reviewer asks why a conservative substitution flipped your call, that is your answer.

The server exposes three modules and they answer three different questions:

  • Predict ranks a set of peptides you already have. This is the one you run on an MHC class II shortlist.
  • Scan generates overlapping peptides across a whole antigen and finds inducing regions. Run this before you have a shortlist, when you are deciding which part of a protein is worth mining at all.
  • Design generates single-residue mutants and looks for the minimum mutations needed to raise inducing potency. If you use it, re-run your MHC class II binding prediction on every mutant you keep, because a mutation that improves the cytokine score can destroy the binding that made the peptide a candidate.

The home page is also the cleanest one-line statement of what the tool is for. It “aims to predict and design the peptides from protein sequences having the capacity to induce IFN-gamma release from CD4+ T cells”, and it notes that “The release of IFN-gamma is the major arm of the Th1 response critical for the control of intracellular pathogens, such as Mycobacterium tuberculosis”.

Want the guided, hands-on version?

Our live Molecular Modeling & MD Simulations cohort bootcamp takes you from zero to running real docking and MD workflows, with a portfolio project for your grad-school applications.

Join the waitlist (free) →

Should you use IFNepitope or IFNepitope2?

This is a real choice, not a version upgrade you should take automatically.

IFNepitope2 is host specific. It exposes the same Predict, Design and Scan modules but under separate human and mouse routes, and its reported performance is given per host: the machine learning models reached AUROC 0.89 for human and 0.83 for mouse, and the hybrid model reached AUROC 0.90 and 0.85 respectively. Its code is on GitHub.

The practical rule is straightforward. If your next experimental step is a mouse challenge study, use IFNepitope2 and pick the mouse model, because a human-trained classifier is answering a question you are not asking. If your work is human facing and you want a method that most published pipelines will recognise, the original IFNepitope is still the common reference and its Biology Direct paper is the one most reviewers know.

One caution on quoting IFNepitope2. AUROC is not accuracy. An AUROC of 0.90 does not mean the model is right 90% of the time, and writing it that way in a thesis is the kind of error an examiner will find. Report it as AUROC and name the host.

Where does the cytokine screen belong in your pipeline?

After MHC binding, after antigenicity, after toxicity, and before conservancy. Running it earlier wastes the screen on peptides you were going to discard anyway, and running it later means you carry dead candidates into construct assembly.

The clearest worked example is our own dengue study, published in Frontiers in Immunology in 2022 and listed on our research page. Its CD4+ funnel is documented number by number:

  • NetMHCIIpan 4.0 predicted 13,501 CD4+ epitope peptides across DENV 1 to 4 against seven HLA-DR alleles (DRB1*0101, DRB1*0401, DRB1*0701, DRB5*0101, DRB1*1501, DRB1*0901 and DRB1*1302), at the default 2% and 10% rank thresholds for strong and weak binders.
  • Binding affinity cut that to 137.
  • Antigenicity cut it to 51.
  • Toxicity screening cut it to 31.
  • The cytokine screen cut 31 down to 10.
  • Conservancy across all four serotypes left 3.

The paper states it directly: “Out of 51 antigenic epitopes, 31 were screened to be non-toxic, and this number was further reduced down to 10 that were IFN-gamma inducing epitopes.”

Look at the arithmetic. The cytokine step removed more than two thirds of what reached it, which makes it the hardest single filter in that funnel. A pipeline that skips it is not saving a step, it is shipping a construct built from peptides two thirds of which were pushing in an unknown direction.

The ten surviving epitopes are named in the paper, so you can see what real output looks like rather than a made-up example: EGKIVGLYGNGVVTT, REGKIVGLYGNGVVT, GKIVGLYGNGVVTTS, TFTMRLLSPVRVPNY, SADLSLEKAAEVSWE, ATFTMRLLSPVRVPN, SSADLSLEKAAEVSW, KREKKLGEFGKAKG, KATYETDVDLGSGTR and VLRGFKKEISNMLN. Notice the overlaps: several are frame-shifted versions of the same region, which is what a sliding window produces and a reminder to deduplicate by region before you count your hits.

That study goes on to in vivo work on the assembled construct. That is a claim about the construct, not about any single predicted epitope, and you should be equally careful in your own write-up. A predicted inducer is a prediction until someone runs the assay.

Our monkeypox study ran the same step, and its methods sentence is the short answer to anyone who asks whether this is standard practice: “The webservers AllergenFP, ToxinPred and IFNepitope were used to determine the antigenic, toxic, allergic and interferon-gamma activation potential of the epitopes.”

What do you do with a strong binder that is also IL-10 inducing?

You drop it from a prophylactic construct, and you say in your methods that you dropped it and why.

This situation feels wrong the first time you hit it. The peptide passed every filter you ran: it is a strong MHC class II binder, it is antigenic, it is non-toxic, it is conserved. Then IL10pred flags it. The instinct is to keep it because four filters beat one.

Resist that. The four filters and the fifth are not voting on the same question. The first four say the peptide will be presented and will not hurt anyone. IL10pred says the response it triggers may suppress the immune reaction you are trying to build. In a construct with ten or fifteen epitopes, one immunosuppressive peptide can work against the rest of the design.

Two caveats keep this honest. A negative interferon gamma call is not the same as a bad epitope; it may simply be a Th2 epitope, and if antibody is your protective mechanism you may want it. And an IL-10 score that sits just past your own threshold on an otherwise excellent candidate is worth reporting as a limitation rather than deleting silently.

How accurate are these predictions?

Accurate enough to rank, not accurate enough to measure. The honest way to see this is to look at what happens to IFNepitope when the negative set changes.

The Biology Direct paper reports three datasets. On the main dataset, “Our best model based on the hybrid approach achieved maximum prediction accuracy of 82.10% with MCC of 0.62”. On the IFNgOnly dataset the hybrid reached “maximum accuracy of 81.39% with 0.57 MCC”. On the IFNrandom dataset, where the negatives are random peptides, the best result was “maximum accuracy of 73.4% and sensitivity of 69.18%”.

That spread is the teaching point. The same method loses roughly nine points of accuracy when the comparison set changes, which tells you the model is partly learning what MHC class II binders look like rather than what interferon gamma induction looks like. IL4pred shows the same pattern from the other side: its hybrid of amino acid pairs plus motif information reached 75.76% accuracy with MCC 0.51, while the plain SVM composition model managed MCC of only 0.29 to 0.31. IL10pred’s Random Forest on dipeptide composition reached 81.24% accuracy with MCC 0.59, against 72.30% with MCC 0.41 for amino acid composition alone and 67.15% with MCC 0.31 for an N-terminal binary profile.

Three consequences for how you write this up. Report MCC alongside accuracy, because MCC survives class imbalance and accuracy does not. Never present a server score as a probability of induction. And say in your limitations section that these are classifiers trained on a few thousand peptides, so the output is a ranked shortlist for an assay and not a result.

One more useful fact from the IL4pred paper: peptide length does not contribute to IL-4 inducing potential. If you are wondering whether to trim a 15-mer before screening, the answer for IL-4 is that it will not help.

What goes wrong, and how do you fix it?

What you seeWhat is actually happeningFix
Your shortlist collapses to almost nothing after the cytokine screenYou ran it before antigenicity and toxicity, so it consumed the full MHC binder listRe-order the pipeline. Binding, then antigenicity and toxicity, then cytokine, then conservancy
Nearly every peptide comes back non-inducingYou fed 9-mer CD8 epitopes to a tool trained on MHC class II bindersThese servers are built around class II binders. Screen your CD4+ set here and keep your CD8+ set on its own track
A good epitope is negative for interferon gamma and you discard itA negative IFN-gamma call is not a quality judgementCheck IL4pred before you drop it. It may be a Th2 epitope, which matters if antibody is your mechanism
You quote a server score as a percentage chance of inductionSVM and Random Forest scores are not calibrated probabilitiesReport the score, your threshold and the model. Never convert it to a probability
You write “IFNepitope2 is 90% accurate”You quoted an AUROC as an accuracyWrite “AUROC 0.90 for the human host” and cite the 2024 paper
The IL10pred home page shows unrelated placeholder textThe page ships with unreplaced template filler. The tool and the reference are realUse the server, but quote the Scientific Reports paper rather than the home page
Two near-identical peptides get opposite callsThe hybrid model checked motifs first and one peptide matched a known motifExpected behaviour. Report which model produced each call

Once your shortlist survives all of this, the next steps are conservancy analysis, population coverage, and then construct assembly, structure prediction and validation, docking against a TLR and MD simulation of the complex. The full sequence is laid out in our immunoinformatics pillar guide, and if you are still deciding which computational skills to build first, start from the computational biology skills roadmap.

Frequently asked questions

Do I need to screen B-cell epitopes for cytokine induction?

No. These four servers are built around MHC class II binders and CD4+ responses. B-cell epitope prediction is a separate track with its own filters, and feeding linear B-cell epitopes into a cytokine classifier produces output that means nothing.

Which threshold should I use on the IFNepitope score?

There is no published universal cutoff, and the servers let you set your own. Choose one before you look at your results, apply it to every candidate, and report it in your methods. A page that quotes a fixed threshold as standard practice is guessing.

Can I screen for IFN-gamma, IL-4 and IL-10 on the same peptide set?

Yes, and you should when your construct’s mechanism is not obvious. Run all three, then read them together. A peptide that is IFN-gamma positive and IL-10 negative is a clean Th1 candidate. One that is positive for both needs a decision you should record.

Is C-ImmSim a substitute for these tools?

No, and they answer different questions. These servers classify individual peptides by sequence. An immune simulation models the response to an assembled construct over time. They sit at different points in the pipeline and neither replaces the other.

My supervisor wants experimental validation. What can I claim from these results?

Claim that you screened and ranked candidates, and name the tool, the model and your threshold. Do not claim that any epitope induces a cytokine. That claim needs an assay, and the gap between a prediction and a measurement is exactly what a good examiner will probe.

Where do these tools fit if I am starting from scratch?

They come late. Begin with target antigen selection, then epitope prediction, then the safety screens, then this one.

Who wrote this

This guide was written by the StemSkills Lab team, whose members have more than ten years of combined work in sequence and structural bioinformatics, drug discovery and design, and multiscale molecular modeling, and who have published immunoinformatics vaccine-design studies including the dengue and monkeypox papers cited above. Primary sources for every number in this article: Dhanda SK, Vir P, Raghava GP, “Designing of interferon-gamma inducing MHC class-II binders”, Biology Direct 2013, 8:30; Dhall A, Patyal S, Raghava GP, “A hybrid method for discovering interferon-gamma inducing peptides in human and mouse”, Scientific Reports 2024, 14:26859; Dhanda SK, Gupta S, Vir P, Raghava GP, “Prediction of IL4 inducing peptides”, Clinical and Developmental Immunology 2013:263952; Nagpal G et al., “Computer-aided designing of immunosuppressive peptides based on IL-10 inducing potential”, Scientific Reports 2017, 7:42851; and the team’s own dengue multi-epitope vaccine study and monkeypox multi-epitope vaccine study.

Want the guided, hands-on version?

Our live Molecular Modeling & MD Simulations cohort bootcamp takes you from zero to running real docking and MD workflows, with a portfolio project for your grad-school applications.

Join the waitlist (free) →

Get StemSkills certified, free
Take a free assessment and earn a verifiable certificate you can download and add to your LinkedIn profile.
Browse free certifications

Keep going

Epitope Conservancy Analysis: A Step-by-Step Guide Run an epitope conservancy analysis across strains and variants, pick a defensible identity threshold, and learn the inverse… Multi-Epitope Vaccine Structure: Model, Refine, Validate Model a vaccine construct with no template, read the low pLDDT at your linkers correctly, refine it, and… How to Select Target Antigens for Reverse Vaccinology Start a vaccine design project right. Five filters, in order, that cut a whole proteome down to the…
See live workshops