How to Refine a Homology Model That Failed Validation

Protein structure refinement improves side-chain packing, removes steric clashes and corrects local backbone geometry. It cannot fix a wrong fold or a bad sequence alignment. Run a server such as ModRefiner on your model, then re-score it in MolProbity and SAVES, and report every metric before and after, not only the one that improved.
This guide sits between two steps the rest of this site already covers. Building a homology model in SWISS-MODEL and building one in MODELLER give you coordinates. Validating a protein structure tells you what the Ramachandran, MolProbity, ERRAT, Verify3D and ProSA numbers mean. This page is the step in between: your model came back with bad numbers, and you need to know what to actually run. It is part of our computational biology skills roadmap, written by the StemSkills Lab team, who have spent more than 10 years in sequence and structural bioinformatics, drug discovery and design, and multiscale molecular modeling.
One warning before any of it. Refinement makes geometry scores better. That is not the same thing as making the model more correct, and the gap between those two statements is where most student projects go wrong. Every server state, form field and command below was probed on 6 October 2026 and is reported as it behaved on that day.
What does protein structure refinement actually change?
Refinement adjusts atom positions inside the fold you already have. It does not search for a different fold. In practice it does three things: it rebuilds and repacks side chains, it pulls apart atoms that overlap, and it corrects bond angles and torsions in short stretches of backbone that came out distorted.
The primary papers are explicit about this scope. ModRefiner was built to convert reduced models into full-atom models and then relax them, and Xu and Zhang tested it on 261 proteins, reporting more accurate side-chain positions, better hydrogen-bonding networks and fewer atomic overlaps (Xu D, Zhang Y, Biophysical Journal 2011;101(10):2525-34). FG-MD, from the same group, uses short molecular dynamics guided by structural fragments and targets steric clashes, torsion angles and hydrogen bonding (Zhang J, Liang Y, Zhang Y, Structure 2011;19(12):1784-95). Both describe local, physical clean-up.
What refinement will not do is move a misthreaded loop onto the right register, repair a model built on a template that shares the wrong domain architecture, or rescue an alignment that slipped by five residues. Those are errors in the model’s topology, and no amount of energy minimisation reaches them. If that is your situation, go back to template selection or compare your options against AlphaFold and experimental structures.
When should you refine, and when should you re-model instead?
Use the validation numbers you already have as the decision rule. The useful split is between scores that report local geometry and scores that report whether the fold itself is plausible.
- Refine. Your ProSA Z-score sits in the range of native structures of similar length, Verify3D passes, and the problems are a handful of Ramachandran outliers, a MolProbity clashscore that is too high, or rotamer outliers. These are local defects inside a sound fold.
- Re-model. ProSA puts your model outside the native range, Verify3D fails over long stretches, or ERRAT shows errors spread across the whole chain rather than clustered. A model that is globally wrong needs a better template or a better alignment, not a refinement run.
- Rebuild the region first. The bad geometry is concentrated in one or two loops, or the model has gaps. Model those regions before any global refinement. Our guide to missing residues and loops covers that step, and it matters here for a practical reason given below.
There is a reason to be strict about this. Sobolev and colleagues revisited the Ramachandran Z score and showed that outlier counts alone can hide bad main-chain geometry. In their set of pathological examples, many structures had more than 90% of residues in the favoured region and all but one had under 2% outliers, yet nearly all of them scored below the Rama-Z threshold of minus three that marks geometrically improbable backbone geometry (Sobolev OV et al., Structure 2020;28(11):1249-1258.e2). The authors put it plainly: zero unexplained outliers as a gold standard “can be misleading if deviations from expected distributions are not considered”. A model that reads 100% favoured after four refinement runs has optimised a metric.
Which refinement tools are available right now?
Server availability is the real obstacle for students, because the tools that papers and protocols name have been going offline and moving hosts. Every row below was checked on 6 October 2026 by loading the submission page and reading its form.
| Tool | What it actually changes | State on 6 October 2026 | What you upload | Account or e-mail | Best for |
|---|---|---|---|---|---|
| ModRefiner | Builds and relaxes full-atom structure from a C-alpha trace, main-chain or full-atom model; side chains and backbone both flexible | Live. Submission form loads and accepts input | One single-chain PDB file, pasted or uploaded; optional reference structure | E-mail and a Zhang group password, both mandatory | A single-chain model with clashes and poor side-chain geometry |
| FG-MD | Short fragment-guided molecular dynamics; removes steric clashes, improves torsions and hydrogen bonding | Live. Submission form loads | One PDB file | E-mail and a Zhang group password, both mandatory | Physical realism after the structure is already close |
| EM-Refiner | Monte Carlo refinement guided by a cryo-EM density map | Live, but not applicable to a plain homology model | A PDB file and an MRC density map plus its resolution | Only when you have experimental cryo-EM density | |
| ModLoop | Rebuilds named loop segments using the MODELLER loop routine; leaves the rest alone | Live. Submission form loads | PDB file plus the loop segments you name | A free MODELLER licence key; e-mail optional | One or two bad loops, not a whole chain |
| GalaxyRefine2 | Iterative optimisation of local regions and overall structure together | Live today, after being unreachable on 5 October 2026 | PDB file with no missing residues | Models needing both local and global improvement, when the server is up | |
| Local energy minimisation in GROMACS | Relieves clashes and strained bonds under a force field; no conformational search | Always available once installed | Your own structure and topology | None | Removing clashes before molecular dynamics, reproducibly |
| YASARA minimisation server | Energy minimisation under the YASARA force field | Live. Page loads and accepts a PDB ID or file | A PDB ID or any PDB file | A quick second opinion, when the data is not confidential |
Two cautions that come straight off those pages. The YASARA server states that results are placed in a public download area, so unpublished thesis structures should not go there, and you need the free YASARA View to open what it returns. EM-Refiner looks like a general refinement server from its name and is not one: its form requires a density map, so a student with only a homology model has nothing to give it.
GalaxyRefine, the 2013 version, is also reachable today at galaxy.seoklab.org/refine. Its method was assessed in CASP10, where according to Heo and colleagues it “showed the best performance in improving the local structure quality” (Heo L, Park H, Seok C, Nucleic Acids Research 2013;41:W384-8). The newer GalaxyRefine2 is described in Lee GR et al., Nucleic Acids Research 2019;47(W1):W451-W455. Treat both as useful when up and unavailable when down, and keep a second route ready.
Want the guided, hands-on version?
Our live Molecular Modeling & MD Simulations cohort bootcamp takes you from zero to running real docking and MD workflows, with a portfolio project for your grad-school applications.
How do you refine a model with ModRefiner, step by step?
ModRefiner is the route to learn first, because it has stayed reachable while others moved or went dark, and because its input requirements are the simplest. Here is the submission as the live form presents it.
- Register once. The form has a mandatory password field. Request one from the Zhang group registration page before you try to submit, or the job will not go through. This is the single most common reason a first submission fails silently.
- Reduce your file to one chain. The web server supports single-chain models only. Strip extra chains, waters and HETATM ligands, and keep standard residues. The standalone 64-bit Linux build handles dimers if you need that.
- Provide the structure. Either paste the coordinates into the PDB-format text box or upload the file in the
INI_FILEfield. ModRefiner accepts a C-alpha trace, a main-chain model or a full-atom model. - Leave the reference structure empty unless you mean it. The optional second upload drives your model towards a reference structure, and only a C-alpha trace is required for it. Use it when you have a trusted related structure and want the refinement pulled towards it. Note that the strength setting described in older write-ups is commented out of the current page, so do not plan a protocol around tuning it.
- Enter your e-mail and password, then submit. Results are e-mailed, so a typo in the address loses the job with no error on screen. The page links an example output under job M42528 so you can see the format before your own arrives.
- Re-validate the returned model. Send it back through MolProbity and SAVES, and re-run ProSA. Record everything, which is the subject of the last section.
If you would rather keep the whole thing on your own machine, the same page offers standalone builds for Linux (32-bit and 64-bit), macOS and Windows. A local binary removes the queue, the e-mail round trip and the single-chain limit.
How do you refine only a bad loop?
Global refinement does not rebuild loops. It relaxes what is there. If one stretch of your model is physically impossible, the correct move is to rebuild that stretch specifically, then refine globally afterwards.
ModLoop does exactly this. It runs the loop modelling routine from MODELLER, which predicts loop conformations by satisfying spatial restraints, and the method is described in Fiser A, Sali A, Bioinformatics 2003;19(18):2500-1. The form asks for your PDB file, the loop segments you want rebuilt, a name for the model, and a MODELLER licence key, which is free for academic use. The e-mail field is optional on this server, unlike the Zhang group ones.
Order matters. Rebuild loops, then refine. Refining first wastes the run, because the refinement will faithfully relax a loop that is in the wrong place. If your model has genuine gaps rather than bad loops, fill them first using the approach in our guide to missing residues and loops, and note that GalaxyRefine2’s own page requires submitted structures to have no missing residues at all.
Can you refine locally instead of using a server?
Yes, and for anything you intend to simulate later this is the better habit. Energy minimisation under a force field relieves clashes and strained bonds without any queue, any e-mail, or any dependence on a host staying up. It performs no conformational search, so it will not repack a badly placed side chain the way ModRefiner will, but it reliably produces a structure that will start a simulation.
The mechanics are already covered on this site rather than repeated here: follow the GROMACS energy minimisation tutorial, and if the model is headed for a simulation, work through preparing a protein for molecular dynamics and assigning protonation states first. One point worth stating: the YASARA minimisation server exists as a hosted equivalent, and its method is documented in Krieger E et al., Proteins 2009;77 Suppl 9:114-22, a CASP8 paper on improving physical realism and side-chain accuracy in homology models.
How do you prove the refinement helped?
Record every metric before and after in one table, and publish the table even when a column gets worse. A single improved number proves nothing, because refinement protocols trade one score against another.
| Metric | Source | Before | After | Better? |
|---|---|---|---|---|
| Ramachandran favoured % | MolProbity | fill in | fill in | fill in |
| Ramachandran outliers % | MolProbity | fill in | fill in | fill in |
| Clashscore | MolProbity | fill in | fill in | fill in |
| Rotamer outliers % | MolProbity | fill in | fill in | fill in |
| ERRAT overall quality factor | SAVES | fill in | fill in | fill in |
| Verify3D pass fraction | SAVES | fill in | fill in | fill in |
| ProSA Z-score | ProSA-web | fill in | fill in | fill in |
| C-alpha RMSD to input | PyMOL | 0.00 | fill in | n/a |
The last row is the one students skip and reviewers ask about. Measure how far the refinement moved your model from where it started, by superimposing the two coordinate files; our guide to superimposing two proteins in PyMOL covers the command. A large displacement with only a small score gain is a warning, not a result. Where a Rama-Z score is available from Phenix or PDB-REDO, add it, because it catches over-restrained geometry that favoured-region percentages miss.
Two practical notes on the validation servers. MolProbity is reachable at http://molprobity.biochem.duke.edu, and on 6 October 2026 the https address failed on a self-signed certificate while the http address answered normally, so check the scheme before concluding the service is down. ProSA-web is at prosa.services.came.sbg.ac.at.
Troubleshooting: real failures and what to do about them
A refinement link from a paper or protocol lands somewhere unexpected. This is now common. On 5 October 2026 both GalaxyRefine addresses redirected to a status page; on 6 October 2026 both answered normally. The 3Drefine server at sysbio.rnet.missouri.edu now redirects to a different host that returns 404, so it is effectively unavailable whatever a protocol says. Check the server yourself on the day you work, and never wait on a host that is answering with a status page or an error.
You submitted to ModRefiner or FG-MD and nothing came back. Both forms require an e-mail address and a password, and results arrive only by e-mail. A missing registration or a mistyped address produces no visible error. Register first, check the address character by character, and look in your spam folder before resubmitting.
The server rejects your file. The usual causes are several chains in one file, HETATM ligands or waters left in, non-standard residues, or missing atoms. The ModRefiner web server takes single-chain models only, and GalaxyRefine2 requires no missing residues. Clean the file to one chain of standard residues with complete atoms, and model gaps before submitting.
Ramachandran improved but ERRAT or Verify3D got worse. That is a trade, not a success. Report both directions. If the fold-level scores degrade while local geometry improves, the refinement is smoothing geometry at the cost of the model’s agreement with sequence-structure compatibility, and you should prefer the input model.
Geometry now reads near-perfect but the model is still wrong. Check the template and the alignment, not the refinement settings. A structure can show over 90% favoured and under 2% outliers and still have improbable main-chain geometry, as Sobolev and colleagues documented. If the template was wrong, the fix is a better template.
Refinement moved side chains in your binding site after you had defined the docking box. Refine first, then re-run pocket detection and re-define the site. Patching afterwards invalidates the grid you already set. If docking is the destination, go through protein and ligand preparation for docking after refinement, not before.
A loop is still unusable after a global refinement run. Expected. Global refinement relaxes, it does not rebuild. Use ModLoop or the MODELLER loop routine on that segment, then refine globally.
Frequently asked questions
Does refining a homology model make it more accurate?
Sometimes, but the scores you measure are geometry scores, not accuracy. Refinement reliably improves side-chain packing, clashes and local backbone geometry. Whether the model is closer to the true structure depends on the template and alignment, which refinement cannot change.
How many times should I refine a model?
Once, then evaluate. Repeated resubmission until a single metric reads perfectly is metric optimisation, not modelling. If one pass does not bring the model into an acceptable range, the problem is usually the template or the alignment rather than the refinement protocol.
Is GalaxyRefine still available in 2026?
It was reachable on 6 October 2026, both the original server and GalaxyRefine2, after being unreachable the previous day. Availability has been intermittent, so check the submission page yourself before planning a protocol around it and keep ModRefiner or local minimisation as a fallback.
Can I refine an AlphaFold model the same way?
You can, but read the confidence scores first. Low-confidence tails and loops are better trimmed or rebuilt than refined. Our guide to interpreting pLDDT and PAE scores explains which regions are worth keeping, and predicting structures with AlphaFold covers the upstream step.
What do I write in my methods section about refinement?
Name the tool and its version or server, the date you ran it, the exact input you gave it, and the full before and after validation table including metrics that worsened. Cite the tool’s primary paper. Reviewers accept a modest, honestly reported improvement and reject an unexplained perfect score.
Do I need refinement before molecular dynamics?
You need clash-free starting coordinates, which a local energy minimisation provides as part of normal simulation setup. A separate server refinement is optional and is most useful when side-chain geometry is poor and the structure will be used for docking or interpretation rather than only as a simulation start point.
The honest version of this workflow is short: decide from your validation numbers whether the fold is sound, rebuild bad loops before anything global, run one refinement pass on a clean single-chain file, and report every metric in both directions. A model that improves modestly and is described accurately will survive a thesis defence. One that reads perfectly for reasons you cannot explain will not.
Want the guided, hands-on version?
Our live Molecular Modeling & MD Simulations cohort bootcamp takes you from zero to running real docking and MD workflows, with a portfolio project for your grad-school applications.
