Assembling multidomain protein structures through analogous global structural alignments

X Zhou, J Hu, C Zhang, G Zhang… - Proceedings of the …, 2019 - National Acad Sciences
Proceedings of the National Academy of Sciences, 2019National Acad Sciences
Most proteins exist with multiple domains in cells for cooperative functionality. However,
structural biology and protein folding methods are often optimized for single-domain
structures, resulting in a rapidly growing gap between the improved capability for tertiary
structure determination and high demand for multidomain structure models. We have
developed a pipeline, termed DEMO, for constructing multidomain protein structures by
docking-based domain assembly simulations, with interdomain orientations determined by …
Most proteins exist with multiple domains in cells for cooperative functionality. However, structural biology and protein folding methods are often optimized for single-domain structures, resulting in a rapidly growing gap between the improved capability for tertiary structure determination and high demand for multidomain structure models. We have developed a pipeline, termed DEMO, for constructing multidomain protein structures by docking-based domain assembly simulations, with interdomain orientations determined by the distance profiles from analogous templates as detected through domain-level structure alignments. The pipeline was tested on a comprehensive benchmark set of 356 proteins consisting of 2–7 continuous and discontinuous domains, for which DEMO generated models with correct global fold (TM-score > 0.5) for 86% of cases with continuous domains and for 100% of cases with discontinuous domain structures, starting from randomly oriented target-domain structures. DEMO was also applied to reassemble multidomain targets in the CASP12 and CASP13 experiments using domain structures excised from the top server predictions, where the full-length DEMO models showed a significantly improved quality over the original server models. Finally, sparse restraints of mass spectrometry-generated cross-linking data and cryo-EM density maps are incorporated into DEMO, resulting in improvements in the average TM-score by 6.3% and 12.5%, respectively. The results demonstrate an efficient approach to assembling multidomain structures, which can be easily used for automated, genome-scale multidomain protein structure assembly.
National Acad Sciences