Changes in information on the number of human proteoforms, post-translational modification (PTM) events, alternative splicing (AS), single-amino acid polymorphisms (SAP) associated with protein-coding genes in the neXtProt database have been retrospectively analyzed. In 2016, our group proposed three mathematical models for predicting the number of different proteins (proteoforms) in the human proteome. Eight years later, we compared the original data of the information resources and their contribution to the prediction results, correlating the differences with new approaches to experimental and bioinformatic analysis of protein modifications. The aim of this work is to update information on the status of records in the databases of identified proteoforms since 2016, as well as to identify trends in changes in the quantities of these records. According to various information models, modern experimental methods may identify from 5 to 125 million different proteoforms: the proteins formed due to alternative splicing, the implementation of single nucleotide polymorphisms at the proteomic level, and post-translational modifications in various combinations. This result reflects an increase in the size of the human proteome by 20 or more times over the past 8 years.
Sarygina E.V., Kozlova A.S., Ponomarenko E.A., Ilgisonis E.V. (2024) The human proteome size as a technological development function. Biomeditsinskaya Khimiya, 70(5), 364-373.
Sarygina E.V. et al. The human proteome size as a technological development function // Biomeditsinskaya Khimiya. - 2024. - V. 70. -N 5. - P. 364-373.
Sarygina E.V. et al., "The human proteome size as a technological development function." Biomeditsinskaya Khimiya 70.5 (2024): 364-373.
Sarygina, E. V., Kozlova, A. S., Ponomarenko, E. A., Ilgisonis, E. V. (2024). The human proteome size as a technological development function. Biomeditsinskaya Khimiya, 70(5), 364-373.
References
Aebersold R., Agar J.N., Amster I.J., Baker M.S., Bertozzi C.R., Boja E.S., Costello C.E., Cravatt B.F., Fenselau C., Garcia B.A., Ge Y., Gunawardena J., Hendrickson R.C., Hergenrother P.J., Huber C.G., Ivanov A.R., Jensen O.N., Jewett M.C., Kelleher N.L., Kiessling L.L., Krogan N.J., Larsen M.R., Loo J.A., Ogorzalek Loo R.R., Lundberg E., MacCoss M.J., Mallick P., Mootha V.K., Mrksich M., Muir T.W., Patrie S.M., Pesavento J.J., Pitteri S.J., Rodriguez H., Saghatelian A., Sandoval W., Schlüter H., Sechi S., Slavoff S.A., Smith L.M., Snyder M.P., Thomas P.M., Uhlén M., van Eyk J.E., Vidal M., Walt D.R., White F.M., Williams E.R., Wohlschlager T., Wysocki V.H., Yates N.A., Young N.L., Zhang B. (2018) How many human proteoforms are there? Nat. Chem. Biol., 14(3), 206–214. CrossRef Scholar google search
Zhang F., Chen J.Y. (2016) A method for identifying discriminative isoform-specific peptides for clinical proteomics application. BMC Genomics, 17(Suppl 7), 522. CrossRef Scholar google search
Prabakaran S., Lippens G., Steen H., Gunawardena J. (2012) Post-translational modification: Nature’s escape from genetic imprisonment and the basis for dynamic information encoding. Wiley Interdiscip. Rev. Syst. Biol. Med., 4(6), 565–583. CrossRef Scholar google search
Schlüter H., Apweiler R., Holzhütter H.G., Jungblut P.R. (2009) Finding one’s way in proteomics: A protein species nomenclature. Chem. Cent. J., 3, 11. CrossRef Scholar google search
Smith L.M., Kelleher N.L., Consortium for Top Down Proteomics (2013) Proteoform: A single term describing protein complexity. Nat. Methods, 10(3), 186–187. CrossRef Scholar google search
Semba R.D., Enghild J.J., Venkatraman V., Dyrlund T.F., van Eyk J.E. (2013) The human eye proteome project: Perspectives on an emerging proteome. Proteomics, 13(16), 2500–2511. CrossRef