Principles of Protein Structural Organization - H. Schultz 1982
Models, Depiction, and Documentation of Protein Structures
Protein Structure Data
Covalent Structure
In many cases, either only the covalent Structure or only the Spatial Structure of a protein is known. Information on Cell/13.html">Protein Structure may include the known Amino Acid Sequence (covalent structure), the known three-dimensional (geometric) structure, or both sets of data simultaneously. At present, significantly more covalent structures are known than spatial ones. On the other hand, there are more than a dozen geometric structures whose Amino acid sequences have not yet been established [80, 124, 217, 221, 236, 261, 303, 306, 307, 310, 313, 316]. All deciphered covalent structures are compiled in the Atlas of Protein Sequence and Structure [20]. Data on known spatial structures are summarized in Table 5.2.
Since Proteins consist of linear polypeptide chains, documenting a covalent structure is straightforward: it can be represented as a linear sequence of amino acid residues written in shorthand using single-letter symbols (Table 1.1). Such notation is easy to read. For convenience, each residue is assigned a sequential number along the chain.
Typically, the same numbering scheme is used for the sequences of homologous proteins. Aligning homologous sequences conveys significantly more information when presented in a unified format, as shown in Fig. 7.1, a. Here, identical numbering is used for all sequences; however, this is only possible if amino acid gaps, or deletions, are permitted. Difficulties arise when subsequent studies reveal homologous sequences containing one or more additional residues. When publishing these data, the established numbering scheme is usually preserved, and the extra residues are assigned the number of the preceding residue (e.g., 27) followed by letters in alphabetical order (e.g., 27A, 27B, ...). Every few years, all sequences in the Atlas of Protein Sequence and Structure [145] are renumbered, and these additional positions become standardized. However, altering familiar residue numbering causes significant inconvenience in reading and writing protein structures. It would likely be more practical to keep all numbering schemes unchanged for a substantial period after each renumbering revision, such as 10 years.
Class="center">
Fig. 7.1. Representation of amino acid sequences.
a — alignment of positions for residues 27–41 in several cytochrome c sequences from [145]. Dashes indicate gaps (deletions). b — sequence Variability histogram for cytochrome c [408]. Variability is determined by the number of different Amino Acids occurring at a given position, normalized by the frequency of The amino acid usually observed at that position.
A low substitution frequency of residues at a given position indicates the Structural and functional importance of that position. Comparison of related protein structures reveals that the sensitivity of a given position in the sequence to Amino Acid Substitutions varies dramatically. This observation is illustrated in Fig. 7.1, b. A high substitution frequency at a specific position indicates that the residue at this site is not crucial for the protein's structure or function. Since most substitutions occur on the protein surface, the pattern of such structural variability can also provide some insight into whether a given residue is located On the surface or in the interior of the protein. This is particularly useful when the spatial STRUCTURE OF THE protein is not yet known.

Fig. 7.2. Representation of Disulfide Bonds in proteins. a — the S—S bonding pattern in wheat germ agglutinin, repeating four times within a single polypeptide chain. Invariant Gly residues are highlighted. Four identical spatial forms were also found in the three-dimensional structure [316]. b — the S—S bonding pattern in human serum albumin [409]. The bonds repeat, though not quite precisely; the amino acid sequence also exhibits repeating segments. The three-dimensional structure of the protein is not yet known.
The arrangement of disulfide bridges reveals evolutionary relationships. Many proteins, particularly extracellular ones, contain disulfide bridges that covalently link distant PARTS OF THE polypeptide chain. Methods for depicting such bonds are shown in Fig. 7.2. S—S bridges tend to be conserved during evolution and are therefore characteristic of a given family of homologous sequences. Furthermore, disulfide bridges can be used to identify structural repeats within a single chain, as seen in wheat germ agglutinin (Fig. 7.2, a), whose structure was elucidated from an electron density map, and human serum albumin (Fig. 7.2, b), whose structure was determined by chemical methods.
Last update: 06/08/2026
Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.
What was processed:
- elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
- editorial organization of content;
- standardization of terminology in accordance with academic sources;
- verification of factual statements against the original source text.
All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.