Principles of Protein Structure Organization - H. Schulz 1982

Prediction of secondary structure from amino acid sequence

As a first approximation, the Introduction/11.html">Secondary Structure of a chain segment is a function of its constituent Amino Acids. The number of known Amino acid sequences far exceeds the number of known three-dimensional structures (see, for example, the Dayhoff Atlas [20]). Since the Amino Acid Sequence contains complete information regarding Cell/13.html">Protein Structure (Section 1.2, and also [177]), it is theoretically possible to determine the Spatial Structure directly from The amino acid sequence without X-ray crystallography. A logical first step in this direction would be a more thorough investigation of the relationship between the two lower Levels of Protein structural Organization (Fig. 5.1)—namely, between the amino acid sequence and the secondary structure. Although this organization is hierarchical in nature, the hierarchy is by no means strictly one-to-one, in the sense that The formation of a secondary structure in a given segment of the polypeptide chain depends not only on the amino acid sequence of that segment, but also on the Influence of other segments located at a distance along the chain*.

Initial data on this relationship were obtained using synthetic homopolymers. The first correlation between amino acid sequence and secondary structure was established by Blout et al. [328]. Based on experiments with synthetic homopolymers—Polypeptides such as poly-Glu, poly-Lys, etc.—they investigated the helix-forming and helix-disrupting tendencies of amino acid residues of seven types (Table 6.1). Residues that adopted an a-helical conformation were classified as helix-forming, whereas those that did not were classified as helix-disrupting.

* Viewed in this light, the functional dependence under consideration is somewhat analogous to the Fourier transform, in which every point in the primary space contributes to a given point in the transformed space, and vice versa.

Davies [329] applied these data to native Globular Proteins and discovered a clear anticorrelation between the helical content determined by optical rotatory dispersion (ORD) measurements [330] and the content of amino acid residues (Ser + Val + Cys + Thr + Ile). As seen from Table 6.1, the first three of these are classified as helix-disrupting residues by Blout et al. [328]. The Thr and Ile residues were chosen because of their structural similarity to Ser and Val, respectively.

Class="center">Table 6.1 Tendency to incorporate into an a-helixa


Blout et al., 1960 [328]

Kotelchuck and Scheraga, 1968 [363]

Lewis et al., 1970 [368]

Robson and Pain, 1971 [346]

Chou and Fasman, 1974 [340]

Finkelstein and Ptitsyn,

1976 [371]

A

Ala

(H)

H

I

+0.09

1.45

1.08

C

Cys

C

H

I

+0.03

0.77

0.95

D

Asp

H

C

B

—0.02

0.98

0.85

E

Glu

H

H

H

+0.12

1.53

1.15

F

Phe

(H)

H

H

+0.03

1.12

1.10

G

Gly

—

—

B

—0.05

0.53

0.55

H

His

(H)

H

I

+0.08

1.24

1.00

I

Ile

(C)

H

H

+0.07

1.00

1.05

K

Lys

(H)

C

I

—0.03

1.07

1.15

L

Leu

H

H

H

+0.11

1.34

1.25

M

Met

H

H

H

+0.10

1.20

1.15

N

Asn

(C)

C

I

—0.4

0.73

0.85

P

Pro

—


B

—

0.59

—

Q

Gln

(H)

H

I

+0.07

1.17

0.95

R

Arg

(H)

H

I

+0.02

0.79

1.05

S

Ser

C

C

B

—0.07

0.79

0.75

T

Thr

(C)

H

I

—0.01

0.82

0.75

V

Val

C

H

I

0.04

1.14

0.95

W

Trp

(H)

C

H

+0.10

1.14

1.10

Y

Tyr

(H)

C

H

—0.02

0.61

1.10

a These properties were derived from observed frequencies of occurrence. Amino acid residues are listed in alphabetical order according to their single-letter symbols. The symbols "H", "I", and "B" are used to designate helix-forming, indifferent, and helix-breaking residues, respectively. The letter "C" denotes a random coil.

Correlation analysis expands as new experimental data become available. Davies' initial success stimulated further research into correlations within globular proteins. However, the Prospects of such work remained limited as long as the experimental database—i.e., the number of known three-dimensional protein structures—remained small. Therefore, the early results of Guzzo [331], Prothero [332], Havenstein [333], Cook [334], Periti et al. [335], Danile [336], and Lowe et al. [337] are primarily of historical interest. As the experimental database expands, the accuracy of correlation improves significantly.

The proposed correlation Methods, or, as they are commonly called, "amino acid sequence-based secondary structure prediction methods" (or simply "prediction methods"), can be divided into two categories: probabilistic and physicochemical. The first category comprises methods that establish regularities based exclusively on statistical analysis of initial X-Ray Diffraction data. Physicochemical methods additionally (or exclusively) utilize other structural information. Obviously, the distinction between these two categories cannot be sharply defined, and intermediate cases may exist whose assignment to either the first or the second category is arbitrary.

This chapter will discuss all currently employed methods, since none of the modern approaches appears to have a clear advantage over the others. On the other hand, even methods that currently seem entirely unsatisfactory may contain ideas useful for the further improvement of the approach as a whole.



Last update: 06/08/2026

Editorial and Educational Adaptation: This material has been compiled based on the primary/original source text. The project team performed an editorial review, corrected technical inaccuracies, structured sections, and adapted the content for an educational format.

What was processed:

  • elimination of formatting defects (OCR errors, structural breaks, corrupted characters);
  • editorial organization of content;
  • standardization of terminology in accordance with academic sources;
  • verification of factual statements against the original source text.

All mentions of the author, publication year, and origin of the primary text have been preserved in accordance with the source.