A-level Chemistry/WJEC/Module 4/Peptides
Amino Acids are an important class of organic compounds that contain both the amino (-NH2) and carboxyl (-COOH) groups. Of these acids, 20 serve as the building blocks of proteins. Here is a list of the "standard" amino acids, with some of their systematic names:
- alanine ((2S)-2-AminoPropanoic acid)
- arginine
- asparagine
- aspartic acid ((2S)-2-AminoButaneDioic acid)
- cysteine ((2R)-2-Amino-3-SulfanylPropanoic acid)
- glutamic acid ((2S)-2-AminoPentaneDioic acid)
- glutamine
- glycine (AminoEthanoic Acid)
- histidine
- isoleucine ((2S, 3S)-2-Amino-3-MethylPentanoic acid)
- leucine ((2S)-2-Amino-4-MethylPentanoic acid)
- lysine ((2S)-2,6-DiAmino-Hexanoic acid)
- methionine
- phenylalanine
- proline
- serine ((2S)-2-Amino-3-HydroxyPropanoic acid)
- threonine ((2S, 3R)-2-Amino-3-HydroxyButanoic acid)
- tryptophan
- tyrosine
- valine ((2S)-2-Amino-3-MethylButanoic acid)
19 of the 20 are constructed according to a general formula:
- H2NCHRCOOH
As the formula shows, the amino and carboxyl groups are both attached to a single carbon atom, which is called the alpha carbon atom. For this reason, these are known as alpha (α) amino acids. Attached to the carbon atom is a variable group (R); it is in their R groups that the molecules of the 20 standard amino acids differ from one another. In the simplest of the acids, glycine, the R consists of a single hydrogen atom. Other amino acids have more complex R groups that contain carbon as well as hydrogen and may include oxygen, nitrogen, or sulfur, as well. Proline has an unusual structure where the R group is also attached to the amine, forming a cyclic structure and making it the only secondary amine in the standard amino acids.



When a living cell makes protein, the carboxyl group of one amino acid reacts with the amino group of another to form a peptide bond. The carboxyl group of the second amino acid similarly reacts with the amino group of a third, and so on, until a long chain is produced. This chainlike molecule, which may contain from 50 to several hundred amino acid subunits, is called a polypeptide. A protein may be formed of a single polypeptide chain, or it may consist of several such chains held together by weak intermolecular bonds. Each protein is formed according to a precise set of instructions contained within DNA — the genetic material of the cell. These instructions determine which of the 20 standard amino acids are to be incorporated into the protein, and in what sequence. The R groups of the amino acid subunits determine the final shape of the protein and its chemical properties; An extraordinary variety of proteins can be produced from the same 20 subunits.
Proteins have roles such as enzymes (e.g. amylase, pepsin), hormones (e.g. insulin), receptors (e.g. the insulin receptor) and structural fibres (e.g. silk and keratin).

The standard amino acids serve as raw materials for the manufacture of many other cellular products, including hormones and pigments. In addition, several of these amino acids are key intermediates in cellular metabolism.
Stereoisomerism
[edit | edit source]The L-amino acid structures follow the "CORN" rule: Putting the H-atom at the front, and the CO, R and N groups behind the alpha-carbon, the CO-R-N groups follow a clockwise pattern.

The R/S notation is almost the same: The H-atom has lowest priority and is placed behind the alpha-carbon. The priorities are usually N > CO > R which are anticlockwise, and gives the (2S) designation. In cysteine, however, the R group contains sulfur so the priority is -NH2 > -CH2SH > -COOH and L-cysteine is designated (2R).
Two amino acids have chiral centres in their sidechains; isoleucine (2S, 3S) and threonine (2S, 3R).
Ionic forms in solution
[edit | edit source]At neutral pH, amino acids exist in the form of zwitterions. The acid and amine groups essentially neutralise one another:
- H2NCHRCOOH ⇌ H3N+CHRCOO–
At low pH, the amino acids accept a proton and become cations. At high pH, they release a proton and become anions:
- Low pH : H3N+CHRCOOH ⇌ H3N+CHRCOO– + H+ ⇌ H2NCHRCOO– + 2 H+ : High pH

The sidechains of several amino acids can donate or accept protons. Acidic amino acids are aspartic acid and glutamic acid. Basic amino acids are lysine, arginine and histidine. At neutral pH, the acidic amino acids tend to form anions and the basic amino acids tend to form cations.
Dipeptides
[edit | edit source]Two amino acids can combine in a condensation reaction to make a dipeptide:
- H3N+CHR1COO– + H3N+CHR2COO– → H3N+CHR1CO-NHCHR2COO– + H2O
The reaction is usually shown without zwitterions:
- H2NCHR1COOH + H2NCHR2COOH → H2NCHR1CO-NHCHR2COOH + H2O
Peptides are recorded, and synthesised naturally, in the sequence N to C. There is a second possible product in the reaction above:
- H2NCHR2COOH + H2NCHR1COOH → H2NCHR2CO-NHCHR1COOH + H2O
Primary Structure
[edit | edit source]The primary structure of a protein is not really a structure, it is the sequence of amino acids that make up its polypeptide chain. The sequence is listed from the end of the polypeptide with free -NH2, to the end with free -COOH. This is the same way that the protein is assembled in living cells.
- H2NCHR1COOH
- H2NCHR1CO-NHCHR2COOH
- H2NCHR1CO-NHCHR2CO-NHCHR3COOH
- H2NCHR1CO-NHCHR2CO-NHCHR3CO-NHCHR4COOH
The genetic code can specify 20 amino acids, so a dipeptide, with 2 amino acids, can have 202 = 400 possible primary sequences. A polypeptide with 99 amino acids could have 2099 = 6.34 x 10128 possible primary sequences.
One form of the enzyme α-amylase, from human saliva, has the 496 amino acid primary sequence:
EYSSN TQQGR TSIVH LFEWR WVDIA LECER YLAPK GFGGV QVSPP NENVA IHNPF RPWWE RYQPV SYKLC TRSGN EDEFR NMVTR CNNVG VRIYV DAVIN HMCGN AVSAG TSSTC GSYFN PGSRD FPAVP YSGWD FNDGK CKTGS GDIEN YNDAT QVRDC RLSGL LDLAL GKDYV RSKIA EYMNH LIDIG VAGFR IDASK HMWPG DIKAI LDKLH NLNSN WFPEG SKPFI YQEVI DLGGE PIKSS DYFGN GRVTE FKYGA KLGTV IRKWN GEKMS YLKNW GEGWG FMPSD RALVF VDNHD NQRGH GAGGA SILTF WDARL YKMAV GFMLA HPYGF TRVMS SYRWP RYFEN GKDVN DWVGP PNDNG VTKEV TINPD TTCGN DWVCE HRWRQ IRNMV NFRNV VDGQP FTNWY DNGSN QVAFG RGNRG FIVFN NDDWT FSLTL QTGLP AGTYC DVISG DKING NCTGI KIYVS DDGKA HFSIS NSAED PFIAI HAESK L
Secondary Structure
[edit | edit source]In 1951 Linus Pauling (1901-1994) deduced, using models, that polypeptide chains could fold into two repeating structures, held together by hydrogen bonds. These regular structures are largely independent of the amino acid sidechains.
Alpha helix
[edit | edit source]Polypeptide chains can fold into a helical structure where the C=O of one peptide bond forms a hydrogen bond to the N-H of a peptide bond 4 amino acids further along the chain. This structure can repeat indefinitely.

Beta sheet
[edit | edit source]Two or more polypeptide chains can fold into an extended structure where the C=O of one peptide bond forms a hydrogen bond to the N-H of a peptide bond on a second chain. The chains in the beta sheet can be loops of one polypeptide chain, or parts of entirely different chains.
Tertiary Structure
[edit | edit source]The full 3D structure of a protein is called its tertiary structure. The secondary structures are folded together by a combination of hydrophobic interactions, dipole-dipole forces, ionic bonds and (sometimes) covalent bonds. Hydrophobic interactions are where non-polar amino acid sidechains tend to be buried in the centre of the protein, away from the surrounding water.
-
Human salivary alpha-amylase structure, showing its alpha helices and beta sheets. There is a single polypeptide chain, coloured from blue (N-terminus) to red (C-terminus). Calcium and chloride ions are shown as yellow and green spheres, respectively.]]
-
Alpha-amylase wireframe model, showing every atom.