Jump to content

A-level Chemistry/WJEC/Module 4/Peptides

From Wikibooks, open books for an open world

Amino Acids are an important class of organic compounds that contain both the amino (-NH2) and carboxyl (-COOH) groups. Of these acids, 20 serve as the building blocks of proteins. Here is a list of the "standard" amino acids, with some of their systematic names:

  • alanine ((2S)-2-AminoPropanoic acid)
  • arginine
  • asparagine
  • aspartic acid ((2S)-2-AminoButaneDioic acid)
  • cysteine ((2R)-2-Amino-3-SulfanylPropanoic acid)
  • glutamic acid ((2S)-2-AminoPentaneDioic acid)
  • glutamine
  • glycine (AminoEthanoic Acid)
  • histidine
  • isoleucine ((2S, 3S)-2-Amino-3-MethylPentanoic acid)
  • leucine ((2S)-2-Amino-4-MethylPentanoic acid)
  • lysine ((2S)-2,6-DiAmino-Hexanoic acid)
  • methionine
  • phenylalanine
  • proline
  • serine ((2S)-2-Amino-3-HydroxyPropanoic acid)
  • threonine ((2S, 3R)-2-Amino-3-HydroxyButanoic acid)
  • tryptophan
  • tyrosine
  • valine ((2S)-2-Amino-3-MethylButanoic acid)

19 of the 20 are constructed according to a general formula:

H2NCHRCOOH

As the formula shows, the amino and carboxyl groups are both attached to a single carbon atom, which is called the alpha carbon atom. For this reason, these are known as alpha (α) amino acids. Attached to the carbon atom is a variable group (R); it is in their R groups that the molecules of the 20 standard amino acids differ from one another. In the simplest of the acids, glycine, the R consists of a single hydrogen atom. Other amino acids have more complex R groups that contain carbon as well as hydrogen and may include oxygen, nitrogen, or sulfur, as well. Proline has an unusual structure where the R group is also attached to the amine, forming a cyclic structure and making it the only secondary amine in the standard amino acids.

The 20 amino acids specified by the genetic code.
Leucine is a typical amino acid, with a primary amine group.
Proline is the only standard amino acid with a secondary amine group.

When a living cell makes protein, the carboxyl group of one amino acid reacts with the amino group of another to form a peptide bond. The carboxyl group of the second amino acid similarly reacts with the amino group of a third, and so on, until a long chain is produced. This chainlike molecule, which may contain from 50 to several hundred amino acid subunits, is called a polypeptide. A protein may be formed of a single polypeptide chain, or it may consist of several such chains held together by weak intermolecular bonds. Each protein is formed according to a precise set of instructions contained within DNA — the genetic material of the cell. These instructions determine which of the 20 standard amino acids are to be incorporated into the protein, and in what sequence. The R groups of the amino acid subunits determine the final shape of the protein and its chemical properties; An extraordinary variety of proteins can be produced from the same 20 subunits.

Proteins have roles such as enzymes (e.g. amylase, pepsin), hormones (e.g. insulin), receptors (e.g. the insulin receptor) and structural fibres (e.g. silk and keratin).

Oxytocin is a 9-amino acid peptide hormone which is involved in human social bonding, the beginning of childbirth and lactation.

The standard amino acids serve as raw materials for the manufacture of many other cellular products, including hormones and pigments. In addition, several of these amino acids are key intermediates in cellular metabolism.

Stereoisomerism

[edit | edit source]

The L-amino acid structures follow the "CORN" rule: Putting the H-atom at the front, and the CO, R and N groups behind the alpha-carbon, the CO-R-N groups follow a clockwise pattern.

The CORN rule for L-amino acids.

The R/S notation is almost the same: The H-atom has lowest priority and is placed behind the alpha-carbon. The priorities are usually N > CO > R which are anticlockwise, and gives the (2S) designation. In cysteine, however, the R group contains sulfur so the priority is -NH2 > -CH2SH > -COOH and L-cysteine is designated (2R).

Two amino acids have chiral centres in their sidechains; isoleucine (2S, 3S) and threonine (2S, 3R).

Ionic forms in solution

[edit | edit source]

At neutral pH, amino acids exist in the form of zwitterions. The acid and amine groups essentially neutralise one another:

H2NCHRCOOH ⇌ H3N+CHRCOO

At low pH, the amino acids accept a proton and become cations. At high pH, they release a proton and become anions:

Low pH : H3N+CHRCOOH ⇌ H3N+CHRCOO + H+ ⇌ H2NCHRCOO + 2 H+ : High pH
Glycine at a range of pH values.

The sidechains of several amino acids can donate or accept protons. Acidic amino acids are aspartic acid and glutamic acid. Basic amino acids are lysine, arginine and histidine. At neutral pH, the acidic amino acids tend to form anions and the basic amino acids tend to form cations.

Dipeptides

[edit | edit source]

Two amino acids can combine in a condensation reaction to make a dipeptide:

H3N+CHR1COO + H3N+CHR2COO → H3N+CHR1CO-NHCHR2COO + H2O

The reaction is usually shown without zwitterions:

H2NCHR1COOH + H2NCHR2COOH → H2NCHR1CO-NHCHR2COOH + H2O

Peptides are recorded, and synthesised naturally, in the sequence N to C. There is a second possible product in the reaction above:

H2NCHR2COOH + H2NCHR1COOH → H2NCHR2CO-NHCHR1COOH + H2O

Primary Structure

[edit | edit source]

The primary structure of a protein is not really a structure, it is the sequence of amino acids that make up its polypeptide chain. The sequence is listed from the end of the polypeptide with free -NH2, to the end with free -COOH. This is the same way that the protein is assembled in living cells.

H2NCHR1COOH
H2NCHR1CO-NHCHR2COOH
H2NCHR1CO-NHCHR2CO-NHCHR3COOH
H2NCHR1CO-NHCHR2CO-NHCHR3CO-NHCHR4COOH

The genetic code can specify 20 amino acids, so a dipeptide, with 2 amino acids, can have 202 = 400 possible primary sequences. A polypeptide with 99 amino acids could have 2099 = 6.34 x 10128 possible primary sequences.

Most calculators can't calculate 2099 but there is a trick using the relationship between standard form numbers and log10. Note that log10(2099) = 99 x log10(20) = 128.8020. This means 2099 = 10128.8020 = 100.8020 x 10128 = 6.34 x 10128.

One form of the enzyme α-amylase, from human saliva, has the 496 amino acid primary sequence:

EYSSN TQQGR TSIVH LFEWR WVDIA LECER YLAPK GFGGV QVSPP NENVA IHNPF RPWWE RYQPV SYKLC TRSGN EDEFR NMVTR CNNVG VRIYV DAVIN HMCGN AVSAG TSSTC GSYFN PGSRD FPAVP YSGWD FNDGK CKTGS GDIEN YNDAT QVRDC RLSGL LDLAL GKDYV RSKIA EYMNH LIDIG VAGFR IDASK HMWPG DIKAI LDKLH NLNSN WFPEG SKPFI YQEVI DLGGE PIKSS DYFGN GRVTE FKYGA KLGTV IRKWN GEKMS YLKNW GEGWG FMPSD RALVF VDNHD NQRGH GAGGA SILTF WDARL YKMAV GFMLA HPYGF TRVMS SYRWP RYFEN GKDVN DWVGP PNDNG VTKEV TINPD TTCGN DWVCE HRWRQ IRNMV NFRNV VDGQP FTNWY DNGSN QVAFG RGNRG FIVFN NDDWT FSLTL QTGLP AGTYC DVISG DKING NCTGI KIYVS DDGKA HFSIS NSAED PFIAI HAESK L

Secondary Structure

[edit | edit source]

In 1951 Linus Pauling (1901-1994) deduced, using models, that polypeptide chains could fold into two repeating structures, held together by hydrogen bonds. These regular structures are largely independent of the amino acid sidechains.

Alpha helix

[edit | edit source]

Polypeptide chains can fold into a helical structure where the C=O of one peptide bond forms a hydrogen bond to the N-H of a peptide bond 4 amino acids further along the chain. This structure can repeat indefinitely.

A section of polypeptide, folded into an alpha-helix, showing the hydrogen bonds in red. The "ribbon" illustration is a common way to show the presence of an alpha helix.

Beta sheet

[edit | edit source]

Two or more polypeptide chains can fold into an extended structure where the C=O of one peptide bond forms a hydrogen bond to the N-H of a peptide bond on a second chain. The chains in the beta sheet can be loops of one polypeptide chain, or parts of entirely different chains.

Tertiary Structure

[edit | edit source]

The full 3D structure of a protein is called its tertiary structure. The secondary structures are folded together by a combination of hydrophobic interactions, dipole-dipole forces, ionic bonds and (sometimes) covalent bonds. Hydrophobic interactions are where non-polar amino acid sidechains tend to be buried in the centre of the protein, away from the surrounding water.