IB Biology SL Topic 1 — Nucleic Acids Paper 1 & 2 Core idea ~11 min read

The Basis of the Genetic Code

DNA has only four letters, yet it has to spell out every protein in every living thing. The trick is that the letters are read in groups of three. Once you see why it has to be three, the rest of this topic falls into place.

📚 What you need to know

The information is in the order

Look back at the structure of DNA and one thing stands out: the sugar is always the same and the phosphate is always the same. The only part that varies from one nucleotide to the next is the base.

So all of the information a cell holds is stored in one thing — the order of A, T, C and G along the strand. A gene is simply a section of that order which codes for one polypeptide.

Think of the backbone as the paper and the bases as the letters printed on it. The paper is identical everywhere; what makes one page a recipe and another a poem is only the order of the letters.

Why the code is read in threes

Here is the problem the cell has to solve. There are 20 different amino acids that proteins are built from, but DNA only has 4 different bases. So how many bases does it take to name 20 different things?

Why three bases and not one or two? There are 20 amino acids, so the code must be able to name at least 20 things 1 base 2 bases 3 bases 4 16 64 combinations combinations combinations far too few still too few plenty of room Two bases would give only 16 combinations – not enough for 20 amino acids. Three bases give 64 combinations, so most amino acids have more than one codon.
Each extra base multiplies the options by four, because any of the four bases can sit in the new position. That is why the jump from 16 to 64 is so big.

Two bases would only give 4 × 4 = 16 combinations, which is short of 20. Three bases give 4 × 4 × 4 = 64. That is comfortably more than enough, and it explains something students often find odd: several different codons can code for the same amino acid.

Number of possible codons 4 × 4 × 4 = 43 = 64 codons for 20 amino acids

Codons and amino acids

A codon is a group of three bases. Reading a gene means starting at one end and taking the bases three at a time, without skipping or overlapping. Each codon is then matched to one amino acid, and the amino acids are joined in that exact order to build a polypeptide.

From base sequence to amino acid chain Take the bases three at a time – no skipping, no overlapping read the base sequence in groups of three A T G C C A G T T T A C codon 1 codon 2 codon 3 codon 4 1 2 3 4 amino acids, joined in this order to make a polypeptide Three bases make one codon, and one codon codes for one amino acid. The order of bases sets the order of amino acids, which sets the protein’s shape and job.
Twelve bases here give four codons and therefore four amino acids. Divide the number of bases by three and you have the number of amino acids.

Which strand gets read?

DNA has two strands, but they do not both carry the message. Only one of them, the coding strand, holds the base sequence that gets read to build the protein. The other strand acts as the template that the coding sequence is copied from.

A neat way to remember why we still need both strands: the second strand is the backup copy. Because of complementary base pairing, if you have one strand you can always rebuild the other exactly. That is the whole basis of DNA replication.

Sequence changes and their effect

Because each codon is read as a fixed block of three, changing even one base can change the codon, which can change the amino acid, which can change the shape of the finished protein. And since a protein’s shape decides its job, a change in shape can stop it working.

That said, a base change does not always cause a problem. With 64 codons for only 20 amino acids, some changes land on a different codon that still codes for the same amino acid, so nothing changes at all.

The chain of cause and effect base sequence → codon → amino acid order → protein shape → protein function

The code is universal

Here is the striking part. The same codon means the same amino acid in a bacterium, a mushroom, an oak tree and a human being. There are a handful of tiny exceptions, but the code is essentially universal.

Two big consequences come out of that, and both are common exam questions:

Conserved sequences. Over long stretches of time, mutations change base sequences. But some sequences have stayed almost identical across wildly different species — these are conserved sequences. They tend to be genes for jobs that every cell depends on, such as the proteins involved in reading DNA and building proteins, and the histone proteins that package DNA. If a sequence is that important, almost any change to it is harmful, so those changes do not get passed on.

Coding and non-coding sequences

Not every part of the genome codes for a protein. Coding sequences are the parts that do. Non-coding sequences do not code for proteins, but many of them are far from useless — they include regions that control when genes are switched on and off.

Worked examples

WORKED EXAMPLE

A coding sequence contains 900 bases. How many amino acids will the polypeptide it codes for contain?

Step 1: recall the rule 3 bases = 1 codon = 1 amino acid. Step 2: divide 900 ÷ 3 = 300 300 amino acids bases to amino acids, divide by 3 – amino acids to bases, multiply by 3
WORKED EXAMPLE

A polypeptide is 146 amino acids long. What is the minimum number of bases needed in the sequence that codes for it?

Step 1: each amino acid needs its own codon 146 × 3 = 438 bases Step 2: think about where the reading stops a stop codon is also needed to mark the end, and it is 3 bases long too 438 + 3 = 441 bases 438 bases code for the amino acids; 441 including a stop codon read the question – if it just says “codes for the amino acids”, 438 is the answer
WORKED EXAMPLE

A human gene for a hormone is inserted into a bacterium, and the bacterium makes the human hormone correctly. Explain why this is possible.

Step 1: name the property The genetic code is universal. Step 2: say what that means The same triplet of bases codes for the same amino acid in bacteria as in humans. Step 3: link it to the result So the bacterium reads the inserted gene in exactly the same way a human cell would. Same code = same amino acid order = the same protein three short linked sentences score far better than one long vague one

💡 Exam tip

⚠ Common mix-up

Up next: Nucleic Acid Structure & Function — we will put DNA and RNA side by side, meet the three types of RNA, and see just how much information one cell’s DNA can hold.

Want this explained one-to-one?

Book a free session with an experienced IB Biology tutor and get your trickiest topics made simple.

Book a Free Session →