Lesson 6 of 6 / Evolution, species and phylogeny
Compare aligned sequences
Why must sequences be aligned before differences are counted?
In this lesson: Explain genome sequences and multiple sequence alignment in classification.
About 7 min
The key ideaCompare homologous positions across suitable sequences; many independent regions give stronger relationship evidence than one short match.
Explore the idea
Compare corresponding sequence positions
Compare homologous positions, one column at a time. A and B differ at 1 of 6 aligned positions. A and B have the fewest differences in this example, supporting a closer relationship in this small dataset.
Nucleotide data provide many quantifiable characters. The same underlying sequence can be compared even when organisms look very different; visible resemblance is not the only evidence available.
Original hypothetical sequences. These rows are not a claim that the shown DNA translates into the separate protein example. Real phylogenetic reconstruction uses suitable homologous regions, many more positions and an explicit evolutionary model; it need not treat every substitution or gap as equally informative. Repeated substitutions and different gene histories can complicate simple distance counting.
Explanation
A multiple sequence alignment arranges homologous DNA or amino-acid sequences so corresponding positions can be compared. Gaps may represent inferred insertions or deletions. Without alignment, one insertion can make all following bases appear mismatched even when most are homologous.
Closely related lineages often share more sequence similarity because less time has elapsed since common ancestry, but rates vary among genes and lineages. A conserved protein may have too few differences for resolving close relationships, while a rapidly changing sequence can obscure very old relationships through repeated substitutions.
Molecular methods provide many quantifiable characters, can compare organisms with very different external forms and are less dependent on subjective judgement of a few anatomical traits. Genome-scale evidence also allows multiple regions to be compared, reducing reliance on one atypical gene history.
Sequence similarity is evidence, not an automatic tree algorithm with no assumptions. Compare homologous regions, use adequate sampling and consider convergence, selection and different gene histories. Amino-acid comparison can reveal conserved protein function, while nucleotide comparison can retain synonymous differences not visible at the protein level.
Step by step
- 1
Check sequence identity
Compare homologous genes or regions.
- 2
Align and count
Treat a gap consistently with the supplied model.
- 3
Infer cautiously
A short sequence supports a hypothesis with limited confidence.
Worked example
Work through the evidence
Aligned sequences are A: ATGCCA, B: ATGCTA, C: AAGTTA. Which pair has the fewest differences?
One way to explain it
A and B differ at one of six positions; A and C differ at three, and B and C at two. A-B similarity supports closer relationship in this small dataset, with limited confidence.
Why this answer works
- Count position by position after alignment.
- Six bases are too few to make a robust whole-genome claim.
Is this true? "A matching six-base sequence proves two organisms are the same species."
Short matches can occur by chance or conservation; species relationships need broader evidence.