Skip to content

Reading Sanger Sequencing Results After Cloning: A Practical Walkthrough

how to interpret sequencing results after cloningMay 21, 2026

Reading Sanger Sequencing Results After Cloning: A Practical Walkthrough

Your colony PCR was clean, the diagnostic digest banded right, and you sent four colonies for Sanger. The traces came back, you opened the alignment, and now you have a screenful of mismatches, ambiguous calls, and what looks like a small insertion 250 bp into the read. Are these real, or are they sequencing artifact? Which colonies do you bank, which do you discard, and which do you re-sequence with a different primer?

This post walks through the analysis workflow for Sanger reads after a cloning experiment: how to read the chromatogram, what counts as a clean alignment, the four common mismatch patterns and what each one means, and when to give up on a clone versus sequence it from the other direction. It assumes you have an aligned trace open in a plasmid editor — SnapGene, Benchling, A Plasmid Editor (ApE), or PlasmidStudio — not just a FASTA dump.

Step 1: Look at the trace before you look at the alignment

Open the chromatogram (.ab1 file) before the alignment view. The trace tells you whether the sequencing reaction itself worked. Three things to check, in order:

  • Peak height and shape: clean single peaks with consistent height across the read. Heights vary by base (A and G are usually taller than C and T on most platforms), but within a base type, peak heights should be reasonably even. A trace where peaks collapse to baseline noise after 200 bp suggests low template concentration or primer-template mismatch.
  • Quality scores: most sequencing services include Phred or platform-equivalent quality scores. Q≥30 across the readable region is the practical bar; trim ends with Q<20 before analysis. Addgene’s sequencing troubleshooting writeup covers the quality-score conventions in more depth.
  • The first 30–50 bp: these are almost always low-quality because of the sequencing primer itself. Don’t panic about garbage right after the primer — that’s expected. The readable region typically starts 30–50 bp downstream of the primer 3′ end and runs to 800–1000 bp.

If the trace is clean — even peaks, high Q across the middle of the read, no premature dropoff — the sequencing worked and any mismatches you see in the alignment are real signal, not noise. If the trace is messy throughout, the alignment doesn’t tell you about your construct; it tells you about the sequencing reaction. Re-submit or re-sequence with fresh template before drawing conclusions.

Step 2: Check coverage before you check identity

A Sanger read covers about 800–1000 bp of usable sequence per primer. For inserts longer than that, you need multiple reads from primers spaced along the construct, or a whole-plasmid read from Plasmidsaurus / Primordium / Elim. Before celebrating a clean alignment, confirm:

  • The entire insert is covered, ideally with two primers (one forward, one reverse) so every base is read at least once. Junctions — the vector-insert boundaries on both sides — must be in the readable region.
  • The vector-insert junctions are bridged. If you cloned via two restriction enzymes, both junction regions matter; if you used Gibson, every overlap region matters. A clean alignment in the middle of the insert with no coverage at a junction means the junction could be wrong and you wouldn’t know.
Tip The single most common verification failure is declaring a clone “correct” based on Sanger reads that don’t span the junctions. A whole-plasmid nanopore read from Plasmidsaurus — about $15 per construct and 24-hour turnaround — is now usually cheaper than two Sanger reads and gives complete coverage. For new constructs, default to whole-plasmid sequencing; reserve Sanger for re-verifying a specific region.

Step 3: Read the four common mismatch patterns

When you see a discrepancy between the read and the reference, the visual pattern in the chromatogram tells you what kind of error it is. The four patterns and what each one means:

Pattern A: Clean single peak, wrong base

The trace shows one tall, well-formed peak, but the base it calls doesn’t match the reference. This is a point mutation in the clone — the sequencing is correct, the construct has the mutation. If this is in a coding region, translate the affected codon to determine whether it’s silent, missense, or nonsense. PCR-introduced mutations (polymerase error, ~1 in 106 bp for Q5 or Phusion, more for Taq) cluster in inserts that were amplified before cloning. Synthesis errors (commercial gene synthesis) cluster in synthesized fragments. Either way, this mismatch is real, and the clone is what it is — bank it or discard it based on whether the mutation matters.

Pattern B: Two overlapping peaks at one position

Two roughly equal peaks at the same base call. The base call will often be an IUPAC ambiguity code (R, Y, S, W, K, M, N). This usually means you sequenced a mixed population — two different inserts in the same colony. Common causes: two colonies picked together, contamination in the mini-prep, or in rare cases, a heteroduplex from incomplete repair after transformation.

The diagnostic: where do the double peaks start? If they appear partway through the read and continue to the end (downstream of a small indel), see Pattern C. If they appear throughout the read at multiple positions, you have two different inserts. Streak the colony to single colonies and re-pick.

Pattern C: Clean read, then shifts to overlapping peaks at a specific position and stays mixed

The read is clean for the first N bases and then transitions abruptly to mixed peaks for the remainder. This is the indel signature in a mixed template: an insertion or deletion in one of two co-sequenced populations causes the frame to shift, and from the indel position downstream, every base call is a superposition of the two sequences (which are now offset by the indel size).

If you align both sequences manually, you can usually recover both, but in practice this signature means you have a mixed colony — the same fix as Pattern B (streak to single, re-pick, re-sequence). The position where the mixing starts is usually a known indel-prone site: a repetitive region, a homopolymer run, or a Gibson overlap junction.

Pattern D: Long stretch of N’s or low-quality calls in the middle of an otherwise clean read

The read is clean before and after, but a 20–100 bp window in the middle is ambiguous or unreadable. This is usually a secondary structure stall — the template has a hairpin or strong GC-rich region that Sanger polymerase struggles through, and the trace quality drops locally. The sequence is probably correct; the read is just unreliable in that window. Re-sequence the same region from the opposite direction. If the reverse read covers the same window cleanly, the construct is fine.

Step 4: When to re-sequence vs. when to give up

A clone is bankable when:

  • Every base of the insert is covered by a high-quality read
  • Both vector-insert junctions are confirmed
  • All apparent mismatches are explained — either confirmed real (and tolerable) or shown to be sequencing artifact via a second read

Re-sequence when the trace is messy in a region you care about, when junctions aren’t covered, or when a Pattern D stall obscures a critical sequence. Switching to the reverse primer is the cheapest first move. If the reverse direction also stalls, you’re probably looking at a stable secondary structure — consider whole-plasmid nanopore sequencing, which doesn’t share Sanger’s GC-rich / hairpin failure modes.

Give up on the clone when:

  • Multiple confirmed point mutations cluster in your coding sequence (suggests a polymerase error landed in a critical region — pick a different colony)
  • A frame-shifting indel is confirmed (the protein is wrong and won’t fix itself)
  • Pattern B mixed reads persist after restreaking (sometimes the strain has a real instability problem — switch hosts, e.g., from DH5α to Stbl3 or NEB Stable, especially for repeat-containing constructs)
Common Mistake Banking a clone based on a single forward read because “the insert sequence matched.” A single Sanger read that doesn’t cover the 3′ junction can hide a bad ligation, a partial overlap in Gibson assembly, or a missing residue at the C-terminus of a tagged fusion. The cost of one extra reverse read is trivial compared to discovering a few weeks later that your tag isn’t actually fused to your protein. The same junction-fidelity discipline drives verifying insert orientation after ligation.

What to record before you move on

Once a clone is verified, save:

  • The reference sequence with annotations (vector + insert + junctions)
  • The aligned trace files (.ab1, not just FASTA) — the chromatograms are the verifiable record; FASTA loses the peak-height context
  • The colony number / glycerol stock identifier
  • Which primers were used for sequencing, with their binding positions on the construct

The .ab1 files are the only thing that lets a labmate (or future-you) verify your conclusion later. A construct documented only with a FASTA and a screenshot of an alignment view loses the chromatogram evidence, and any future dispute about whether the construct is “really” correct ends up being adjudicated by re-sequencing. Keep the traces.

If your plasmid editor supports it, save the reference and the read together as a single annotated record — storing the trace alongside the annotated reference means the next person who opens the construct sees the same evidence you used to decide it was correct, and the same junction-bridging logic you used to verify your double-digest cloning strategy still holds up months later.

Try PlasmidStudio

AI-assisted plasmid design with automated validation. Start free — $0 to sign up.

Get started free