Almost every research peptide spec sheet carries a line that looks like gibberish to a newcomer: GEPPPGKPADDAGLV. Or Ac-Nle-cyclo[Asp-His-D-Phe-Arg-Trp-Lys]-NH2. Or AEDG.
That line is the single most information-dense thing on the page. It is the compound's identity claim — the specific molecule the manufacturer says is in the vial, and the thing every analytical test on the COA is ultimately measured against. It also quietly predicts a surprising amount: how the molecule will behave on an HPLC column, whether UV quantitation will work at all, which degradation pathways it is exposed to, and whether it could have been made by a bacterium or had to be built by chemistry.
This is a reading guide. No dosing, no handling — just how to decode the string.
The Alphabet: One Letter, Three Letters, Same Thing
There are twenty amino acids the ribosome routinely uses, and each has a three-letter abbreviation and a one-letter code. Gly/G is glycine, Pro/P is proline, Trp/W is tryptophan, and so on. The two systems are interchangeable; which one gets used is mostly a matter of length and habit.
Short peptides are usually written three-letter, because the individual residues matter and are being discussed. Epithalon is Ala-Glu-Asp-Gly — four residues, four hyphens. Written one-letter it collapses to AEDG, which is how you will most often see it.
Longer sequences go one-letter for readability. BPC-157 as GEPPPGKPADDAGLV is fifteen residues; spelling that out three-letter would be a paragraph.
The hyphens are not optional in three-letter notation. They separate residues. One-letter notation runs the residues together and relies on the reader knowing each character is one amino acid.
Direction: N-Terminus on the Left, Always
A peptide is a chain, and the chain has a direction. By universal convention, sequences are written N-terminus first (left) and C-terminus last (right).
The N-terminus carries a free amino group; the C-terminus carries a carboxyl group. Peptide bonds link the carboxyl of one residue to the amino group of the next, so the chain is directional the way a sentence is — reversing it produces a different molecule, not the same one backwards.
This convention is why position numbering also runs left to right, and why "residue 2" or "position 8" always counts from the N-terminal end.
It matters more than it sounds. The termini are the exposed positions that exopeptidases attack, which is why so many stabilizing modifications on this shelf cluster at the two ends of the chain rather than in the middle — a point covered in detail in the peptidase and clearance explainer.
The Ends Are Part of the Molecule
Here is the most common beginner error: treating the sequence letters as the whole story and ignoring what is written before and after them.
-NH2 at the end means C-terminal amidation. The carboxyl has been converted to an amide. Ipamorelin, oxytocin and SS-31 are all written with a terminal -NH2. The unamidated version — the "free acid" — is a different compound, about 1 Da heavier, and in many neuropeptides the amide is required for receptor activity rather than being a bolt-on stability feature.
Ac- at the front means N-terminal acetylation, adding about 42 Da. Thymosin Alpha-1 and the Ac-LKKTETQ heptapeptide sold as TB-500 both carry it.
pGlu- means pyroglutamate, an N-terminal glutamine that has cyclized onto itself, losing about 17 Da. Gonadorelin — GnRH — begins this way.
How much do the ends matter? Consider Melanotan 2 and PT-141. They share the same acetylated, cyclized core. The published difference between them sits at the C-terminus: the amide versus the free acid. That roughly 1 Da difference is associated with a distinctly different receptor-activity profile in the literature, and PT-141 — under its INN, bremelanotide — went on to become an approved drug while MT-2 did not. A one-dalton terminal difference is not a rounding error.
What the Letters Cannot Show You
The one-letter alphabet was designed for ribosome-built proteins. It systematically fails to represent several things this catalog is full of.
Chirality. Amino acids come in L and D forms — mirror images. The ribosome builds L only, so plain letters imply L by default. A D-residue must be flagged explicitly: D-Ala in Dermorphin at position 2, D-Arg at position 1 of SS-31, and the entire chain of FOXO4-DRI, where "DRI" stands for D-retro-inverso — reversed sequence, all-D residues. If a spec sheet writes such a sequence without the D-prefixes, the sequence line is wrong.
Non-proteinogenic residues. Anything outside the standard twenty gets written three-letter because it has no one-letter code. Aib (α-aminoisobutyric acid) sits at position 8 of semaglutide and at the N-terminus of ipamorelin. Nle is norleucine. Dmt in SS-31 is 2',6'-dimethyltyrosine. 2-Nal is 2-naphthylalanine. Seeing three-letter codes mixed into an otherwise one-letter string is a signal, not a typo.
Rings. cyclo[...] brackets mark a covalent bridge between residue side chains — in Melanotan 2 and PT-141, a lactam bridge between an Asp and a Lys. Disulfides between two cysteines are a separate kind of ring, usually annotated as a bond between positions: oxytocin's Cys1–Cys6 disulfide is what makes it a cyclic nonapeptide rather than a floppy string.
Side-chain attachments. The C18 fatty diacid on semaglutide hangs off a lysine side chain via a spacer. No letter conveys it.
The rule: the letters describe composition and order; everything else is written in the annotations around them. If a vendor's sequence line has no annotations on a compound you know to be modified, the line has been simplified — which is exactly the kind of simplification that makes a mass spectrum "not match."
Numbers, Fragments, and Why Position 1 Isn't Always First
Peptides derived from larger parent proteins usually keep the parent's numbering. That is why you see:
- GLP-1(7-36)amide — residues 7 through 36 of the parent, C-terminally amidated. The mature hormone starts at 7, which is why semaglutide's Aib substitution is called Aib8 even though it is the second residue of the actual molecule.
- hGH 176-191 — the C-terminal region of growth hormone that gives HGH Fragment 176-191 and AOD-9604 their names.
- ACTH(4-7) — the fragment underlying Semax (MEHFPGP), where the last three residues are an added Pro-Gly-Pro tail rather than parent sequence.
- Thymosin β4 residues 17-23 — the actin-binding stretch corresponding to Ac-LKKTETQ.
Two traps here. First, a number in a name may be parent-protein coordinates, catalog index, or residue count — a distinction worked through in the nomenclature primer. Second, a fragment is not its parent. Sharing residues 17-23 does not mean sharing the parent's evidence file, which is the central issue in the BPC-157 vs TB-500 comparison.
Reading Properties Straight Off the Sequence
Once decoded, the sequence predicts practical things.
Aromatics decide whether UV quantitation works. Trp (W) and Tyr (Y) absorb at 280 nm. BPC-157, Epithalon (AEDG), KPV and Selank (TKPRPGP) contain neither — so A280 quantitation is noise for them, and they are detected weakly at the low-UV wavelengths HPLC uses instead. MOTS-c carries both and quantifies fine. This is the mechanism behind the caveats in the net peptide content and chromatogram-reading pieces.
Met (M) and Cys (C) are the oxidation-prone residues. Scan for them and you know the compound's main oxidative liability before reading a single stability study.
Asn (N) and Asp (D) next to Gly (G) are the classic instability motifs — Asn-Gly for deamidation, Asp-Gly for aspartimide formation during synthesis. Epithalon's AEDG contains an Asp-Gly at its C-terminal end, which is precisely the motif chemists watch during assembly.
Arg (R) and Lys (K) density predicts a low net-peptide-content number on an otherwise excellent lot, because basic residues bind more counterion. Composition, not quality.
Length predicts the manufacturing route. Under roughly 50 residues is comfortable solid-phase synthesis; past about 80 the balance shifts to recombinant expression — which is why the 83-residue IGF-1 LR3 cannot be a synthetic peptide and the 191-residue HGH 191AA certainly cannot. And any D-residue forces the chemical route regardless of length, because no ribosome builds D-amino acids. See the synthesis-routes piece for what each route's impurity menu looks like.
Why This Is the Identity Claim
The sequence is what analytical testing is checked against. A mass spectrum confirms that the observed mass matches the theoretical mass of the fully annotated molecule — acetyl included, amide included, disulfide included. If the vendor's theoretical mass was calculated from the bare letters while the actual product is acetylated, the two will disagree by 42 Da and nothing is actually wrong except the paperwork.
And the reverse limitation holds. Intact mass confirms composition, not order. A scrambled sequence, an epimer, or a mis-paired disulfide can all carry the correct mass. Confirming the sequence itself requires MS/MS fragmentation or amino acid analysis, as laid out in HPLC vs mass spec.
FAQ
Does a longer sequence mean a more potent or more advanced compound? No. Length reflects the parent molecule and design approach, nothing else. AEDG is four residues and HGH 191AA is 191; the number describes size, not activity, evidence, or quality.
If two vendors list slightly different sequences for the same product name, which is right? That question cannot be answered from the sequences alone, and it is a real occurrence — TB-500 is the standing example, listed by different suppliers as either a 43-residue full-length protein or a 7-residue fragment with a roughly 4,000 Da mass difference. The sequence line is a claim; a named mass on an identity test is what supports it.
Why do some sequences show three-letter codes for a few residues and one-letter for the rest? Because the three-letter ones have no one-letter code — they are outside the standard twenty. Mixed notation almost always signals a non-proteinogenic residue such as Aib, Nle or a D-form, and those residues usually exist for a specific engineering reason.
Everything above is a reading skill, not a laboratory procedure. Start with the sequence, expand every annotation, and the rest of the library becomes considerably easier to navigate.
This article is educational and for the laboratory research community. Trulogic Labs products are sold for laboratory and research use only and are not for human consumption.