A peptide sequence is the whole identity of the molecule. Two peptides with the same formula and the same mass are different compounds if the residues are in a different order, which is why a certificate confirming "identity by LC-MS" is confirming the mass is consistent with a sequence, not that the sequence is correct in isolation. This post covers the two notations you will meet, what the termini and the modification prefixes mean, and how to check a stated molecular weight with nothing but the table below and a calculator.
Why are there two notations?
The three-letter codes came first and read naturally: Gly-Pro-Glu is obviously glycine, proline, glutamate. They are unambiguous and still standard for short peptides and for anything being discussed in prose.
The one-letter set was standardised by the IUPAC-IUB Commission on Biochemical Nomenclature in 1968, published simultaneously in Biochemistry and the Biochemical Journal, for a practical reason: protein sequences had become long enough that three-letter codes were unreadable and unprintable at scale. A 191-residue hormone is 764 characters in three-letter notation and 191 in one-letter.
Both are current. Short research peptides are usually written three-letter, longer ones one-letter, and a certificate may use either.
The full amino acid table
| Amino acid | 3-letter | 1-letter | Residue mass (Da) | Side chain |
|---|---|---|---|---|
| Alanine | Ala | A | 71.08 | Nonpolar |
| Arginine | Arg | R | 156.19 | Basic |
| Asparagine | Asn | N | 114.10 | Polar, deamidation-prone |
| Aspartate | Asp | D | 115.09 | Acidic |
| Cysteine | Cys | C | 103.14 | Polar, oxidation-prone |
| Glutamate | Glu | E | 129.12 | Acidic |
| Glutamine | Gln | Q | 128.13 | Polar, deamidation-prone |
| Glycine | Gly | G | 57.05 | Nonpolar |
| Histidine | His | H | 137.14 | Basic |
| Isoleucine | Ile | I | 113.16 | Nonpolar |
| Leucine | Leu | L | 113.16 | Nonpolar |
| Lysine | Lys | K | 128.17 | Basic |
| Methionine | Met | M | 131.19 | Nonpolar, oxidation-prone |
| Phenylalanine | Phe | F | 147.18 | Aromatic |
| Proline | Pro | P | 97.12 | Nonpolar, rigid |
| Serine | Ser | S | 87.08 | Polar |
| Threonine | Thr | T | 101.10 | Polar |
| Tryptophan | Trp | W | 186.21 | Aromatic |
| Tyrosine | Tyr | Y | 163.18 | Aromatic |
| Valine | Val | V | 99.13 | Nonpolar |
The one-letter codes are not all first initials, because several amino acids share one. The 1968 rules assigned the initial to the most frequent residue and made the rest phonetic or arbitrary: R for a*r*ginine, K for lysine (the letter next to L), W for tryptophan (the double ring), Y for t*y*rosine, Q for glutamine, N for asparagine.
Which end is which?
Sequences are written N-terminus first, left to right. The N-terminus carries a free amino group, the C-terminus a free carboxyl group. This is a convention agreed for consistency, and it matters: Gly-Pro-Glu and Glu-Pro-Gly are different compounds with identical formulas.
This is why KPV — lysine-proline-valine — is specified in that order. Written backwards it would describe a molecule that is not the C-terminal fragment of alpha-MSH and has no relationship to the literature on it.
What do the prefixes and suffixes mean?
The 1972 IUPAC-IUB recommendations cover symbols for derivatives. Three appear constantly in research peptides:
- Ac- at the front: the N-terminus is acetylated. Adds 42.01 Da and removes the positive charge that a free amino group carries.
- -NH2 at the end: the C-terminus is amidated rather than a free acid. Subtracts 0.98 Da from the free-acid mass and removes the negative charge. Many natural signalling peptides are amidated, and the amidated form is often far more active.
- D- before a residue: the D-stereoisomer rather than the natural L form. Same mass, different molecule, usually chosen to resist enzymatic cleavage.
Modified GRF (1-29) — the CJC-1295 no-DAC backbone — is a good example of why this notation matters. It is GHRH (1-29) with four substitutions, and writing it without those substitutions describes plain sermorelin.
How do you check a molecular weight yourself?
Add the residue masses from the table, then add 18.02 for the water molecule released when the chain was formed. That is the free-acid, unmodified mass.
Worked example, Epithalon, sequence Ala-Glu-Asp-Gly:
- Ala 71.079 + Glu 129.116 + Asp 115.089 + Gly 57.052 = 372.335
- Plus water: 372.335 + 18.015 = 390.35 Da
PubChem lists epithalon at 390.35. The arithmetic checks to the second decimal. It checks for the rest of the short bioregulators too: KPV calculates to 342.44 against PubChem's 342.43, Vilon to 275.30, Pinealon to 418.41 against 418.40, Vesugen to 390.39.
If your calculation misses a stated mass by a recognisable amount, the difference usually names the modification:
| Discrepancy | Likely explanation |
|---|---|
| −0.98 Da | C-terminal amide rather than free acid |
| +42.01 Da | N-terminal acetylation |
| +16.00 Da | An oxidised methionine — a degradation product, not a design feature |
| +1.00 Da | A deamidated asparagine or glutamine |
| Large, variable | Counter-ion, usually TFA or acetate, included in the quoted mass |
That last row catches people out. A quoted "molecular weight" that is much heavier than the arithmetic may be quoting the salt rather than the free peptide. Our post on TFA and acetate salts covers why that matters for anything mass-based.
What can a sequence not tell you?
It does not tell you the three-dimensional structure, and it does not tell you purity. A certificate stating a sequence and a matching observed mass has confirmed the molecule is consistent with that sequence; it has not proven the order of residues, which requires sequencing rather than mass measurement. For practical research use the mass match plus a single clean HPLC peak is the standard evidence, and our post on reading purity figures explains what that number covers.
Frequently asked questions
Why do some sequences use B, Z or X?
They are ambiguity codes from the same IUPAC-IUB rules. B means aspartate or asparagine, Z means glutamate or glutamine, and X means any or unknown. They appear in sequences determined by methods that cannot distinguish the amide from the acid, and should not appear on a synthetic peptide certificate.
Does the one-letter code include selenocysteine?
Yes, U, added long after the original rules. It is vanishingly rare in synthetic research peptides.
Why is proline listed as rigid?
Its side chain loops back and bonds to the backbone nitrogen, which locks the local geometry. This is why proline-rich tails such as the Pro-Gly-Pro extension on Semax and Selank resist peptidase cleavage — the enzyme cannot fit the conformation it needs.
Where can I check a sequence and mass independently?
PubChem, which publishes formula, molecular weight and structure for most catalogued peptides. Every product page on this site links its PubChem CID directly so you can compare our stated figures against the public record.
References
- IUPAC-IUB Commission on Biochemical Nomenclature. A one-letter notation for amino acid sequences: tentative rules. Biochemistry 1968;7(8):2703-2705. doi.org/10.1021/bi00848a001
- IUPAC-IUB Commission on Biochemical Nomenclature. A one-letter notation for amino acid sequences. Biochemical Journal 1969;113(1):1-4. doi.org/10.1042/bj1130001
- IUPAC-IUB Commission on Biochemical Nomenclature. Symbols for amino-acid derivatives and peptides: recommendations. Biochemical Journal 1972;126(4):773-780. doi.org/10.1042/bj1260773
Every product mentioned is sold for laboratory research use only and is not for human or animal use. Nothing on this page describes or recommends use of the material sold here in humans or animals.




