Research chemicals onlyFor laboratory research use only. Not for human consumption, therapeutic, or diagnostic purposes.
    Optimized Aminos
    Free BAC water $50+ · Free US shipping $100+

    Research Blog

    How to Read a Peptide's Amino Acid Sequence: Single-Letter Codes Explained

    Published

    A plain guide to the single-letter amino acid code system used on peptide COAs and in published literature, including directionality and modification notation.

    For laboratory and research use only. Not for human consumption.

    Every research peptide is, at its core, a defined chain of amino acids — and the shorthand researchers use to write that chain out is the single-letter amino acid code. If you've looked at a Certificate of Analysis, a published paper's methods section, or a peptide's technical sheet and seen a string like "GHK" or "H-Aib-His-D-Phe-Arg-Trp-Gly-NH2," this article explains how to read it: what each letter stands for, how directionality works, and what special notation (lowercase letters, brackets, terminal caps) usually means.

    Key Facts

    • The single-letter amino acid code is an IUPAC-IUBMB standard assigning one letter to each of the 20 standard amino acids.
    • Sequences are conventionally written N-terminus to C-terminus, left to right.
    • Some letter assignments (like N for asparagine, Q for glutamine) don't match the first letter of the name, because that letter was already taken by a different amino acid.
    • Lowercase letters in a sequence typically denote D-amino acids rather than the standard L-amino acid form.
    • Terminal notations like "-NH2" (amidation) or "Ac-" (acetylation) describe end-group modifications that fall outside the core 20-letter alphabet.

    The 20-Letter Alphabet

    The 20 standard amino acids that make up naturally occurring peptides and proteins each have a three-letter abbreviation and a single-letter code, standardized by the International Union of Pure and Applied Chemistry and the International Union of Biochemistry and Molecular Biology (IUPAC-IUBMB). The single-letter system exists purely for compactness: writing out "glycine-histidine-lysine" is far more cumbersome in a table or a sequence-alignment tool than writing "GHK."

    • G Glycine, A Alanine, V Valine, L Leucine
    • I Isoleucine, P Proline, F Phenylalanine, W Tryptophan
    • M Methionine, S Serine, T Threonine, C Cysteine
    • Y Tyrosine, N Asparagine, Q Glutamine, D Aspartate
    • E Glutamate, K Lysine, R Arginine, H Histidine

    Why Some Letters Look Mismatched

    Most codes are intuitive — G for glycine, A for alanine, L for leucine. But a handful don't match the first letter of the name, and there's a specific reason: with 20 amino acids competing for 26 letters, several names share a first letter. Aspartate and asparagine both start with "A," so asparagine was assigned "N" (from a later letter in "asparagiNe"). Glutamate and glutamine share "G" territory with glycine, so glutamine became "Q." Isoleucine and leucine both want "L," so isoleucine took "I." These reassignments are memorization points, not logic puzzles — the codes are fixed by the standard, not derivable from first principles.

    Reading Directionality: N-Terminus to C-Terminus

    A peptide chain has a direction, because it's built from amino acids linked through peptide bonds between a carboxyl group and an amino group. One end of the chain retains a free amino group — the N-terminus — and the other end retains a free carboxyl group — the C-terminus. By universal convention, sequences are written and read left to right, N-terminus first. So in a tripeptide written as "GHK," glycine is the N-terminal residue and lysine is the C-terminal residue. This convention matters when comparing a written sequence to a structural diagram or a mass-spectrometry fragmentation pattern, since fragment ions are also typically labeled relative to N- and C-terminal position.

    Beyond the Basic Alphabet: Modifications and Special Notation

    Many research peptides aren't simple chains of the 20 standard L-amino acids — they include structural modifications that require notation beyond the basic single-letter code. Recognizing these conventions is part of reading a sequence correctly:

    Lowercase Letters: D-Amino Acids

    Standard amino acids in biological peptides are almost universally the L-stereoisomer (a specific three-dimensional configuration). Some engineered research peptides substitute a D-amino acid (the mirror-image stereoisomer) at a specific position, often to increase resistance to enzymatic degradation in in vitro assays. When this is denoted in running text, a lowercase letter (or a "D-" prefix on the three-letter code, e.g., "D-Phe") flags the substitution — a capital letter denotes the standard L-form.

    Terminal Caps: Amidation and Acetylation

    A sequence ending in "-NH2" indicates C-terminal amidation — the free carboxyl group is capped with an amide group instead, a modification reported in stability literature to affect resistance to carboxypeptidase activity in certain in vitro degradation assays. A sequence beginning with "Ac-" indicates N-terminal acetylation, capping the free amino group. Both are common, checkable modifications listed explicitly on a technical sheet or COA.

    Non-Standard Residues and Brackets

    Some sequences include residues outside the standard 20, such as Aib (2-aminoisobutyric acid) or side-chain attachments like a fatty-acyl chain used in some long-acting peptide designs. These are typically written out with their abbreviation directly in the sequence (not forced into a single letter) or flagged with brackets/parentheses describing the attachment point and chemical group.

    Cross-Checking a Sequence Against a COA

    Reading sequence notation isn't just an academic exercise — it's a practical verification tool. A Certificate of Analysis for a research peptide should state the sequence (or reference the compound's defined structure) alongside the molecular weight confirmed by mass spectrometry. A researcher who can read single-letter notation can cross-check that the stated sequence is internally consistent with the reported molecular weight, an extra layer of identity verification beyond trusting the label alone. Our testing and COA documentation page shows how this identity information is presented alongside purity data for each production lot.

    For a broader vocabulary of terms that come up alongside sequence notation — residue, terminus, isoform, and more — see our 150-term peptide chemistry glossary. For how these sequences are physically assembled in a lab, see how research peptides are made: solid-phase synthesis explained. And for a worked example of how a single side-chain modification changes a peptide's chemistry entirely, see GHK-Cu: the chemistry of a copper-binding peptide, which builds directly on the GHK tripeptide sequence introduced above.

    Frequently Asked Questions

    What is the single-letter amino acid code system?

    The single-letter amino acid code is a standardized notation, adopted by IUPAC-IUBMB, that assigns one letter of the alphabet to each of the 20 standard amino acids so that a peptide or protein sequence can be written compactly as a single string of letters, such as GHK for glycine-histidine-lysine.

    Why do some amino acids share similar-looking letter codes?

    Because 20 amino acids need distinct letters and several share a first letter (like aspartate and asparagine, both starting with "A"), the system assigns some amino acids a letter from elsewhere in their name (asparagine is N, glutamine is Q) to avoid duplicate codes.

    How do you read directionality in a written peptide sequence?

    By convention, a peptide sequence is written from the N-terminus (the end with a free amino group) to the C-terminus (the end with a free carboxyl group), so the leftmost letter in a sequence like "H-Gly-His-Lys-OH" or "GHK" represents the N-terminal residue.

    What do lowercase letters or brackets in a sequence usually indicate?

    Lowercase letters in a written sequence commonly denote D-amino acids (the mirror-image stereoisomer of the standard L-amino acid), while brackets or parentheses typically flag a non-standard modification, such as an acetyl group, an amide cap, or a side-chain attachment like a fatty-acid acylation.

    For laboratory and research use only. Not for human consumption.

    Related research compounds

    Compounds referenced in this article, available as research-grade lyophilized peptides with third-party tested COA.

    Continue reading

    Related references