For laboratory and research use only. Not for human or veterinary use. Not a drug, supplement, or medical device.

What Actually Counts as a Peptide? Three Places the Line Gets Drawn

A single research catalog will happily file a four-residue tetrapeptide, a 191-residue folded protein, a dinucleotide, and an indoleamine under the same shelf heading. In casual usage they all become "peptides." That is convenient for navigation and useless for reasoning, because almost every practical question a researcher asks downstream — how was this made, which analytical test means anything, how is it cleared, does receptor vocabulary even apply — is answered by the classification, not by the shelf.

So it is worth asking the apparently trivial question directly: what is a peptide, where does it stop being one, and who decides? The answer is that the line gets drawn in at least three different places, by three different communities, for three different reasons — and none of them is wrong.

The One Thing That Is Not Ambiguous

A peptide is a chain of amino acids joined by amide bonds — the peptide bond, formed between the carboxyl group of one residue and the amino group of the next, releasing water. A chain of n residues contains n − 1 peptide bonds.

That bond is unusually rigid. The carbonyl-to-nitrogen linkage carries partial double-bond character, which locks the six atoms around it into a plane and removes free rotation. Conformational freedom lives in the angles on either side, not in the bond itself. This is why a peptide backbone is not a floppy string — it is a series of planar plates connected by hinges.

Everything else about the definition is convention. The bond is chemistry; the boundary is bookkeeping.

Line One: The Textbook Convention (~50 Residues)

The most commonly taught boundary places peptides at roughly 2 to 50 amino acids, with anything longer called a protein. At an average residue mass near 110 Da, that puts the line somewhere around 5 kDa. You will also see the intermediate terms oligopeptide (a handful of residues) and polypeptide (a long chain, used interchangeably with protein in much of the literature).

The convention is genuinely soft. Different textbooks place it anywhere from about 40 to 50 residues, and some reserve "protein" for chains of 100 or more. Insulin, at 51 residues across two chains, is conventionally called a protein hormone; GLP-1, at around 30, is unambiguously a peptide. Nothing chemically distinguishes residue 50 from residue 51.

What the convention is really gesturing at is not length but capability — which brings us to the second line.

Line Two: The Structural Line (Does It Fold?)

The reason length matters at all is that extra chain buys higher-order structure. A protein has a primary sequence, local secondary structure (helices, sheets), and a tertiary fold — a specific three-dimensional arrangement that is itself the functional unit.

Most short peptides do not have this. In aqueous solution a 7-mer or a 16-mer is largely disordered, sampling many conformations rather than holding one. Structure, when it appears, is often inducedLL-37 adopts its amphipathic helix on contact with an anionic membrane surface rather than carrying it around in dilute buffer. Small cyclic peptides are the interesting exception: oxytocin is only nine residues but a Cys1–Cys6 disulfide constrains it into a defined ring, which is why disulfide integrity is a real analytical question for a molecule that short.

At the other end of the same catalog, IGF-1 LR3 (83 residues, three disulfides) and HGH 191AA (191 residues, a four-helix bundle) are folded proteins in every meaningful sense.

This line has direct analytical consequences. For a short unstructured peptide, sequence plus mass plus purity comes close to describing identity. For a folded protein, it does not — misfolded and disulfide-scrambled species are isomers, identical in composition and mass, invisible to intact-mass spectrometry, and untouched by a purity percentage. That is a different documentation problem, discussed further in reading an HPLC chromatogram and mass spectrum.

Line Three: The Regulatory Line, and It Is a Hard Number

Here the fuzziness stops. Under 21 CFR 600.3, as amended by the FDA's final rule Definition of the Term "Biological Product" (published February 21, 2020; effective March 23, 2020), a protein means any alpha amino acid polymer with a specific defined sequence that is greater than 40 amino acids in size.

Forty. Not "about fifty." The rule implemented changes from the BPCI Act and the Further Consolidated Appropriations Act, 2020, which removed the old statutory parenthetical excluding "any chemically synthesized polypeptide" from the protein category — meaning a molecule's classification now turns on what it is, not on how it was manufactured.

The consequence is structural. Greater than 40 residues → protein → regulated as a biological product under a Biologics License Application. Forty or fewer → regulated as a drug under the FD&C Act, with the New Drug Application and Abbreviated New Drug Application pathways available. There is also a counting rule: where subunits are associated in a manner that occurs in nature, FDA counts the total residues across all subunits — so a heterodimeric glycoprotein like HCG (a 92-residue alpha subunit with a 145-residue beta subunit) is far past the line.

Run the shelf against that number and the sorting is unfamiliar. Semaglutide at 31 residues and tirzepatide at 39 both sit on the drug side. Full-length thymosin beta-4, at 43 residues, sits three residues onto the protein side — while the Ac-LKKTETQ heptapeptide circulated under the same TB-500 name sits nowhere near it, a naming problem covered in the BPC-157 vs TB-500 comparison.

Why That Number Is In the News Right Now

Because semaglutide and tirzepatide fall on the drug side of the 40-residue line, generic versions can in principle be pursued through the ANDA pathway — and the technical standards for doing so are actively being rewritten.

On July 28, 2026, FDA published a batch of 17 revised draft product-specific guidances for peptide products, with the Federal Register availability notice appearing July 29, 2026 and a comment deadline of September 28, 2026. The list spans semaglutide, tirzepatide, liraglutide, teriparatide, pegcetacoplan, vosoritide, dasiglucagon, glucagon and calcitonin salmon. The revisions address five areas: submission of recombinantly, synthetically or semi-synthetically produced peptides as ANDAs, innate immune response testing, impurity thresholds, higher order structure assessment, and biological activity assessment.

At the same time FDA withdrew its May 2021 guidance ANDAs for Certain Highly Purified Synthetic Peptide Drug Products That Refer to Listed Drugs of rDNA Origin — the document that had covered glucagon, liraglutide, nesiritide, teriparatide and teduglutide — stating it no longer reflects the agency's current scientific thinking, with a reissued version planned.

Two things are worth noticing, neither of which is a claim about any research compound. First, an agency drawing a line at a round number for administrative clarity has produced a boundary that determines an entire manufacturing and evidence regime. Second, "higher order structure assessment" appearing on a peptide guidance list is line two and line three colliding: chains short enough to be regulated as drugs are nonetheless expected to be characterized for structure. These are drafts in a comment period, not settled requirements.

The Things On the Shelf That Are Not Peptides At All

Classification is not pedantry when the item in question contains no amide-linked amino acids whatsoever:

  • NAD+ is a dinucleotide — two nucleotides joined through a phosphate bridge. Not a peptide.
  • AICAR is a nucleoside/nucleotide-class molecule, a purine biosynthesis intermediate that acts as a prodrug for an AMP mimetic. Not a peptide.
  • 5-Amino-1MQ is a small-molecule methylquinolinium enzyme inhibitor. Not a peptide.
  • Melatonin is an indoleamine, N-acetyl-5-methoxytryptamine. Not a peptide.
  • GHK-Cu is a tripeptide — but as supplied it is a coordination complex with copper(II), so metal stoichiometry is an analytical axis entirely separate from sequence purity.
  • HMG is not a single molecule at all but a standardized biological fraction, which is why a purity percentage has no clean referent for it.

The practical rule: a test only means something if it can detect a failure mode the molecule is capable of. Amino acid analysis and peptide mapping lines on a small-molecule COA are template, not characterization. This is the same argument developed in peptide synthesis routes and impurity profiles.

Where the Boundary Is Blurred On Purpose

Modern peptide chemistry spends most of its effort making molecules that are built from amino acids but are no longer peptides in the biosynthetic sense — peptidomimetics, designed to keep the binding surface while shedding the liabilities.

Non-proteinogenic residues such as Aib (α-aminoisobutyric acid) have no genetic codon. D-amino acids cannot be installed by a ribosome — whether engineered, as in SS-31 and the all-D retro-inverso FOXO4-DRI, or naturally post-translationally epimerized, as in dermorphin. Side-chain fatty-acid conjugation, lactam bridges, N-terminal acetylation and C-terminal amidation all sit outside the plain twenty-letter alphabet.

The useful reframing: "peptide" increasingly describes a molecule's ancestry rather than its chemistry. A compound can be entirely amino-acid-derived and still be uncopyable by any cell — which, as it happens, is exactly what forces the chemical synthesis route and sets the impurity menu that follows.

The Payoff: What the Classification Predicts

Before reading anything else about an unfamiliar item, place it. Five things follow almost automatically:

  1. Manufacturing route. Under ~50 residues is solid-phase synthesis territory; past ~80 generally means recombinant expression; any D-residue or synthetic side chain forces a chemical step regardless of length.
  2. Which analytics matter. Short and unstructured → mass, sequence, purity, counterion, water. Folded or recombinant → add folding, host cell protein, host DNA, endotoxin. Not a peptide → different panel entirely.
  3. Clearance. Essentially every 1–10 kDa chain is freely filtered by the kidney and subject to peptidase attack; small molecules and nucleosides do not obey those rules. See the peptidase and clearance explainer.
  4. Whether receptor vocabulary applies. Several catalog compounds have no established receptor and act on intracellular partners — affinity, efficacy and selectivity are the wrong words for those.
  5. What kind of file could even exist. Drug pathway, biologic pathway, or no regulatory file at all — which is where most of this shelf sits.

FAQ

Is a longer peptide a more serious or more potent compound? No. Length predicts manufacturing route, structural complexity and regulatory category. It predicts nothing about potency, and nothing about how much evidence exists. A four-residue compound and a 191-residue protein can sit at opposite ends of the evidence hierarchy in either direction.

If FDA says over 40 residues is a protein, is a 43-residue research compound "a biologic"? The definition classifies molecules for regulatory purposes; it does not confer any status on an unapproved research compound. The useful takeaway is comparative: it tells you which regulatory machinery a molecule of that size would encounter, and it explains why manufacturing and characterization expectations differ so sharply across a catalog whose members span 4 to 191 residues.

Why does any of this matter when reading a COA? Because a certificate of analysis is only as meaningful as the match between the tests run and the molecule's actual failure modes. Chirality, folding and metal stoichiometry are each invisible to a purity percentage, and each is relevant to a different subset of the shelf. More on that in /quality/ and across the library.

This article is educational and for the laboratory research community. Trulogic Labs products are sold for laboratory and research use only and are not for human consumption.

Explore the reference library

Every compound we discuss has a full educational monograph — mechanism, research findings, and pharmacokinetics. For laboratory research use only.

Browse the Peptide Library