Skip to main content
Science

What Is Synthetic Biology? Designing Cells Instead of Discovering Them

Synthetic biology treats DNA as something you write, not just something you edit. How the design–build–test cycle works, the cell built with only 473 genes, what the field has actually delivered — and the 149 genes nobody can explain.

13 min read
Share:
A macro photograph of a petri dish showing rounded bacterial colonies in amber, cream and green growing across a dark green agar surface
Guy Grandjean on Unsplash

Most of biology's history has been observational. You find an organism, work out what it does, and describe it. Synthetic biology inverts that. You decide what you want a cell to do, write the DNA sequence that should produce it, build that DNA chemically, and put it into a cell to find out whether you were right.

That inversion — from discovering biology to specifying it — is the whole field in one sentence. Everything else is consequence.

Editor's note: This is an educational explainer, not medical advice. Several applications described here are approved medicines and several are active research; decisions about any treatment belong with a qualified clinician. The methods below are well established, and the frontier sections are sourced at the end.

The Definition That Actually Distinguishes It

A workable definition: synthetic biology aims to produce cells with new or improved biological functions that do not already exist in nature, using the methods of engineering — specification, standardised parts, modular assembly and iterative testing.

The operative word is engineering. A civil engineer does not discover a bridge. They specify a load, select components with known properties, assemble them to a design, and test whether the result holds. Synthetic biology attempts the same discipline with genes, promoters, and metabolic pathways as the components.

That ambition is why the field talks about parts, circuits and chassis — vocabulary borrowed wholesale from electronics. A "part" is a characterised piece of DNA with predictable behaviour. A "chassis" is the host cell you install it in, typically E. coli or brewer's yeast.

Synthetic Biology Is Not Gene Editing

This is the most common confusion, and it is worth settling early, because the two do different things.

Two-column comparison diagram. The left column, in amber, describes gene editing with CRISPR as changing what is already there: it starts from a living organism's genome, cuts, disables or swaps specific letters, makes edits that are usually small and targeted, and aims to fix or alter an existing trait — the verdict being revision of an existing text. The right column, in blue, describes synthetic biology as specifying then building: it starts from a design written as sequence, DNA is chemically synthesised rather than cut, whole pathways or genomes are assembled, and the goal is a function that did not exist before — the verdict being composition of a new text. A band notes the tools overlap but the intent does not
Editing revises a text that already exists. Synthetic biology composes a new one — then uses editing tools freely while doing it.

CRISPR gene editing is a precision instrument for altering DNA that is already in a cell: cut here, disable that, swap this letter. It is revision.

Synthetic biology starts from a sequence written on a screen. The DNA is chemically synthesised rather than cut, and the unit of work is often an entire metabolic pathway or, at the limit, an entire genome. It is composition.

In practice the two are inseparable — synthetic biology uses editing tools constantly. What distinguishes it is the direction of work: design first, then build to specification.

The Method: Design, Build, Test

The field's defining process is a loop, and it is the reason synthetic biology produces results in domains where rational design in biology has historically failed.

Pipeline diagram of the design-build-test cycle in four numbered stages. One, design: specify the DNA sequence for the function you want, often assembled from characterised parts. Two, build: chemically synthesise the DNA and assemble it into a host cell such as yeast or E. coli. Three, test: measure what the cell actually does, where most designs fail and the failure is the data. Four, learn: feed the result back into the next design, increasingly the step where AI models sit. A band notes this is why the field is called engineering rather than discovery
Most designs fail at the test step. The method's strength is that failure is cheap and informative, so the loop can run many times.

The honest part of this diagram is stage three. Most designs do not work. Biology is not modular in the way electronics are: a promoter that behaves predictably in one context misbehaves in another, and cells respond to engineering by mutating, silencing inserted genes, or simply dying.

What makes the field viable is that the loop got cheap. DNA synthesis costs collapsed, automation made building and screening thousands of variants routine, and so a method that depends on many failed attempts became affordable.

The Landmark: A Cell With Only 473 Genes

If you want one experiment that captures both the field's power and its limits, it is the minimal cell.

In 2010, researchers at the J. Craig Venter Institute built JCVI-syn1.0, the first cell running on a chemically synthesised genome. Then they asked a harder question: what is the smallest genome that can still support a living, self-replicating cell?

The answer, published in 2016, was JCVI-syn3.0 — a working bacterium with 531,560 base pairs and just 473 genes, the smallest genome of any organism that can be grown in laboratory media. Everything non-essential had been stripped away.

And here is the detail that should be quoted far more often than it is: 149 of those 473 essential genes have no known biological function.

The cell will not live without them. Remove any one and it fails. Nobody can say what they do.

That is a remarkable and clarifying result. We can now build a genome from chemicals and boot a living cell with it, while remaining unable to explain roughly a third of the minimum instruction set. The engineering has outrun the understanding.

Writing a Whole Eukaryotic Genome

Bacteria are the easy case. The harder target is a eukaryote — a cell with a nucleus, like ours.

That is the goal of Sc2.0, an international consortium that has spent well over a decade rewriting the genome of brewer's yeast, Saccharomyces cerevisiae, chromosome by chromosome. The consortium has now produced functional synthetic versions of all 16 native yeast chromosomes, plus a 17th "neochromosome" built from scratch to hold all the cell's nuclear tRNA genes — a structure no organism has ever had.

Crucially, this is not transcription of the original. The synthetic genome is a redesign: stop codons recoded, non-functional sequence removed, and genetic watermarks inserted so synthetic chromosomes can be identified. Yeast has also been given a deliberate mechanism for large-scale genome rearrangement, letting researchers scramble the design and observe what survives.

A fully synthetic yeast cell running the complete set at once has not yet been assembled, but with every chromosome built and validated it is now a consolidation problem rather than an open scientific question.

What Synthetic Biology Has Actually Produced

The field is often discussed in the future tense. Its record is longer than that suggests.

Table diagram listing five results of synthetic biology with status: recombinant insulin, a human gene expressed in a microbe, approved 1982; artemisinic acid, an antimalarial precursor made in yeast, up to 25 grams per litre; JCVI-syn3.0, a cell with only 473 working genes, built 2016; Sc2.0 yeast, all 16 chromosomes synthesised, chromosomes done; and designed proteins, sequences no organism has ever used, AI-driven and active. A band notes that of JCVI-syn3.0's 473 essential genes, 149 have no known biological function
A long record of delivered results, and one unresolved puzzle sitting inside the most celebrated of them.

Recombinant insulin is the field's ancestor and its most consequential product. Approved in 1982, it was made by expressing the human insulin gene in a microorganism — replacing insulin extracted from animal pancreases and transforming diabetes care worldwide.

Artemisinin is the clearest demonstration of engineered metabolism. The frontline antimalarial is derived from sweet wormwood, a plant crop with volatile supply and price. Researchers engineered brewer's yeast with genes drawn from several organisms to produce artemisinic acid, the chemical precursor, reaching demonstrated titres of up to 25 grams per litre — an industrial-scale fermentation route to a plant compound.

Designed proteins are the current frontier. Rather than screening natural proteins for useful behaviour, researchers now design sequences that no organism has ever used, aiming at a specified structure and function.

The same platform logic underpins mRNA vaccines: once the sequence is the product, manufacturing becomes a fixed process and the design becomes software.

Where AI Fits

The design step was always the field's weakest link. Predicting what a novel sequence will do inside a living cell is extraordinarily hard, which is why so much of synthetic biology has proceeded by building many variants and screening them.

Machine-learning models have changed the balance. Reviews of the field describe a transition from experience-based to data-driven design, in which the next generation of proteins is engineered to specification rather than discovered by screening — models proposing sequences predicted to fold into a target structure and perform a target function.

This is the same underlying shift described in our guide to machine learning: a domain where the search space is far too large to explore exhaustively becomes tractable when a model can rank candidates before anyone builds them. The loop still runs. AI makes stage one much less blind.

The Biosecurity Problem Is Real, and Already Regulated

A technology that lets you order synthesised DNA to specification has an obvious dual-use problem, and the field has not been naive about it.

Providers of synthetic double-stranded DNA operate under screening frameworks — in the United States, the Department of Health and Human Services' Screening Framework Guidance for Providers of Synthetic Double-Stranded DNA — which asks suppliers to screen both sequence orders and the customers placing them against sequences of concern. Major projects layer their own requirements on top: everyone working on Sc2.0, for instance, is trained in biosafety and dual-use issues.

The honest assessment is that this is a meaningful control rather than a complete one. Screening depends on recognising sequences of concern, on suppliers participating, and on the synthesis capability remaining concentrated among providers who screen. Each of those assumptions weakens as synthesis gets cheaper and more distributed, and AI-assisted design tools raise the additional question of what a model should refuse to propose.

What Is Genuinely Hard

Three obstacles have proven durable, and any realistic view of the field has to sit with them.

Biology resists modularity. Parts do not compose reliably. A genetic circuit that works in one strain frequently fails in another, because the host's own metabolism interacts with everything installed in it.

Living systems evolve away from your design. An engineered cell under selection pressure will often mutate to disable the expensive pathway you added, because not making your product is cheaper for the cell.

Understanding lags construction. The 149 unexplained genes in the minimal cell are the clean example. We can build genomes we cannot fully read.

The Bottom Line

Synthetic biology is the application of engineering method to living systems: specify a function, write the DNA, build it, test it, and iterate. It differs from gene editing in direction rather than tooling — composition instead of revision.

Its record is real and mostly unglamorous: a medicine in every pharmacy since 1982, an industrial fermentation route to an antimalarial, a bacterium running on a synthesised genome, and a yeast whose every chromosome has been rewritten. Its limits are equally real, and the most instructive of them is that the minimal cell contains 149 essential genes nobody can explain.

That gap between what can be built and what is understood is the most interesting thing about the field. For more, follow our biology hub and medical research coverage.

Frequently Asked Questions

What is synthetic biology in simple terms?

It is engineering applied to living systems. Instead of studying what organisms already do, researchers decide what they want a cell to do, write the DNA sequence intended to produce it, synthesise that DNA chemically, and install it in a host cell to test the result. The aim is biological functions that do not exist in nature, built from characterised parts to a specification.

How is synthetic biology different from CRISPR gene editing?

CRISPR edits DNA that is already present in a cell — cutting, disabling or swapping specific sequences. Synthetic biology starts from a design written as sequence and builds the DNA chemically, often assembling entire pathways or genomes. Editing is revision of an existing text; synthetic biology is composition of a new one. In practice, synthetic biology uses editing tools routinely.

What is the minimal cell, and why does it matter?

JCVI-syn3.0, built in 2016 at the J. Craig Venter Institute, is a working bacterium with 531,560 base pairs and only 473 genes — the smallest genome of any organism that can be grown in laboratory media. It matters twice over: it proves a living cell can run on a designed, chemically synthesised genome, and it revealed that 149 of those essential genes have no known biological function.

Has anyone built a synthetic version of a complex organism's genome?

Not a complete cell yet, but the groundwork is done. The international Sc2.0 consortium has produced functional synthetic versions of all 16 chromosomes of brewer's yeast, plus a 17th artificial neochromosome carrying all the nuclear tRNA genes. The synthetic genome is a redesign rather than a copy, with recoded stop codons and non-functional sequence removed. Assembling the full set into one cell remains in progress.

What products has synthetic biology actually delivered?

Several, over four decades. Recombinant human insulin, approved in 1982, was produced by expressing a human gene in a microorganism. Engineered brewer's yeast produces artemisinic acid, the precursor to the antimalarial artemisinin, at demonstrated titres up to 25 grams per litre. More recently, AI-assisted protein design has begun producing sequences no organism has ever used.

Is synthetic biology dangerous?

It carries genuine dual-use risk, which is why DNA synthesis is screened. Providers in the United States work under the HHS Screening Framework Guidance for Providers of Synthetic Double-Stranded DNA, checking both orders and customers against sequences of concern, and major projects add their own biosafety training. That control is meaningful but incomplete: it depends on suppliers participating and on synthesis capability staying concentrated, and both assumptions weaken as the technology cheapens.

Why do so many synthetic biology designs fail?

Because biology is not modular in the way engineering assumes. A genetic part that behaves predictably in one strain often misbehaves in another, since the host cell's own metabolism interacts with everything added to it. Engineered cells also evolve away from the design, frequently mutating to disable an added pathway because not producing the product is metabolically cheaper.

Sources

BiologyMedicine Research#synthetic biology#biology#genetics#biotechnology#genomics
Share:
A glass vial of Moderna COVID-19 vaccine standing on a clinical worktop beside a stack of alcohol prep pads, its label showing the multiple-dose formulation

ScienceGuide

How mRNA Vaccines Actually Work

An mRNA vaccine doesn't deliver a drug — it delivers instructions your own cells read once and throw away. Here's what really happens, and why the same trick now targets cancer.

Aug 22, 202621 min
Artist's concept of the three planets orbiting the pulsar PSR B1257+12, the first planets ever confirmed outside our solar system, with the pulsar's twisted magnetic fields glowing blue and an aurora lighting the night side of the nearest world

ScienceGuide

How We Find Exoplanets — and Why We've Barely Seen Any

We have confirmed more than 6,000 planets around other stars. We have directly photographed about 98 of them. Here is how the other 98% were found, and why that method leaves our map of the galaxy badly distorted.

Aug 15, 202615 min