In one year, the quality manager at a cold-smoked salmon plant in Chile has recovered Listeria monocytogenes six times: three times from the same floor drain, once from a slicer frame, once in the raw fish intake area and once in finished product. Is this one resident strain hiding in a harbourage site, or separate introductions arriving with raw fish? Serotyping cannot tell. Whole genome sequencing can, and public-health laboratories now use it to link patients, foods and factories. This guide answers the practical whole genome sequencing food safety questions: how genomes are compared, what a match means and where the method stops.
In short
- Whole genome sequencing (WGS) reads almost the complete DNA of a bacterial isolate, so isolates from patients, foods and factories can be compared down to single-base differences.
- Genomes are compared by SNP analysis, which counts single-base differences, or by cgMLST, which compares versions of a fixed set of core genes.
- WGS has replaced PFGE as the main subtyping method in PulseNet and many other surveillance networks, revealing smaller and more dispersed outbreaks.
- A close genomic match shows that isolates are closely related; proving that a food caused illness also needs epidemiological and traceback evidence.
- In factories, WGS separates persistent strains from repeated introductions and helps select and confirm strains for challenge tests and validation studies.
What is whole genome sequencing?
Whole genome sequencing is the laboratory process of determining nearly the complete DNA sequence of an organism, in food safety usually a bacterial isolate, which is a pure culture of a single strain grown from a sample. A Listeria monocytogenes genome is about 3 million base pairs long, and WGS reads it many times over to build an accurate sequence. The workflow has five stages:
- Isolate. Culture the organism from the food, swab or clinical sample and confirm the species.
- Prepare DNA. Extract and purify the DNA and convert it into a sequencing library.
- Sequence. Short-read instruments give highly accurate reads of a few hundred bases; long-read instruments read many thousands of bases and resolve repeats and plasmids.
- Assemble and check. Reads are assembled or mapped to a reference, and coverage, contamination and species identity are checked.
- Analyse. Software assigns a subtype, predicts serotype, screens for virulence and antimicrobial resistance genes and compares the genome with others in a database.
How did WGS replace PFGE in outbreak surveillance?
WGS replaced pulsed-field gel electrophoresis (PFGE) because it can separate unrelated isolates that PFGE could not tell apart. PFGE cuts the chromosome with a restriction enzyme and compares the pattern of large DNA fragments on a gel; common patterns were shared by many unrelated isolates, so apparent clusters were often coincidence.
PulseNet International, a network of national and regional laboratory networks in more than 80 countries, built global foodborne disease surveillance on PFGE and has been moving its members to WGS. PulseNet in the United States made WGS its primary subtyping method in 2019. Regulators also sequence food and environmental isolates; the US Food and Drug Administration’s GenomeTrakr network shares such genomes publicly, where they can be compared with clinical isolates.
| Feature | PFGE | WGS |
|---|---|---|
| What is compared | Banding pattern of large DNA fragments | Nearly the complete DNA sequence |
| Resolution | Low to moderate; unrelated isolates often share a pattern | Down to single-base differences |
| Extra information | None | Serotype prediction, virulence and resistance genes, plasmids |
| Sharing between laboratories | Gel images need strict standardisation | Digital sequence data with standard schemes such as cgMLST |
| Outbreaks detected | Mainly larger, concentrated clusters | Also small clusters spread over months or years |
The 2017 to 2018 listeriosis outbreak in South Africa, which WHO described as the largest recorded, showed what this enables. WHO reported 1,060 laboratory-confirmed cases and 216 deaths among cases with a known outcome. Sequencing showed that most clinical isolates belonged to one sequence type, ST6, and linked them to ready-to-eat processed meat from a single production facility.
How are genomes compared: SNP analysis or cgMLST?
Two approaches dominate. SNP analysis counts single-nucleotide polymorphisms, positions where one DNA base differs between genomes, after aligning reads to a closely related reference genome. Core genome multilocus sequence typing (cgMLST) compares a fixed set of genes shared by almost all strains of a species and counts how many carry a different allele, meaning a different version of that gene’s sequence.
Classical multilocus sequence typing (MLST) uses just seven housekeeping genes to assign a sequence type (ST), and related STs form clonal complexes. It describes lineages well but is too coarse for outbreak work.
| Feature | SNP analysis | cgMLST |
|---|---|---|
| Unit of difference | Single base | Whole gene (allele) |
| Resolution | Highest | Slightly lower: several SNPs within one gene count as one allele difference |
| Standardisation | Depends on reference genome and pipeline | Shared allele names, portable between laboratories |
| Typical use | Fine detail within a cluster | Surveillance databases and international comparison |
For L. monocytogenes, a widely used cgMLST scheme from the Institut Pasteur covers 1,748 core genes and groups isolates no more than 7 allele differences apart into one cluster, originally called a cgMLST type (CT) and now a genetic cluster in its LIN code nomenclature. That grouping is a naming convention, not proof of a common source. No universal threshold defines “the same strain”; interpretation depends on the organism’s diversity, the time between isolates and the pipeline. The Advanced Food Science course works through cluster interpretation of this kind.
What does a genomic match mean in an outbreak investigation?
A close genomic match means the isolates share a recent common ancestor; it does not, on its own, prove that a particular food caused particular illnesses. Investigators combine genomic, epidemiological and traceback evidence before naming a source.
WGS has changed which outbreaks are found. Clusters of two or three cases spread over months or years, invisible to PFGE, are now detected, and food or environmental isolates in public databases can be matched to past clinical cases. It also supports source attribution, the estimation of how much human illness comes from each food or animal reservoir. Traps remain: highly clonal lineages can give near-identical isolates without a recent shared source, two factories buying from one supplier can both carry the matching strain, and a strain persisting in one plant for years slowly diversifies.
For a food business the implication is direct. An environmental isolate sequenced today may be matched to illnesses that happened years ago, so persistent Listeria in a ready-to-eat plant is a public-health and legal risk as well as a hygiene problem.
How does WGS help investigate Listeria persistence?
WGS separates a persistent strain, the same strain recovered repeatedly from a plant over months or years, from repeated introductions of different strains. That distinction decides the corrective action. A harbourage site is a protected niche, such as a hollow roller, cracked seal or drain, where bacteria survive cleaning; persistence points to one, while repeated introductions point to raw materials, people or traffic.
| Isolate | Month | Sampling site | cgMLST result (illustrative data) |
|---|---|---|---|
| 1 | 1 | Floor drain, slicing room | Cluster A (reference isolate) |
| 2 | 3 | Same floor drain | Cluster A, 1 allele from isolate 1 |
| 3 | 5 | Slicer frame | Cluster A, 2 alleles from isolate 1 |
| 4 | 7 | Raw fish intake area | Different lineage, hundreds of alleles away |
| 5 | 9 | Same floor drain | Cluster A, 3 alleles from isolate 1 |
| 6 | 12 | Finished product | Cluster A, 2 alleles from isolate 1 |
Worked example
Isolates 1, 2, 3, 5 and 6 are within 3 alleles of isolate 1 across 11 months: one strain is resident in the plant and has reached product. Isolate 4 is unrelated, a separate introduction with raw fish.
The investigation should focus on the drain and everything that links it to the slicer: foot traffic, trolley wheels, hoses and splashing during cleaning. Corrective action means opening up the drain and nearby equipment, repairing or replacing what cannot be cleaned, then verifying with repeated negative results and sequencing any further isolates to confirm the strain has gone.
WGS can also reveal genes associated with tolerance to quaternary ammonium sanitisers, such as bcrABC and qacH. These raise tolerance to low concentrations, so check dosing, contact time and access to the harbourage before changing chemistry.
Can WGS support process validation?
Yes, indirectly. WGS does not measure heat or pressure resistance, but it improves the strains and evidence that validation studies depend on.
- Strain selection. Challenge tests and inactivation studies use cocktails of several strains. Genomes confirm that the strains are genuinely different, cover relevant lineages and can include the plant’s own persistent strain.
- Identity checks. Sequencing confirms that culture collection strains are what their labels say and that survivors recovered at the end of a study are the inoculated strains, not contaminants.
- Surrogates. A surrogate is a harmless organism with process resistance equal to or greater than the target pathogen, used where the pathogen cannot enter a factory, such as Enterococcus faecium NRRL B-2354 for Salmonella in low-moisture foods. Genome data confirm identity and screen for virulence or transferable resistance genes.
- Unusual resistance. Some genetic elements, such as the locus of heat resistance found in some Escherichia coli strains, are linked to markedly higher heat tolerance. Finding one is a reason to test the strain, not a substitute for testing.
- Cultures. Starter cultures and probiotics can be screened for acquired resistance and toxin genes; EFSA asks for WGS data when it assesses microorganisms intentionally used in the food chain.
The working rule is that genotype suggests and phenotype decides. D values, z values and growth limits must still be measured, as taught in Food Science for Industry Professionals.
What are the limits of whole genome sequencing food safety evidence?
WGS evidence is only as good as the isolate, the database and the interpretation behind it. Five limits matter most:
- No isolate, no genome. Many clinical laboratories now use culture-independent diagnostic tests, PCR panels that diagnose patients quickly but produce no isolate unless the sample is also cultured. Metagenomic sequencing straight from samples is developing but does not yet match isolate WGS for resolution.
- Thresholds depend on context. SNP and allele cut-offs differ by organism, lineage and pipeline.
- Genes are not behaviour. A resistance or virulence gene may be silent or truncated; confirm with phenotypic tests.
- Databases see only what is submitted. Countries with the most sequencing capacity dominate public databases, so the absence of a match is weak evidence.
- Results have consequences. Agree in advance how the business will respond, cooperate with authorities and protect consumers if its isolates match clinical cases.
Frequently asked questions
Is whole genome sequencing the same as PCR testing?
No. PCR detects or identifies an organism by amplifying a short, chosen stretch of DNA, which makes it fast and suited to screening. Whole genome sequencing reads nearly the entire genome of a pure isolate, so strains can be compared with each other at single-base resolution. Plants typically use culture or PCR to find Listeria, then send selected isolates for WGS to learn where they came from.
How many SNPs or alleles mean two isolates are the same strain?
There is no universal number. Closely related isolates of foodborne pathogens often differ by single digits to low tens of SNPs or alleles, but the meaningful range depends on the organism, the lineage, the time between isolations and the analysis pipeline. A strain persisting in a factory for years can drift further apart. Treat thresholds as triggers for investigation and interpret them with epidemiological and site evidence.
Can WGS prove that a food caused an outbreak?
Not by itself. A close genomic match shows that isolates from patients and from a food or factory share a recent common ancestor, which makes a common source likely and directs the investigation. Proof of cause also needs epidemiological evidence, such as what patients ate, and traceback or product testing that shows how the food reached them. The three lines of evidence together support a causal conclusion.
Should a food company sequence its own environmental isolates?
For ready-to-eat producers with recurring Listeria findings, it is often worthwhile. Sequencing shows whether positives are one persistent strain or several introductions, which directs corrective action and shows whether a deep clean actually removed the strain. Freeze isolates from the environmental monitoring programme so they can be sequenced later, and agree in advance how results will be used and shared with authorities when required.
What is the difference between serotyping, MLST and cgMLST?
Serotyping groups isolates by surface antigens, and most human Listeria isolates fall into just a few serotypes. Classical MLST uses the sequences of seven housekeeping genes to assign a sequence type, such as ST6, which is useful for describing lineages. cgMLST compares hundreds to thousands of core genes (1,748 in a widely used Listeria scheme) and resolves individual strains within a lineage, the level needed for outbreak and persistence work.
Next step. Advanced Food Science covers genomic epidemiology, SNP and cgMLST interpretation, source attribution and antimicrobial resistance, alongside microbial stress physiology, biofilms and the predictive models used in process validation. It ends with a proctored final assessment and an ASC certificate. To compare levels and topics, see all eleven food science and technology courses.
Sources. FAO, Applications of Whole Genome Sequencing in Food Safety Management, technical background paper (FAO, 2016); WHO, Disease Outbreak News and WHO Regional Office for Africa reports on listeriosis in South Africa (2018); A. Moura et al., Whole genome-based population biology and epidemiological surveillance of Listeria monocytogenes, Nature Microbiology (2016); EFSA, Statement on the requirements for whole genome sequence analysis of microorganisms intentionally used in the food chain, EFSA Journal (2021, updated 2024); M. P. Doyle, F. Diez-Gonzalez and C. Hill (eds), Food Microbiology: Fundamentals and Frontiers, 5th edn (ASM Press, 2019).
This article is general guidance and is not a substitute for the applicable standard, your national legislation or the advice of a qualified food safety professional.