Reading and writing DNA are now infrastructure, and the interesting choices are about fit rather than capability. Read length decides which variants you can see at all, error profile decides what you can call, and cost per sample decides how many you can run. This guide catalogs 26 methods across six classes, with the read lengths, accuracies, throughputs and costs that decide between them. Instrument specifications move faster than anything else in this collection, so the numbers here are bands rather than datasheet values.
Sequencing by synthesis is the chemistry behind the instruments that have produced most of the world's sequence data. DNA fragments are attached to a flow cell surface and amplified in place into clusters of identical copies, so that each cluster gives a signal strong enough to read. Sequencing then proceeds one base at a time: fluorescently labeled nucleotides with a reversible chemical block are washed in, one is incorporated at each cluster, the flow cell is imaged, and the label and block are cleaved so the next cycle can run. Every cluster is read in parallel, and a modern flow cell holds billions of them, which is where the throughput comes from. Reads are short because the chemistry accumulates errors: molecules within a cluster gradually fall out of step with each other, and the signal degrades until base calling becomes unreliable, which happens somewhere around 250 to 300 cycles.
Strengths & weaknessesThe strengths are accuracy, throughput and cost. Raw per-base accuracy above 99.9% is routine, and the errors that do occur are mostly substitutions rather than insertions or deletions, which is the easier error type to handle in variant calling. Throughput on the largest instruments is measured in terabases per run, and cost per gigabase is the lowest of any method. The ecosystem is the real moat: two decades of tools, pipelines, reference datasets and trained people all assume this data. The weaknesses follow from the read length. Short reads cannot span repetitive regions, so they miss structural variation, cannot phase variants that sit far apart, and align ambiguously in segmental duplications and the several percent of the genome that stayed unresolved until long reads arrived. Amplification introduces bias against regions of extreme base composition and creates duplicate reads. The instruments are expensive and the largest ones only make economic sense at high, steady volume.
When to useUse short-read sequencing by synthesis as the default for counting applications and for small variants: whole genome and exome sequencing for single-nucleotide variants, RNA quantification, targeted panels, and any experiment where the answer comes from how many reads fall somewhere rather than from what a single molecule looks like end to end. It is the right choice when cost per sample matters and the question does not involve structure. Reach for a long-read platform when you need structural variants, phasing, repeat expansions, full-length transcript isoforms, or assembly of a new genome. Many projects now run both, using short reads for depth and accuracy and long reads for structure, and that combination is more often the right answer than either alone.
Key numbersRead lengths typically 100–300 bases, most commonly 150 bases paired-end · raw accuracy above 99.9%, with errors mostly substitutions · output from a few gigabases on a benchtop instrument to terabases on the largest · a 30-fold coverage human genome takes roughly 90–100 gigabases · cost per gigabase has fallen by several orders of magnitude since 2007 and continues to fall with each instrument generation · run times from several hours to a couple of days · instrument capital from tens of thousands to over a million dollars.
Failure modesThe characteristic failures are all consequences of short reads and amplification. Repetitive regions produce ambiguous alignment, so variants there are either missed or called wrongly, and this is silent: the pipeline returns a confident answer for the regions it can handle and simply says nothing about the ones it cannot. Amplification bias depletes coverage in regions of very high or very low base composition, which is why some clinically important regions have persistently poor coverage. PCR duplicates inflate apparent depth and can make a sequencing error look like a real variant if not removed. Index hopping on patterned flow cells assigns reads to the wrong sample in multiplexed runs, which matters most when looking for rare variants and is controlled with unique dual indexes. Coverage uniformity, not raw accuracy, is usually what limits a clinical assay.
ExamplesThe Illumina instrument line, which has dominated the market for over a decade, from benchtop sequencers through the highest-throughput production instruments; the UK Biobank and All of Us whole-genome programs, which produced hundreds of thousands of genomes on this chemistry; essentially every large population genomics dataset; and the RNA sequencing and single-cell workflows that use it as the readout.
Economic profileThe clearest example in life sciences of a razor-and-blade model, and one of the most profitable. Instruments are sold at moderate margin and consumables at high margin, with the flow cell as the recurring purchase, and the installed base plus the software ecosystem makes switching genuinely expensive even when a competitor's specifications are better. That position held for roughly fifteen years and has recently come under real pressure from several new entrants, which has moved list prices more in the last few years than in the decade before. For a buyer, the important consequence is that cost per gigabase quotes are contract-dependent and the list price is not the price; for anyone building a business on sequencing, the input cost is falling but the supplier concentration is a genuine risk.
Avidity sequencing keeps the cluster-and-image architecture of conventional short-read sequencing and changes how a base is identified. Instead of incorporating one labeled nucleotide per cycle and reading its color, it uses large multivalent substrates: a polymer core carrying many copies of the same nucleotide, which binds the sequencing complex through many simultaneous contacts at once. Because binding strength grows steeply with the number of contacts, the correct base binds far more tightly than an incorrect one, and the discrimination between right and wrong is much better than single-molecule binding gives. The bound complex can then be imaged at low reagent concentration, which reduces the background that limits conventional chemistry. Separating the recognition step from the incorporation step is the structural change: the accuracy comes from binding, and the chain extension happens afterwards with unlabeled nucleotides.
Strengths & weaknessesThe strengths are accuracy and reagent economics. Reported raw error rates are lower than conventional sequencing by synthesis by a meaningful margin, which reduces the coverage needed to make a confident call and therefore the cost of an answer rather than the cost of a base. Because the detection step uses very low nucleotide concentrations, reagent consumption per base is low. The chemistry also allows the extension step to be optimized separately from detection, which gives room to improve. The weaknesses are ecosystem and scale rather than chemistry. Read lengths are in the same short range as conventional platforms, so none of the structural limitations of short reads are addressed. Tooling, reference datasets and pipeline validation are far thinner than for the incumbent, which matters most in clinical settings where an assay has to be validated against established data. Installed base and service infrastructure are small by comparison, and throughput per instrument is below the largest production sequencers.
When to useUse avidity sequencing where short reads are the right answer and cost per sample or per confident call is the deciding factor, particularly in a laboratory that is not locked into an existing validated pipeline. It is a strong option for whole-genome sequencing at population scale, for RNA quantification, and for laboratories willing to trade ecosystem maturity for price. It is a weaker choice when an assay has to be validated against a body of existing data generated on another platform, or when the workflow depends on tools that assume a specific data format and error profile. As with every short-read platform, it does not address structural variation, phasing or repeat expansions, so it competes with other short-read instruments rather than with long reads.
Key numbersRead lengths in the same 100–300 base range as conventional short-read sequencing · reported raw accuracy better than conventional sequencing by synthesis, with vendor claims of substantially reduced error rates · multivalent binding gives far stronger discrimination between correct and incorrect bases than single-molecule binding · reagent concentrations during detection are orders of magnitude below conventional chemistry · commercially launched in the 2020s and gaining share · instrument and consumable pricing has been positioned aggressively below the incumbent.
Failure modesShares the structural limitations of all short-read methods: repetitive regions align ambiguously, structural variants are invisible, and phasing over any distance is impossible. Amplification is still used to make clusters, so amplification bias against extreme base composition and duplicate reads remain. The platform-specific risks are practical rather than chemical. Pipelines tuned to another platform's error profile can behave unexpectedly, and a variant caller trained on one error model applied to another will misestimate confidence. Reference datasets and benchmarking material generated on the incumbent platform do not transfer cleanly, so a laboratory switching platforms has real revalidation work, and underestimating that work is the most common way the cost advantage fails to materialize.
ExamplesThe Element Biosciences AVITI platform, which introduced the chemistry commercially and has been adopted by genome centers and service providers looking for an alternative to the incumbent; population sequencing programs that have run comparisons across platforms; and service laboratories that have added it alongside existing instruments rather than replacing them, which is the common adoption pattern.
Economic profileThe most consequential thing about this platform may be its effect on prices rather than its own market share. A credible short-read competitor with better raw accuracy forced list price movement in a market that had been stable for a decade, which benefits every buyer regardless of what they purchase. The business model is the same razor-and-blade structure, and the strategic question is whether an alternative can accumulate enough ecosystem, pipeline validation and clinical acceptance to displace an incumbent whose real moat was never the chemistry. History in instrument markets suggests that takes longer than the specification advantage alone would predict.
This approach attacks sequencing cost from the physical side rather than the chemical one. Instead of clusters on a patterned flow cell, DNA is amplified onto beads which are then packed at very high density onto a circular substrate that spins under the optics, so imaging is continuous rather than a step-and-repeat scan of discrete tiles. The chemistry also departs from conventional sequencing by synthesis: rather than adding one reversibly blocked base per cycle, mostly natural nucleotides are used with only a fraction labeled, and the sequence is inferred from signal intensity across flows of single bases. Removing the reversible terminator removes a slow and expensive chemical step. The result is very high throughput at low reagent cost, aimed squarely at applications where the requirement is an enormous number of bases rather than the highest possible per-base accuracy.
Strengths & weaknessesThe strength is cost per gigabase at scale, which has been the platform's entire proposition and has been genuinely disruptive to pricing expectations for large sequencing programs. Continuous imaging of a spinning substrate uses the optics far more efficiently than tile-by-tile scanning. The weakness is the error profile. Inferring sequence from flow intensities rather than from discrete terminated additions makes homopolymer runs difficult, because the signal from four identical bases in a row has to be distinguished from three or five by intensity alone, and that is exactly the error mode that flow-based chemistries have always struggled with. Insertion and deletion errors in homopolymers are consequently more common than on terminator-based platforms, which matters for clinical variant calling and less for counting applications. Read lengths remain short, so none of the structural limitations are addressed.
When to useUse this platform where the question is counting and the volume is large: RNA quantification, single-cell sequencing, methylation surveys, and very large population whole-genome programs where cost per sample dominates. It suits high-throughput production settings with steady volume rather than laboratories running occasional diverse projects. Avoid it where homopolymer accuracy matters, which includes clinical variant calling in genes with homopolymer runs and any application where insertions and deletions are the variants of interest. As always with short reads, structural variation, phasing and repeat expansions require a long-read platform regardless of which short-read instrument is chosen.
Key numbersRead lengths in the short-read range, comparable to other platforms in this class · very high throughput per run, aimed at the highest-volume production applications · reagent cost per base is low because mostly natural nucleotides replace reversible terminators · homopolymer regions carry a higher insertion and deletion error rate than terminator chemistries · adoption has concentrated in large genome centers and high-volume service providers rather than in clinical laboratories.
Failure modesThe signature failure is homopolymer miscalling, and it is systematic rather than random, which makes it worse than an equivalent rate of random error: the same run of identical bases will be miscalled the same way every time, so extra coverage does not fix it. Pipelines have to account for this explicitly, and variant callers tuned for substitution-dominated error profiles will misbehave. The short-read limitations all apply: ambiguous alignment in repeats, invisible structural variation, no phasing. Because the platform is optimized for scale, small runs are economically inefficient, and a laboratory with variable or low volume will not see the cost advantage that justifies the platform.
ExamplesThe Ultima Genomics platform, which introduced this architecture with an explicit goal of driving whole-genome cost sharply below prevailing prices; adoption by large genome centers and by companies running very high volumes of single-cell and methylation sequencing; and use in population-scale research programs where the total base requirement is the binding constraint.
Economic profileA pure cost-per-base play, and it has moved the market's expectations even where it has not won the sale. The strategy is to serve the highest-volume buyers, whose economics are dominated by consumable cost and who have the pipeline sophistication to handle a different error profile, and to leave clinical and low-volume segments to incumbents. That is a coherent position and it targets the segment where switching costs are lowest relative to the saving. The broader effect on the industry is the same as any credible competitor entering a long-stable market: prices move for everyone, and buyers benefit whether or not they switch.
DNA nanoball sequencing replaces the amplification step that most short-read platforms use with a different one that does not compound errors. A DNA fragment is circularized, then copied round and round by rolling circle amplification to produce a long single strand containing hundreds of tandem copies of the original, which collapses into a compact ball a few hundred nanometers across. Because every copy is made from the original circle rather than from a previous copy, errors do not accumulate the way they do in PCR, where a mistake early in amplification is propagated to everything downstream. The nanoballs are loaded onto a patterned array, one per spot, and sequenced by a combinatorial probe-anchor chemistry. The architecture gives high signal density with low reagent consumption, and it avoids the duplicate reads and exponential error propagation of PCR-based cluster generation.
Strengths & weaknessesThe strengths are error containment, low reagent use and cost. Linear rather than exponential amplification means the error profile is cleaner in a specific way that matters for detecting low-frequency variants, since a PCR error made in an early cycle looks like a real low-frequency variant and a rolling circle error does not propagate. Duplicate rates are very low because each nanoball comes from one original molecule. Cost per gigabase is competitive with any platform, and the highest-throughput instruments in this line are among the largest-output sequencers available. The weaknesses are ecosystem and geopolitics rather than chemistry. Read lengths are short with all the usual consequences. Tooling and clinical validation are less developed outside China than for the incumbent. Most consequentially, the platform's availability in the US market has been affected by patent litigation and by legislation aimed at Chinese biotechnology companies, which is a supply risk unrelated to performance.
When to useUse nanoball sequencing where short reads are appropriate and cost matters, and particularly for low-frequency variant detection where the linear amplification advantage is real, such as liquid biopsy and somatic variant calling at low allele fraction. It is widely used outside the US and is a genuine alternative to the market leader on both cost and specifications. The decisive consideration for a US-based laboratory is usually supply and regulatory risk rather than technical merit: patent disputes and legislative attention have made availability uncertain in ways that a laboratory building a clinical service cannot ignore. Where that risk is acceptable or does not apply, the technical case is strong.
Key numbersRead lengths in the short-read range, typically 100–200 bases · rolling circle amplification is linear, so errors do not propagate exponentially as in PCR · duplicate read rates are very low because each nanoball derives from a single original molecule · the highest-throughput instruments in this family are among the largest-output sequencers available · widely deployed in China and internationally, with restricted availability in the US market · cost per gigabase competitive with any platform.
Failure modesAll the short-read limitations apply: repeats align ambiguously, structural variants are invisible, phasing is impossible over any distance. Platform-specific issues are mostly practical. Error profiles differ from the incumbent's, so variant callers and quality thresholds need retuning, and a laboratory migrating an assay has real revalidation work. Reference and benchmarking datasets are more available for the incumbent. The most significant practical failure mode has nothing to do with the chemistry: a clinical service built on an instrument whose market availability depends on ongoing litigation and legislation carries a continuity risk that should be assessed explicitly rather than discovered.
ExamplesThe MGI and Complete Genomics instrument lines, which use this chemistry across a range from benchtop to very high throughput; large national sequencing programs outside the US built on the platform; and the extensive patent litigation with the market incumbent, which has shaped where the instruments can be sold and is as important to understanding the platform's position as any specification.
Economic profileAggressively priced and technically credible, and its commercial trajectory has been shaped as much by law and policy as by performance. The platform has taken substantial share in markets where it can compete freely and has been constrained in the largest market by litigation and legislative scrutiny of Chinese biotechnology. For buyers, the effect of its existence has been to put sustained pressure on prices everywhere. For anyone planning a sequencing-dependent business, the lesson generalizes beyond this platform: supplier concentration in sequencing is a real strategic exposure, and the alternatives to the incumbent each carry their own non-technical risk.
Nanopore sequencing reads a DNA molecule by pulling it through a protein pore set in a membrane and measuring the change in ionic current as bases pass through. Each combination of bases sitting in the pore's narrow constriction blocks the current by a characteristic amount, and a neural network converts the resulting current trace into sequence. Nothing is copied, so there is no amplification and no incorporation chemistry, and the read continues for as long as the molecule keeps threading through: read length is set by the length of DNA you managed to extract, not by the instrument. That is the defining property. It also means the native molecule is what gets read, so chemical modifications such as methylation change the current signal and are detected directly, without bisulfite conversion or any separate assay.
Strengths & weaknessesThe strengths are read length, direct modification detection, real-time output and portability. Reads of tens of kilobases are routine and reads over a megabase have been reported, which resolves repeats, structural variants and phasing that short reads cannot touch. Methylation comes free with the sequence. Data streams as the run proceeds, so an answer can arrive in minutes rather than after a run completes, and the smallest devices run off a laptop in the field. The weaknesses are accuracy and throughput. Raw per-read accuracy has improved enormously across pore and basecaller generations but remains below short-read platforms, and the residual errors concentrate in homopolymers and in specific sequence contexts, which is a systematic error that coverage does not fully cure. Throughput per flow cell is well below the largest short-read instruments, so very large projects cost more. Yield depends heavily on DNA quality, and a poorly extracted sample gives short reads regardless of the platform's capability.
When to useUse nanopore sequencing whenever the question involves structure: structural variants, repeat expansions, phasing, full-length transcript isoforms, or assembling a genome without a reference. It is the right tool for rapid clinical answers where time matters, including same-day whole-genome sequencing in critical care, and for any setting where a laboratory is not available, from field surveillance to space stations. Direct methylation detection makes it attractive for epigenetics without a separate library preparation. Use short reads instead when the answer comes from counting, when very large sample numbers dominate cost, or when the assay must call single-nucleotide variants at the highest confidence. Invest in high-molecular-weight extraction before blaming the platform for short reads, because sample preparation is usually the limiting factor.
Key numbersRead length limited by DNA fragment length rather than by chemistry, with tens of kilobases routine and megabase reads reported · raw per-read accuracy has improved across successive pore and basecaller generations and remains below short-read platforms · errors concentrate in homopolymers and specific sequence contexts · methylation detected directly from the current signal, with no bisulfite conversion · data available in real time as the run proceeds · devices range from a pocket-sized flow cell to high-throughput benchtop instruments · very low instrument capital at the small end.
Failure modesThe dominant practical failure is sample preparation. Read length is set by fragment length, so ordinary extraction methods that shear DNA give short reads on a long-read instrument, and laboratories new to the platform routinely conclude the technology underperforms when the extraction is at fault. Homopolymer and context-specific errors are systematic, so they persist at high coverage and require basecallers and variant callers matched to the current chemistry. Pore blocking and membrane degradation reduce active pores over a run, so yield falls with time and with dirty samples. Because basecalling is a machine-learning step, results change when the basecaller is updated, which is a real reproducibility consideration: the same raw data reanalyzed later can give different calls, and clinical use requires pinning versions.
ExamplesOxford Nanopore's device range from the pocket-sized MinION to high-throughput PromethION instruments; the telomere-to-telomere human genome assembly, which used ultra-long nanopore reads to resolve regions no short-read technology could; rapid whole-genome sequencing in neonatal intensive care, where turnaround measured in hours changes clinical management; Ebola and SARS-CoV-2 field surveillance sequencing; and use aboard the International Space Station.
Economic profileA genuinely different business model from the razor-and-blade instrument market: low capital cost, flow cells sold in a range from very cheap to production scale, and no large instrument purchase required to start. That has made sequencing accessible to laboratories and countries that could not justify a production sequencer, which is a real democratizing effect and a large part of the platform's strategic significance. Cost per gigabase remains above the highest-throughput short-read platforms, so it competes on capability and on turnaround rather than on bulk price. The company's independence in a market otherwise consolidated around a single supplier is itself commercially important.
Single-molecule real-time sequencing watches one polymerase copy one DNA molecule, detecting each nucleotide as it is incorporated by the flash of fluorescence from a label that is cleaved away immediately afterwards. The polymerase sits at the bottom of a well small enough that only the nucleotide being incorporated is illuminated. On its own this gives long reads with a fairly high raw error rate, and the important development was circular consensus: the DNA fragment is circularized, the polymerase goes round it many times, and the resulting repeated passes over the same molecule are collapsed into a consensus. Because the errors are random rather than systematic, averaging many passes drives accuracy up sharply. The result, called HiFi, is reads of roughly 15 to 25 kilobases at accuracy comparable to short-read sequencing, which is the combination that made long reads acceptable for clinical variant calling.
Strengths & weaknessesThe strength is having both length and accuracy at once, which nothing else offers. Reads long enough to phase variants and span most repeats, at per-base accuracy that supports confident single-nucleotide variant calling, means one dataset answers questions that previously required two platforms. Methylation is detected directly from polymerase kinetics with no separate assay. Assemblies from HiFi data are the current standard for reference-quality genomes. The weaknesses are cost, throughput and the length ceiling. Circular consensus spends sequencing capacity on repeated passes over the same molecule, so usable output per run is well below what the raw chemistry generates and well below short-read platforms, which makes cost per gigabase substantially higher. Read lengths, while long, are shorter than nanopore's, because the fragment must be short enough for the polymerase to circle it several times. Instruments are expensive and input DNA quality requirements are demanding.
When to useUse HiFi when you need long reads and high accuracy in the same dataset: reference genome assembly, clinical structural variant detection where a confident call is required, phasing across a gene, and comprehensive variant detection where missing something is worse than the cost. It has become the standard for rare disease genome sequencing in programs that can afford it, because a single assay finds small variants, structural variants, repeat expansions and methylation together. Choose nanopore instead when you need reads longer than about 25 kilobases, when turnaround time is critical, or when cost per sample is the constraint. Choose short reads when the question is counting and structure does not matter. Sample quality matters as much as it does for nanopore.
Key numbersRead lengths typically 15–25 kilobases for HiFi mode · accuracy comparable to short-read platforms after circular consensus, from averaging many passes over the same molecule · errors in the raw signal are random rather than systematic, which is why consensus works so well · methylation detected directly from polymerase kinetics · output per run well below short-read platforms, so cost per gigabase is several times higher · instrument capital in the high hundreds of thousands · the standard input for reference-grade genome assembly.
Failure modesThe main trap is the trade-off between read length and accuracy, which is set by how many passes the polymerase makes: longer inserts mean fewer passes and lower consensus accuracy, so a library pushed toward maximum length quietly loses the accuracy that justified the platform. Input DNA quality is critical, and degraded samples produce short inserts and poor yield. Cost per gigabase is high enough that projects sized on short-read intuitions come out several times over budget, which is the most common planning error. The platform reads through most but not all difficult regions, so the longest repeat arrays and some centromeric sequence still need ultra-long reads from another technology, and assuming HiFi resolves everything is a mistake.
ExamplesPacBio's Revio and related instruments; the Human Pangenome Reference Consortium assemblies, which combined HiFi with ultra-long nanopore reads to produce reference-quality diploid genomes; rare disease programs that replaced exome sequencing with long-read genome sequencing and found diagnoses that short reads had missed; and the telomere-to-telomere reference work, where HiFi provided the accurate backbone and nanopore the ultra-long scaffolding.
Economic profilePositioned as the premium long-read platform, sold on the value of a complete and confident answer rather than on cost per base. That works well in rare disease diagnostics, where a diagnosis after years of testing is worth a great deal, and in reference genomics, where quality is the product. It works poorly against short reads for any counting application and against nanopore where turnaround or capital cost dominates. The competitive dynamic between the two long-read platforms has been good for buyers, with both improving accuracy and throughput quickly, and the practical result is that many large projects now use both rather than choosing.
Sequencing by expansion solves a physical problem with nanopore sensing. DNA bases are about 0.34 nanometers apart, which is far smaller than the sensing region of any practical pore, so several bases sit in the constriction at once and the current signal is a blur of all of them that has to be deconvolved computationally. This chemistry sidesteps that by never sequencing the DNA itself. The template is first copied into a surrogate polymer, an expandable molecule in which each original base is represented by a much larger reporter unit separated by a long spacer. Threading that surrogate through a pore presents one reporter at a time, well separated from its neighbors, so the signal is a clean sequence of discrete events rather than an overlapping blur. Reading the surrogate instead of the original trades an extra enzymatic conversion step for a much easier measurement.
Strengths & weaknessesThe strengths, as claimed and demonstrated in early data, are speed and signal quality. Because reporters are spaced far apart, they can be translocated quickly while still being resolved, which supports very high per-pore throughput. The signal is discrete rather than convolved, so basecalling is a simpler problem than for direct nanopore sensing and homopolymers are not intrinsically ambiguous. The weaknesses are that it is new and that the conversion step is an additional place for things to go wrong. The template must be copied into the surrogate faithfully, and any error or incompleteness in that step is indistinguishable from a sequencing error downstream. Read lengths reported so far are shorter than direct nanopore reads. There is essentially no independent validation, no ecosystem, no clinical precedent, and no track record of how the platform behaves on difficult samples. Specifications from a platform that has not shipped broadly should be treated as targets.
When to useThere is no established use case yet, which is the honest position for a platform at this stage. The reason to track it is that if the claims hold it changes the cost and throughput assumptions that everything else on this sheet is priced against, and it comes from an organization with the resources to commercialize at scale. For a laboratory, the practical advice is to wait for independent benchmarking on standard reference materials before making plans that depend on it, and to be skeptical of comparisons run by the vendor on samples of their choosing. For anyone building a business whose economics depend on sequencing cost, this is a platform whose arrival would be worth modeling as a scenario.
Key numbersBases in native DNA are about 0.34 nanometers apart, well below the resolution of a practical nanopore, which is the problem this chemistry exists to solve · the surrogate polymer spaces reporters far enough apart to be read individually · reported translocation speeds are much faster than direct nanopore sequencing while retaining signal resolution · read lengths reported so far are shorter than direct nanopore reads · announced in the mid-2020s and not yet broadly deployed · no independent benchmarking at the time of writing.
Failure modesThe conversion step is the structural risk: any infidelity in copying the template into the surrogate produces errors that look exactly like sequencing errors and cannot be distinguished from them after the fact, so the accuracy of the whole method is bounded by a step that direct sequencing does not have. Beyond that, the failure modes of a new platform are the ordinary ones and they are the ones that matter: unvalidated performance on degraded or low-input samples, absent pipeline tooling, no reference datasets, and specifications quoted under favorable conditions. Any laboratory adopting early should plan for substantial internal validation work and should not assume that published specifications will hold on their sample types.
ExamplesRoche's sequencing by expansion platform, announced in the mid-2020s as the company's return to the sequencing market after previously exiting it; early demonstration data presented at industry conferences; and the broader set of nanopore-adjacent approaches that attempt to make single-molecule electrical sensing easier by changing what is threaded through the pore rather than by improving the pore.
Economic profileImpossible to assess directly, and strategically interesting regardless. A large diagnostics company entering sequencing with a differentiated architecture is a meaningful competitive event in a market that has been effectively a duopoly at the high end and a monopoly for most of its history. Whether or not this specific platform succeeds, the pattern of the last few years, with several credible new entrants using genuinely different physics, has already changed pricing behavior. For buyers that is good; for anyone whose business model assumes sequencing costs stay where they are, it is a reason to model further declines.
Sanger sequencing works by deliberately breaking DNA synthesis. A polymerase copies the template in the presence of a small proportion of chain-terminating nucleotides, each labeled with a different color, so the reaction produces a population of fragments that all start at the same place and stop at every possible position. Separating those fragments by size in a capillary and reading the color at each length gives the sequence directly. It was the method that sequenced the first human genome, and although it was displaced for anything at scale, it remains the default for reading one thing at a time. Two properties keep it alive: a single reaction reads 700 to 900 bases of very high quality with no library preparation and no informatics, and the resulting trace is a directly interpretable picture that a person can look at and judge.
Strengths & weaknessesThe strengths are simplicity, turnaround and interpretability. For one amplicon, one plasmid or one clone, the answer arrives the same day for a few dollars with no library preparation, no multiplexing, no alignment and no variant caller. The trace is human-readable, which matters when a result has to be defended. Accuracy in the middle of a read is very high, which is why it has been the confirmatory method for clinical variant calling for decades. The weaknesses are throughput and sensitivity. One reaction reads one molecule population, so cost per base is orders of magnitude above any next-generation platform and sequencing more than a handful of targets stops making sense quickly. It cannot detect a variant present in less than roughly 15–20% of the molecules, so low-frequency somatic variants and mosaicism are invisible. The first 20 to 40 bases after the primer are unreliable, and read quality degrades toward the end.
When to useUse Sanger sequencing when you have a small number of specific things to check: verifying a plasmid or a cloned construct, confirming a variant found by another method, checking an edited cell line at a known site, or genotyping a handful of samples. It is the default for sequence verification in molecular biology and remains the confirmatory standard in many clinical laboratories. The crossover point is worth knowing: below roughly ten to twenty targets, Sanger is cheaper and faster than preparing a sequencing library, and above that it is not. Do not use it to look for low-frequency variants, to survey anything large, or to detect structural variation, and be aware that its detection floor around 15–20% allele fraction means a negative Sanger result does not exclude a low-level variant.
Key numbersRead length 700–900 bases per reaction, with the first 20–40 bases unreliable · very high accuracy in the middle of the read · detection limit roughly 15–20% allele fraction, so low-frequency variants are invisible · cost per reaction of a few dollars, which is very high per base and very low per answer · turnaround same day, with no library preparation · the crossover against next-generation sequencing sits at roughly ten to twenty targets · the method used for the original human genome project.
Failure modesThe characteristic failure is a mixed trace. If the template contains two different sequences, from a heterozygous insertion or deletion, from a mixed clone, or from a poorly specific PCR, the trace becomes a superposition of two sequences after the difference and is unreadable without deconvolution software. Indels are the usual cause and the reason a clean-looking failure is often misdiagnosed as a bad reaction. The detection floor is the other trap: a laboratory using Sanger to confirm variants can report a negative when a real variant is present below 15%, which matters in somatic testing and mosaicism. Poor template quality, primer dimers and secondary structure all degrade traces, and homopolymer runs can slip.
ExamplesPlasmid and construct verification, which is still overwhelmingly done this way; clinical confirmation of variants found by next-generation panels, long the standard practice; genotyping of edited cell lines at a known target site; the Human Genome Project, completed on capillary Sanger instruments; and the commercial services that will sequence a tube of plasmid overnight for a few dollars, which is what keeps the method ubiquitous.
Economic profileA mature commodity service business with very thin margins and enormous volume, sustained by the fact that molecular biology still generates a constant stream of single constructs that need checking. Instrument sales have declined for years while sample volumes have not, because the work moved to service providers who run capillary instruments at scale. The interesting recent development is competitive pressure from cheap targeted next-generation sequencing and from nanopore-based plasmid sequencing services, both of which read a whole plasmid in one pass rather than requiring several primer walks, and which are eroding the one application that kept Sanger indispensable.
An amplicon panel amplifies a defined set of regions by PCR and sequences only those, rather than sequencing everything and discarding most of it. Hundreds to thousands of primer pairs are pooled into one or a few multiplex reactions, the products are barcoded and sequenced, and coverage lands where the primers put it. The economics are the point: sequencing a 50-gene panel at very high depth costs a fraction of sequencing a whole genome at low depth, and the depth is what lets the assay detect variants present in a small fraction of the molecules, which is what somatic and liquid biopsy testing require. Amplicon panels are also the fastest targeted method, since the workflow is PCR plus sequencing with no hybridization step, and turnaround measured in hours matters for infectious disease and for time-sensitive oncology testing.
Strengths & weaknessesThe strengths are depth, speed, cost and low input requirements. Coverage of thousands of times over a small region is affordable, which pushes the detection limit for low-frequency variants down to a fraction of a percent when combined with molecular barcodes. Input DNA requirements are low, which matters for small biopsies and degraded samples. Workflow is short. The weaknesses come from PCR. Primers amplify unevenly, so coverage varies across the panel and some regions consistently underperform, which has to be characterized and reported rather than assumed away. Regions where a primer binding site carries a variant will amplify poorly or not at all, causing allele dropout that produces a false negative in exactly the sample that has the variant. Panels detect only what they target, so a variant outside the design is invisible by construction, and copy number and structural variants are difficult. Redesigning a panel to add genes is a real revalidation exercise.
When to useUse an amplicon panel when the gene set is known, the required depth is high, and turnaround or cost per sample matters. It is the standard approach for somatic oncology panels, liquid biopsy, infectious disease typing including viral genome sequencing, and inherited disease panels where the relevant genes are well established. Use hybrid capture instead when the target region is large, when uniform coverage matters more than speed, or when the design needs to change often, since capture handles large and evolving target sets better. Use whole genome or exome sequencing when the answer might be outside the panel, which is the situation panels handle worst and where their apparent cost advantage disappears if a second test is needed.
Key numbersPanels typically cover from a handful to a few thousand targets · depth of hundreds to thousands of times over the target region is routine and affordable · detection limits below 1% allele fraction with molecular barcodes, against 15–20% for Sanger · input requirements as low as tens of nanograms, and lower for optimized designs · turnaround measured in hours for the fastest workflows · cost per sample well below whole genome sequencing, which is the entire argument.
Failure modesAllele dropout is the failure that matters, and it is insidious because it produces a confident wrong answer: a variant under a primer binding site prevents amplification of that allele, so the assay reports the sample as homozygous reference or homozygous variant when it is neither. Coverage non-uniformity means some targets are always weakly covered, and unless minimum coverage is enforced per target rather than averaged, a region can be reported as negative when it was never adequately read. PCR errors introduced in early cycles look identical to real low-frequency variants, which is why molecular barcodes are essential for any application below a few percent allele fraction. Amplicon boundaries also make copy number and structural variant calling unreliable.
ExamplesSomatic oncology panels used routinely in clinical pathology; circulating tumor DNA assays for treatment selection and monitoring; the ARTIC protocol for SARS-CoV-2 genome sequencing, which used tiled amplicons and was deployed globally during the pandemic; inherited cardiomyopathy and cancer predisposition panels; and 16S ribosomal RNA amplicon sequencing for microbial community profiling.
Economic profileThe workhorse of clinical sequencing economics. Panels are what make routine molecular pathology affordable, since they turn an expensive sequencing run into a cheap per-sample test by only reading what matters, and reimbursement structures have grown up around them. The commercial tension is that panels have to be revalidated whenever the gene content changes, while the clinically actionable gene list keeps growing, so laboratories face repeated revalidation costs and a slow drift toward exome or genome sequencing, where content changes are a reanalysis rather than a new assay. That drift is real but slow, because reimbursement and turnaround still favor panels.
Hybrid capture enriches a chosen part of the genome by fishing for it. A sequencing library is made from the whole sample, then mixed with a pool of biotinylated probes complementary to the regions of interest. The probes hybridize to matching fragments, streptavidin beads pull them out, everything else is washed away, and the captured library is amplified and sequenced. Because enrichment happens after library construction rather than through amplification of specific targets, the fragments retain their original ends, which preserves the information that molecular barcodes and duplicate detection depend on and makes copy number analysis work properly. Probe pools scale to very large target sets: exome capture covers all roughly 20,000 protein-coding genes, and custom panels can cover anything from a few genes to tens of megabases without redesigning a primer scheme.
Strengths & weaknessesThe strengths are coverage uniformity, scale and design flexibility. Coverage is more even than amplicon panels achieve, so fewer regions fall below the minimum depth an assay requires. Target sets can be very large, and adding content means adding probes to a pool rather than rebalancing a multiplex PCR. Original fragment ends are preserved, so duplicate reads can be identified properly and copy number can be called from read depth. Off-target reads still carry usable low-coverage information about the rest of the genome. The weaknesses are time, cost and input. The workflow includes an overnight hybridization, so turnaround is a day or more longer than an amplicon panel. Cost per sample is higher, from the probe pool and the extra steps. Input DNA requirements are greater, which is a real constraint on small biopsies. Capture efficiency falls in regions of extreme base composition and in repetitive sequence, where probes bind poorly or promiscuously.
When to useUse hybrid capture when the target region is large, when uniform coverage matters, or when the content will change over time. Exome sequencing, comprehensive cancer panels of hundreds of genes, and any assay where copy number is part of the answer are the natural applications. It is also the better choice when molecular barcode-based error correction is needed on native fragment ends. Use amplicon panels instead when input is scarce, turnaround must be measured in hours, or the target set is small and stable, since capture's advantages do not repay its cost at small scale. For very large target sets the calculation flips again: once a panel covers a substantial fraction of the exome, whole genome sequencing can be cheaper than capturing.
Key numbersTarget sets from a few genes to the full exome of roughly 20,000 genes and beyond · coverage uniformity better than amplicon panels, with fewer targets falling below minimum depth · workflow includes an overnight hybridization, adding a day or more against amplicon methods · input requirements higher than amplicon panels, typically tens to hundreds of nanograms · on-target rates commonly 50–80% depending on design and target size · cost per sample higher than amplicon panels and well below whole genome sequencing for small target sets.
Failure modesCapture efficiency varies with sequence, so regions of extreme base composition and repetitive sequence are consistently under-covered, and a panel's weak spots are systematic rather than random. Probes can cross-hybridize to similar sequences elsewhere, pulling in pseudogene copies that then misalign and generate spurious variant calls, which is a well-known problem in genes with close paralogs. Off-target capture wastes sequencing capacity and is the main determinant of how much data a sample needs. Because the workflow is long, sample tracking errors have more opportunities to occur, and unique dual indexes are important. As with any enrichment, variants outside the design are invisible, and expanding coverage later means reprocessing samples rather than reanalyzing data.
ExamplesClinical exome sequencing, which is the largest single application; comprehensive genomic profiling panels in oncology covering hundreds of genes with copy number and fusion detection; circulating tumor DNA assays that combine capture with molecular barcodes to reach very low detection limits; methylation capture panels; and target enrichment for non-human genomics where a reference exists but whole genome sequencing is unaffordable at the sample numbers required.
Economic profileSits in the middle of the targeted sequencing market and is squeezed from both sides. Amplicon panels are cheaper and faster for small stable target sets, and falling whole-genome costs erode the case for capture at the large end, since capturing an increasingly large fraction of the genome eventually costs more than sequencing all of it. The probe pool is the recurring consumable and where suppliers make their margin. The long-term direction is clear even if the timing is not: as sequencing gets cheaper, enrichment becomes harder to justify, and the industry drifts toward sequencing everything and filtering computationally, which also removes the revalidation burden that comes with changing an assay's content.
A genotyping microarray measures which version of a chosen set of variants a sample carries, without sequencing anything. Hundreds of thousands to millions of short probes are printed or bead-loaded onto a chip, each complementary to a specific known variant site. Sample DNA is fragmented, labeled and hybridized, and the fluorescence at each probe reports which allele is present. The critical difference from sequencing is that an array only sees what was designed into it: it is an interrogation of known positions, not a discovery method. That constraint is also its strength, because measuring a million pre-selected positions is far cheaper than sequencing a genome, and for common variation it is enough. Imputation extends the reach considerably, using population reference panels to infer the genotypes of millions of unmeasured variants from the correlation structure of the ones that were measured.
Strengths & weaknessesThe strengths are cost and throughput. Arrays cost a small fraction of even the cheapest whole-genome sequencing, run on simple instruments, and produce a compact well-behaved dataset that needs no alignment or variant calling. For genome-wide association studies across hundreds of thousands of people, they remain the economically rational choice. Copy number can be inferred from signal intensity. The weaknesses follow directly from measuring only known sites. Rare and novel variants are invisible, which makes arrays unsuitable for diagnosing rare disease, where the causal variant is usually rare by definition. Probe performance varies and some sites fail consistently. Ancestry bias is a serious and underappreciated problem: arrays were designed around variation cataloged mostly in European-ancestry populations, so both direct coverage and imputation accuracy are worse for other ancestries, which propagates into polygenic scores and association studies.
When to useUse arrays when you need common variant genotypes across very large numbers of samples and cost per sample dominates: genome-wide association studies, polygenic risk scoring, population biobanks, agricultural breeding programs, and consumer ancestry testing. They remain the right tool for that work. Do not use them for rare disease diagnosis, for anything requiring novel variant discovery, or for populations whose variation is poorly represented in the array design and the imputation reference. Check imputation accuracy for the specific ancestries in your cohort rather than accepting a global figure, because the difference between populations is large enough to change conclusions. As sequencing costs fall, the crossover point where low-coverage whole genome sequencing beats an array keeps moving, and it is worth recalculating rather than assuming.
Key numbersHundreds of thousands to a few million variant positions per array · cost per sample a small fraction of whole genome sequencing, which is the entire rationale · measures only designed positions, so novel and rare variants are invisible · imputation extends coverage to millions of additional variants using population reference panels · imputation accuracy is substantially lower for ancestries under-represented in reference panels · copy number inferred from intensity rather than measured directly · used for the great majority of published genome-wide association studies.
Failure modesThe dominant systematic failure is ancestry bias, and it is not a technical artifact but a design consequence: probes were chosen from variant catalogs built mostly from European-ancestry cohorts, so an array measures other populations less well and imputes them worse, and polygenic scores built on that data transfer poorly. This has produced published findings that do not replicate across populations. Beyond that, probes at sites where a nearby variant disrupts hybridization give wrong genotypes, batch effects between chip lots can masquerade as association signals if cases and controls were run separately, and copy number inference from intensity is noisy. The most common analysis error is treating imputed genotypes as measured ones without carrying their uncertainty through.
ExamplesThe UK Biobank genotyping of roughly half a million participants, which underpins a very large fraction of modern human genetics; consumer ancestry and health testing services; agricultural genomic selection programs in livestock and crops, where arrays are used at enormous scale for breeding decisions; and the genome-wide association studies of the 2007 to 2020 period, essentially all of which used arrays.
Economic profileA mature, low-margin, high-volume business, and one of the few places in genomics where the incumbent's cost advantage comes from manufacturing scale rather than from chemistry. The strategic question is how long arrays remain rational as sequencing prices fall. Low-coverage whole genome sequencing with imputation already competes on cost in some settings and gives a superset of the information, including novel variants and better performance across ancestries. Consumer genomics has been the most visible commercial application and also the most volatile, since the business model depends on repeat engagement that a one-time test does not naturally produce.
Optical genome mapping produces a picture of genome structure without reading a single base. Very long DNA molecules, hundreds of kilobases to megabases, are labeled at every occurrence of a specific short sequence motif, then stretched out in nanochannels and imaged. What you get per molecule is the pattern of distances between labels, which is a fingerprint that can be aligned to a reference map. Because the molecules are enormous compared with any sequencing read, rearrangements that span hundreds of kilobases are visible directly rather than inferred from breakpoint evidence. The method replaces karyotyping, fluorescence in situ hybridization and chromosomal microarray with one assay, and it detects balanced translocations and inversions, which microarrays cannot see at all because no material is gained or lost.
Strengths & weaknessesThe strengths are the size of what it can see and its coverage of balanced events. Molecules in the megabase range span rearrangements that no sequencing read approaches, and detection of balanced translocations and inversions fills a genuine gap between microarray, which misses them entirely, and karyotyping, which sees them at very low resolution. Resolution is far better than karyotyping and the workflow is simpler than running three separate cytogenetic assays. The weaknesses are that it produces no sequence. There is no base-level information, so point mutations are invisible and a breakpoint is localized to a region rather than to a base, which matters when the exact junction is clinically relevant. Ultra-high-molecular-weight DNA extraction is required and is genuinely demanding, which limits which sample types work and rules out most degraded or fixed material. Repetitive regions with dense or sparse labeling are harder. The installed base is small and the ecosystem is thin.
When to useUse optical genome mapping when the question is structural and large: hematological malignancy characterization, constitutional disorders where a balanced rearrangement is suspected, repeat expansion sizing, and any case where karyotyping plus fluorescence in situ hybridization plus microarray is the current standard and could be replaced by one assay. It is a strong complement to short-read sequencing, which handles the small variants it cannot see, and the pair covers most of the diagnostic space. Do not use it as a discovery method for point mutations or as a standalone comprehensive test. Confirm that your sample types can yield ultra-high-molecular-weight DNA before planning a service around it, because extraction is the practical gate and fixed or degraded samples will not work.
Key numbersMolecules typically hundreds of kilobases to megabases in length, far beyond any sequencing read · resolution for structural variants substantially better than karyotyping · detects balanced translocations and inversions, which chromosomal microarray cannot see · produces no base-level sequence, so point mutations are invisible · requires ultra-high-molecular-weight DNA, which rules out most fixed and degraded samples · replaces three conventional cytogenetic assays with one workflow · small installed base relative to sequencing platforms.
Failure modesSample preparation is the recurring failure, and it is unforgiving: the assay depends on molecules that ordinary extraction methods shear, so a laboratory that has not solved ultra-high-molecular-weight extraction for its sample types will get poor results and may attribute them to the platform. Fixed tissue generally does not work. Regions with unusually dense or sparse label sites are poorly resolved, so coverage of the genome is not uniform in the way sequencing coverage is. Breakpoints are localized to a window rather than a base, so a result often needs sequencing follow-up to characterize the junction, which means the assay is frequently a first step rather than a final answer. Interpretation requires cytogenetics expertise that not every molecular laboratory has.
ExamplesAdoption in hematological malignancy workups, where it has been compared directly against the standard combination of karyotype, fluorescence in situ hybridization and microarray and found comparable or better for structural findings; constitutional genetic testing for balanced rearrangements; repeat expansion sizing in disorders where the expansion is too large for sequencing to size accurately; and the Bionano platform, which is the main commercial implementation.
Economic profileA niche instrument business competing against both an entrenched conventional workflow and an improving alternative. The value proposition against cytogenetics is real, consolidating three assays into one with better resolution, and reimbursement and laboratory practice change slowly, which has made adoption gradual. The longer-term pressure comes from long-read sequencing, which is closing the gap on structural variant detection while also providing base-level sequence, and which would make a structure-only assay harder to justify. The durable case is where molecule length beyond even long reads is what matters, which is a real but narrow requirement.
Single-cell RNA sequencing measures gene expression one cell at a time instead of averaging across a tissue. The dominant approach partitions cells into droplets, each containing one cell and one bead carrying millions of copies of a barcode unique to that droplet. Inside the droplet the cell is lysed and its messenger RNA is captured and tagged with that barcode plus a molecular identifier unique to each original transcript. Everything is then pooled and sequenced together, and the barcodes sort the reads back into cells computationally. The result is a matrix of tens of thousands of cells by tens of thousands of genes. Bulk RNA sequencing tells you the average expression of a tissue, which for a mixture of cell types is a number that may describe no cell in it; single-cell resolves the actual populations, which is why it displaced bulk methods for any question about heterogeneity.
Strengths & weaknessesThe strengths are resolution and discovery. Cell types and states can be identified without knowing in advance what to look for, rare populations become visible, and developmental trajectories can be reconstructed from a single snapshot. The technique has reshaped immunology, developmental biology and tumor biology in about a decade. The weaknesses are cost, sparsity and sample handling. Only a fraction of each cell's transcripts is captured, so the data are sparse and most genes in most cells read as zero, which makes distinguishing genuine absence from a capture failure a persistent analytical problem. Cost per sample is high, which pushes experiments toward too few biological replicates and toward treating cells as independent samples when they are not. Tissue dissociation is the biggest practical problem: making a single-cell suspension kills fragile cell types, biases the population toward robust ones, and induces stress responses that appear in the data as real biology.
When to useUse single-cell RNA sequencing when the question is about composition or heterogeneity: which cell types are present, how their proportions change, what states exist within a population, and which cells express a target. It is the right tool for immune profiling, tumor microenvironment work, and developmental atlases. Use bulk RNA sequencing when the sample is homogeneous or when the question is about overall expression with many replicates, because bulk is far cheaper and better powered for differential expression across conditions. Use single-nucleus sequencing when the tissue does not dissociate well, which includes brain, muscle, heart and frozen archival material. Budget for biological replicates rather than for more cells, since cells within a sample are not independent observations and analyses that treat them as such produce confident nonsense.
Key numbersTypically thousands to tens of thousands of cells per run, with high-throughput protocols reaching hundreds of thousands · a few thousand genes detected per cell, out of roughly 20,000 expressed, so the matrix is mostly zeros · capture efficiency of transcripts per cell is a fraction rather than a majority · cost per sample in the hundreds to low thousands of dollars, plus sequencing · doublet rates rise with loading concentration and are typically a few percent · sequencing depth of tens of thousands of reads per cell is a common target.
Failure modesDissociation bias is the failure that most often invalidates conclusions, because it is invisible in the data: cell types that do not survive dissociation are simply absent, and a paper reporting their proportions is describing the dissociation protocol rather than the tissue. Stress and immediate-early gene induction during dissociation appear as genuine transcriptional states. Ambient RNA from lysed cells contaminates every droplet, creating apparent low-level expression of markers in cells that do not express them. Doublets, two cells in one droplet, look like hybrid cell types and have fooled published analyses. And the statistical error of treating thousands of cells from three mice as thousands of independent samples inflates significance dramatically; the unit of replication is the animal or the donor, not the cell.
ExamplesThe Human Cell Atlas and the many tissue atlases built on this technology; tumor microenvironment studies that revealed the composition of immune infiltrates and their relationship to checkpoint inhibitor response; developmental trajectory reconstructions across organisms; the 10x Genomics Chromium platform, which dominates commercially; and combinatorial indexing methods, which barcode cells through rounds of splitting and pooling rather than in droplets and reach much higher cell numbers at lower cost.
Economic profileA large and profitable consumables business built on a proprietary microfluidic cartridge, and one that has been the subject of extensive patent litigation, which is worth noting because it has shaped which alternatives are available. Cost per experiment remains high enough to constrain experimental design, and that constraint distorts the science toward underpowered studies. The trend is toward higher cell numbers at lower cost per cell, through combinatorial indexing and through cheaper sequencing, and the practical effect is that experiments are gradually shifting from a few samples with many cells toward many samples with fewer cells each, which is the statistically correct direction.
Single-cell ATAC sequencing measures which parts of the genome are physically open in each cell, using an enzyme that inserts sequencing adapters preferentially into accessible chromatin. Open regions are where transcription factors can bind, so the profile is a readout of regulatory state rather than of expression, and it identifies which regulatory elements are active in which cells. Multiome methods measure accessibility and gene expression in the same cell, which is the important advance: rather than profiling two populations separately and correlating them statistically, you observe both layers in one nucleus and can link a regulatory element directly to the gene it controls in that specific cell type. Because these methods work on nuclei rather than whole cells, they also solve a practical problem, since nuclei can be isolated from frozen and difficult tissue that will not survive dissociation into intact cells.
Strengths & weaknessesThe strengths are mechanism and sample compatibility. Expression tells you what a cell is doing; accessibility tells you something about why, and measuring both in one nucleus removes the inference step that separate assays require. Working from nuclei makes brain, heart, muscle, kidney and frozen archival tissue accessible, which are exactly the tissues where whole-cell dissociation fails. The weaknesses are sparsity and cost. Accessibility data are much sparser than expression data, because each cell has only two copies of each genomic region to sample rather than many transcripts, so a single cell's profile is nearly binary and almost all analysis depends on aggregating cells into groups first. That aggregation limits how much can be said about rare populations. Multiome kits cost more per sample than either assay alone, analysis is considerably harder, and nuclei preparations lose cytoplasmic RNA, which changes the expression profile relative to whole-cell data.
When to useUse single-cell ATAC when the question is regulatory: which enhancers are active, which transcription factors are driving a state, or how chromatin changes during differentiation or disease. Use multiome when you need to link regulation to expression within a cell type, which is the main reason to pay for it. Use nucleus-based methods whenever the tissue does not dissociate cleanly or the material is frozen, which is a practical rather than a scientific reason and is often the deciding one. Prefer whole-cell RNA sequencing when the question is only about expression and cell type composition, since it is cheaper, denser and easier to analyze. Plan for aggregation-based analysis from the start, and do not expect single-cell resolution of accessibility in rare populations.
Key numbersAccessibility data are near-binary per cell, since only two genomic copies are available to sample per region · typically a few thousand to tens of thousands of nuclei per run · fragments per cell in the low thousands, far sparser than transcript counts · multiome kits cost more per sample than either single assay · works on frozen and archival tissue where whole-cell dissociation fails · analysis requires aggregating cells into groups before most inference is possible.
Failure modesSparsity is the structural problem and it drives most analytical errors: treating a per-cell accessibility profile as a measurement rather than as a very noisy sample leads to overinterpretation, and clustering on accessibility alone is far less stable than clustering on expression. Nuclei preparations differ systematically from whole cells because cytoplasmic transcripts are lost, so multiome expression data are not directly comparable to whole-cell single-cell data, and combining datasets across those protocols introduces batch effects that look like biology. Ambient DNA contamination affects accessibility the way ambient RNA affects expression. And the biggest interpretive trap is assuming that an accessible region is an active one, since accessibility is necessary for regulation and not sufficient.
ExamplesBrain atlases built on single-nucleus multiome data, where whole-cell dissociation is impossible and the regulatory question is central; immune cell differentiation studies linking enhancer accessibility to lineage commitment; tumor studies identifying transcription factor programs driving resistance states; and the commercial multiome kits that made the paired assay routine rather than a specialist protocol.
Economic profileA premium extension of the single-cell consumables business, sold at a higher price per sample on the argument that paired measurement answers questions neither assay answers alone. Adoption has been slower than for single-cell RNA sequencing because the analysis burden is real and the biological payoff is less immediate: expression data yield cell types straight away, while accessibility data yield regulatory hypotheses that need follow-up. The most durable driver of adoption has been practical rather than scientific, since nucleus-based protocols unlock frozen biobank material that whole-cell methods cannot use, and that material is abundant and already collected.
Sequencing-based spatial transcriptomics keeps track of where in a tissue each transcript came from. A tissue section is placed on a slide covered in capture probes, each carrying a barcode that encodes its position on the slide. The tissue is permeabilized, messenger RNA diffuses down onto the surface and is captured, and the position barcode is incorporated during library construction. Sequencing then reports both what the transcript was and where it was. The critical specification is the size of each capture spot, because that determines whether you are measuring a single cell or a neighborhood. Early implementations used spots of roughly 55 micrometers, which contain several to tens of cells depending on tissue, so the measurement is a local average. Newer designs use much smaller features, down to sub-micrometer, which approach or reach single-cell resolution and change what the data can support.
Strengths & weaknessesThe strength is that it is unbiased: the whole transcriptome is captured without choosing genes in advance, so it is a discovery method in a way that imaging-based spatial methods are not. Tissue architecture is preserved, which matters enormously for tumors, brain and any tissue where position is functional. It works on standard sections and integrates with existing histology. The weaknesses are resolution and capture efficiency. At the older spot sizes, each measurement mixes several cells, so cell-type composition has to be inferred computationally by deconvolution against a single-cell reference, and those inferences are only as good as the reference. Capture efficiency is lower than for single-cell methods, so the data are sparser still. Cost per section is high and the area covered per slide is small, which limits how much tissue can be surveyed. Diffusion of transcripts before capture blurs the spatial signal at fine scales.
When to useUse sequencing-based spatial methods when you need whole-transcriptome coverage with spatial context and do not know in advance which genes matter. Tumor microenvironment mapping, tissue atlases and discovery work in structured tissue are the natural applications. Use imaging-based spatial methods instead when you know your gene panel, need genuine single-molecule and single-cell resolution, or need to work with fixed archival tissue where sequencing-based capture performs poorly. If spot sizes in your chosen implementation are larger than a cell, plan the deconvolution approach and the matching single-cell reference before running samples, because the analysis depends on it and retrofitting a reference is worse than designing for one.
Key numbersCapture spot sizes have fallen from roughly 55 micrometers, covering several to tens of cells, down to sub-micrometer features in newer implementations · whole transcriptome captured without gene selection, which is the defining advantage · capture efficiency lower than single-cell methods, so data are sparse · tissue area per slide is limited, typically a few square millimeters to a few square centimeters · cost per section in the hundreds to low thousands of dollars plus sequencing · transcript diffusion before capture limits effective resolution below the nominal feature size.
Failure modesThe most common error is treating spot-level data as cell-level data. At multi-cell spot sizes every measurement is a mixture, and reporting that a spot expresses two markers does not mean any cell expresses both, which has generated a great deal of overstated colocalization. Deconvolution can recover composition but inherits every bias of the reference dataset, including its dissociation bias. Permeabilization time is a critical and tissue-specific parameter: too short and capture fails, too long and transcripts diffuse and the spatial signal smears. Section quality, RNA integrity and tissue thickness all affect results substantially. And because tissue area per slide is small, sampling a tumor at one location and generalizing is a real risk that the method's expense encourages.
ExamplesThe Visium platform and its higher-resolution successors, which made spatial transcriptomics routine; Slide-seq and related bead-based methods that reached near-cellular resolution earlier; tumor studies mapping immune exclusion and spatial niches; developmental atlases where position is the biology; and the growing set of studies pairing spatial with single-cell data from the same samples, which is now the standard design.
Economic profileAn expensive consumables business in rapid technical flux, which makes purchasing decisions difficult: resolution has improved by more than an order of magnitude across a few product generations, so an instrument or workflow bought today may be superseded quickly. Competition between sequencing-based and imaging-based approaches is active and has not resolved, and they are converging on similar capabilities from different directions. For a laboratory, the practical implication is to favor service providers over capital purchases until the field settles, and for an investor, the platform question is whether either approach establishes a durable advantage before the other closes the gap.
Imaging-based spatial methods find individual transcript molecules in intact tissue by looking at them. A panel of probes is hybridized to chosen target genes, and the identity of each gene is encoded as a sequence of fluorescent signals read across many successive rounds of imaging, so that a manageable number of color channels can distinguish hundreds to thousands of genes through combinatorial coding. Each detected spot is one molecule at a known location with sub-cellular precision. Because the tissue is never dissociated and never sequenced, the measurement preserves true single-cell boundaries and even sub-cellular localization, which sequencing-based spatial methods cannot reach. The trade-off is fundamental to the approach: you must decide which genes to measure before you start, so this is a measurement method rather than a discovery method.
Strengths & weaknessesThe strengths are resolution and sensitivity. Individual molecules are localized to within a fraction of a cell, cell boundaries are real rather than inferred from capture spots, and detection efficiency per targeted transcript is considerably higher than capture-based methods achieve. It works well on formalin-fixed archival tissue, which unlocks enormous existing sample collections that sequencing-based capture handles poorly. Sub-cellular localization is a genuinely new kind of information. The weaknesses are panel selection and throughput. Measuring only chosen genes means the experiment can only answer questions anticipated in advance, and a panel that misses the relevant biology cannot be rescued by reanalysis. Imaging many rounds over large tissue areas is slow and generates very large image datasets. Cell segmentation, dividing the image into cells, is a persistent source of error, and molecules assigned to the wrong cell produce apparent co-expression that is an artifact.
When to useUse imaging-based spatial profiling when you know which genes matter, need genuine single-cell or sub-cellular resolution, or are working with fixed archival tissue. Validating a hypothesis generated from single-cell or sequencing-based spatial data is the classic application, and the two approaches pair well in that order: discover with sequencing, confirm and localize with imaging. It is the better choice for clinical and translational work on retrospective cohorts, where the samples are fixed and the gene set is defined. Use sequencing-based methods when the gene set is unknown. Treat cell segmentation quality as a primary quality metric rather than a preprocessing detail, since it determines whether the single-cell claims hold at all.
Key numbersPanels of hundreds to a few thousand genes, chosen in advance · single-molecule detection with sub-cellular localization · detection efficiency per targeted transcript higher than capture-based spatial methods · combinatorial encoding across many imaging rounds is what allows a few color channels to distinguish thousands of genes · works on formalin-fixed paraffin-embedded tissue, unlike most capture-based approaches · imaging time and data volume scale with area and with the number of rounds · cost per section high, comparable to sequencing-based spatial.
Failure modesCell segmentation is the dominant error source. Assigning molecules to cells requires drawing boundaries in an image, and in dense tissue those boundaries are ambiguous, so molecules are routinely assigned to neighbors. The consequence is apparent co-expression of markers that belong to adjacent cell types, which looks like a novel hybrid population and is an artifact. Optical crowding in highly expressed genes causes spots to overlap and be undercounted, which compresses dynamic range at the top end. Autofluorescence in fixed tissue creates false spots. Registration errors across imaging rounds corrupt the combinatorial code and produce misidentified genes. And panel design is a one-way decision: the genes not chosen are permanently absent from the dataset.
ExamplesThe MERFISH and seqFISH methods that established combinatorial single-molecule imaging; the Xenium and CosMx commercial platforms that made it accessible to non-specialist laboratories; brain atlases mapping cell types in situ at single-molecule resolution; tumor immunology studies on archival fixed cohorts where no fresh tissue exists; and the growing practice of pairing a discovery single-cell experiment with an imaging-based validation on the same cohort.
Economic profileA capital-intensive instrument market with high-margin consumable panels, competing directly with sequencing-based spatial from the opposite technical direction. The competitive picture is genuinely unsettled: imaging methods are adding genes toward whole-transcriptome coverage while sequencing methods are shrinking features toward single-cell resolution, and each is approaching the other's advantage. For laboratories the sensible position is to use service providers rather than buy, since the technology generation cycle is short. The archival tissue compatibility is the most durable commercial advantage, because the world's pathology archives are enormous, fixed, and unusable by most other methods.
A sequencer cannot read raw DNA. Library preparation converts a sample into fragments of the right size carrying the adapter sequences the instrument needs to bind, prime and index them. Two chemistries dominate. Ligation-based preparation shears DNA mechanically or enzymatically, repairs the ends, adds a single overhanging base, and ligates adapters on. Tagmentation uses a transposase loaded with adapter sequences, which cuts the DNA and inserts adapters in one step, collapsing several enzymatic reactions into one and cutting hands-on time substantially. Ligation gives more even and predictable fragmentation and works over a wider input range; tagmentation is faster, needs less input, and introduces a small sequence bias at insertion sites. Both then usually amplify, which is where most library artifacts originate, and PCR-free protocols exist for applications where that bias is unacceptable.
Strengths & weaknessesThe strengths of good library preparation are invisible when it works: even coverage, faithful representation of the original molecules, and low duplicate rates. Tagmentation in particular has made library construction fast enough and cheap enough to run thousands of samples, which changed what experiments are affordable. The weaknesses are that this step silently determines what the sequencer can see. Amplification introduces bias against fragments of extreme base composition, so coverage drops in regions that may be exactly the ones of interest, and it creates duplicates that inflate apparent depth. Fragment size distribution caps insert size and therefore how much information paired reads carry. Adapter dimers consume sequencing capacity and produce useless reads. Every one of these problems is easier to prevent at the bench than to correct computationally, and most sequencing failures attributed to the instrument originate here.
When to useChoose ligation-based preparation when input is plentiful, when coverage uniformity matters most, or when insert size needs tight control. Choose tagmentation when throughput, hands-on time, or low input dominate, which covers most high-volume applications. Choose a PCR-free protocol when input allows and coverage uniformity is critical, which is the case for high-quality whole genome sequencing and for anything where regions of extreme base composition matter clinically. Match the fragment size to the read length: paying for long inserts that the read length cannot span wastes information, and inserts shorter than the read length cause the two reads of a pair to overlap and read the same bases twice. Quantify libraries accurately before pooling, because unequal pooling wastes capacity on samples that did not need it.
Key numbersLigation-based preparation works over a wide input range from nanograms to micrograms; tagmentation works from picogram to nanogram inputs · tagmentation collapses fragmentation, end repair and adapter addition into one step, cutting hands-on time substantially · PCR amplification introduces coverage bias against fragments of very high or very low base composition · PCR-free protocols require higher input, typically hundreds of nanograms to a microgram · adapter dimers produce short useless reads and are removed by size selection · insert size should be matched to read length to avoid wasted or redundant sequencing.
Failure modesAlmost every downstream problem traces back here. Amplification bias produces systematically low coverage in regions of extreme base composition, and because the pipeline reports what it saw rather than what it missed, those regions look uninformative rather than untested. Duplicate reads inflate apparent depth and, if not removed, make a PCR error look like a supported variant. Adapter dimers can consume a large fraction of a run if size selection fails. Incorrect library quantification leads to over- or under-clustering, which reduces quality or wastes capacity. Index hopping between multiplexed samples assigns reads to the wrong sample, which matters most for low-frequency variant detection and is controlled with unique dual indexes. Cross-contamination between samples during preparation is common and is the reason negative controls belong in every batch.
ExamplesNextera and related tagmentation kits, which made high-throughput library construction practical; TruSeq and other ligation-based chemistries, still preferred where uniformity matters; PCR-free whole genome protocols used in reference-quality sequencing and in large population programs; automated library preparation on liquid handlers, which is how genome centers achieve consistency at scale; and the single-cell workflows that embed a specialized library preparation inside a larger protocol.
Economic profileLibrary preparation kits are a substantial and profitable consumables market, and one where the instrument vendors compete directly with independent suppliers. Cost per library has fallen and matters increasingly as sequencing itself gets cheaper: when sequencing was the dominant cost, library preparation was a rounding error, and at current prices it can be a comparable line item for small genomes and targeted assays. That has driven miniaturization onto acoustic liquid handlers and nanoliter-scale reactions, which cut reagent cost by an order of magnitude for laboratories with the automation to run them, and it is one of the few places where capital investment in automation pays back quickly.
A molecular barcode, usually called a unique molecular identifier, is a short random sequence attached to each original DNA fragment before any amplification. Every copy made from that fragment carries the same barcode, so after sequencing you can group reads by barcode and know they came from one starting molecule. That grouping does two things. It lets you count original molecules instead of reads, which removes amplification bias from quantification. And it lets you build a consensus across the copies of each original molecule, so that an error introduced during PCR or sequencing appears in only some copies and is voted out, while a real variant appears in all of them. Duplex sequencing extends this by tagging both strands of the original double helix separately and requiring agreement between them, which removes errors that occurred on one strand before amplification.
Strengths & weaknessesThe strengths are sensitivity and honest counting. Consensus calling across barcode families drops the effective error rate by orders of magnitude, which is what makes detecting a variant present in one molecule in ten thousand possible at all. For expression quantification, counting molecules rather than reads removes a major source of bias. Duplex methods reach error rates low enough to measure somatic mutation in normal tissue, which no other approach can do. The weaknesses are cost and depth. Building a consensus requires sequencing every original molecule several times over, so a large fraction of the sequencing output is spent confirming what you already read, and duplex sequencing is dramatically more expensive still because it needs both strands represented multiple times. Barcode collisions occur when two different molecules receive the same barcode, and barcodes themselves accumulate sequencing errors, both of which need handling.
When to useUse molecular barcodes whenever the variant of interest may be present at low frequency: circulating tumor DNA, minimal residual disease monitoring, somatic mosaicism, mutagenicity testing, and any assay claiming detection below a few percent allele fraction. They are effectively mandatory for those applications, and an assay claiming sub-percent sensitivity without them should be treated skeptically. Use them for RNA quantification wherever amplification is involved, which includes essentially all single-cell methods. Use duplex sequencing when the requirement is measuring genuinely rare mutations in normal tissue or reaching the lowest achievable error rates, and budget for it accordingly, because the sequencing cost per informative molecule is very high.
Key numbersA barcode is typically 8–16 random bases attached before amplification · consensus across a barcode family reduces effective error rates by orders of magnitude against raw sequencing · detection limits below 0.1% allele fraction are achievable with barcodes, against roughly 1% without and 15–20% for Sanger · duplex sequencing requires both strands of the original molecule and reaches the lowest error rates available · several-fold to tens-of-fold sequencing redundancy is needed per original molecule, which is the cost · barcode collisions rise as the number of input molecules approaches the barcode space.
Failure modesThe commonest mistake is claiming a detection limit that the input allows. Sensitivity is bounded by how many original molecules were in the sample, not by sequencing depth: if a plasma sample contains 3,000 genome equivalents, no amount of sequencing detects a variant present at one in ten thousand, because the molecule is not there. Assays are routinely oversold on this point. Barcode collisions assign two original molecules to one family and corrupt the consensus. Errors in the barcode itself split one family into two, which looks like extra molecules. Insufficient family size means consensus is built from too few reads to vote reliably, and family size distribution should be reported rather than an average. Contamination and index hopping both create apparent low-frequency variants that barcoding does not fix.
ExamplesCirculating tumor DNA assays for treatment selection and minimal residual disease detection, which depend entirely on this technique; duplex sequencing used to measure somatic mutation accumulation in normal human tissue and to detect low-frequency mutagenic effects in toxicology; single-cell RNA sequencing, where every commercial protocol includes molecular identifiers; and error-corrected sequencing used to measure the fidelity of gene editing and of DNA synthesis.
Economic profileBarcoding is cheap to add and expensive to use, because the cost lands in sequencing depth rather than in reagents. That structure has shaped an entire diagnostics segment: liquid biopsy assays are priced on the sequencing burden their sensitivity requires, and the competitive question between vendors is how few molecules and how little depth can achieve a clinically acceptable limit of detection. As sequencing prices fall, assays that were uneconomic become viable, which is why minimal residual disease monitoring has moved from research to commercial reality over the last few years, and it is one of the clearest examples of falling sequencing cost creating a new market rather than just reducing an old cost.
Long-read sequencers read whatever length of molecule they are given, so read length is set by extraction rather than by the instrument. Ordinary DNA extraction methods, designed decades ago for short-read applications, shear DNA into fragments of tens of kilobases through pipetting, vortexing and column binding, which is invisible and harmless for short-read work and destroys the entire value proposition of a long-read platform. High-molecular-weight extraction avoids shear at every step: cells are lysed gently, mixing is done by slow inversion rather than pipetting, wide-bore tips are used, and DNA is bound to magnetic disks or recovered by precipitation rather than pulled through a silica column. For the longest reads, cells are embedded in agarose plugs so the DNA is never handled in free solution at all, and size selection removes the short fragments that would otherwise dominate the sequencing pores.
Strengths & weaknessesThe strength is that it is the cheapest way to improve long-read data, usually by a large margin. Doubling read length by fixing extraction costs almost nothing compared with buying more sequencing, and it improves assembly contiguity and structural variant detection more than any downstream change. The weakness is that it is slow, manual and sample-dependent. Plug-based methods take days. Yields are lower than column methods because gentle handling recovers less. It cannot rescue material that is already degraded, so archival, fixed and long-stored samples are permanently limited regardless of technique. Blood and cultured cells work well; solid tissue is harder; formalin-fixed tissue is essentially hopeless for long reads. Automation is limited precisely because the thing being avoided, mechanical stress, is what automated liquid handling applies.
When to useUse high-molecular-weight extraction whenever the downstream method is long-read sequencing or optical mapping, which is to say whenever molecule length is the point. Treat it as part of the assay rather than as sample preparation, and validate it on your actual sample types before committing to a platform, because the read lengths in a vendor's specification were achieved on their choice of material. Use size selection to remove short fragments when maximum read length matters, accepting the yield loss, since short molecules occupy sequencing capacity that long ones would use better. Do not attempt long-read sequencing on fixed or degraded material and expect long reads: choose short-read methods for those samples rather than paying long-read prices for short-read data.
Key numbersOrdinary column-based extraction typically yields fragments of tens of kilobases, capping long-read length at that regardless of instrument · gentle methods routinely reach hundreds of kilobases, and agarose plug methods reach megabases · size selection removes short fragments at a cost in total yield · plug-based protocols take days against hours for column methods · fresh blood and cultured cells give the best results; solid tissue is harder; formalin-fixed material is unsuitable · read length improvements from better extraction typically exceed anything achievable by changing sequencing chemistry.
Failure modesThe dominant failure is not recognizing that extraction is the limiting step. A laboratory gets shorter reads than expected, concludes the sequencer underperforms, and optimizes the wrong thing, sometimes for months. Mechanical shear is cumulative and invisible: every pipetting step, every vortex, every column passage shortens the distribution a little, and the damage does not show up until sequencing. Contaminants that ordinary extraction tolerates, including residual protein, polysaccharide and phenol, inhibit long-read library preparation and block pores, so purity requirements are stricter than for short-read work. Freeze-thaw cycles shear DNA. And samples collected and stored under protocols designed for short-read sequencing are usually already too degraded, which is a planning failure that cannot be fixed at the bench.
ExamplesThe telomere-to-telomere human genome assembly, which depended on ultra-long reads that required specialized extraction to obtain; reference genome projects across many species, where extraction protocol development is often a substantial part of the work; clinical long-read programs that had to redesign sample collection before sequencing could work; and the commercial extraction kits and instruments developed specifically for this purpose, which is a small but growing segment.
Economic profileA small consumables and instrument market with outsized influence on the value of a much larger one, since it determines whether an expensive long-read platform delivers what it promised. The cost of getting extraction right is trivial next to the cost of the sequencing it enables, which makes it one of the highest-return investments in a genomics workflow and one that is routinely underfunded because it looks like sample preparation rather than like the assay. For a service provider, extraction expertise is a genuine differentiator that is hard to copy quickly, because it is process knowledge rather than a purchasable kit.
DNA methylation is a chemical mark on cytosine that a sequencer cannot see, because a methylated cytosine and an ordinary one are read as the same base. Methylation library preparation makes the mark visible by converting the DNA chemically before sequencing. Bisulfite treatment, the long-standing method, deaminates unmethylated cytosines to uracil, which reads as thymine, while methylated cytosines are protected and still read as cytosine; comparing the result to the reference reveals which cytosines were marked. The treatment is harsh: it fragments DNA severely and destroys a large fraction of the sample. Enzymatic conversion achieves the same read-out through a two-enzyme reaction that protects methylated cytosines and then deaminates the rest, without the acid and heat, which preserves far more intact DNA and gives more even coverage. Long-read platforms sidestep the whole problem by detecting methylation directly from the raw signal.
Strengths & weaknessesThe strength of conversion-based methods is that they give single-base resolution methylation across the genome on ordinary short-read instruments, with a large body of established protocols and reference data. Enzymatic conversion in particular gives good coverage uniformity from low input. The weaknesses are damage and complexity. Bisulfite conversion degrades most of the input DNA, so it needs more material and yields libraries with poor uniformity, and the converted genome has drastically reduced sequence complexity because most cytosines have become thymines, which makes alignment harder and less accurate. Incomplete conversion is indistinguishable from methylation and inflates apparent methylation levels, so conversion controls are essential. The converted data cannot be aligned with ordinary tools. Long-read direct detection avoids all of this but costs more per base and has its own accuracy considerations.
When to useUse enzymatic conversion in preference to bisulfite for new work: it is gentler, needs less input, and gives better coverage uniformity, and the reason bisulfite persists is inertia and comparability with existing datasets rather than performance. Use bisulfite when comparability with a large existing bisulfite dataset is required. Use long-read direct detection when you also want structural information or phasing, or when you want methylation without any conversion chemistry at all, which is increasingly the right answer for whole-genome methylation. Use targeted methylation panels rather than whole-genome methods when the regions of interest are known, which is the case for most clinical applications and cuts cost dramatically. Always include conversion controls, because incomplete conversion silently inflates every methylation estimate in the run.
Key numbersBisulfite treatment degrades a large fraction of input DNA and requires correspondingly more material · enzymatic conversion works from low input with better coverage uniformity · conversion reduces the effective sequence complexity of the genome, since most cytosines become thymines, which complicates alignment · conversion efficiency must be measured with spike-in controls, since incomplete conversion reads as methylation · long-read platforms detect methylation directly with no conversion · targeted methylation panels cost a fraction of whole-genome methylation sequencing.
Failure modesIncomplete conversion is the classic and most consequential failure, because it produces a systematic overestimate of methylation that looks entirely plausible and is invisible without spike-in controls. Over-conversion, where methylated cytosines are converted too, produces the opposite error. Bisulfite-induced fragmentation biases the library toward regions that survive, which is not random with respect to base composition, so coverage is uneven in a way that correlates with the biology being measured. Alignment of converted reads is harder and produces more mapping errors, particularly in repetitive regions. And PCR after conversion amplifies the strand bias introduced by the chemistry, so the two strands of a locus can give different answers, which is why strand-specific analysis matters.
ExamplesWhole-genome bisulfite sequencing, the long-standing reference method for methylome mapping; enzymatic methyl sequencing kits, which have taken substantial share for new work; methylation arrays, which remain the cheapest way to survey hundreds of thousands of defined sites and underpin most epigenetic clock work; targeted methylation panels used in multi-cancer early detection tests, which is the largest clinical application; and direct methylation calling on long-read platforms, now routine.
Economic profileA specialized preparation market that has been substantially reshaped by the clinical arrival of methylation-based cancer detection, which turned methylation from a research measurement into the basis of large commercial diagnostic programs. That has pulled investment into targeted panels and away from whole-genome methods, since a clinical test needs a few hundred informative regions rather than the whole methylome. The technical trend is toward avoiding conversion entirely, either through enzymatic methods that are gentler or through long-read platforms that read the mark directly, and bisulfite's long dominance is ending for reasons of data quality rather than cost.
Almost every synthetic DNA molecule in the world is made by phosphoramidite chemistry, a four-step cycle developed in the early 1980s and largely unchanged since. Synthesis runs on a solid support in a column, building the chain one base at a time from the three-prime end. Each cycle removes a protecting group from the growing chain, couples the next activated nucleotide, caps any chains that failed to react so they cannot participate later, and oxidizes the new linkage to its stable form. The cycle takes a few minutes and repeats once per base. Coupling efficiency per step is very high but not perfect, and because the yield of full-length product is the per-step efficiency raised to the power of the length, the arithmetic imposes a hard practical ceiling: at 99% per step, a 100-mer comes out at roughly 37% full length, and the rest is truncated material that has to be removed or tolerated.
Strengths & weaknessesThe strengths are maturity, quality and availability. The chemistry is completely characterized, instruments are widely available, oligos arrive in days from many suppliers, and modifications such as labels, phosphorothioate backbones and unnatural bases are routine, which matters enormously because therapeutic oligonucleotides depend on them. Per-oligo quality is high and purification methods are established. The weaknesses are length, cost at scale and waste. Length is capped around 200 bases by yield arithmetic, and quality degrades well before that. The process is column-based, so each distinct sequence needs its own column and its own reagent volumes, which makes making a thousand different oligos a thousand times the cost of making one. It uses large volumes of acetonitrile and other organic solvents, generating hazardous waste that is a real environmental and cost consideration at industrial scale, and it is the main driver of interest in enzymatic alternatives.
When to useUse column phosphoramidite synthesis for anything needing high quality per sequence at modest sequence count: PCR primers, probes, guide RNAs, therapeutic oligonucleotides, and any oligo requiring chemical modification. It remains the only practical route to modified backbones and unnatural chemistry, which is why the entire antisense and siRNA drug industry runs on it. Use array-based synthesis instead when you need many different sequences and can tolerate lower per-oligo quality and tiny quantities, which is the case for gene assembly and library construction. Do not plan on sequences much beyond 150 to 200 bases from a single synthesis; longer constructs are assembled from shorter pieces, which is a separate step with its own error considerations.
Key numbersCoupling efficiency per cycle is very high, typically above 99%, and full-length yield is that raised to the power of the length · at 99% per step a 100-mer is roughly 37% full length; at 99.5% it is about 61% · practical length ceiling around 150–200 bases, with quality falling well before that · cycle time a few minutes per base · cost per base falls sharply with synthesis scale, and cost per oligo is dominated by fixed setup rather than by length at small scale · large solvent volumes per synthesis, which is the main environmental cost.
Failure modesTruncated sequences are the characteristic impurity, and capping is what makes them tolerable: without it, a chain that failed to couple in one cycle would resume in the next and produce a deletion mutant that is nearly the right length and nearly impossible to separate, which is far worse than a clean truncation. Depurination under the acidic deprotection step damages the growing chain and worsens with length. Sequences with strong secondary structure or long homopolymer runs couple poorly and give lower yields. For therapeutic oligonucleotides, the impurity profile is the regulatory issue rather than the yield: every truncation and modification-related species has to be characterized and controlled, and that analytical burden grows with length.
ExamplesEvery PCR primer and probe in routine use; the approved antisense and siRNA drugs, all made by this chemistry at kilogram scale with modified backbones; guide RNAs for CRISPR experiments; the commercial oligo suppliers that deliver custom sequences in days; and the large-scale manufacturing plants built through the 2020s to meet demand from oligonucleotide therapeutics, which turned a laboratory reagent business into a pharmaceutical manufacturing one.
Economic profileTwo very different businesses run on the same chemistry. Research oligos are a commodity with thin margins, fast turnaround and many suppliers, competing on price and delivery time. Therapeutic oligonucleotide manufacturing is a pharmaceutical business with high margins, long qualification cycles and few qualified suppliers, and demand from approved siRNA and antisense drugs has driven substantial capacity investment. Solvent consumption and hazardous waste are a genuine cost and regulatory pressure at that scale, which is the strongest practical argument for enzymatic synthesis and the reason serious money is going into replacing a chemistry that has worked well for forty years.
Array synthesis runs the same phosphoramidite chemistry as a column, but on a chip with tens of thousands to millions of independently addressed features, each building a different sequence in parallel. Which feature receives which base at each step is controlled either by light, using photolabile protecting groups and a micromirror array, or by electrochemistry, generating deprotecting acid at chosen electrodes. Reagent volumes per sequence collapse from microliters to picoliters, and the cost per base falls by three to four orders of magnitude against column synthesis. The output is a pool: all the sequences mixed together in tiny total quantity, at unequal abundances, with a higher error rate than column material. That output form defines the applications, since a pool is exactly what you want for building genes from fragments or for making a library, and useless when you need one pure sequence.
Strengths & weaknessesThe strengths are cost per sequence and parallelism. Making 100,000 different oligos on a chip costs a small fraction of making one on a column, which is what made synthetic biology, large-scale library construction and DNA data storage economically conceivable. It is the enabling technology for gene synthesis, since a gene is assembled from array-derived fragments. The weaknesses are quantity, purity and uniformity. Each sequence is present in femtomole quantities, so anything needing material has to amplify first, which introduces bias and errors. Error rates per base are higher than column synthesis, so downstream assembly needs error correction or sequence verification. Abundance across the pool is uneven, often by orders of magnitude, so some designed sequences are effectively absent. Modified chemistry is limited compared with column synthesis, and separating one sequence from a pool requires amplification with specific primers, which is a design constraint on the whole pool.
When to useUse array pools whenever you need many different sequences and small amounts of each: gene assembly fragments, mutagenesis and variant libraries, CRISPR guide libraries, hybridization capture probes, and DNA data storage. It is the right tool any time the number of distinct sequences is the cost driver. Design the pool with orthogonal primer binding sites so subsets can be amplified separately, because retrieving one gene's worth of fragments from a pool of 50,000 is a design problem that has to be solved before synthesis, not after. Use column synthesis when you need one sequence at usable quantity, high purity, or with chemical modification. Expect to sequence-verify anything assembled from array material, since the error rate makes verification a required step rather than a precaution.
Key numbersTens of thousands to millions of distinct sequences per chip · cost per base three to four orders of magnitude below column synthesis · quantity per sequence in the femtomole range, requiring amplification before use · error rates per base higher than column synthesis, typically by several-fold · abundance across a pool varies by orders of magnitude between sequences · practical length per oligo of roughly 150–300 bases, similar to column but with more error · limited support for modified backbones and labels.
Failure modesUneven representation is the failure that most often derails a library experiment: sequences that came off the chip at low abundance are underrepresented or absent, and a screen then reports nothing about genes that were never really present, which looks like a negative result. Measuring the actual pool composition by sequencing before use is the fix and is frequently skipped. Amplification of the pool introduces further bias and can amplify a subset preferentially, particularly where sequences differ in base composition or secondary structure. Cross-hybridization between similar sequences during retrieval pulls out the wrong fragments. And the higher error rate means assembled constructs carry mutations at a rate that requires clonal verification, so treating array material as if it were column-grade is a reliable way to build a gene with a silent error in it.
ExamplesCommercial gene synthesis, essentially all of which assembles genes from array-derived fragments; CRISPR guide libraries for genome-wide screens; deep mutational scanning libraries covering every amino acid substitution in a protein; hybridization capture probe panels; large-scale DNA data storage demonstrations, which depend entirely on array cost per base; and the oligo pool products sold by several suppliers as a standard catalog item.
Economic profileThe cost structure that made synthetic biology possible. Reducing the price of a distinct sequence by four orders of magnitude changed which experiments are conceivable, and the entire field of high-throughput library-based biology rests on it. Commercially, array synthesis is mostly sold indirectly, embedded in gene synthesis and library products rather than as chips, and the suppliers with proprietary array platforms have a genuine cost advantage in those downstream markets. The main competitive pressure now comes from enzymatic synthesis, which promises comparable parallelism without organic solvents and with potentially longer and cleaner products, and which is the technology most likely to reset this market.
Enzymatic synthesis builds DNA with an enzyme in water instead of with phosphoramidite chemistry in organic solvent. The enzyme used is terminal deoxynucleotidyl transferase, which naturally adds nucleotides to the end of a DNA strand without needing a template. Left alone it would add many bases at once, so control comes from blocking each added nucleotide so that exactly one goes on per cycle, then removing the block before the next. Two designs exist: attaching the nucleotide to the enzyme itself, so one enzyme adds one base and then stalls until it is cleaved away, or using a chemically blocked nucleotide with a small reversible group. Either way the cycle is add, wash, deblock, wash, in aqueous buffer at mild pH. The attraction is that DNA is much more stable in these conditions than under the repeated acid exposure of phosphoramidite chemistry, which is what limits conventional synthesis length.
Strengths & weaknessesThe strengths are potential length, environmental profile and compatibility with biology. Removing acidic deprotection removes depurination, which is the damage that caps phosphoramidite length, so longer products should be achievable in principle. The process runs in water with no acetonitrile and no hazardous solvent waste, which matters increasingly at industrial scale and is a genuine regulatory and cost advantage. Aqueous mild conditions are compatible with enzymatic assembly and with benchtop instruments that do not need solvent handling. The weaknesses are that the promise is largely unrealized commercially. Per-cycle efficiency has been harder to push to phosphoramidite levels than expected, and since yield is efficiency to the power of length, a small deficit compounds badly. Modified backbones such as phosphorothioates, which the therapeutic oligonucleotide industry depends on, are not straightforward enzymatically. Products and error rates from shipping systems have not yet clearly beaten the incumbent.
When to useToday, use enzymatic synthesis where its practical advantages apply rather than where its theoretical ones do: benchtop instruments that let a laboratory make its own oligos overnight without solvent infrastructure, and applications where avoiding hazardous waste matters. Continue to use phosphoramidite chemistry for anything requiring modified backbones, high purity per sequence, or established regulatory precedent, which covers all therapeutic oligonucleotides. Watch per-cycle efficiency as the number that decides whether the technology displaces the incumbent, since it determines achievable length, and treat length claims without a stated efficiency and full-length yield as marketing. For a laboratory considering a benchtop synthesizer, the question is turnaround time against outsourcing, not cost per base, because at low volume the fixed costs dominate either way.
Key numbersRuns in aqueous buffer at mild pH with no organic solvent, against large acetonitrile volumes for phosphoramidite chemistry · removing acidic deprotection removes depurination, which is what caps conventional synthesis length · per-cycle efficiency must approach phosphoramidite's very high figures to compete, since full-length yield is efficiency to the power of length · modified backbones including phosphorothioates are difficult enzymatically · benchtop systems deliver oligos overnight without solvent infrastructure · commercial products exist but have not displaced phosphoramidite for any major application.
Failure modesThe compounding arithmetic is unforgiving and is where most optimism has foundered: an enzymatic cycle a fraction of a percent less efficient than a chemical one gives dramatically less full-length product at 200 bases, so "comparable efficiency" is not comparable at length. Incomplete deblocking leaves chains unable to extend, producing truncations, and the capping step that makes phosphoramidite truncations clean is not straightforward to replicate enzymatically, so deletion mutants are a bigger risk. Enzyme activity varies with sequence, particularly with secondary structure and homopolymers. Scaling from a demonstration to reproducible manufacturing has been the practical barrier, and several well-funded efforts have taken far longer than projected.
ExamplesDNA Script's benchtop synthesizers, which put overnight oligo production in a laboratory without solvent handling; Ansa Biotechnologies and Molecular Assemblies, pursuing longer products and different blocking strategies; Twist and other established suppliers investing in enzymatic routes alongside their array businesses; and the DNA data storage efforts, which need synthesis cost far below anything current chemistry offers and are a major driver of interest in the approach.
Economic profileA technology that has attracted substantial investment on a clear thesis, replacing a forty-year-old chemistry with a cleaner and potentially longer one, and has consistently taken longer to deliver than expected. The environmental argument is strengthening as oligonucleotide therapeutics scale and solvent waste becomes a real manufacturing cost. The market it would unlock, DNA data storage, needs cost reductions of several orders of magnitude that no current approach achieves. For an investor, the discriminating question is per-cycle efficiency and full-length yield at a stated length, not cost per base in isolation, because those two numbers determine whether the technology can address any of the applications that matter.
No synthesis chemistry makes a gene directly, because the length ceiling is a few hundred bases and a gene is one to several kilobases. Genes are therefore built by joining short synthetic fragments. Several methods do this and they differ mainly in whether the joins are designed or arbitrary. Polymerase cycling assembly anneals overlapping oligos and extends them into a full-length product in one reaction. Gibson assembly joins fragments with matching end sequences using three enzymes in a single isothermal step, and because the overlaps are designed into the fragments, the joins leave no scar and the method handles several fragments at once. Golden Gate assembly uses restriction enzymes that cut outside their recognition site, so the cut leaves a chosen four-base overhang and many fragments can be joined in a defined order in one reaction, which suits standardized part libraries. Yeast assembly uses homologous recombination in vivo and handles the largest constructs.
Strengths & weaknessesThe strengths are that these methods work reliably at the kilobase scale and are cheap in reagent terms. Gibson and Golden Gate are seamless, so no unwanted sequence is left at the junctions, which matters when the construct has to encode exactly what was designed. Golden Gate's defined overhangs make ordered multi-fragment assembly routine and are the basis of standardized part collections. Yeast recombination assembles constructs of tens to hundreds of kilobases that no in vitro method handles. The weaknesses are error accumulation and sequence-dependent failure. Every base came from synthesis with an error rate, and assembly does not correct errors, so a longer construct is more likely to carry one, which makes verification mandatory rather than optional. Repetitive sequence causes misassembly, since overlaps become ambiguous. Constructs toxic to the host will not clone. Success rates vary enough with sequence that gene synthesis providers quote turnaround ranges rather than fixed times.
When to useUse Gibson assembly for joining a few fragments seamlessly when the sequences are arbitrary and you control the ends, which covers most cloning. Use Golden Gate when building from a standardized part library or assembling many fragments in a defined order, which is the synthetic biology workhorse. Use polymerase cycling assembly to build a gene directly from an oligo pool, which is what commercial gene synthesis does at the first stage. Use yeast recombination for anything above about 20 kilobases. In every case, plan verification into the workflow, and design out repeats and problematic sequence before ordering, since a construct that will not assemble is usually a design problem rather than a protocol problem and is much cheaper to fix on the screen.
Key numbersSynthesis fragments are a few hundred bases; genes are one to several kilobases, which is why assembly exists · Gibson handles several fragments in one isothermal reaction with seamless joins · Golden Gate uses four-base designed overhangs, supporting ordered assembly of many fragments at once · yeast recombination assembles tens to hundreds of kilobases · error rate of the final construct is inherited from synthesis and accumulates with length · repetitive sequence is the main cause of misassembly · commercial gene synthesis turnaround is typically days to a few weeks depending on difficulty.
Failure modesRepeats are the reliable killer: any sequence that appears twice makes overlaps ambiguous, and the assembly produces deletions, inversions or nothing at all. This is why long repetitive constructs, including many natural regulatory regions and anything with tandem repeats, are quoted at higher prices or refused by synthesis providers. Secondary structure prevents annealing and extension. High or low base composition regions fail disproportionately. Constructs encoding something toxic to the cloning host are selected against, so the colonies that grow are the ones carrying mutations that broke the toxic element, which is a particularly deceptive failure because it produces plasmid that sequences cleanly at the vector and wrongly at the insert. Carrying synthesis errors through unverified is the other common outcome.
ExamplesCommercial gene synthesis, which uses polymerase cycling assembly from array-derived oligos followed by cloning and verification; the MoClo and Golden Braid standardized Golden Gate part systems used across plant and microbial synthetic biology; the synthetic yeast genome project, which assembled whole designer chromosomes by yeast recombination; and the routine molecular biology use of Gibson assembly, which largely displaced restriction cloning for arbitrary constructs.
Economic profileGene synthesis has become a competitive commodity service priced per base, with surcharges for repetitive or structured sequence. Prices have fallen far enough that ordering a designed gene is normally cheaper and faster than cloning one from a template, which has quietly changed how molecular biology is done. Suppliers compete on price, turnaround and how much difficult sequence they will accept, and that last point is where the genuine technical differentiation lies, since easy genes are easy for everyone.
Synthetic DNA arrives with errors, and verification is how you find out which molecules are correct. Two distinct jobs sit under this heading. Error correction reduces the error rate of a synthesized population before cloning, usually by melting and reannealing the DNA so that strands carrying different errors pair with each other, creating mismatches that a mismatch-binding enzyme then cuts or binds, allowing the damaged molecules to be removed. That improves the population but does not guarantee any individual molecule. Clonal verification does the rest: individual molecules are isolated, usually by transformation into bacteria so that one colony equals one molecule, and each candidate is sequenced until a perfect one is found. Verification used to mean Sanger sequencing with primer walks, and increasingly means sequencing whole plasmids in one pass on a nanopore or short-read platform, which is cheaper and catches rearrangements that primer walking misses.
Strengths & weaknessesThe strength is that these steps convert an unreliable synthetic product into a defined reagent, which is what makes gene synthesis usable at all. Enzymatic error correction can reduce error rates several-fold for very little cost, which materially raises the fraction of clones that are perfect and therefore reduces how many have to be screened. Whole-plasmid sequencing has made verification cheap enough to do routinely rather than selectively. The weaknesses are cost in time and the limits of what each method sees. Screening clones is a numbers game that scales badly with construct length, since the probability of a perfect clone falls with every base. Error correction reduces but does not eliminate errors and works poorly on some error types, particularly insertions and deletions in homopolymers. Sanger primer walking, still common, misses large rearrangements entirely because it only reads where the primers point.
When to useVerify everything that will be used more than once or that anything depends on, which in practice means every construct. Use whole-plasmid sequencing rather than Sanger primer walks: it costs about the same, reads the whole molecule including the vector backbone, and catches the rearrangements and insertions that walking cannot see. Use enzymatic error correction when assembling long constructs from synthetic fragments, because raising the fraction of perfect clones from a few percent to a few tens of percent changes the screening burden from painful to routine. For anything above a few kilobases, expect to screen multiple clones and plan the timeline accordingly. And verify after any step that could change the sequence, including passaging, since constructs with a fitness cost drift in culture.
Key numbersProbability of a perfect clone falls with construct length, since it is the per-base accuracy raised to the power of the length · enzymatic error correction typically reduces error rates several-fold in one round · whole-plasmid sequencing reads the entire construct in one pass, against Sanger's 700–900 bases per primer · verification cost per construct is now low enough to be routine · error correction works poorly on insertions and deletions in homopolymer runs · clones needing screening rise sharply above a few kilobases.
Failure modesThe most consequential failure is verifying only part of the construct. Sanger primer walks read where the primers point and report clean sequence there while a large deletion, an inversion, or an insertion sits somewhere unread, which is why whole-plasmid methods have displaced walking. Toxic constructs select for mutants during cloning, so the clones that grow well are frequently the broken ones, and a laboratory that screens only the healthiest colonies finds them enriched for the mutations that relieved the toxicity. Repeated sequences cause misassembly that short-read verification cannot resolve, because reads do not span the repeat. And constructs that were verified once and then propagated for months are not still verified, since selection acts continuously on anything with a fitness cost.
ExamplesCommercial gene synthesis providers, who run error correction and clonal verification as standard and quote a guaranteed sequence; nanopore-based whole-plasmid sequencing services that return a full construct sequence overnight for a few dollars, which have largely replaced Sanger primer walking; enzymatic mismatch cleavage kits used in laboratory-scale gene assembly; and the routine practice of resequencing working stocks of important constructs, which catches drift that would otherwise be discovered as an unreproducible experiment.
Economic profileVerification is a small cost that prevents large ones, and its economics improved dramatically when whole-plasmid sequencing arrived: reading an entire construct for a few dollars removed the reason to verify selectively. The service market for plasmid sequencing has grown quickly for exactly that reason and has taken share from Sanger providers. For gene synthesis companies, error correction and verification are where much of the real process know-how sits, since anyone can order oligos but delivering a guaranteed perfect multi-kilobase construct at a competitive price depends on how efficiently the errors are removed.
Benchtop synthesizers put DNA writing inside the laboratory rather than at a supplier. A researcher enters a sequence in the afternoon and has oligos or a short gene the next morning, without shipping, customs or a supplier queue. The instruments use either miniaturized phosphoramidite chemistry with cartridge-contained reagents or enzymatic synthesis, which is better suited to a benchtop because it avoids the solvent handling and hazardous waste that make chemical synthesis a facilities question. The reason this matters beyond convenience is biosecurity. Commercial synthesis providers screen orders against sequences of concern and screen customers, a practice that has been voluntary in most jurisdictions and is increasingly the subject of policy. An instrument that synthesizes DNA on a bench moves that control point from a supplier who screens to a device that may not, which is the central governance question the technology raises.
Strengths & weaknessesThe strengths are turnaround and control. Same-day or overnight availability changes the pace of design-build-test cycles substantially, and for laboratories in places where importing reagents is slow or unreliable it can be the difference between doing the work and not. Sequences never leave the institution, which matters for confidential constructs. The weaknesses are cost per base and quality. At the volumes a single laboratory uses, an instrument plus consumables rarely beats a commercial supplier on price, because suppliers run at enormous scale and compete hard. Product quality and length are generally below what a specialist supplier delivers. The biosecurity concern is real rather than hypothetical: distributed synthesis capability without screening removes the chokepoint that current governance relies on, and screening implemented in device firmware can in principle be circumvented in ways a supplier's process cannot.
When to useUse a benchtop synthesizer when turnaround time genuinely limits your work and the volume justifies the instrument, which is most often true in high-iteration protein engineering and synthetic biology groups running many design cycles. It is also worth considering where supply chains are slow or where sequence confidentiality is a requirement. Do not buy one to save money at typical laboratory volumes, because the arithmetic usually favors ordering. Whatever the setting, treat screening as part of the purchase decision: check what the instrument screens, whether that screening can be disabled, and what the institution's own obligations are, because the responsibility that a supplier used to carry moves to the operator.
Key numbersOvernight or same-day turnaround against days to weeks for ordering · cost per base at single-laboratory volumes is generally above commercial suppliers, which run at far greater scale · product length and purity typically below specialist supplier output · enzymatic instruments avoid organic solvent handling, which is what makes a benchtop deployment practical · commercial synthesis screening has been largely voluntary, coordinated through industry consortia · policy attention to synthesis screening has increased substantially in the 2020s.
Failure modesOn the technical side, benchtop output is usually less pure and shorter than supplier material, so constructs built from it need more verification, and a laboratory that treats instrument output as equivalent to ordered oligos will find the difference during assembly rather than before. Consumable cartridges tie the instrument to one supplier at that supplier's pricing, which is the same razor-and-blade exposure as any instrument purchase. On the governance side, the failure mode is institutional rather than technical: assuming screening happens because it used to. If an instrument does not screen, or screening is optional, the obligation sits with the operator and the institution, and discovering that after the fact is a much worse position than deciding it in advance.
ExamplesDNA Script's enzymatic benchtop synthesizers, the most established commercial instruments; Telesis Bio and similar systems that combine synthesis with assembly to produce constructs rather than oligos; the International Gene Synthesis Consortium's voluntary screening framework, which covers a large share of commercial synthesis; and the successive policy efforts in the US and elsewhere to make screening of both sequences and customers a condition of federal funding.
Economic profileA small instrument market with an outsized policy footprint. The commercial case is turnaround rather than cost, which limits the addressable market to groups whose iteration speed is genuinely constrained, and that is a smaller set than the marketing implies. The more consequential dynamic is regulatory: as synthesis screening moves from voluntary industry practice toward a funding or legal requirement, the compliance burden falls differently on centralized suppliers, who already screen at scale, and on distributed instruments, which have to implement it per device. How that resolves will shape whether benchtop synthesis becomes common laboratory equipment or stays a specialist purchase.
No methods match the current filters.
Try clearing a facet or broadening the search.
Terms that show up in the method explorer and are not obvious from outside the field. Numbers are typical values, not specifications.
| Term | What it means |
|---|---|
| Adapter | A short synthetic sequence attached to every fragment in a library so the instrument can bind, prime and index it. Adapters that ligate to each other instead of to sample DNA form dimers, which are short, sequence efficiently, and can consume a large share of a run if size selection fails. |
| Allele dropout | Failing to amplify one of the two copies of a region, usually because a variant sits under a PCR primer's binding site. The assay then reports the sample as homozygous when it is not. It is dangerous because it produces a confident wrong answer rather than a failure, and it happens preferentially in exactly the samples that carry variants. |
| Ambient RNA | Free-floating RNA from cells that broke during sample handling, which gets captured in every droplet of a single-cell experiment. It creates apparent low-level expression of markers in cells that do not express them, and it is a common source of implausible cell-type assignments. |
| Bisulfite conversion | Chemical treatment that turns unmethylated cytosines into something read as thymine while leaving methylated ones alone, making methylation visible to an ordinary sequencer. It is harsh, destroying much of the input DNA, and incomplete conversion is indistinguishable from methylation, which is why spike-in controls are essential. |
| Circular consensus | Sequencing the same circularized molecule many times round and collapsing the passes into one high-accuracy read. It works because the errors are random rather than systematic, so averaging removes them. It is what turned long reads from a structure-only tool into something accurate enough for clinical variant calling. |
| Coupling efficiency | The fraction of growing chains that successfully add a base in one synthesis cycle. Full-length yield is this number raised to the power of the length, which is why 99% per step leaves only about 37% full-length product at 100 bases. This single piece of arithmetic governs the whole synthesis field. |
| Coverage | How many times each position was read. Depth is what most people quote, and uniformity is usually what limits an assay: regions that consistently read poorly get reported as negative rather than as untested, so the pipeline sounds confident about the parts it could handle and silent about the rest. |
| Deconvolution | Working out what mixture of cell types produced a measurement that averaged several cells together, using a single-cell dataset as reference. It is how spatial data at multi-cell resolution is interpreted, and it inherits every bias of the reference, including which cells survived dissociation to be in it. |
| Depurination | Loss of a base from the DNA backbone under acidic conditions, which happens at every deprotection step of chemical synthesis and gets worse with length. It is the specific damage that caps phosphoramidite synthesis at a couple of hundred bases, and avoiding it is the main technical argument for enzymatic synthesis. |
| Dissociation | Breaking a tissue into single cells so it can be measured one cell at a time. Fragile cell types do not survive, so they are absent from the results with no indication they existed, and the stress of the process induces gene expression that appears in the data as real biology. It is the most common way a single-cell experiment describes the protocol rather than the tissue. |
| Doublet | Two cells captured in one droplet and barcoded as if they were one. The combined profile looks like a cell type expressing two lineages at once, which has been reported as a novel hybrid population more than once. Rates rise with how heavily the system is loaded. |
| Duplex sequencing | Tagging both strands of an original DNA molecule separately and requiring them to agree before calling a variant. It removes errors that happened on one strand before amplification and reaches the lowest error rates available, which is what makes measuring rare mutations in normal tissue possible. It is also dramatically expensive per informative molecule. |
| Gibson and Golden Gate assembly | Two ways of joining DNA fragments seamlessly in one reaction. Gibson uses matching end sequences designed into the fragments; Golden Gate uses restriction enzymes that cut outside their recognition site, leaving chosen four-base overhangs so many parts can be joined in a defined order. Both leave no unwanted sequence at the junctions. |
| Homopolymer | A run of the same base repeated. Any chemistry that infers length from signal intensity rather than counting discrete additions struggles to tell four identical bases from five, so insertion and deletion errors concentrate here. The errors are systematic, which means more coverage does not fix them. |
| Hybrid capture | Fishing chosen regions out of a whole-genome library with complementary biotinylated probes and magnetic beads. It handles very large and changing target sets better than PCR-based enrichment and preserves the original fragment ends, which is what molecular barcodes and duplicate detection depend on. |
| Imputation | Inferring genotypes at positions that were never measured, using the correlation structure of variants in a population reference panel. It extends an array of a million sites to tens of millions of usable variants. Accuracy depends on how well the reference represents the sample's ancestry, and it is substantially worse for under-represented populations. |
| Index hopping | Reads from one multiplexed sample being assigned to another, caused by free adapters swapping during cluster generation. It is usually a fraction of a percent, which is irrelevant for germline genotyping and very relevant when hunting for variants at similar frequencies. Unique dual indexes are the standard control. |
| Molecular barcode | A short random sequence attached to each original fragment before amplification, so every copy of that fragment can be recognized as coming from one starting molecule. It lets you count molecules instead of reads and build a consensus that votes out amplification and sequencing errors. It is what makes detection below a percent possible. |
| On-target rate | The share of sequencing that lands where the enrichment was aimed. The rest is wasted capacity, so this number determines how much data a sample actually needs. It falls with smaller target sets and rises with careful probe design. |
| Oligo | A short synthetic DNA sequence, typically under 200 bases. Made one sequence per column when quality and quantity matter, or tens of thousands at a time on an array when the number of distinct sequences is what matters. An array pool costs a tiny fraction per sequence and delivers a tiny fraction of the material. |
| PCR duplicates and PCR-free | Duplicates are multiple reads from the same original molecule created during amplification. They inflate apparent depth, and if not removed they can make a single early PCR error look like a well-supported variant. PCR-free protocols skip amplification entirely, giving much more even coverage at the cost of needing more input DNA. |
| Phasing | Working out which variants sit on the same copy of a chromosome. It matters clinically because two damaging variants in one gene mean something different depending on whether they are on the same copy or on opposite ones. Short reads cannot phase across any real distance, which is one of the clearest arguments for long reads. |
| Segmentation | Drawing cell boundaries in an image so that detected molecules can be assigned to cells. In dense tissue the boundaries are genuinely ambiguous, so molecules get assigned to neighbors, which produces apparent co-expression of markers from adjacent cell types. It is the main artifact in imaging-based spatial data. |
| Size selection | Removing fragments outside a chosen length range, usually with magnetic beads. It cleans adapter dimers out of a library and, for long-read work, removes short molecules that would otherwise occupy sequencing capacity that long ones would use better. It always costs yield. |
| Structural variant | A large change to the genome: a deletion, duplication, inversion or translocation, typically anything above about 50 bases. Short reads mostly infer these from indirect evidence and miss balanced events entirely, since nothing is gained or lost. Long reads and optical mapping see them directly. |
| Tagmentation | Using a transposase preloaded with adapters to cut DNA and insert adapters in a single step, collapsing several enzymatic reactions into one. It made library preparation fast and cheap enough to run thousands of samples, at the cost of a small sequence bias at the insertion sites. |
| Reversible terminator | A nucleotide carrying a chemical block that stops the chain after exactly one base is added, so each cycle adds one letter that can be imaged before the block is removed. It is what makes sequencing by synthesis count bases discretely rather than infer them from signal intensity, which is why homopolymers are not a problem for that chemistry. |
| Ultra-long and high-molecular-weight DNA | Very long intact DNA molecules, hundreds of kilobases and up. Long-read platforms read whatever length they are given, so read length is set by extraction rather than by the instrument. Ordinary extraction shears DNA invisibly, which is why laboratories new to long reads often blame the sequencer for their pipetting. |
Two questions settle most of this. Are you counting things or looking at structure? And do you know in advance what you are looking for? Counting favors short reads, structure needs long ones, and knowing your targets in advance makes everything cheaper. Most disappointing sequencing experiments come from getting one of those two answers wrong at the design stage, not from the instrument.
| Factor | Why it matters |
|---|---|
| Read length | Decides what is visible at all. A read shorter than a repeat cannot span it, so short reads simply do not see structural variants, repeat expansions, or which variants sit on the same chromosome. No amount of depth fixes this; it is a property of the molecule, not the coverage. |
| Error profile, not error rate | Substitution errors are easy to handle statistically; insertion and deletion errors in homopolymers are systematic and survive high coverage, because the same run is miscalled the same way every time. Ask which errors a platform makes, not just how many. |
| Input DNA quality | The most underestimated variable. Long-read platforms read whatever length they are given, so ordinary extraction that shears DNA turns a long-read instrument into an expensive short-read one. Fixed and archival material caps what is achievable regardless of instrument. |
| Amplification | Introduces bias against extreme base composition, creates duplicates, and makes early PCR errors look like real low-frequency variants. It is why molecular barcodes exist and why PCR-free protocols are worth the extra input where coverage uniformity matters. |
| Coverage uniformity | Usually more limiting than raw accuracy in a clinical assay. Regions that consistently under-cover are reported as negative rather than as untested, so a pipeline returns a confident answer about the regions it can handle and stays silent about the rest. |
| Targeted or untargeted | Panels and probes only find what was designed in. That makes them cheap and fast, and it makes a negative result meaningless outside the design. Untargeted methods cost more and can be reanalyzed as knowledge changes without touching the sample. |
| Detection limit against input | Sensitivity is bounded by how many original molecules were in the sample, not by sequencing depth. If a plasma tube holds 3,000 genome equivalents, nothing detects a variant at one in ten thousand. Assays are routinely oversold on this point. |
| Coupling efficiency, for synthesis | Full-length yield is per-step efficiency raised to the power of the length, which is why a chemistry a fraction of a percent worse per step is dramatically worse at 200 bases. This single arithmetic fact governs the entire synthesis field. |
| Repeats, for synthesis and assembly | Repetitive sequence makes assembly overlaps ambiguous and is the main reason a construct fails or is quoted at a premium. It is a design problem, and it is far cheaper to fix on the screen than at the bench. |
| Factor | Why it matters |
|---|---|
| Cost per answer, not per base | The number that matters is what it costs to reach a confident conclusion. A platform with a better error rate needs less coverage, so a higher price per gigabase can still be cheaper per call. Panels look cheap until a second test is needed because the answer was outside the design. |
| Instrument utilization | High-throughput sequencers only make economic sense at steady high volume, because a partially filled flow cell costs the same as a full one. A laboratory with variable demand is usually better served by a smaller instrument or a service provider. |
| Ecosystem and validation | The incumbent's real moat has never been chemistry. Pipelines, reference datasets, trained staff and clinical validation all assume one data type, and switching platforms carries revalidation work that is routinely underestimated when comparing list prices. |
| Supplier concentration | Sequencing has been close to a monopoly for most of its history, and the alternatives each carry non-technical risk, from patent litigation to legislation. Anyone building a business on sequencing input costs should treat supply as a strategic exposure. |
| Library preparation as a cost line | When sequencing was expensive, library preparation was a rounding error. At current prices it can be a comparable line item for small genomes and targeted assays, which is what makes miniaturization onto liquid handlers pay back quickly. |
| Revalidation on content change | Adding genes to a panel is a new assay requiring validation; adding genes to an exome or genome analysis is a reanalysis. That difference pushes clinical laboratories slowly toward untargeted sequencing as actionable gene lists keep growing. |
| Specifications go stale | Instrument specifications in this field move faster than anything else in these sheets, and several platforms have changed accuracy or throughput by large factors across a couple of product generations. Treat any number older than a year as a band, not a value. |
| Screening obligations, for synthesis | Commercial synthesis providers screen orders and customers, largely voluntarily. Benchtop instruments move that control point into the laboratory, and the obligation moves with it, which is a purchase-decision question rather than an afterthought. |
The pattern that repeats across this sheet is that what happens before the instrument sets a ceiling the instrument cannot raise. Read length on a long-read platform is set by extraction, not by chemistry, and laboratories routinely spend months optimizing the wrong end of the workflow. Detection limit in a liquid biopsy is set by how many genome equivalents were in the tube. What a single-cell experiment sees is set by which cells survived dissociation, and the cell types that died are simply absent from the results with no indication that they ever existed. Coverage uniformity is set by library chemistry and amplification. In every one of these cases the instrument reports confidently on what it received and says nothing about what it never saw, which is why a clean-looking dataset is not evidence that the sample preparation worked. If a result is surprising, check the preparation before the analysis.
Sequencing cost per base has fallen by something like six orders of magnitude since 2001, far outpacing semiconductor cost curves over the same period, and it has fallen fastest whenever a credible competitor appeared. Synthesis has not followed. Phosphoramidite chemistry is essentially the same process it was in the mid-1980s, the length ceiling around 200 bases has barely moved, and the cost reductions that did arrive came from parallelism on arrays rather than from better chemistry. The gap matters strategically: many of the ideas that motivate synthetic biology, including DNA data storage and large-scale genome writing, need synthesis costs several orders of magnitude below where they are, and no incremental improvement to the existing chemistry gets there. That is the case for enzymatic synthesis, and it is also why the field has consistently overestimated how quickly writing would catch up with reading.
Pick the method from the question, not the specification sheet. Counting and small variants mean short reads; structure, phasing and repeats mean long reads; a known target list means a panel and a fraction of the cost. Then check the two things that quietly cap the result: how many original molecules the sample contains, which bounds sensitivity absolutely, and what the preparation does to the material before the instrument sees it. For synthesis, remember that yield is efficiency to the power of length, that repeats are the usual reason a construct fails, and that anything assembled needs whole-construct verification rather than a primer walk.
The durable positions here have been in ecosystems rather than in chemistry. Incumbents held share for a decade after competitors matched their specifications, because pipelines, reference data and clinical validation are harder to replace than an instrument. That is worth remembering both when assessing a new platform's prospects and when estimating your own switching cost.
These seven are the instruments a laboratory actually chooses between. Read length is listed first because it decides which questions are answerable at all, and cost is banded rather than quoted because list prices move every year and real prices depend on contract volume. The three tables after this one cover how to target a subset of the genome, how to add spatial or single-cell resolution, and how to write DNA rather than read it.
| Platform | Read length | Error profile | Cost | Pick it when |
|---|---|---|---|---|
| Sequencing by synthesis | 100–300 b | Above 99.9%, mostly substitutions | Very low per base | The question is counting or small variants and volume matters. The default, and its real advantage is the ecosystem: two decades of pipelines, reference datasets and trained people all assume this data type. |
| Avidity | 100–300 b | Better than conventional sequencing by synthesis | Very low | Short reads are right and you are not locked into a validated pipeline. Better accuracy means fewer reads per confident call. Thinner tooling and reference data, so budget revalidation if migrating an assay. |
| High-density bead | Short | Higher indel rate in homopolymers | Very low, aimed at scale | Enormous volume of counting work: RNA quantification, single-cell, methylation surveys. Homopolymer errors are systematic, so extra coverage does not fix them, which rules it out for some clinical calling. |
| DNA nanoball | 100–200 b | Linear amplification, so errors do not compound; very low duplicates | Very low | Low-frequency variant detection, where not propagating early amplification errors is a genuine advantage. Technically strong; the deciding factor for US laboratories is usually supply and litigation risk rather than performance. |
| Nanopore | Tens of kb to megabases | Below short reads; systematic in homopolymers | Low, with very low capital at the small end | Structure, phasing, repeat expansions, or a same-day answer. Reads the native molecule, so methylation comes free. Read length is set by your extraction, not by the instrument, which is where most disappointment originates. |
| HiFi | 15–25 kb | Comparable to short reads after circular consensus | Moderate, several times short reads per base | You need length and accuracy in one dataset: reference assembly, clinical structural variants, comprehensive rare disease sequencing. Pushing insert length trades away the accuracy that justified the platform. |
| Optical mapping | Hundreds of kb to megabases | No sequence at all | Moderate | Large structural rearrangements, especially balanced translocations that microarrays cannot see. Replaces three cytogenetic assays with one. Needs ultra-high-molecular-weight DNA, so fixed material is out. |
Sequencing everything is often the wrong answer, and the three targeting approaches trade against each other on speed, uniformity and how easily the content can change. The crossover points move every year as sequencing gets cheaper.
| Approach | Scale | Turnaround | Input | Pick it when |
|---|---|---|---|---|
| Sanger | 1 target | Same day | Low | You have a handful of things to check: a plasmid, a clone, an edited site, a variant to confirm. Below roughly ten to twenty targets it beats preparing a library. Blind under 15–20% allele fraction. |
| Amplicon panel | Tens to a few thousand targets | Hours | Very low, works on small biopsies | Known gene set, high depth, fast turnaround, low input. The workhorse of clinical molecular pathology. Watch for allele dropout, where a variant under a primer site produces a confident wrong answer. |
| Hybrid capture | Tens of genes to the whole exome | A day or more longer | Higher | Large or changing target sets, or when coverage uniformity and copy number matter. Preserves native fragment ends, so molecular barcodes and duplicate detection work properly. |
| Genotyping array | Hundreds of thousands to millions of known sites | Days | Low | Common variants across very large cohorts, where cost per sample dominates. Sees only designed positions, so it is useless for rare disease, and imputation accuracy is substantially worse for under-represented ancestries. |
| Whole genome | Everything | Days | Moderate | The answer might be anywhere, or the content will change and you want reanalysis rather than a new assay. Increasingly cheaper than capturing a large fraction of the genome. |
Four ways to stop averaging over a tissue. The choice turns on whether you know your gene list in advance and whether the tissue will dissociate, and those two questions usually settle it before any specification comparison.
| Method | Resolution | Gene coverage | Sample type | Pick it when |
|---|---|---|---|---|
| Single-cell RNA-seq | Single cell, no position | Whole transcriptome, sparse | Fresh, must dissociate | Composition and heterogeneity are the question. Dissociation bias is the failure that invalidates conclusions, because cell types that do not survive are simply absent with no trace in the data. |
| Single-nucleus and multiome | Single nucleus | Transcriptome plus accessibility | Frozen and difficult tissue | The tissue will not dissociate, or you need to link regulation to expression in the same cell. Accessibility data are near-binary per cell, so almost all analysis requires aggregating first. |
| Sequencing-based spatial | Spot, from tens of micrometers down to sub-micrometer | Whole transcriptome | Fresh frozen sections | You need spatial context and do not know which genes matter. At larger spot sizes every measurement mixes several cells, so colocalization claims need deconvolution rather than inspection. |
| Imaging-based spatial | Single molecule, sub-cellular | Hundreds to a few thousand chosen genes | Works on fixed archival tissue | You know the gene list, need true single-cell boundaries, or the samples are fixed. Cell segmentation errors create apparent co-expression, which is the artifact most often mistaken for a novel population. |
Synthesis methods differ mainly in how many distinct sequences you get and how much of each. Nothing makes a gene directly, so anything above a few hundred bases is assembled from pieces and then has to be verified.
| Method | Sequences | Quantity each | Length | Pick it when |
|---|---|---|---|---|
| Column phosphoramidite | One per column | Usable to industrial | To roughly 150–200 b | You need one sequence at quality and quantity, or any chemical modification. The only practical route to phosphorothioate backbones, which is why every approved oligonucleotide drug uses it. |
| Array pools | Tens of thousands to millions | Femtomoles, must amplify | 150–300 b, higher error | Many different sequences and small amounts of each: gene assembly fragments, guide libraries, capture probes. Uneven representation across the pool is the failure that quietly ruins screens. |
| Enzymatic | One to many depending on format | Varies | Comparable to chemical so far | Solvent-free operation matters, or a benchtop instrument fits the workflow. Judge it on per-cycle efficiency and full-length yield, since yield is efficiency to the power of length. |
| Assembly | Joins fragments into genes | Clonal after transformation | Kilobases to hundreds of kb | Anything longer than synthesis allows. Gibson for a few arbitrary fragments, Golden Gate for ordered multi-part builds, yeast recombination above about 20 kb. Repeats are the usual reason it fails. |
j and k work from anywhere on the page. The arrow keys move between entries once one is selected, so they still scroll normally the rest of the time.