Skip to content

Barcode clashing

By the nature of linked read technologies, there will (almost always) be more DNA fragments than unique barcodes for them. As a result, it’s common for barcodes to reappear in sequences. This is referred to as clashing or convolution and often needs to be dealt with so as to not incorrectly associate fragments with each other.

There are approaches to overcome barcode clashing (deconvolve/deconvolute) in linked read data by determining which sequences are sharing barcodes by chance rather than because they originated from the same DNA molecule. The likelihood of it happening is usually low, but it’s not impossible. When the algorithm determines that barcodes are being shared by unrelated molecules, the barcode typically gets suffixed with a hyphenated number (e.g. BX:Z:ATACG becomes BX:Z:ATACG-1). As of this writing, there are only a few pieces of software dedicated to deconvoluting linked-read data and do so with varying degrees of success and computational resource requirements. Similarly, linked-read software is variable in its flexibility towards accepting the deconvolved-barcode format.

Rather than incorrectly assume that all sequences/alignments with the same barcode originated from the same orignal DNA molecule, linked-read aware programs include a threshold parameter to determine a “cutoff” distance between alignments with the same barcode. This parameter can be interpreted as “if a barcode appears more than X base pairs away from the same barcode (on the same contig), then we’ll consider them as originating from different molecules.” If this threshold is lower, then you are being more strict and indicating that alignments sharing barcodes must be closer together to be considered originating from the same DNA molecule. Conversely, a higher threshold indicates you are being more lax and indicating barcodes can be further away from each other and still be considered originating from the same DNA molecule. A threshold of 50kb-150kb is considered a decent balance, but you should choose larger/smaller values if you have evidence to support them.

Molecule origin is determined by the distance between alignments with the same barcode relative to the specified threshold

Alignment distance Inferred origin
less than threshold same molecule
greater than threshold different molecules