Software
Linked reads came on the scene in 2016 with the release of 10X Genomics Chromium platform and it was met with considerable enthusiasm. 10X themselves released the LongRanger software platform that became the go-to for processing their linked reads. During this early and eager adoption of linked-read technology, independent research groups also began writing their own tools to do cool things with linked-read data. Then, due to litigation, the 10X Chromium linked-read chemistry was discontinued. At the time, there were no published alternatives (TELL-seq and stLFR were published in 2020, haplotagging in 2021), so the initial enthusiasm regardling linked reads and their software infrastructure quickly waned as researchers pivoted towards existing technologies.
Although 3 years (2016-2019) doesn’t seem like much, that brief period resulted in quite a few linked-read specific bioinformatics tools, including new assemblers, scaffolders, aligners, phasers, and SV callers. However, many of those tools are specific to 10X-style data (or another technology-specific format), or were deprecated/abandoned following the discontinuation of 10X Genomics linked-reads. There is an Awesome-Linked-Reads repository maintained by Pontus Höjer that lists known linked-read software. Scanning it, you will notice there are many options and that most of them haven’t been maintained in 2 or more years.
What this means
Section titled “What this means”There are now several methods to generate linked-read data from sample tissue or data simulation, so linked-read data is accessible again. The community would benefit significantly from renewed bioinformatic and algorithmic interest in linked-reads at all levels: assembly, scaffolding, phasing, imputation, alignment, and especially variant calling.
Our role
Section titled “Our role”By creating a fully open-source BLink-seq chemistry, we’ve made linked reads more accessible than ever. We also created the BLink-seq organization as stewards and promotors of linked-read chemistry, regardless of platform. We’ve provided a knowledge base on this website of all things linked-read. We created the LASTQ specification to unify linked-read data into a single platform-agnostic and future-proof format. We’ve been developing Harpy, the go-to software suite for linked-read data for the last several years. We created the Djinn linked-read data format converter. We wrote Mimick, the linked-read sequence simulator. We are reviving the Lariat barcode-aware aligner as Arachne. Despite creating our own linked-read chemistry, each of these tools remains open and compatible with the current popular chemistries1. Our role is to empower the research community who wants to use linked-read sequencing.
Footnotes
Section titled “Footnotes”-
The current most popular linked-read chemistries are stLFR, TELL-seq, and haplotagging. There are others like LinkPrep, Illumina SLR, and LoopSeq, but there are few publications featuring their adoption. Regardless, our hope is a unified data format that makes the chemistry just an implementation detail. ↩