A greedy algorithm for aligning DNA sequences.
TL;DR: A new greedy alignment algorithm is introduced with particularly good performance and it is shown that it computes the same alignment as does a certain dynamic programming algorithm, while executing over 10 times faster on appropriate data.
read more
Abstract: For aligning DNA sequences that differ only by sequencing errors, or by equivalent errors from other sources, a greedy algorithm can be much faster than traditional dynamic programming approaches and yet produce an alignment that is guaranteed to be theoretically optimal. We introduce a new greedy alignment algorithm with particularly good performance and show that it computes the same alignment as does a certain dynamic programming algorithm, while executing over 10 times faster on appropriate data. An implementation of this algorithm is currently used in a program that assembles the UniGene database at the National Center for Biotechnology Information.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Figures

FIG. 3. Three cases for nding a point of distance d on diagonal k. The R -values give x -coordinates of the last positions on each diagonal with D -value d ¡ 1. We can move right from a d ¡ 1, or along a diagonal from a d ¡ 1, or up from a d ¡ 1, in order to nd a point that can be reached with d differences. The furthest of these points along diagonal k must have D -value d . The rst two moves raise the x -coordinate by 1. See line 10 of Figure 4. 
FIG. 2. A dynamic-programming X-drop algorithm. ![FIG. 1. Pruning of small S-values. The white boxes on antidiagonal k ¡ 1 designate points where S(i, j ) . ¡ 1. S(i, j ) on antidiagonal k is computed for i 2 [L , U 1 1], after which entries whose scores are too small are reset to ¡ 1, and any such positions at the ends of the white boxes are pruned away giving L 0 and U 0.](/figures/fig-1-pruning-of-small-s-values-the-white-boxes-on-1bdgzrug.png)
FIG. 1. Pruning of small S-values. The white boxes on antidiagonal k ¡ 1 designate points where S(i, j ) . ¡ 1. S(i, j ) on antidiagonal k is computed for i 2 [L , U 1 1], after which entries whose scores are too small are reset to ¡ 1, and any such positions at the ends of the white boxes are pruned away giving L 0 and U 0. 
FIG. 5. An algorithm to determine an optimal edit script.
Citations
Fast and accurate short read alignment with Burrows–Wheeler transform
18 May 2009
TL;DR: Burrows-Wheeler Alignment tool (BWA), a new read alignment package that is based on backward search with Burrows–Wheeler Transform (BWT), to align short sequencing reads against a large reference sequence such as the human genome, allowing mismatches and gaps is implemented.
8.5K
A core gut microbiome in obese and lean twins
Peter J. Turnbaugh,Micah Hamady,Tanya Yatsunenko,Brandi L. Cantarel,Alexis E. Duncan,Ruth E. Ley,Mitchell L. Sogin,William J. Jones,Bruce A. Roe,Jason P. Affourtit,Michael Egholm,Bernard Henrissat,Andrew C. Heath,Rob Knight,Jeffrey I. Gordon +14 more
TL;DR: The faecal microbial communities of adult female monozygotic and dizygotic twin pairs concordant for leanness or obesity, and their mothers are characterized to address how host genotype, environmental exposure and host adiposity influence the gut microbiome.
Primer-BLAST: A tool to design target-specific primers for polymerase chain reaction
TL;DR: A new software tool called Primer-BLAST is presented to alleviate the difficulty in designing target-specific primers and combines BLAST with a global alignment algorithm to ensure a full primer-target alignment and is sensitive enough to detect targets that have a significant number of mismatches to primers.
5.8K
A Draft Sequence of the Neandertal Genome
Richard E. Green,Johannes Krause,Adrian W. Briggs,Tomislav Maricic,Udo Stenzel,Martin Kircher,Nick Patterson,Heng Li,Weiwei Zhai,Markus Hsi-Yang Fritz,Nancy F. Hansen,Eric Durand,Anna-Sapfo Malaspinas,Jeffrey D. Jensen,Tomas Marques-Bonet,Tomas Marques-Bonet,Can Alkan,Kay Prüfer,Matthias Meyer,Hernán A. Burbano,Jeffrey M. Good,Jeffrey M. Good,Rigo Schultz,Ayinuer Aximu-Petri,Anne Butthof,Barbara Höber,Barbara Höffner,Madien Siegemund,Antje Weihmann,Chad Nusbaum,Eric S. Lander,Carsten Russ,Nathaniel Novod,Jason P. Affourtit,Michael Egholm,Christine Verna,Pavao Rudan,Dejana Brajković,Željko Kućan,Ivan Gušić,Vladimir B. Doronichev,Liubov V. Golovanova,Carles Lalueza-Fox,Marco de la Rasilla,Javier Fortea,Antonio Rosas,Ralf Schmitz,Philip L. F. Johnson,Evan E. Eichler,Daniel Falush,Ewan Birney,James C. Mullikin,Montgomery Slatkin,Rasmus Nielsen,Janet Kelso,Michael Lachmann,David Reich,David Reich,Svante Pääbo +58 more
TL;DR: The genomic data suggest that Neandertals mixed with modern human ancestors some 120,000 years ago, leaving traces of Ne andertal DNA in contemporary humans, suggesting that gene flow from Neand Bertals into the ancestors of non-Africans occurred before the divergence of Eurasian groups from each other.
ABySS: A parallel assembler for short read sequence data
Jared T. Simpson,Kim Wong,Shaun D. Jackman,Jacqueline E. Schein,Steven J.M. Jones,Inanc Birol +5 more
TL;DR: ABySS (Assembly By Short Sequences), a parallelized sequence assembler, was developed and assembled 3.5 billion paired-end reads from the genome of an African male publicly released by Illumina, Inc, representing 68% of the reference human genome.
References
Basic Local Alignment Search Tool
TL;DR: A new approach to rapid sequence comparison, basic local alignment search tool (BLAST), directly approximates alignments that optimize a measure of local similarity, the maximal segment pair (MSP) score.
98.8K
Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
Stephen F. Altschul,Thomas L. Madden,Alejandro A. Schäffer,Jinghui Zhang,Zheng Zhang,Webb Miller,David J. Lipman +6 more
TL;DR: A new criterion for triggering the extension of word hits, combined with a new heuristic for generating gapped alignments, yields a gapped BLAST program that runs at approximately three times the speed of the original.
Database resources of the National Center for Biotechnology Information
David L. Wheeler,Deanna M. Church,Ron Edgar,Scott Federhen,Wolfgang Helmberg,Thomas L. Madden,Joan Pontius,Gregory D. Schuler,Lynn M. Schriml,Edwin Sequeira,Tugba O. Suzek,Tatiana Tatusova,Lukas Wagner +12 more
TL;DR: In addition to maintaining the GenBank(R) nucleic acid sequence database, the National Center for Biotechnology Information (NCBI) provides data analysis and retrieval resources for the data in GenBank and other biological data made available through NCBI’s website.
Deciphering the biology of Mycobacterium tuberculosis from the complete genome sequence
Stewart T. Cole,Roland Brosch,Julian Parkhill,Thierry Garnier,Carol Churcher,David Harris,Stephen V. Gordon,Karin Eiglmeier,S. Gas,Clifton E. Barry,Fredj Tekaia,K. Badcock,D. Basham,D. Brown,Tracey Chillingworth,R. Connor,Robert L. Davies,K. Devlin,Theresa Feltwell,S. Gentles,N. Hamlin,S. Holroyd,T. Hornsby,Kay Jagels,Anders Krogh,J. McLean,Sharon Moule,Lee Murphy,K. Oliver,J. Osborne,Michael A. Quail,Marie-Adèle Rajandream,Jane Rogers,S. Rutter,K. Seeger,Jason Skelton,Rob Squares,S. Squares,John Sulston,K. Taylor,Sally Whitehead,Bart Barrell +41 more
TL;DR: The complete genome sequence of the best-characterized strain of Mycobacterium tuberculosis, H37Rv, has been determined and analysed in order to improve the understanding of the biology of this slow-growing pathogen and to help the conception of new prophylactic and therapeutic interventions.
Base-calling of automated sequencer traces using Phred. I. accuracy assessment
TL;DR: In this article, a base-calling program for automated sequencer traces, phred, with improved accuracy was proposed. But it was not shown to achieve a lower error rate than the ABI software, averaging 40%-50% fewer errors in the data sets examined independent of position in read, machine running conditions, or sequencing chemistry.