A greedy algorithm for aligning DNA sequences.
TL;DR: A new greedy alignment algorithm is introduced with particularly good performance and it is shown that it computes the same alignment as does a certain dynamic programming algorithm, while executing over 10 times faster on appropriate data.
read more
Abstract: For aligning DNA sequences that differ only by sequencing errors, or by equivalent errors from other sources, a greedy algorithm can be much faster than traditional dynamic programming approaches and yet produce an alignment that is guaranteed to be theoretically optimal. We introduce a new greedy alignment algorithm with particularly good performance and show that it computes the same alignment as does a certain dynamic programming algorithm, while executing over 10 times faster on appropriate data. An implementation of this algorithm is currently used in a program that assembles the UniGene database at the National Center for Biotechnology Information.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Figures

FIG. 3. Three cases for nding a point of distance d on diagonal k. The R -values give x -coordinates of the last positions on each diagonal with D -value d ¡ 1. We can move right from a d ¡ 1, or along a diagonal from a d ¡ 1, or up from a d ¡ 1, in order to nd a point that can be reached with d differences. The furthest of these points along diagonal k must have D -value d . The rst two moves raise the x -coordinate by 1. See line 10 of Figure 4. 
FIG. 2. A dynamic-programming X-drop algorithm. ![FIG. 1. Pruning of small S-values. The white boxes on antidiagonal k ¡ 1 designate points where S(i, j ) . ¡ 1. S(i, j ) on antidiagonal k is computed for i 2 [L , U 1 1], after which entries whose scores are too small are reset to ¡ 1, and any such positions at the ends of the white boxes are pruned away giving L 0 and U 0.](/figures/fig-1-pruning-of-small-s-values-the-white-boxes-on-1bdgzrug.png)
FIG. 1. Pruning of small S-values. The white boxes on antidiagonal k ¡ 1 designate points where S(i, j ) . ¡ 1. S(i, j ) on antidiagonal k is computed for i 2 [L , U 1 1], after which entries whose scores are too small are reset to ¡ 1, and any such positions at the ends of the white boxes are pruned away giving L 0 and U 0. 
FIG. 5. An algorithm to determine an optimal edit script.
Citations
Fast and accurate short read alignment with Burrows–Wheeler transform
Heng Li,Richard Durbin +1 more
TL;DR: Burrows-Wheeler Alignment tool (BWA) is implemented, a new read alignment package that is based on backward search with Burrows–Wheeler Transform (BWT), to efficiently align short sequencing reads against a large reference sequence such as the human genome, allowing mismatches and gaps.
BLAST+: architecture and applications.
Christiam Camacho,George Coulouris,Vahram Avagyan,Ning Ma,Jason S. Papadopoulos,Kevin Bealer,Thomas L. Madden +6 more
TL;DR: The new BLAST command-line applications, compared to the current BLAST tools, demonstrate substantial speed improvements for long queries as well as chromosome length database sequences.
BEDTools: a flexible suite of utilities for comparing genomic features
Aaron R. Quinlan,Ira M. Hall +1 more
- 28 Jan 2010
TL;DR: A new software suite for the comparison, manipulation and annotation of genomic features in Browser Extensible Data (BED) and General Feature Format (GFF) format, BEDTools, which allows the user to compare large datasets with both public and custom genome annotation tracks.
12.3K
Database resources of the National Center for Biotechnology Information
David L. Wheeler,Deanna M. Church,Ron Edgar,Scott Federhen,Wolfgang Helmberg,Thomas L. Madden,Joan Pontius,Gregory D. Schuler,Lynn M. Schriml,Edwin Sequeira,Tugba O. Suzek,Tatiana Tatusova,Lukas Wagner +12 more
TL;DR: In addition to maintaining the GenBank(R) nucleic acid sequence database, the National Center for Biotechnology Information (NCBI) provides data analysis and retrieval resources for the data in GenBank and other biological data made available through NCBI’s website.
BLAT—The BLAST-Like Alignment Tool
TL;DR: How BLAT was optimized is described, which is more accurate and 500 times faster than popular existing tools for mRNA/DNA alignments and 50 times faster for protein alignments at sensitivity settings typically used when comparing vertebrate sequences.
References
AvGI, an index of genes transcribed in the salivary glands of the ixodid tick Amblyomma variegatum
Vishvanath Nene,Dan Lee,John Quackenbush,Robert A. Skilton,Stephen Mwaura,Malcolm J. Gardner,Richard P. Bishop +6 more
TL;DR: Random clones from a cDNA library made from mRNA purified from dissected salivary glands of feeding female Amblyomma variegatum ticks were subjected to single pass sequence analysis to construct a gene index called AvGI, which represents an electronic knowledge base, which can be used to launch investigations of the biology of the salivaries of this tick species.
97
Alignments without low-scoring regions.
TL;DR: The results indicate that computing an optimal alignment under this constraint is very expensive, however, less rigorous conditions on the alignment can be guaranteed by quite efficient algorithms.
Methylation of messenger RNA in Escherichia coli
TL;DR: It is concluded that neither T4 mRNA nor E. coli mRNA is methylated; the level of methylation, if greater than zero, must be less than one base in 3500.
46
PipTools: a computational toolkit to annotate and analyze pairwise comparisons of genomic sequences.
Laura Elnitski,Cathy Riemer,Hanna Petrykowska,Liliana Florea,Scott Schwartz,Webb Miller,Ross C. Hardison +6 more
TL;DR: The utility of the toolkit is illustrated using annotation of a pairwise comparison of the mouse MHC class II and class III regions with orthologous human sequences to identify conserved, noncoding sequences that are DNase I hypersensitive sites in chromatin of mouse cells.
45
A tool for aligning very similar DNA sequences.
TL;DR: The alignment scoring scheme is designed to model sequencing errors, rather than evolutionary processes, and can align a 100 kb sequence to a 1 megabase sequence in a few seconds on a workstation, provided that there are very few differences between the shorter sequence and some region in the longer sequence.
34