EXMOTIF: efficient structured motif extraction.
Yongqiang Zhang,Mohammed J. Zaki +1 more
TL;DR: An efficient algorithm, called EXMOTIF, that given some sequence(s), and a structured motif template, extracts all frequent structured motifs that have quorum q.
read more
Abstract: Extracting motifs from sequences is a mainstay of bioinformatics. We look at the problem of mining structured motifs, which allow variable length gaps between simple motif components. We propose an efficient algorithm, called EXMOTIF, that given some sequence(s), and a structured motif template, extracts all frequent structured motifs that have quorum q. Potential applications of our method include the extraction of single/composite regulatory binding sites in DNA sequences. EXMOTIF is efficient in terms of both time and space and is shown empirically to outperform RISO, a state-of-the-art algorithm. It is also successful in finding potential single/composite transcription factor binding sites. EXMOTIF is a useful and efficient tool in discovering structured motifs, especially in DNA sequences. The algorithm is available as open-source at: http://www.cs.rpi.edu/~zaki/software/exMotif/
.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
SMOTIF: efficient structured pattern and profile motif search
Yongqiang Zhang,Mohammed J. Zaki +1 more
TL;DR: An efficient algorithm, called SMOTIF, to solve the structured motif search problem, which can effectively search for LTR retrotransposons and is well suited to searching for motifs with long range gaps.
Advances in enzyme biotechnology
Pratyoosh Shukla,Brett I. Pletschke +1 more
- 01 Jan 2013
TL;DR: Industrial enzyme applications in biorefineries for starchy materials and the role of enzymes and proteins in plant - microbe interaction are studied.
49
Establishing relationships among patterns in stock market data
Dietmar H. Dorr,Anne Denton +1 more
- 01 Mar 2009
TL;DR: This work introduces an algorithm for capturing the relationships among similar, contiguous subsequences, identifies patterns based on the similarity among sequences, captures the sequence-subsequence relationships among patterns in the form of a directed acyclic graph (DAG), and determines pattern conglomerates that allow the application of additional meta-analyses and mining algorithms.
37
MoTeX-II: structured MoTif eXtraction from large-scale datasets
TL;DR: MoTeX-II, a word-based high-performance computing tool for structured MoTif eXtraction from large-scale datasets, uses state-of-the-art algorithms for solving the fixed-length approximate string matching problem and shows that it matches or outperforms these tools in terms of runtime efficiency.
31
Mining Loosely Structured Motifs from Biological Data
TL;DR: This paper focuses on the mining of loosely structured motifs, i.e., of more general kinds of motif where several "exceptions" may be tolerated in pattern repetitions, and an algorithm exploiting data structures conceived to efficiently handle pattern variabilities is presented and analyzed.
22
References
Tandem repeats finder: a program to analyze DNA sequences
TL;DR: A new algorithm for finding tandem repeats which works without the need to specify either the pattern or pattern size is presented and its ability to detect tandem repeats that have undergone extensive mutational change is demonstrated.
SPADE: An Efficient Algorithm for Mining Frequent Sequences
TL;DR: SPADE is a new algorithm for fast discovery of Sequential Patterns that utilizes combinatorial properties to decompose the original problem into smaller sub-problems, that can be independently solved in main-memory using efficient lattice search techniques, and using simple join operations.
Finding composite regulatory patterns in DNA sequences
Eleazar Eskin,Pavel A. Pevzner +1 more
- 01 Jul 2002
TL;DR: This paper presents a MITRA (MIsmatch TRee Algorithm) approach for discovering composite signals and demonstrates that MITRA performs well for both monad and composite patterns by presenting experiments over biological and synthetic data.
YMF: a program for discovery of novel transcription factor binding sites by statistical overrepresentation
Saurabh Sinha,Martin Tompa +1 more
TL;DR: The program YMF identifies good candidates for DNA binding sites by searching for statistically overrepresented motifs and enumerates all motifs in the search space and is guaranteed to produce those motifs with greatest z-scores.
Mining periodic patterns with gap requirement from sequences
TL;DR: The complexity of the mining problem is shown and why traditional mining algorithms are computationally infeasible is discussed, and practical algorithms for solving the problem are proposed and proposed.
146