Boyer–Moore string search algorithm

Topic Tools

Papers published on a yearly basis

Papers

Journal Article•10.1137/0206024•

Fast Pattern Matching in Strings

[...]

Donald E. Knuth, James Morris, Vaughan R. Pratt

01 Jun 1977-SIAM Journal on Computing

TL;DR: An algorithm is presented which finds all occurrences of one given string within another, in running time proportional to the sum of the lengths of the strings, showing that the set of concatenations of even palindromes, i.e., the language $\{\alpha \alpha ^R\}^*$, can be recognized in linear time.

...read moreread less

Abstract: An algorithm is presented which finds all occurrences of one given string within another, in running time proportional to the sum of the lengths of the strings. The constant of proportionality is low enough to make this algorithm of practical use, and the procedure can also be extended to deal with some more general pattern-matching problems. A theoretical application of the algorithm shows that the set of concatenations of even palindromes, i.e., the language $\{\alpha \alpha ^R\}^*$, can be recognized in linear time. Other algorithms which run even faster on the average are also considered.

...read moreread less

3,461 citations

Journal Article•10.1145/360825.360855•

Efficient string matching: an aid to bibliographic search

[...]

Alfred V. Aho¹, Margaret J. Corasick¹•Institutions (1)

Bell Labs¹

01 Jun 1975-Communications of The ACM

TL;DR: A simple, efficient algorithm to locate all occurrences of any of a finite number of keywords in a string of text that has been used to improve the speed of a library bibliographic search program by a factor of 5 to 10.

...read moreread less

Abstract: This paper describes a simple, efficient algorithm to locate all occurrences of any of a finite number of keywords in a string of text. The algorithm consists of constructing a finite state pattern matching machine from the keywords and then using the pattern matching machine to process the text string in a single pass. Construction of the pattern matching machine takes time proportional to the sum of the lengths of the keywords. The number of state transitions made by the pattern matching machine in processing the text string is independent of the number of keywords. The algorithm has been used to improve the speed of a library bibliographic search program by a factor of 5 to 10.

...read moreread less

3,450 citations

Journal Article•10.1145/359842.359859•

A fast string searching algorithm

[...]

Robert S. Boyer¹, J. Strother Moore²•Institutions (2)

SRI International¹, PARC²

01 Oct 1977-Communications of The ACM

TL;DR: The algorithm has the unusual property that, in most cases, not all of the first i.” in another string, are inspected.

...read moreread less

Abstract: An algorithm is presented that searches for the location, “il” of the first occurrence of a character string, “pat,” in another string, “string.” During the search operation, the characters of pat are matched starting with the last character of pat. The information gained by starting the match at the end of the pattern often allows the algorithm to proceed in large jumps through the text being searched. Thus the algorithm has the unusual property that, in most cases, not all of the first i characters of string are inspected. The number of characters actually inspected (on the average) decreases as a function of the length of pat. For a random English pattern of length 5, the algorithm will typically inspect i/4 characters of string before finding a match at i. Furthermore, the algorithm has been implemented so that (on the average) fewer than i + patlen machine instructions are executed. These conclusions are supported with empirical evidence and a theoretical analysis of the average behavior of the algorithm. The worst case behavior of the algorithm is linear in i + patlen, assuming the availability of array space for tables linear in patlen plus the size of the alphabet.

...read moreread less

2,710 citations

Book Chapter•10.1007/978-3-642-16321-0_27•

Extracting powers and periods in a string from its runs structure

[...]

Maxime Crochemore¹, Costas S. Iliopoulos², Marcin Kubica³, Jakub Radoszewski³, Wojciech Rytter³, Tomasz Waleń³ - Show less +2 more•Institutions (3)

King's College London¹, Curtin University², University of Warsaw³

11 Oct 2010

TL;DR: This paper uses Lyndon words and introduces the Lyndon structure of runs as a useful tool when computing powers and presents an efficient algorithm for testing primitivity of factors of a string and computing their primitive roots.

...read moreread less

Abstract: A breakthrough in the field of text algorithms was the discovery of the fact that the maximal number of runs in a string of length n is O(n) and that they can all be computed in O(n) time. We study some applications of this result. New simpler O(n) time algorithms are presented for a few classical string problems: computing all distinct kth string powers for a given k, in particular squares for k = 2, and finding all local periods in a given string of length n. Additionally, we present an efficient algorithm for testing primitivity of factors of a string and computing their primitive roots. Applications of runs, despite their importance, are underrepresented in existing literature (approximately one page in the paper of Kolpakov & Kucherov, 1999). In this paper we attempt to fill in this gap. We use Lyndon words and introduce the Lyndon structure of runs as a useful tool when computing powers. In problems related to periods we use some versions of the Manhattan skyline problem.

...read moreread less

444 citations

Journal Article•10.1145/79173.79184•

A very fast substring search algorithm

[...]

Daniel M. Sunday¹•Institutions (1)

East Stroudsburg University of Pennsylvania¹

01 Aug 1990-Communications of The ACM

TL;DR: A substring search algorithm that is faster than the Boyer-Moore algorithm and does not depend on scanning the pattern string in any particular order is described.

...read moreread less

Abstract: This article describes a substring search algorithm that is faster than the Boyer-Moore algorithm. This algorithm does not depend on scanning the pattern string in any particular order. Three variations of the algorithm are given that use three different pattern scan orders. These include: (1) a “Quick Search” algorithm; (2) a “Maximal Shift” and (3) an “Optimal Mismatch” algorithm.

...read moreread less

444 citations

...

Expand

Year	Papers
2025	2
2024	1
2023	1
2022	6
2021	3
2020	12

Topic Tools

Papers published on a yearly basis

Papers

Fast Pattern Matching in Strings

Efficient string matching: an aid to bibliographic search

A fast string searching algorithm

Extracting powers and periods in a string from its runs structure

A very fast substring search algorithm

Related Topics (5)

Performance Metrics