Learning string-edit distance
TL;DR: The stochastic model allows us to learn a string-edit distance function from a corpus of examples and is applicable to any string classification problem that may be solved using a similarity function against a database of labeled prototypes.
read more
Abstract: In many applications, it is necessary to determine the similarity of two strings. A widely-used notion of string similarity is the edit distance: the minimum number of insertions, deletions, and substitutions required to transform one string into the other. In this report, we provide a stochastic model for string-edit distance. Our stochastic model allows us to learn a string-edit distance function from a corpus of examples. We illustrate the utility of our approach by applying it to the difficult problem of learning the pronunciation of words in conversational speech. In this application, we learn a string-edit distance with nearly one-fifth the error rate of the untrained Levenshtein distance. Our approach is applicable to any string classification problem that may be solved using a similarity function against a database of labeled prototypes.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Entity resolution: theory, practice & open challenges
Lise Getoor,Ashwin Machanavajjhala +1 more
- 01 Aug 2012
TL;DR: This tutorial brings together perspectives on ER from a variety of fields, including databases, machine learning, natural language processing and information retrieval, to provide, in one setting, a survey of a large body of work.
382
Record linkage: similarity measures and algorithms
Nick Koudas,Sunita Sarawagi,Divesh Srivastava +2 more
- 27 Jun 2006
TL;DR: This tutorial provides a comprehensive and cohesive overview of the key research results in the area of record linkage methodologies and algorithms for identifying approximate duplicate records, and available tools for this purpose.
Ecological Patterns of nifH Genes in Four Terrestrial Climatic Zones Explored with Targeted Metagenomics Using FrameBot, a New Informatics Tool
Qiong Wang,John F. Quensen,Jordan A. Fish,Tae Kwon Lee,Yanni Sun,James M. Tiedje,James R. Cole +6 more
TL;DR: To accurately detect and correct frameshifts caused by indel sequencing errors, FrameBot was developed, a tool for frameshift correction and nearest-neighbor classification, and its accuracy was compared to that of two other rapid frameshIFT correction tools.
315
Toward an Epistemology ofWikipedia
TL;DR: To improve Wikipedia, it is suggested that to clarify what the authors' epistemic values are and to better understand why Wikipedia works as well as it does, changes to Wikipedia need to be identified.
280
Iterative record linkage for cleaning and integration
Indrajit Bhattacharya,Lise Getoor +1 more
- 13 Jun 2004
TL;DR: Results are presented that illustrate the power and feasibility of making use of join information when comparing records and the need to make multiple passes over the data to correctly find all duplicates.
References
Error bounds for convolutional codes and an asymptotically optimum decoding algorithm
TL;DR: The upper bound is obtained for a specific probabilistic nonsequential decoding algorithm which is shown to be asymptotically optimum for rates above R_{0} and whose performance bears certain similarities to that of sequential decoding algorithms.
7.6K
The viterbi algorithm
Jr. G.D. Forney
- 01 Mar 1973
TL;DR: This paper gives a tutorial exposition of the Viterbi algorithm and of how it is implemented and analyzed, and increasing use of the algorithm in a widening variety of areas is foreseen.
6.5K