TL;DR: This work seeks to establish the relative importance of each step of mid-level feature extraction through a comprehensive cross evaluation of several types of coding modules and pooling schemes and shows how to improve the best performing coding scheme by learning a supervised discriminative dictionary for sparse coding.
Abstract: Many successful models for scene or object recognition transform low-level descriptors (such as Gabor filter responses, or SIFT descriptors) into richer representations of intermediate complexity. This process can often be broken down into two steps: (1) a coding step, which performs a pointwise transformation of the descriptors into a representation better adapted to the task, and (2) a pooling step, which summarizes the coded features over larger neighborhoods. Several combinations of coding and pooling schemes have been proposed in the literature. The goal of this paper is threefold. We seek to establish the relative importance of each step of mid-level feature extraction through a comprehensive cross evaluation of several types of coding modules (hard and soft vector quantization, sparse coding) and pooling schemes (by taking the average, or the maximum), which obtains state-of-the-art performance or better on several recognition benchmarks. We show how to improve the best performing coding scheme by learning a supervised discriminative dictionary for sparse coding. We provide theoretical and empirical insight into the remarkable performance of max pooling. By teasing apart components shared by modern mid-level feature extractors, our approach aims to facilitate the design of better recognition architectures.
TL;DR: In this article, the authors studied the problem of distributed average consensus in sensor networks with quantized data and random link failures. But their work was restricted to the case where the quantizer range is unbounded.
Abstract: The paper studies the problem of distributed average consensus in sensor networks with quantized data and random link failures. To achieve consensus, dither (small noise) is added to the sensor states before quantization. When the quantizer range is unbounded (countable number of quantizer levels), stochastic approximation shows that consensus is asymptotically achieved with probability one and in mean square to a finite random variable. We show that the mean-squared error (mse) can be made arbitrarily small by tuning the link weight sequence, at a cost of the convergence rate of the algorithm. To study dithered consensus with random links when the range of the quantizer is bounded, we establish uniform boundedness of the sample paths of the unbounded quantizer. This requires characterization of the statistical properties of the supremum taken over the sample paths of the state of the quantizer. This is accomplished by splitting the state vector of the quantizer in two components: one along the consensus subspace and the other along the subspace orthogonal to the consensus subspace. The proofs use maximal inequalities for submartingale and supermartingale sequences. From these, we derive probability bounds on the excursions of the two subsequences, from which probability bounds on the excursions of the quantizer state vector follow. The paper shows how to use these probability bounds to design the quantizer parameters and to explore tradeoffs among the number of quantizer levels, the size of the quantization steps, the desired probability of saturation, and the desired level of accuracy ? away from consensus. Finally, the paper illustrates the quantizer design with a numerical study.
TL;DR: This paper reports the recent exploration of the layer-by-layer learning strategy for training a multi-layer generative model of patches of speech spectrograms and shows that the binary codes learned produce a logspectral distortion that is approximately 2 dB lower than a subband vector quantization technique over the entire frequency range of wide-band speech.
Abstract: This paper reports our recent exploration of the layer-by-layer learning strategy for training a multi-layer generative model of patches of speech spectrograms. The top layer of the generative model learns binary codes that can be used for efficient compression of speech and could also be used for scalable speech recognition or rapid speech content retrieval. Each layer of the generative model is fully connected to the layer below and the weights on these connections are pretrained efficiently by using the contrastive divergence approximation to the log likelihood gradient. After layer-bylayer pre-training we “unroll” the generative model to form a deep auto-encoder, whose parameters are then fine-tuned using back-propagation. To reconstruct the full-length speech spectrogram, individual spectrogram segments predicted by their respective binary codes are combined using an overlapand-add method. Experimental results on speech spectrogram coding demonstrate that the binary codes produce a logspectral distortion that is approximately 2 dB lower than a subband vector quantization technique over the entire frequency range of wide-band speech. Index Terms: deep learning, speech feature extraction, neural networks, auto-encoder, binary codes, Boltzmann machine
TL;DR: This paper introduces residual vector quantization based approaches that are appropriate for unstructured vectors that are compared to two state-of-the-art methods, spectral hashing and product quantization, on both structured and unstructuring datasets.
Abstract: A recently proposed product quantization method is efficient for large scale approximate nearest neighbor search, however, its performance on unstructured vectors is limited. This paper introduces residual vector quantization based approaches that are appropriate for unstructured vectors. Database vectors are quantized by residual vector quantizer. The reproductions are represented by short codes composed of their quantization indices. Euclidean distance between query vector and database vector is approximated by asymmetric distance, i.e., the distance between the query vector and the reproduction of the database vector. An efficient exhaustive search approach is proposed by fast computing the asymmetric distance. A straight forward non-exhaustive search approach is proposed for large scale search. Our approaches are compared to two state-of-the-art methods, spectral hashing and product quantization, on both structured and unstructured datasets. Results show that our approaches obtain the best results in terms of the trade-off between search quality and memory usage.
TL;DR: The method is evaluated, which is based on high-dimensional data from industrial processes and based on benchmark datasets from the Internet and compared with well-known batch-training methods, in terms of accuracy and complexity of the fuzzy systems.
Abstract: In this paper, we deal with a novel data-driven learning method [sparse fuzzy inference systems (SparseFIS)] for Takagi-Sugeno (T-S) fuzzy systems, extended by including rule weights. Our learning method consists of three phases: The first phase conducts a clustering process in the input/output feature space with iterative vector quantization and projects the obtained clusters onto 1-D axes to form the fuzzy sets (centers and widths) in the antecedent parts of the rules. Hereby, the number of clusters = rules is predefined and denotes a kind of upper bound on a reasonable granularity. The second phase optimizes the rule weights in the fuzzy systems with respect to least-squares error measure by applying a sparsity-constrained steepest descent-optimization procedure. Depending on the sparsity threshold, weights of many or a few rules can be forced toward 0, thereby, switching off (eliminating) some rules (rule selection). The third phase estimates the linear consequent parameters by a regularized sparsity-constrained-optimization procedure for each rule separately (local learning approach). Sparsity constraints are applied in order to force linear parameters to be 0, triggering a feature-selection mechanism per rule. Global feature selection is achieved whenever the linear parameters of some features in each rule are (near) 0. The method is evaluated, which is based on high-dimensional data from industrial processes and based on benchmark datasets from the Internet and compared with well-known batch-training methods, in terms of accuracy and complexity of the fuzzy systems.
TL;DR: In this article, a properly normalized object is first described by a set of depth-buffer views captured on the surrounding vertices of a given unit geodesic sphere, and then represent each view as a word histogram generated by the vector quantization of the view's salient local features.
Abstract: This paper presents a novel 3D shape retrieval method, which uses Bag-of-Features and an efficient multi-view shape matching scheme. In our approach, a properly normalized object is first described by a set of depth-buffer views captured on the surrounding vertices of a given unit geodesic sphere. We then represent each view as a word histogram generated by the vector quantization of the view’s salient local features. The dissimilarity between two 3D models is measured by the minimum distance of their all (24) possible matching pairs. This paper also investigates several critical issues including the influence of the number of views, codebook, training data, and distance function. Experiments on four commonly-used benchmarks demonstrate that: 1) Our approach obtains superior performance in searching for rigid models. 2) The local feature and global feature based methods are somehow complementary. Moreover, a linear combination of them significantly outperforms the state-of-the-art in terms of retrieval accuracy.
TL;DR: Results show that this method features stronger a watermark in comparison with conventional QIM and, as a result, has better performance while it does not suffer from the drawbacks of a previously proposed logarithmic quantization algorithm.
Abstract: In this paper, a novel arrangement for quantizer levels in the Quantization Index Modulation (QIM) method is proposed. Due to perceptual advantages of logarithmic quantization, and in order to solve the problems of a previous logarithmic quantization-based method, we used the compression function of ? -Law standard for quantization. In this regard, the host signal is first transformed into the logarithmic domain using the ? -Law compression function. Then, the transformed data is quantized uniformly and the result is transformed back to the original domain using the inverse function. The scalar method is then extended to vector quantization. For this, the magnitude of each host vector is quantized on the surface of hyperspheres which follow logarithmic radii. Optimum parameter ? for both scalar and vector cases is calculated according to the host signal distribution. Moreover, inclusion of a secret key in the proposed method, similar to the dither modulation in QIM, is introduced. Performance of the proposed method in both cases is analyzed and the analytical derivations are verified through extensive simulations on artificial signals. The method is also simulated on real images and its performance is compared with previous scalar and vector quantization-based methods. Results show that this method features stronger a watermark in comparison with conventional QIM and, as a result, has better performance while it does not suffer from the drawbacks of a previously proposed logarithmic quantization algorithm.
TL;DR: A snapshot of the recent VQ codebook generation schemes is presented, which include mean-distance-ordered partial codebook search (MPS), enhance LBG (ELBG), neural network based techniques, genetic-based algorithms, principal component analysis (PCA) approaches, tabu search (TS) schemes, and more.
Abstract: One of the key roles of Vector Quantization (VQ) is how to generate a good codebook such that the distortion between the original image and the reconstructed image is the minimum. In the past years, many improved algorithms of VQ codebook generation approaches have been developed. In this paper, we present a snapshot of the recent de- veloped schemes. The discussed schemes include mean-distance-ordered partial codebook search (MPS), enhance LBG (ELBG), neural network based techniques, genetic-based algorithms, principal component analysis (PCA) approaches, tabu search (TS) schemes,
TL;DR: A new prototype learning algorithm based on the conditional log-likelihood loss (CLL), which isbased on the discriminative model called log- likelihood of margin (LOGM), which yields higher classification accuracies than the MCE, generalized learning vector quantization (GLVQ), soft nearest prototype classifier (SNPC) and the robust soft learning vectorquantization (RSLVQ).
TL;DR: This paper discusses the use of the Scale Invariance Feature Transform (SIFT) features for bare hand gesture recognition, and introduces the Bag-of-features model, which is fed into the multi-class SVM training classifier model to recognize the hand gesture.
Abstract: This paper discusses the use of the Scale Invariance Feature Transform (SIFT) features for bare hand gesture recognition. In the training stage, we can not use SIFT keypoints of training images directly with a multi-class Support Vector Machine (SVM) to build a training classifier model, because of the space incompatibility of the SIFT keypoints for every training image that contains the hand gesture only. Therefore, the Bag-of-features model was introduced. After extracting the keypoints for every training image using the SIFT algorithm, a vector quantization technique is used to unify them. The quantization will map keypoints extracted from every training image into a unified dimensional histogram vector (Bag-of-words) after K-means clustering. This histogram is treated as an input vector for a multi-class SVM to build the training classifier model. In the testing stage, the keypoints are extracted from every image captured from the webcam and fed into the cluster model to map them with one (Bag-of-words) vector, which is finally fed into the multi-class SVM training classifier model to recognize the hand gesture.
TL;DR: The distribution preserving quantization (DPQ) achieves the optimal trade-off between mean square error and bit rate asymptotically and provides a continuum ranging from rate-distortion optimal signal quantization to parametric coding.
Abstract: A new quantization scheme that preserves the probability distribution of the source signal is presented. The distribution preserving quantization (DPQ) achieves the optimal trade-off between mean square error and bit rate asymptotically. It provides a continuum ranging from rate-distortion optimal signal quantization to parametric coding. The method can be used as a core component for scalable coding. Its efficacy is illustrated by applying the scheme to audio coding.
TL;DR: This work presents a monaural speech enhancement method based on sparse coding of noisy speech signals in a composite dictionary, consisting of the concatenation of a speech and interferer dictionary, both being possibly over-complete.
Abstract: The enhancement of speech degraded by non-stationary interferers is a highly relevant and difficult task of many signal processing applications. We present a monaural speech enhancement method based on sparse coding of noisy speech signals in a composite dictionary, consisting of the concatenation of a speech and interferer dictionary, both being possibly over-complete. The speech dictionary is learned off-line on a training corpus, while an environment specific interferer dictionary is learned on-line during speech pauses. Our approach optimizes the trade-off between source distortion and source confusion, and thus achieves significant improvements on objective quality measures like cepstral distance, in the speaker dependent and independent case, in several real-world environments and at low signal-to-noise ratios. Our enhancement method outperforms state-of-the-art methods like multi-band spectral subtraction and approaches based on vector quantization.
TL;DR: A novel batch-image encryption algorithm that combines Vector Quantization (VQ) and additional index-compression process to benefit from their computational efficiency and low transmission bandwidth without affecting the original compression rate is presented.
TL;DR: A new reversible scheme based on locally adaptive coding for VQ-compressed images that has the best compression rate and the highest embedding capacity compared with other reversible VQ embedding methods is proposed.
TL;DR: This paper describes segmentation method consisting of two phases, in the first phase, the MRI brain image is acquired from patients database, and after that Hierarchical Self Organizing Map is applied for image segmentation.
Abstract: Image Segmentation is an important and challenging factor in the medical image segmentation. This paper describes segmentation method consisting of two phases. In the first phase, the MRI brain image is acquired from patients database, In that film artifact and noise are removed. After that Hierarchical Self Organizing Map(HSOM) is applied for image segmentation. The HSOM is the extension of the conventional self organizing map used to classify the image row by row. In this lowest level of weight vector, a higher value of tumor pixels, computation speed is achieved by the HSOM with vector quantization
TL;DR: A novel reversible data-hiding scheme that embeds secret data into a transformed image and achieves lossless reconstruction of vector quantization (VQ) indices is presented.
Abstract: This work presents a novel reversible data-hiding scheme that embeds secret data into a transformed image and achieves lossless reconstruction of vector quantization (VQ) indices. The VQ compressed image is modified by the side-matched VQ scheme to yield a transformed image. Distribution of the transformed image is employed to achieve high embedding capacity and a low bit rate. Moreover, three configurations, under-hiding, normal-hiding, and over-hiding schemes, are utilized to improve the proposed scheme further for various applications. Experimental results demonstrate that the proposed scheme significantly enhances the compression ratio and embedding capacity. Experimental results also show that the proposed scheme achieves the best performance among approaches in literature in terms of the compression ratio and embedding capacity.
TL;DR: In this paper, local descriptors are extracted from an image and the image vector is compressed using a vector quantization algorithm to generate a compressed image vector, which is then concatenated with the compressed sub-vectors to generate the final image vector.
Abstract: Local descriptors are extracted from an image. An image vector is generated having vector elements indicative of parameters of mixture model components of a mixture model representing the extracted local descriptors. The image vector is compressed using a vector quantization algorithm to generate a compressed image vector. Optionally, the compressing comprises splitting the image vector into a plurality of sub-vectors each including at least two vector elements, compressing each sub-vector independently using the vector quantization algorithm, and concatenating the compressed sub-vectors to generate the compressed image vector. Optionally, each sub-vector includes only vector elements indicative of parameters of a single mixture model component, and any sparse sub-vector whose vector elements are indicative of parameters of a mixture model component that does not represent any of the extracted local descriptors is not compressed.
TL;DR: Experimental results show the proposed algorithm is effective, the weight trained from image of Lena is successfully used to other images' compression and reconstruction.
Abstract: A color image compression algorithm based on quaternion neural network approach is proposed. The original RGB based color image of Lena can be firstly modeled as pure imaginary quaternion matrix, i.e. any pixel of R,G,B corresponding to the I,J,K imaginary axis , to ensure the integrity of pixel in the computation. The obtained quaternion matrix can be split up into 8×8 sub-blocks and vector quantization to make up of a new sample set. This sample set then is used to train the quaternion neural network adopting quaternion Generalized Hebbian Algorithm (QGHA), acquiring a quaternion weight coefficient that can get the principal components (PCs), the weight can be used to compress and reconstruct the image. Experimental results show the proposed algorithm is effective, the weight trained from image of Lena is successfully used to other images' compression and reconstruction.
TL;DR: The proposed method, built on the basis of the previous deterministic forecasting method that does not require the overhead of determining the order number, as in other high-order models, utilizes a vector quantization technique to support forecasting if there are no matching historical patterns, which is usually the case with long-term forecasting
TL;DR: The concepts of vector quantization (VQ) and association rules in data mining are employed to propose a robust watermarking technique that achieves effective resistance against several image processings such as blurring, sharpening, adding in Gaussian noise, cropping, and JPEG lossy compression.
TL;DR: This work develops an efficient lossless image compression scheme called super-spatial structure prediction, which is motivated by motion prediction in video coding, attempting to find an optimal prediction of structure components within previously encoded image regions.
Abstract: We recognize that the key challenge in image compression is to efficiently represent and encode high-frequency image structure components, such as edges, patterns, and textures. In this work, we develop an efficient lossless image compression scheme called super-spatial structure prediction. This super-spatial prediction is motivated by motion prediction in video coding, attempting to find an optimal prediction of structure components within previously encoded image regions. We find that this super-spatial prediction is very efficient for image regions with significant structure components. Our extensive experimental results demonstrate that the proposed scheme is very competitive and even outperforms the state-of-the-art lossless image compression methods.
TL;DR: A histogrambased data-dependent estimate is proposed adopting a version of Barron-type histogram-based estimate, with the stipulation of su cient conditions on the partition scheme to make the estimate strongly consistent.
TL;DR: In this paper, the authors present a control system for the process of rotary drilling based on recognition of geomechanical class of rock with the method of vector quantization, and the necessary information about the character of the process are obtained from the signal of the accompanying vibro-acoustic emissions.
Abstract: Rotary drilling belongs to the key methods of rock separation not only in mining, but also in wider areas of geotechnologies 1, 2 . The most efficient separation of rocks is in so-called volume area which can be achieved with the correct choice of work tool and work mode while taking into account the geomechanical properties of rock separation. Theoretical research of rock separation by rotary drilling and subsequent experiments on a drilling stand 1, 3, 4 showed that there exists an optimum – efficient mode of drilling from the viewpoint of specific energy consumption w (J/m), from the viewpoint of the wear of the separation tool, but also from the viewpoint of the speed of drilling v (m/s). These facts led to the idea of efficient control of the process of drilling. Further research showed that in the neighborhood of efficient mode of the process of drilling the accompanying vibro-acoustic signal has identifiable properties 3, 4 . In the contribution are given first partial results in the area of design of control system for the process of rock drilling, based on recognition of geomechanical class of rock with the method of vector quantization. At the same time, the necessary information about the character of the process are obtained from the signal of the accompanying vibro-acoustic emissions.
TL;DR: The proposed algorithm gives less distortion as compared to well known Linde Buzo Gray (LBG) algorithm and Kekre’s Proportionate Error (KPE) Algorithm by introducing new orientation every time to split the clusters.
Abstract: —The paper presents new clustering algorithm. The proposed algorithm gives less distortion as compared to well known Linde Buzo Gray (LBG) algorithm and Kekre’s Proportionate Error (KPE) Algorithm. Constant error is added every time to split the clusters in LBG, resulting in formation of cluster in one direction which is 135 0 in 2-dimensional case. Because of this reason clustering is inefficient resulting in high MSE in LBG. To overcome this drawback of LBG proportionate error is added to change the cluster orientation in KPE. Though the cluster orientation in KPE is changed its variation is limited to ± 45 0 over 135 . The proposed algorithm takes care of this problem by introducing new orientation every time to split the clusters. The proposed method reduces PSNR by 2db to 5db for codebook size 128 to 1024 with respect to LBG. Keywords-component; Vector Quantization; Codebook; Codevector; Encoding; Compression. I. I NTRODUCTION Exhaustive Search (ES) method gives the optimal result at the World Wide Web Applications have extensively grown since last few decades and it has become requisite tool for education, communication, industry, amusement etc. All these applications are multimedia-based applications consisting of images and videos. Images/videos require enormous volume of data items, creating a serious problem as they need higher channel bandwidth for efficient transmission. Further high degree of redundancies is observed in digital images. Thus the need for image compression arises for resourceful storage and transmission. Image compression is classified into two categories, lossless image compression and lossy image compression technique. Vector quantization (VQ) is one of the lossy data compression techniques[1], [2] and has been used in number of applications, like pattern recognition [3], speech recognition and face detection [4], [5], image segmentation [6-9], speech data compression [10], Content Based Image Retrieval (CBIR) [11], [12], Face recognition[13], [14] iris recognition[15], tumor detection in mammography images [29] etc. VQ is a mapping function which maps k-dimensional vector space to a finite set CB = {C
TL;DR: In this article, the authors describe two algorithms for learning distance metrics based on convex optimization, which can be used to measure the dissimilarity between different feature vectors in a multidimensional vector space.
Abstract: The goal of machine learning is to build automated systems that can classify and recognize complex patterns in data. Not surprisingly, the representation of the data plays an important role in determining what types of patterns can be automatically discovered. Many algorithms for machine learning assume that the data are represented as elements in a metric space. For example, in popular algorithms such as nearest-neighbor classification, vector quantization, and kernel density estimation, the metric distances between different examples provide a measure of their dissimilarity [1]. The performance of these algorithms can depend sensitively on the manner in which distances are measured. When data are represented as points in a multidimensional vector space, simple Euclidean distances are often used to measure the dissimilarity between different examples. However, such distances often do not yield reliable judgments; in addition, they cannot highlight the distinctive features that play a role in certain types of classification, but not others. For example, consider two schemes for clustering images of faces: one by age, one by gender. Images can be represented as points in a multidimensional vector space in many ways—for example, by enumerating their pixel values, or by computing color histograms. However the images are represented, different components of these feature vectors are likely to be relevant for clustering by age versus clustering by gender. Naturally, for these different types of clustering, we need different ways of measuring dissimilarity; in particular, we need different metrics for computing distances between feature vectors. This article describes two algorithms for learning such distance metrics based on recent developments in convex optimization.
TL;DR: Numerical results show that even with one bit quantization, the proposed approach achieves a superior mean square deviation performance (with respect to the global linear minimum mean-square error estimate) within a moderate number of iterations.
Abstract: We consider the problem of distributed estimation of a Gauss-Markov random field using a wireless sensor network (WSN), where due to the stringent power and communication constraints, each sensor has to quantize its data before transmission. In this case, the convergence of conventional iterative matrix-splitting algorithms is hindered by the quantization errors. To address this issue, we propose a one-bit adaptive quantization approach which leads to decaying quantization errors. Numerical results show that even with one bit quantization, the proposed approach achieves a superior mean square deviation performance (with respect to the global linear minimum mean-square error estimate) within a moderate number of iterations.
TL;DR: This paper presents an image-based method to effectively address the problem of matching non-rigid shapes in content-based 3D object retrieval and obtains better retrieval performance compared to the state-of-the-art.
Abstract: Matching non-rigid shapes is a challenging research field in content-based 3D object retrieval. In this paper, we present an image-based method to effectively address this problem. Multidimensional Scaling (MDS) and Principal Component Analysis (PCA) are first applied to each object to calculate its canonical form, which is afterward represented by 66 depth-buffer images captured on the vertices of an unit geodesic sphere. Then, each image is described as a word histogram obtained by the vector quantization of the image's salient local features. Finally, a multi-view shape matching scheme is carried out to measure the dissimilarity between two models. Experimental results on the McGill Articulated Shape Benchmark database demonstrate that, our method obtains better retrieval performance compared to the state-of-the-art.
TL;DR: Experiments and simulations show that quantizers jointly designed with channel conditions significantly reduce the EED when compared with quantizers designed separately without reference to channel conditions, which reveals a practical and effective design for noisy-channel quantization as to simplify the channel model by considering a random index assignment.
Abstract: This paper studies the design of vector quantization on noisy channels and its high rate asymptotic performance. Given a tandem source-channel coding system with vector quantization, block channel coding, and random index assignment, a closed-form formula is first derived for computing the average end-to-end distortion (EED) of the system, which reveals a structural factor called the scatter factor of a noisy channel quantizer. Based on this formula, we propose a noisy-channel quantization design method by minimizing the EED. Experiments and simulations show that quantizers jointly designed with channel conditions significantly reduce the EED when compared with quantizers designed separately without reference to channel conditions, which reveals a practical and effective design for noisy-channel quantization as to simplify the channel model by considering a random index assignment. Furthermore, we have presented the high rate asymptotic analysis of the EED for the tandem system, while convergence analysis of the iterative algorithm is included in the Appendix.
TL;DR: This work proposes a scheme for compressing distributions called Type Coding, which offers lower complexity and higher compression efficiency compared to tree-based quantization schemes proposed in prior work, and constructs optimal Entropy Constrained Vector Quantization (ECVQ) code-books and shows that Type Coded comes close to achieving optimal performance.
Abstract: We study different quantization schemes for the Compressed Histogram of Gradients (CHoG) image feature descriptor. We propose a scheme for compressing distributions called Type Coding, which offers lower complexity and higher compression efficiency compared to tree-based quantization schemes proposed in prior work. We construct optimal Entropy Constrained Vector Quantization (ECVQ) code-books and show that Type Coding comes close to achieving optimal performance. The proposed descriptors are 16× smaller than SIFT and perform on par. We implement the descriptor in a mobile image retrieval system and for a database of 1 million CD, DVD and book covers, we achieve 96% retrieval accuracy using only 4 kilobytes of data per query image.
TL;DR: A new method based on the firefly algorithm to construct the codebook of vector quantization is proposed, which gets higher quality than those generated from the LBG and PSO-LBG algorithms, but there are not significantly different to the HBMO- LBG algorithm.
Abstract: The vector quantization (VQ) was a powerful technique in the applications of digital image compression. The traditionally widely used method such as the Linde-Buzo-Gray (LBG) algorithm always generated local optimal codebook. This paper proposed a new method based on the firefly algorithm to construct the codebook of vector quantization. The proposed method uses LBG method as the initial of firefly algorithm to develop the VQ algorithm. This method is called FF-LBG algorithm. The FF-LBG algorithm is compared with the other three methods that are LBG, PSO-LBG and HBMO-LBG algorithms. Experimental results showed that the computation of this proposed FF-LBG algorithm is faster than the PSO-LBG, and the HBMO-LBG algorithms. Furthermore, the reconstructured images get higher quality than those generated from the LBG and PSO-LBG algorithms, but there are not significantly different to the HBMO-LBG algorithm.