Proceedings Article10.1145/1015330.1015332
Solving large scale linear prediction problems using stochastic gradient descent algorithms
Tong Zhang
- 04 Jul 2004
- pp 116
1.4K
TL;DR: Stochastic gradient descent algorithms on regularized forms of linear prediction methods, related to online algorithms such as perceptron, are studied, and numerical rate of convergence for such algorithms is obtained.
read more
Abstract: Linear prediction methods, such as least squares for regression, logistic regression and support vector machines for classification, have been extensively used in statistics and machine learning. In this paper, we study stochastic gradient descent (SGD) algorithms on regularized forms of linear prediction methods. This class of methods, related to online algorithms such as perceptron, are both efficient and very simple to implement. We obtain numerical rate of convergence for such algorithms, and discuss its implications. Experiments on text data will be provided to demonstrate numerical and statistical consequences of our theoretical findings.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Enhanced Machine Learning Sketches for Network Measurements
01 Apr 2023
TL;DR: In this paper , a generic machine learning framework is proposed to reduce the dependence of sketching techniques on network traffic characteristics, such as flow sizes, top- < inline-formula> flows, and number of flows.
•Posted Content
Leveraging Machine Learning to Detect Data Curation Activities
TL;DR: In this paper, a machine learning approach for annotating and analyzing data curation work logs at ICPSR, a large social sciences data archive, is described, which is used to study the relationship between curation and data reuse.
Patent
Determining soft graph correspondence between node sets based on feature representations
U Kang,Ravindranath Konuru,Jimeng Sun,Hanghang Tong +3 more
- 17 Jan 2013
TL;DR: In this article, a method for determining a correspondence between a first node set of a first graph and a second node set in a second graph is presented. But the method is not suitable for the case where the first graph has a fixed number of vertices and the second graph has many vertices.
Facilitating Human-Wildlife Cohabitation through Conflict Prediction
TL;DR: In this article , a case study on human-wildlife conflicts in the Bramhapuri Forest Division in Chandrapur, Maharashtra, India is presented, where a sparse conflict training dataset is used to identify the right features to make accurate predictions of conflicts at the required spatial granularity.
A Multi-parameter Updating Fourier Online Gradient Descent Algorithm for Large-scale Nonlinear Classification
TL;DR: Empirical studies on several benchmark data sets demonstrate that compared with the state-of-the-art online random Fourier feature map methods, the proposed MPU-FOGD can obtain better test accuracy.
References
Principles of neurodynamics. perceptrons and the theory of brain mechanisms
TL;DR: The background, basic sources of data, concepts, and methodology to be employed in the study of perceptrons are reviewed, and some of the notation to be used in later sections are presented.
2.5K
Acceleration of stochastic approximation by averaging
Boris T. Polyak,Anatoli Juditsky +1 more
TL;DR: Convergence with probability one is proved for a variety of classical optimization and identification problems and it is demonstrated for these problems that the proposed algorithm achieves the highest possible rate of convergence.
2.3K
Discriminative training methods for hidden Markov models: theory and experiments with perceptron algorithms
Michael Collins
- 06 Jul 2002
TL;DR: Experimental results on part-of-speech tagging and base noun phrase chunking are given, in both cases showing improvements over results for a maximum-entropy tagger.
Large margin classification using the perceptron algorithm
Yoav Freund,Robert E. Schapire +1 more
- 24 Jul 1998
TL;DR: A new algorithm for linear classification which combines Rosenblatt‘s perceptron algorithm with Helmbold and Warmuth’s leave-one-out method is introduced, which takes advantage of data that are linearly separable with large margins.
•Book
Stochastic Approximation Algorithms and Applications
Harold J. Kushner,George Yin +1 more
- 01 Jan 1997
TL;DR: Applications and issues application to learning, state dependent noise and queueing applications to signal processing and adaptive control mathematical background convergence with probability one, introduction weak convergence methods for general algorithms applications, proofs of convergence rate of convergence averaging of the iterates distributed/decentralized and asynchronous algorithms.
1.2K