Open AccessPosted Content
Implicitly Maximizing Margins with the Hinge Loss.
TL;DR: It is shown that for a linear classifier on linearly separable data with fixed step size, the margin of this modified hinge loss converges to the max-margin at the rate of $\mathcal{O}( 1/t)$, which is fast when compared with the rates of exponential losses such as the logistic loss.
read more
Abstract: A new loss function is proposed for neural networks on classification tasks which extends the hinge loss by assigning gradients to its critical points We will show that for a linear classifier on linearly separable data with fixed step size, the margin of this modified hinge loss converges to the $\ell_2$ max-margin at the rate of $\mathcal{O}( 1/t )$ This rate is fast when compared with the $\mathcal{O}(1/\log t)$ rate of exponential losses such as the logistic loss Furthermore, empirical results suggest that this increased convergence speed carries over to ReLU networks
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Enhanced detection of threat materials by dark-field x-ray imaging combined with deep neural networks
T. Partridge,Alberto Astolfo,Sai Shankar,Fabio A. Vittoria,Marco Endrizzi,Simon R. Arridge,T. Riley-Smith,Ian Haig,David Bate,Alessandro Olivo +9 more
TL;DR: In this article , the authors show that dark-field creates a texture which is characteristic of the imaged material, and that its combination with conventional attenuation leads to an improved discrimination of threat materials.
References
Gradient-based learning applied to document recognition
Yann LeCun,Léon Bottou,Léon Bottou,Yoshua Bengio,Yoshua Bengio,Yoshua Bengio,Patrick Haffner +6 more
- 01 Jan 1998
TL;DR: In this article, a graph transformer network (GTN) is proposed for handwritten character recognition, which can be used to synthesize a complex decision surface that can classify high-dimensional patterns, such as handwritten characters.
53.5K
•Dissertation
Learning Multiple Layers of Features from Tiny Images
Alex Krizhevsky
- 01 Jan 2009
TL;DR: In this paper, the authors describe how to train a multi-layer generative model of natural images, using a dataset of millions of tiny colour images, described in the next section.
•Proceedings Article
Support vector machines for multi-class pattern recognition.
Jason Weston,Chris Watkins +1 more
- 01 Jan 1999
TL;DR: A formulation of the SVM is proposed that enables a multi-class pattern recognition problem to be solved in a single optimisation and a similar generalization of linear programming machines is proposed.
953
On Loss Functions for Deep Neural Networks in Classification
TL;DR: This paper investigates how particular choices of loss functions affect deep models and their learning dynamics, as well as resulting classifiers robustness to various effects, and shows that L1 and L2 losses are justified classification objectives for deep nets, by providing probabilistic interpretation in terms of expected misclassification.
607
•Posted Content
On Loss Functions for Deep Neural Networks in Classification
TL;DR: In this paper, the authors investigate how particular choices of loss functions affect deep models and their learning dynamics, as well as resulting classifiers robustness to various effects, and show that L1 and L2 losses are, quite surprisingly, justified classification objectives for deep nets, by providing probabilistic interpretation in terms of expected misclassification.
361