Proceedings Article10.1109/ICDM.2012.120
Distributed Matrix Completion
Christina Teflioudi,Faraz Makari,Rainer Gemulla +2 more
- 10 Dec 2012
- pp 655-664
122
TL;DR: The DALS, ASGD, and DSGD++ algorithms are novel variants of the popular alternating least squares and stochastic gradient descent algorithms, they exploit thread-level parallelism, in-memory processing, and asynchronous communication.
read more
Abstract: We discuss parallel and distributed algorithms for large-scale matrix completion on problems with millions of rows, millions of columns, and billions of revealed entries. We focus on in-memory algorithms that run on a small cluster of commodity nodes, even very large problems can be handled effectively in such a setup. Our DALS, ASGD, and DSGD++ algorithms are novel variants of the popular alternating least squares and stochastic gradient descent algorithms, they exploit thread-level parallelism, in-memory processing, and asynchronous communication. We provide some guidance on the asymptotic performance of each algorithm and investigate the performance of both our algorithms and previously proposed Map Reduce algorithms in large-scale experiments. We found that DSGD++ outperforms competing methods in terms of overall runtime, memory consumption, and scalability. Using DSGD++, we can factor a matrix with 10B entries on 16 compute nodes in around 40 minutes.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Good Intentions: Adaptive Parameter Servers via Intent Signaling
Alexander Renz-Wieland,Andreas Kieslinger,Robert Gericke,Rainer Gemulla,Zoi Kaoudi,Volker Markl +5 more
TL;DR: This paper proposes a novel intent signaling mechanism that acts as an enabler for adaptivity and naturally integrates into ML tasks, and proposes a fully adaptive, zero-tuning PS called AdaPS based on this mechanism.
Parallel algorithms for tensor completion in the CP format
Lars Karlsson,Daniel Kressner,André Uschmajew +2 more
- 01 Sep 2016
TL;DR: Novel parallel algorithms for tensor completion problems, with applications to recommender systems and function learning, based on a combination of the canonical polyadic tensor format with block coordinate descent methods are proposed.
Provable and efficient algorithms for robust subspace learning and tracking
TL;DR: This thesis develops provable and efficient algorithms for robust subspace learning and tracking, addressing data corruption, missing values, and distributed data, leveraging low-dimensional structures such as sparsity and low-rank representations in high-dimensional data.
To Index or Not to Index: Optimizing Exact Maximum Inner Product Search
Firas Abuzaid,Geet Sethi,Peter Bailis,Matei Zaharia +3 more
- 08 Apr 2019
TL;DR: A hardware-efficient brute-force approach, blocked matrix multiply (BMM), can outperform the state-of-the-art MIPS solvers by over an order of magnitude, for some—but not all—inputs, and a novel MIPS solution is presented, MAX-IMUS, that takes advantage of hardware efficiency and pruning of the search space.
FlexiFaCT: Scalable Flexible Factorization of Coupled Tensors on Hadoop
Alex Beutel,Partha Pratim Talukdar,Abhimanu Kumar,Christos Faloutsos,Evangelos E. Papalexakis,Eric P. Xing +5 more
- 01 Jan 2014
TL;DR: FlexiFaCT provides a distributed, scalable method for decomposing matrices, tensors, and coupled data sets through stochastic gradient descent on a variety of objective functions.
References
Matrix Factorization Techniques for Recommender Systems
TL;DR: As the Netflix Prize competition has demonstrated, matrix factorization models are superior to classic nearest neighbor techniques for producing product recommendations, allowing the incorporation of additional information such as implicit feedback, temporal effects, and confidence levels.
Exact Matrix Completion via Convex Optimization
TL;DR: It is proved that one can perfectly recover most low-rank matrices from what appears to be an incomplete set of entries, and that objects other than signals and images can be perfectly reconstructed from very limited information.
A limited memory algorithm for bound constrained optimization
TL;DR: An algorithm for solving large nonlinear optimization problems with simple bounds is described, based on the gradient projection method and uses a limited memory BFGS matrix to approximate the Hessian of the objective function.
A limited-memory algorithm for bound-constrained optimization
Richard H. Byrd,L. Peihuang,Jorge Nocedal +2 more
- 01 Mar 1996
TL;DR: An algorithm for solving large nonlinear optimization problems with simple bounds is described, based on the gradient projection method and uses a limited-memory BFGS matrix to approximate the Hessian of the objective function.
Collaborative Filtering for Implicit Feedback Datasets
Yifan Hu,Yehuda Koren,Chris Volinsky +2 more
- 15 Dec 2008
TL;DR: This work identifies unique properties of implicit feedback datasets and proposes treating the data as indication of positive and negative preference associated with vastly varying confidence levels, which leads to a factor model which is especially tailored for implicit feedback recommenders.