• Home
  • Agent Gallery
  • Templates
  • Chat with PDF
  • Literature Review
  • AI Writer
  • Find Topics
  • Paraphraser
  • Citation Generator
  • Extract Data
  • AI Detector
  • AI Humanizer
Scispace (Formerly Typeset)
  1. Home
  2. Topics
  3. Coppersmith–Winograd algorithm
  4. 2012
  1. Home
  2. Topics
  3. Coppersmith–Winograd algorithm
  4. 2012
Showing papers on "Coppersmith–Winograd algorithm published in 2012"
Proceedings Article•10.1145/2312005.2312044•
Communication-optimal parallel algorithm for strassen's matrix multiplication

[...]

Grey Ballard1, James Demmel1, Olga Holtz1, Benjamin Lipshitz1, Oded Schwartz1 •
University of California, Berkeley1
25 Jun 2012
TL;DR: In this article, a new parallel algorithm based on Strassen's fast matrix multiplication algorithm is presented, which is communication-optimal and exhibits perfect strong scaling within the maximum possible range.
Abstract: Parallel matrix multiplication is one of the most studied fundamental problems in distributed and high performance computing. We obtain a new parallel algorithm that is based on Strassen's fast matrix multiplication and minimizes communication. The algorithm outperforms all known parallel matrix multiplication algorithms, classical and Strassen-based, both asymptotically and in practice. A critical bottleneck in parallelizing Strassen's algorithm is the communication between the processors. Ballard, Demmel, Holtz, and Schwartz (SPAA '11) prove lower bounds on these communication costs, using expansion properties of the underlying computation graph. Our algorithm matches these lower bounds, and so is communication-optimal. It exhibits perfect strong scaling within the maximum possible range.Benchmarking our implementation on a Cray XT4, we obtain speedups over classical and Strassen-based algorithms ranging from 24% to 184% for a fixed matrix dimension n=94080, where the number of processors ranges from 49 to 7203.Our parallelization approach generalizes to other fast matrix multiplication algorithms.

124 citations

Proceedings Article•10.1145/2312005.2312021•
Brief announcement: strong scaling of matrix multiplication algorithms and memory-independent communication lower bounds

[...]

Grey Ballard1, James Demmel1, Olga Holtz1, Benjamin Lipshitz1, Oded Schwartz1 •
University of California, Berkeley1
25 Jun 2012
TL;DR: A memory-independent communication cost lower bound is obtained on classical and Strassen-based distributed-memory matrix multiplication algorithms that imply that no classical or Strassan-based parallel matrix multiplication algorithm can strongly scale perfectly beyond the ranges already attained by the two parallel algorithms.
Abstract: A parallel algorithm has perfect strong scaling if its running time on $P$ processors is linear in $1/P$, including all communication costs. Distributed-memory parallel algorithms for matrix multiplication with perfect strong scaling have only recently been found. One is based on classical matrix multiplication (Solomonik and Demmel, 2011), and one is based on Strassen's fast matrix multiplication (Ballard, Demmel, Holtz, Lipshitz, and Schwartz, 2012). Both algorithms scale perfectly, but only up to some number of processors where the inter-processor communication no longer scales. We obtain a memory-independent communication cost lower bound on classical and Strassen-based distributed-memory matrix multiplication algorithms. These bounds imply that no classical or Strassen-based parallel matrix multiplication algorithm can strongly scale perfectly beyond the ranges already attained by the two parallel algorithms mentioned above. The memory-independent bounds and the strong scaling bounds generalize to other algorithms.

72 citations

Posted Content•
Strong Scaling of Matrix Multiplication Algorithms and Memory-Independent Communication Lower Bounds

[...]

Grey Ballard1, James Demmel1, Olga Holtz1, Benjamin Lipshitz1, Oded Schwartz1 •
University of California, Berkeley1
14 Feb 2012-arXiv: Data Structures and Algorithms
TL;DR: In this article, the authors obtained a memory-independent communication cost lower bound on classical and Strassen-based distributed-memory matrix multiplication algorithms, which implies that no classical or fast matrix multiplication algorithm can strongly scale perfectly beyond the ranges already attained by the two parallel algorithms mentioned above.
Abstract: A parallel algorithm has perfect strong scaling if its running time on P processors is linear in 1/P, including all communication costs. Distributed-memory parallel algorithms for matrix multiplication with perfect strong scaling have only recently been found. One is based on classical matrix multiplication (Solomonik and Demmel, 2011), and one is based on Strassen's fast matrix multiplication (Ballard, Demmel, Holtz, Lipshitz, and Schwartz, 2012). Both algorithms scale perfectly, but only up to some number of processors where the inter-processor communication no longer scales. We obtain a memory-independent communication cost lower bound on classical and Strassen-based distributed-memory matrix multiplication algorithms. These bounds imply that no classical or Strassen-based parallel matrix multiplication algorithm can strongly scale perfectly beyond the ranges already attained by the two parallel algorithms mentioned above. The memory-independent bounds and the strong scaling bounds generalize to other algorithms.

19 citations

Comparative Study of Strassen's Matrix Multiplication Algorithm

[...]

Juby Mathew, R. Vijaya kumar
1 Jan 2012
TL;DR: The overall finding is that the Strassen’s algorithm is more efficient than conventional algorithm on large size of matrices, however, in scientific computing, memory has to be considered.
Abstract: The main focus of this paper is to compare the execution time complexity and space complexity between Strassen’s algorithm and the conventional algorithm for matrix multiplication. The aim is to design a program, which generates two matrices with various dimensions, and multiplies the two matrices using both the Strassen’s algorithm and the conventional algorithm. The execution time of each algorithm is recorded to evaluate the performance of each algorithm. The programming language used this project is Java. Some of the main achievements in this project are, successfully divide matrices into blocks, the Strassen’s algorithm was applied to each blocks recursively, and the level of recursion was controlled. The overall finding is that the Strassen’s algorithm is more efficient than conventional algorithm on large size of matrices. However, in scientific computing, memory has to be considered. The results show that Strassen’s algorithm needs more memory allocations than the conventional algorithm, due to the fact in design that more arrays need to be created.

5 citations

Communication-Optimal Parallel Algorithm for Strassen's Matrix Multiplication Regular Submission

[...]

Grey Ballard, James Demmel, Olga Holtz, Benjamin Lipshitz, Oded Schwartz 
1 Jan 2012
TL;DR: In this paper, a new parallel algorithm based on Strassen's fast matrix multiplication and minimizing communication is presented, which outperforms all known parallel matrix multiplication algorithms, both asymptotically and in practice.
Abstract: Parallel matrix multiplication is one of the most studied fundamental problems in distributed and high performance computing. We obtain a new parallel algorithm that is based on Strassen’s fast matrix multiplication and minimizes communication. The algorithm outperforms all known parallel matrix multiplication algorithms, classical and Strassen-based, both asymptotically and in practice. A critical bottleneck in parallelizing Strassen’s algorithm is the communication between the processors. Ballard, Demmel, Holtz, and Schwartz (SPAA’11) prove lower bounds on these communication costs, using expansion properties of the underlying computation graph. Our algorithm matches these lower bounds, and so is communication-optimal. It exhibits perfect strong scaling within the maximum possible range. Research supported by Microsoft (Award #024263) and Intel (Award #024894) funding and by matching funding by U.C. Discovery (Award #DIG07-10227). Additional support comes from Par Lab aliates

Tools

SciSpace AgentBiomedical AgentSciSpace RecruitSciSpace for EnterpriseAgent GalleryChat with PDFLiterature ReviewAI WriterFind TopicsParaphraserCitation GeneratorExtract DataAI DetectorAI Humanizer

Learn

ResourcesCompareGuidesLive Workshops

SciSpace

CareersSupportBrowse PapersPricingSciSpace Affiliate ProgramCancellation & Refund PolicyTermsPrivacyData Sources

Directories

PapersTopicsJournalsAuthorsConferencesInstitutionsPublishersCitation StylesWriting templates

Extension & Apps

SciSpace Chrome ExtensionSciSpace Mobile App

Contact

[email protected]
SciSpace

© 2026 | PubGenius Inc. | Suite # 217 691 S Milpitas Blvd Milpitas CA 95035, USA

soc2