Skip to content

Latest commit

 

History

124 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

PieRank: Embedded Large-Scale Sparse Matrix Processing

License

PieRank is licensed under the MIT License.

Introduction

I introduce PieRank, a library aimed at embedded large-scale sparse matrix processing. It enjoys significant advantages over previous state-of-the-art for handling big, sparse data sets in scalability, speed, and usability. My approach is also more general, allowing a wider variety of applications beyond graphs. Inspired by Raspberry Pi(e) for its low cost, I name the library PieRank with the goal of dramatically reducing the hardware requirements for cutting-edge sparse matrix processing. My experiments show PieRank outperforms GraphChi by a wide margin, making it possible for a single embedded device to meet or exceed the scalability and performance of server-class computers on big matrices such as AGATHA 2015, a deep-learning network with 5.8B parameters as the largest instance in the popular SuiteSparse Matrix Collection.

Available as an open-source project on GitHub under the permissive MIT license, PieRank is implemented in plain C++ and compiles and runs on Linux, Mac OS, and Raspbian.

Publication

PieRank has been accepted as a regular paper to the 2023 IEEE International Conference on Big Data. PDF of the paper can be viewed here. List of accepted papers can be found here.

Features

  • Utilizes FlexArray, a late-binding random-access container for flexible data encoding. FlexArray supports high-speed mapping from a nominal integer type to a low-level binary representation that is more memory-efficient for the data.
  • Employs asynchronous, multi-threaded parallel processing. Using a set of disjoint integer-based intervals called ranges, a user can write a PieRank Program (PRP) to process the matrix in parallel.
  • Applies an efficient delta compression technique, called Sketch Compression, for (sorted) integer arrays with guaranteed worst-case constant O(1) decode time, an improvement of the conventional decoding techniques that often require multiple decode rounds to reach the target entry.
  • Provides built-in support for memory-mappable sparse matrices for RAM-oblivious sparse matrix processing
  • Offers a library for sparse matrix processing
  • Contains sample applications PageRank and connected components.
  • Can run graphs with billions of edges, with linear scalability, on an embedded device such as a Raspberry Pi.
  • Compiles and runs on Linux, Mac OS, and Raspbian.

Getting Started

  1. sudo apt install cmake libgflags-dev libgoogle-glog-dev pkg-config
  2. git clone https://github.com/mmichellezhou/pierank.git && cd pierank
  3. git submodule init && git submodule update
  4. mkdir release && cd release
  5. cmake -DCMAKE_BUILD_TYPE=Release ..
  6. make -j64

Experimental Results

For all experiments, I used a Raspberry Pi 4 Model B with 8GB of RAM and a 960GB ADATA SC680 USB 3.2 Gen 2 external SSD with a maximum read (write) speed of 530 (460) MB/s. The OS is 64-bit Raspbian Buster. GCC 8.3.0 was used to generate all binaries, including GraphChi.

$p$# name #rows nnz
1 web-Stanford 281,903 2,312,497
2 amazon-2008 735,323 5,158,388
3 Stanford_Berkeley 683,446 7,583,376
4 333SP 3,712,815 11,108,633
5 cit-Patents 3,774,768 16,518,948
6 wikipedia-20070206 3,566,907 45,030,389
7 soc-LiveJournal1 4,847,571 68,993,773
8 uk-2005 39,459,925 936,364,282
9 twitter-2010 41,652,230 1,468,365,182
10 AGATHA_2015 183,964,077 5,794,362,982

The above table shows the sparse matrices used in the experiments. They are from the SuiteSparse Matrix Collection website, including AGATHA 2015 as the largest instance in the collection. It is a deep-learning network with 5.8B parameters (or 11.6B if one counts each undirected edge as two parameters). Instead of referring to matrices by their names, the rest of the tables all use their problem number mentioned above (e.g., $p$#10 refers to AGATHA_2015).

$p$# mtx size prm size ratio $s^*$ cb rb
1 30,579,901 7,502,539 4.1x 11 2 3
2 70,636,197 16,946,668 4.2x 13 2 3
3 102,294,093 24,138,518 4.2x 8 2 3
4 171,065,574 37,967,050 4.5x 5 1 3
5 261,658,301 57,075,516 4.6x 13 2 3
6 642,663,421 145,792,458 4.4x 16 3 3
7 1,011,609,587 217,282,551 4.7x 6 2 3
8 16,451,551,195 3,903,296,960 4.2x 0 4 4
9 26,141,060,759 5,999,068,381 4.4x 9 3 4
10 104,523,784,790 23,773,593,875 4.4x 5 3 4

The table above is a comparison of encoding sizes, including matrix market file size (mtx size), PieRank matrix file size (prm size), compression ratio (ratio), optimal number of sketch bits (s*), number of bytes for each element of COL_INDEX (cb) and ROW_INDEX (rb) in PieRank’s binary CSC format.

$p$# mtx read (sec) prm read (sec) ratio prm mmap (sec)
1 1.41 0.16 8.8x 0.0006
2 3.52 0.17 20.4x 0.0024
3 4.93 0.14 34.3x 0.0027
4 7.06 0.60 11.7x 0.0090
5 10.61 1.43 7.4x 0.0027
6 28.29 3.97 7.1x 0.0028
7 52.75 3.91 13.5x 0.0062
8 721.53 17.50 41.2x 0.0025
9 896.75 83.93 10.7x 0.0077
10 3,524.49 727.82 4.8x 0.1309

The above table is a comparison of matrix read time in seconds, including matrix market read time, PRM read time, speedup ratio, and PRM memory-map time.

$p$# $GC$ (sec) $PR_\textrm{RAM}$ (sec) $PR_\textrm{RO}$ (sec) ratio
1 0.64 0.06 0.06 10.7x
2 1.23 0.06 0.07 17.6x
3 1.20 0.05 0.05 24.0x
4 3.74 0.21 0.22 17.0x
5 6.10 0.52 0.52 11.7x
6 12.04 1.42 1.43 8.4x
7 15.42 1.57 1.59 9.7x
8 152.03 3.62 4.34 35.0x
9 386.75 44.47 44.60 8.7x
10 2,207.72 MEM 328.11 6.7x

The table above is a comparison of average PageRank iteration time, including GraphChi (GC), in-RAM PieRank (PRRAM), RAM-oblivious PieRank (PRRO), and speedup ratio for PRRO.

$p$# $GC$ (sec) $PR_\textrm{RAM}$ (sec) $PR_\textrm{RO}$ (sec) ratio
1 2.19 0.06 0.06 36.5x
2 3.29 0.06 0.07 47.0x
3 3.88 0.05 0.05 77.6x
4 14.22 0.25 0.25 56.9x
5 12.60 0.60 0.61 20.7x
6 27.54 1.57 1.62 17.0x
7 35.43 1.49 1.58 22.4x
8 195.48 4.36 6.12 31.9x
9 812.55 47.09 46.50 17.5x
10 MEM MEM 350.06 N/A

The above table is a comparison of average connected-components iteration time, including GraphChi (GC), in-RAM PieRank (PRRAM), RAM-oblivious PieRank (PRRO), and speedup ratio for PRRO.

How to Reproduce Experimental Results

Please help yourself to the scripts directory which contains all of the Bash scripts I use to run PieRank, including:

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages