I am comparing CPU/GPU libraries for accelerating computations. Each cell contains a MVP code setup for each library, which is timed consecutively for three things:
1. Preparation: simulating a (flat) matrix by ingesting a pre-filled array with random floating point values from the start and from the end, resulting in two arrays A & B
2. Calculation: running a basic computation per value, in this case ( A[i] ** B[i] ) / B[i]
3. Returning values: finding and displaying the max value of the resulting array.
The first and last step are included to time GPU I/O to see how quick the calculated values are available again for javascript use. In addition, it's a crude measure of accuracy to see how far GPU (and sometimes, webworkers; see the 'thread' method) calculations diverge from native methods.
The last method ('supgpu') is partly designed by me and clearly the quickest for large arrays, in some cases providing a 10x speed up over (for instance) Tensorflow.js. It makes uses of a WebGL2 wrapper that I have tuned to enable parallel computation, currently making use of 256 GPU threads.
I'd like to better understand how one library in particular measures up. GPU.js. this is because I'm the maintainer of it, and I worked deeply with it over the years so that it runs at a pretty good clip. I'm not sure I clearly understand if it is tested and how can you help me to see that?
Sure thing, it was your library that set me off exploring this space! It's in the wgpu.js cell, simply timed with performance.now(). Some more accurate measurements are definitely in the pipeline, as well as exploring alternative algo designs using GPU.js (for instance, finding the max value on the GPU alone).
This first iteration was as much a learning experience as providing MVPs for javascript <-> GPU I/O.
One of the biggest things I'd like to understand is that the initial setup and compile times vs runtime are going to have drastically different results. Another is going to be kernel to kernel computational time. If you use pipeline mode (where one kernel sends values to another), you'll see MASSIVE speedups, because the textures don't need to sync the GPU's processes, and as well there is no round-trip (picture CPU -> GPU -> CPU -> GPU -> CPU) trip per texture. However, I do want to impart my amazement for providing a fantastic playground and all the hard work that went with it. Thank you for the comparisons across the board!
Thank you, glad that you found it helpful:)
Re: your first point, I've already used the pipeline feature between the operation step (ops kernel) and the max finding step (batch kernel), the latter reducing the initial array to a much smaller one of potential maximum values. This is then synced back to the CPU after which the final max is calculated (however, some faster 'reduction' to a single value on the GPU might be possible, but I haven't been able to figure this out for GPU.js just yet).
Re: your second question, the next iteration of this notebook will execute multiple runs per library with more gracious startup/cooldown periods in between runs, as well as more harsh memory/context resets in between runs, comparing both GPU performance timing and IO performance (also aiming to include regl and swissgl).
Right now I was mainly interested in including IO in the benchmarks, as I wanted to investigate how quickly different libraries and methods could respond to fresh data and make the results available for further synchronous compute steps. I would say GPU.js outperforms Tensorflow.js for large arrays for general calculations, with hamster.js (the 'thread' cell) providing a competitive CPU solution by parallelizing computions via workers (8 worker 'threads' seem to be upper end of providing benefits on my system) while remaining easy to use.