Files
OpenCL-examples/example01/README.md
2015-06-23 16:52:01 -04:00

930 B

Example 01

This example compares the timings of adding vectors on the CPU versus adding vectors on the GPU, the latter of which has different implementations.

Compiling

clang++ -std=c++0x -framework OpenCL main.cpp -o main.out

To ignore deprecation warnings, add the flag -Wno-deprecated-declarations.

Run from this directory, as a relative path is used for the OpenCL header file (for now).

About

The code runs the following implementations of adding large vectors (131072 elements; 8 * 32 * 512). The vectors are added together 10000 times.

  • CPU
  • GPU, where 1024 threads are spawned and each thread thus gets 128 elements to calculate; there are two implementations of this:
    • (Version 1) each thread gets 128 sequential elements (thread 0 gets 0-127, 1 gets 128-255, ...)
    • (Version 2) each thread gets 128 elements, but coalescing happens (thread 0 gets 0,128,256..., thread 1 gets 1,129,257...)