User and Reference Guide for the Intel® C++ Compiler 14.0
SIMD-enabled functions (formerly called elemental functions) are a general language construct to express a data parallel algorithm. A SIMD-enabled function is written as a regular C/C++ function, and the algorithm within describes the operation on one element, using scalar syntax. The function can then be called as a regular C/C++ function to operate on an single element or it can be called in a data parallel context, providing many elements to operate on. In Intel® Cilk™ Plus, the data parallel context is provided as an array.
If you are using SIMD-enabled functions and need to link a compiler object file with an object file from a previous version of the compiler (for example, 13.1), you need to use the [Q]vecabi compiler option, specifying the legacy keyword. The default (compat) behavior is compatible with GCC vector function support of both Intel® Cilk™ Plus and OpenMP* 4.0.
When you write a SIMD-enabled function, the compiler generates a short vector form of the function, which can perform your function's operation on multiple arguments in a single invocation. The short vector version may be able to perform multiple operations as fast as the regular implementation performs a single one by utilizing the vector ISA in the CPU. In addition, when invoked from a cilk_for or pragma omp construct, the compiler may assign different copies of the SIMD-enabled functions to different threads (or workers), executing them concurrently. The end result is that your data parallel operation executes on the CPU utilizing both the parallelism available in the multiple cores and the parallelism available in the vector ISA.
If the short vector function is called inside a parallel loop, a cilk_for loop or an auto-parallelized loop that is vectorized, you can achieve both vector-level and thread-level parallelism.
In order for the compiler to generate the short vector function, you need to provide an indication in your code.
Windows* OS:
Use the __declspec(vector (clauses)) declaration, as follows:
__declspec(vector (clauses)) return_type simd_enabled_function_name(arguments)
Linux* OS and OS X*:
Use the __attribute__((vector (clauses))) declaration, as follows:
__attribute__((vector (clauses))) return_type simd_enabled_function_name(arguments)
The clauses for the vector declaration take the following values:
Write the code inside your function using existing C/C++ syntax.
Typically, the invocation of a SIMD-enabled function provides arrays wherever scalar arguments are specified as formal parameters. Use the array notation syntax available in Intel® Cilk™ Plus to provide the arrays succinctly. Alternatively, you can invoke the function from a _Cilk_for loop.
The following examples show how to use SIMD-enabled functions to add two large arrays and store the result in a third array, taking advantage of the parallelism available in both the cores and the vectors in the CPU:
Windows* OS:
|
Example |
|---|
//declaring the function body __declspec((vector)) double ef_add (double x, double y){
return x + y; } //invoking the function using array notations a[:] = ef_add(b[:],c[:]); //operates on the whole extent of the arrays a,b,c a[0:n:s] = ef_add(b[0:n:s],c[0:n:s]); //use the full array notation construct to also specify n as an extend and s as a stride //Use the _Cilk_for construct to invoke the SIMD-enabled function in a data parallel context _Cilk_for (j = 0; j < n; ++j) {
a[j] = ef_add(b[j],c[j]) } |
Linux* OS and OS X*:
|
Example |
|---|
//declaring the function body __attribute__((vector)) double ef_add (double x, double y){
return x + y; } //invoking the function using array notations a[:] = ef_add(b[:],c[:]); //operates on the whole extent of the arrays a,b,c a[0:n:s] = ef_add(b[0:n:s],c[0:n:s]); //use the full array notation construct to also specify n as an extend and s as a stride //Use the _Cilk_for construct to invoke the SIMD-enabled function in a data parallel context _Cilk_for (j = 0; j < n; ++j) {
a[j] = ef_add(b[j],c[j]) } |
Only the calling code using the _Cilk_for calling syntax is able to use all available parallelism. The array notation syntax, as well as calling the SIMD-enabled function from the regular for loop, results in invoking the short vector function in each iteration and utilizing the vector parallelism but the invocation is done in a serial loop, without utilizing multiple cores.
Limitations
The following language constructs are disallowed within SIMD-enabled functions: