User and Reference Guide for the Intel® C++ Compiler 14.0
Causes the loop statements following the pragma to run in parallel when the statements are executed on the target. This pragma is deprecated. This pragma only applies to Intel® Graphics Technology.
#pragma parallel_loop [clause] |
clause |
Can be one of the following:
If you do not specify a clause or you use #pragma parallel_loop collapse (1), then only the outer for loop will run in parallel. |
Use this pragma along with the offload directive to parallelize nested loops. Use the collapse clause when you want to execute a perfect loop nest in parallel.
If the target is not available, the pragma is ignored and the CPU will execute the loop iterations serially. To execute the loop iterations in parallel on the CPU, use another method, such as using the cilk_for keyword.
The loop iterations following this pragma must follow the same rules as the cilk_for loop.
The following pragmas can be in loops following this pragma:
prefetch/noprefetch
simd
The following keywords and pragmas cannot be in the loops following this pragma:
cilk_for
cilk_spawn
cilk_sync
OpenMP* pragmas
The following conditions will cause an error:
The statement following the pragma is not a for loop.
An incorrect array notation expression.
A loop that does not follow the rules for a loop using the cilk_for keyword.
The number of for loops is less than the number specified in the collapse (n) clause.
Example: Isolating loop iterations to make a perfect loop next using the collapse clause for execution on the target |
|---|
bool MatmultLocalsAN::execute_offload(int do_offload) {
int m = m_height, n = m_width, k = m_common;
float (* A)[k] = (float (*)[])m_matA;
float (* B)[n] = (float (*)[])m_matB;
float (* C)[n] = (float (*)[])m_matC;
// Execute the following on the specified target
#pragma offload target(gfx) if (do_offload) \
pin(A: length(m*k)), pin(B: length(k*n)), pin(C: length(m*n))
// The first two loop iterations are an example of a perfect loop nest.
// The remaining loop iterations call other functions outside of the loop.
// Including the remaining loop iterationss will result in an imperfect
// loop nest. To execute the parts of the loop that are in a perfect loop
// nest, use the collapse clause to specify the loop iterations to
// execute in parallel.
#pragma parallel_loop collapse(2)
for (int r = 0; r < m; r += TILE_m) {
for (int c = 0; c < n; c += TILE_n) {
float atile[TILE_m][TILE_k], btile[TILE_n], ctile[TILE_m][TILE_n];
#pragma unroll
ctile[:][:] = 0.0f;
// Including the remaining loop iterations will make an
// imperfect loop nest.
for (int t = 0; t < k; t += TILE_k) {
#pragma unroll
atile[:][:] = A[r:TILE_m][t:TILE_k];
#pragma unroll
for (int rc = 0; rc < TILE_k; rc++) {
btile[:] = B[t+rc][c:TILE_n];
#pragma unroll
for (int rt = 0; rt < TILE_m; rt++) {
ctile[rt][:] += atile[rt][rc] * btile[:];
}
}
}
#pragma unroll
C[r:TILE_m][c:TILE_n] = ctile[:][:];
}
}
return true;
} |