User and Reference Guide for the Intel® C++ Compiler 14.0

parallel_loop

Causes the loop statements following the pragma to run in parallel when the statements are executed on the target. This pragma is deprecated. This pragma only applies to Intel® Graphics Technology.

Syntax

#pragma parallel_loop [clause]

Arguments

clause

Can be one of the following:

collapse

The compiler determines the maximum number of for loops in a perfect loop nest to run in parallel.

collapse (n)

Runs in parallel the n number of for loops in a perfect loop nest, starting with the outer loop. The specified number n is a positive integer and must be equal to or greater than the number of for loops following this pragma.

If you do not specify a clause or you use #pragma parallel_loop collapse (1), then only the outer for loop will run in parallel.

Description

Note

This pragma is deprecated. The replacement options are cilk_for, cilk_spawn, or cilk_sync

Use this pragma along with the offload directive to parallelize nested loops. Use the collapse clause when you want to execute a perfect loop nest in parallel.

If the target is not available, the pragma is ignored and the CPU will execute the loop iterations serially. To execute the loop iterations in parallel on the CPU, use another method, such as using the cilk_for keyword.

The loop iterations following this pragma must follow the same rules as the cilk_for loop.

The following pragmas can be in loops following this pragma:

The following keywords and pragmas cannot be in the loops following this pragma:

The following conditions will cause an error:

Example: Isolating loop iterations to make a perfect loop next using the collapse clause for execution on the target

bool MatmultLocalsAN::execute_offload(int do_offload) {
	
	int m = m_height, n = m_width, k = m_common;
	float (* A)[k] = (float (*)[])m_matA;
	float (* B)[n] = (float (*)[])m_matB;
	float (* C)[n] = (float (*)[])m_matC;
	
	// Execute the following on the specified target
	#pragma offload target(gfx) if (do_offload) \
	pin(A: length(m*k)), pin(B: length(k*n)), pin(C: length(m*n))
	
	// The first two loop iterations are an example of a perfect loop nest. 
	// The remaining loop iterations call other functions outside of the loop.
	// Including the remaining loop iterationss will result in an imperfect 
	// loop nest. To execute the parts of the loop that are in a perfect loop
	// nest, use the collapse clause to specify the loop iterations to 
 // execute in parallel.
	#pragma parallel_loop collapse(2)
	for (int r = 0; r < m; r += TILE_m) {
		for (int c = 0; c < n; c += TILE_n) {
			float atile[TILE_m][TILE_k], btile[TILE_n], ctile[TILE_m][TILE_n];
			
			#pragma unroll
			ctile[:][:] = 0.0f;

			// Including the remaining loop iterations will make an
			// imperfect loop nest.
			for (int t = 0; t < k; t += TILE_k) {
				#pragma unroll
				atile[:][:] = A[r:TILE_m][t:TILE_k];
				#pragma unroll

				for (int rc = 0; rc < TILE_k; rc++) {
					btile[:] = B[t+rc][c:TILE_n];
					#pragma unroll
					for (int rt = 0; rt < TILE_m; rt++) {
						ctile[rt][:] += atile[rt][rc] * btile[:];
					}
				}
			}

			#pragma unroll
			C[r:TILE_m][c:TILE_n] = ctile[:][:];
		}
	}
	return true;
}

See Also


Submit feedback on this help topic