User and Reference Guide for the Intel® C++ Compiler 14.0
This topic only applies to IA-32 architecture targeting Intel® Graphics Technology. Intel® Graphics Technology is a preview feature.
Some algorithms require in-order iterative execution of parallelized code. Consider the outer-most for loop in the following example:
for (int k = 0; k < numNodes; k++) {
#pragma offload target(gfx) …
#pragma parallel_loop
for (unsigned int y = 0; y < numNodes; ++y) {
…
}
}
The offload block is executed numNodes times and there is no CPU-side execution dependent on results of any intermediate iteration of the k loop. The compiler and runtime can further optimize such patterns via hoisting loop invariant offload logic out of the loop, as in the k loop in the example, and enqueueing multiple offload tasks without waiting for each separate offload task to complete, which substantially speeds up such patterns. The requirements are as follows:
The #pragma offload block should be manually or automatically inlined into the loop.
Other than incrementing the loop, the loop should not contain any host-side computation.
|
Intel's compilers may or may not optimize to the same degree for non-Intel microprocessors for optimizations that are not unique to Intel microprocessors. These optimizations include SSE2, SSE3, and SSSE3 instruction sets and other optimizations. Intel does not guarantee the availability, functionality, or effectiveness of any optimization on microprocessors not manufactured by Intel. Microprocessor-dependent optimizations in this product are intended for use with Intel microprocessors. Certain optimizations not specific to Intel microarchitecture are reserved for Intel microprocessors. Please refer to the applicable product User and Reference Guides for more information regarding the specific instruction sets covered by this notice. Notice revision #20110804 |