User and Reference Guide for the Intel® C++ Compiler 14.0
This topic only applies to IA-32 architecture targeting Intel® Graphics Technology. Intel® Graphics Technology is a preview feature.
Intel® Graphics Technology uses multiple SIMD execution units, each capable of simultaneously running several threads by sharing part of its resources. Running many simultaneous threads on each execution unit helps hide the latency of some operations. Significant increase of the number of threads adds to the cost of offloading code, but as a rule, a reasonably large number of threads delivers the best results.
The maximum number of target threads to parallelize loop nests annotated with #pragma parallel_loop can be controlled via the GFX_MAX_THREAD_COUNT environment variable set before a heterogeneous application is started. The default value is -1, which means the runtime automatically determines the maximum thread count. The real thread count for a particular offload execution can be lower than the maximum, and is determined by the offload runtime, depending on the real iteration space for that execution.
The iteration space of the top-most loop within a #pragma parallel_loop section may be insufficient to fully leverage Intel® Graphics Technology parallelism, especially if this loop is also vectorized. The number of iterations of the outermost loop of the offloaded loop nest may be lower than the number of target threads delivering the best performance. However, explicit collapsing of a loop nest into a single loop in source code can be inconvenient. You can define collapse(<N>) in a parallel_loop pragma to parallelize the larger iteration space of N perfectly nested loops under the offload pragma.
The compiler virtually ignores #pragma parallel_loop for CPU code. The corresponding CPU version of the loop nest is not parallelized and runs in a single CPU thread.
|
Intel's compilers may or may not optimize to the same degree for non-Intel microprocessors for optimizations that are not unique to Intel microprocessors. These optimizations include SSE2, SSE3, and SSSE3 instruction sets and other optimizations. Intel does not guarantee the availability, functionality, or effectiveness of any optimization on microprocessors not manufactured by Intel. Microprocessor-dependent optimizations in this product are intended for use with Intel microprocessors. Certain optimizations not specific to Intel microarchitecture are reserved for Intel microprocessors. Please refer to the applicable product User and Reference Guides for more information regarding the specific instruction sets covered by this notice. Notice revision #20110804 |