User and Reference Guide for the Intel® C++ Compiler 14.0

Initiating an Offload on Intel® Graphics Technology

This topic only applies to IA-32 architecture targeting Intel® Graphics Technology. Intel® Graphics Technology is a preview feature.

The code inside loops or loop nests following #pragma offload target(gfx) and in functions qualified with __declspec(target(gfx)) is compiled to both the target and the CPU. The qualified functions can be called for execution on the target from other code executed on the target. The code is executed on the target if the target is present on the system and the if clause is evaluated to true, otherwise it is executed on the CPU (see examples below).

You can place #pragma offload target(gfx) only before a perfect loop nest explicitly marked as parallel by #pragma parallel_loop.

#pragma offload can contain the following clauses when programming for Intel® Graphics Technology:

Note

Using pin substantially reduces the cost of offloading because instead of copying data to or from memory accessible by the target, the pin clause organizes sharing the same memory area between the CPU and the target, which is much faster. For kernels that perform substantial work on a relatively small data size, such as O(N2)), this optimization is not important.

#pragma parallel_loop [collapse(n)] indicates that the underlying perfect nest of one or more loops will be parallelized over the target's threads.

Although by default the compiler builds an application that runs on both the host CPU and target, you can also build the same source code to run on just the CPU, using the Qoffload- compiler option.

Example: Offloading to the Target

unsigned parArrayRHist[256][256],
     parArrayGHist[256][256], parArrayBHist[256][256];

#pragma offload target(gfx) if (do_offload) \
     pin(inputImage: length(imageSize)) \
     out(parArrayRHist, parArrayGHist, parArrayBHist)
#pragma parallel_loop
     for (int ichunk = 0; ichunk < chunkCount; ichunk++){
          …
     }

In the example above, the generated CPU code and the runtime do the following:

Example: Offloading Using parallel_loop collapse

float (* A)[k] = (float (*)[k])matA;
float (* B)[n] = (float (*)[n])matB;
float (* C)[n] = (float (*)[n])matC;

#pragma offload target(gfx) if (do_offload) \
     pin(A: length(m*k)), pin(B: length(k*n)), pin(C: length(m*n))
#pragma parallel_loop collapse(2)
     for (int r = 0; r < m; r += TILE_m) {
          for (int c = 0; c < n; c += TILE_n) {
               …
          }
     }

In the example above:

Optimization Notice

Intel's compilers may or may not optimize to the same degree for non-Intel microprocessors for optimizations that are not unique to Intel microprocessors. These optimizations include SSE2, SSE3, and SSSE3 instruction sets and other optimizations. Intel does not guarantee the availability, functionality, or effectiveness of any optimization on microprocessors not manufactured by Intel. Microprocessor-dependent optimizations in this product are intended for use with Intel microprocessors. Certain optimizations not specific to Intel microarchitecture are reserved for Intel microprocessors. Please refer to the applicable product User and Reference Guides for more information regarding the specific instruction sets covered by this notice.

Notice revision #20110804

See Also


Submit feedback on this help topic