File: single-work-item-barrier.rst

package info (click to toggle)
swiftlang 6.0.3-2
  • links: PTS, VCS
  • area: main
  • in suites: forky, sid, trixie
  • size: 2,519,992 kB
  • sloc: cpp: 9,107,863; ansic: 2,040,022; asm: 1,135,751; python: 296,500; objc: 82,456; f90: 60,502; lisp: 34,951; pascal: 19,946; sh: 18,133; perl: 7,482; ml: 4,937; javascript: 4,117; makefile: 3,840; awk: 3,535; xml: 914; fortran: 619; cs: 573; ruby: 573
file content (58 lines) | stat: -rw-r--r-- 1,804 bytes parent folder | download | duplicates (21)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
.. title:: clang-tidy - altera-single-work-item-barrier

altera-single-work-item-barrier
===============================

Finds OpenCL kernel functions that call a barrier function but do not call
an ID function (``get_local_id``, ``get_local_id``, ``get_group_id``, or
``get_local_linear_id``).

These kernels may be viable single work-item kernels, but will be forced to
execute as NDRange kernels if using a newer version of the Altera Offline
Compiler (>= v17.01).

If using an older version of the Altera Offline Compiler, these kernel
functions will be treated as single work-item kernels, which could be
inefficient or lead to errors if NDRange semantics were intended.

Based on the `Altera SDK for OpenCL: Best Practices Guide
<https://www.altera.com/en_US/pdfs/literature/hb/opencl-sdk/aocl_optimization_guide.pdf>`_.

Examples:

.. code-block:: c++

  // error: function calls barrier but does not call an ID function.
  void __kernel barrier_no_id(__global int * foo, int size) {
    for (int i = 0; i < 100; i++) {
      foo[i] += 5;
    }
    barrier(CLK_GLOBAL_MEM_FENCE);
  }

  // ok: function calls barrier and an ID function.
  void __kernel barrier_with_id(__global int * foo, int size) {
    for (int i = 0; i < 100; i++) {
      int tid = get_global_id(0);
      foo[tid] += 5;
    }
    barrier(CLK_GLOBAL_MEM_FENCE);
  }

  // ok with AOC Version 17.01: the reqd_work_group_size turns this into
  // an NDRange.
  __attribute__((reqd_work_group_size(2,2,2)))
  void __kernel barrier_with_id(__global int * foo, int size) {
    for (int i = 0; i < 100; i++) {
      foo[tid] += 5;
    }
    barrier(CLK_GLOBAL_MEM_FENCE);
  }

Options
-------

.. option:: AOCVersion

   Defines the version of the Altera Offline Compiler. Defaults to ``1600``
   (corresponding to version 16.00).