This page is the linear start for the public Vulkan GPU path in cthreads. Vulkan is a cross-vendor GPU API. cthreads lowers @Gpu kernels to SPIR-V (an intermediate shader form Vulkan compute pipelines consume).
You should already know the CPU basics (types, @Thread, pack / writeback). If not, read the CPU quickstart first.
Install the GPU package (do not also install plain cthreads):
pip install cthreads-gpuMore install detail: install.md.
GitHub renders Mermaid diagrams. Clickable links inside Mermaid are unreliable on GitHub. Use the diagram for shape. Use the numbered list under it for real links.
flowchart TD
A[CPU quickstart]
B[This GPU quickstart]
C[gpu/concepts.md]
D[kernels.md]
E[indexes.md]
F[launch.md]
G[arena.md]
H[sync.md]
I[best_practices.md]
J[examples.md]
K[errors.md]
A --> B
B --> C
C --> D
C --> E
C --> F
F --> G
C --> H
C --> I
C --> J
C --> K
- CPU quickstart - types, compile, launch on CPU
- This page - probe, first
@Gpukernel, join, arena sketch - concepts.md - host vs device, workgroups, writeback
- Then deepen as needed:
- Hub overview: README.md
from cthreads import gpu
# Soft probe. False means this install or machine cannot run @Gpu.
print(gpu.available())
if gpu.available():
print(gpu.device_name())If available() is False, stop and read errors.md. Decorating @Gpu or calling gpu() raises when the path is not usable.
GPU kernels are narrower than CPU kernels today.
Allowed arguments:
int,float,boollistof those scalars
Returns must be -> None. Results come back by writing into list arguments (writeback on join). Locals still need annotated assignment (i: int = ...).
There is no @Threadable, str, or dict on the public GPU path yet.
@Gpu marks a function for compilation to a compute shader. Calling the function as ordinary Python still runs the Python body. Device execution happens only through gpu(fn, ...).
from cthreads.gpu import Gpu, GlobalIdx, gpu
@Gpu
def saxpy(n: int, a: float, x: list[float], y: list[float]) -> None:
# GlobalIdx.x is this invocation's 1D index in the launch grid.
i: int = GlobalIdx.x
# Dispatch may pad past n. Always guard.
if i >= n:
return
y[i] = a * x[i] + y[i]GlobalIdx is an index builtin. For 1D element-wise work, .x is the flat index you usually want. More index forms: indexes.md.
from cthreads.gpu import Gpu, GlobalIdx, gpu
@Gpu
def saxpy(n: int, a: float, x: list[float], y: list[float]) -> None:
i: int = GlobalIdx.x
if i >= n:
return
y[i] = a * x[i] + y[i]
x: list[float] = [1.0, 2.0, 3.0, 4.0]
y: list[float] = [10.0, 20.0, 30.0, 40.0]
n: int = len(x)
a: float = 2.0
# First gpu() in a process may compile SPIR-V (cached afterward).
job = gpu(saxpy, n, a, x, y)
job.join()
print(y) # [12.0, 24.0, 36.0, 48.0]What gpu() does on first use:
- Checks that Vulkan is usable.
- Compiles registered
@Gpufunctions to SPIR-V (then caches them). - Uploads arguments and records a compute dispatch.
- Returns a GpuJob that is already started.
join() waits for the GPU fence. By default it downloads list arguments into the same Python list objects you passed. job.result() is always None on this path.
When you launch many kernels over the same lists, bind them once with GpuArena. That reduces host traffic between passes.
from cthreads.gpu import GpuArena, gpu
with GpuArena() as arena:
arena.bind(x=x, y=y)
for _ in range(100):
# download=False skips writeback until you ask for it.
gpu(saxpy, n, a, x, y).join(download=False)
arena.sync() # copy device lists back into x and y for PythonFull guide: arena.md.
GlobalIdx, ThreadIdx, BlockIdx, and related builtins describe where an invocation sits in the grid. Dispatch rounds up to a workgroup size. Always bounds-check with n.
Guide: indexes.md.
prepare, compile caching, GpuJob.join(download=...), and error cases.
Guide: launch.md.
__sync_threads() and Barrier.arrive_and_wait() meet inside one workgroup. They are not a whole-grid barrier. Workgroup shared memory arrays are not in the public dialect yet. Prefer multiple gpu() launches for global phases.
Guide: sync.md.
Patterns that stay correct: best_practices.md. More samples: examples.md. Troubleshooting: errors.md.
| Mistake | What happens | Fix |
|---|---|---|
Missing if i >= n: return |
Out-of-range writes on padded invocations | Always bounds-check |
Expecting job.result() to hold an array |
Always None |
Read list args after join |
Using dict / @Threadable / str args |
Type error at decorate or compile | Scalars and scalar lists only |
Calling __sync_threads() from plain Python |
Runtime error | Only inside @Gpu bodies |
| Assuming a barrier syncs the whole array | Only one workgroup waits | Use multiple host launches for global phases |
pip install cthreads then expecting @Gpu |
GPU not built into that wheel | Install cthreads-gpu alone |
Read concepts.md for host vs device vocabulary. Then skim kernels.md before you write larger kernels.