diff --git a/.gitignore b/.gitignore
index 06432db..7a4aba2 100644
--- a/.gitignore
+++ b/.gitignore
@@ -62,5 +62,6 @@ libcthreads_kernels.dylib
# <<< cthreads (auto)
todo.md
-# demo artifacts
-*.gif
\ No newline at end of file
+# demo artifacts (allow README showcase gifs)
+*.gif
+!docs/__ressources/*.gif
diff --git a/ReadMe.md b/ReadMe.md
index ec40238..b557388 100644
--- a/ReadMe.md
+++ b/ReadMe.md
@@ -1,11 +1,26 @@

+[](https://colab.research.google.com/github/K-T0BIAS/CThreads/blob/main/colab/Mandelbrot.ipynb)
+
----
-**cthreads** compiles a typed Python subset into C++ so work can run on real OS threads **without the GIL** allowing true concurrency without `multiprocessing`'s process boundaries and pickling tax.
+**cthreads** compiles a typed Python subset into C++ so work can run on real OS threads **without the GIL**, allowing true concurrency without `multiprocessing`'s process boundaries and pickling tax. Optional **Vulkan GPU** kernels (`@Gpu`) ship as a separate install ( **`cthreads-gpu`**) same `import cthreads`, do not install alongside the CPU-only `cthreads` wheel.
+
+Use `@Thread` / `@Threadable` on the CPU path and `@Gpu` when you installed `cthreads-gpu`. The whitelist covers the usual scalars and containers, plus your own Threadable types. Code runs at native speed while you keep a Python-shaped control flow (jobs, pools, sync; GPU: `gpu()` / `join`).
+
+| cthreads SPH | Python + NumPy SPH |
+|:---:|:---:|
+| realtime | linear speedup up to **100× realtime** on this scene |
+|
|
|
+
+**Fig. 1.** Side-by-side Smoothed Particle Hydrodynamics (SPH) fluid demo under identical scene setup (same particle set, forces, timestep, and camera).
+**Left:** dynamics advanced with **cthreads** (`@Thread` kernels on a `ThreadPool`, shared native buffers). **Right:** the same solver expressed as **Python + NumPy** array kernels on the host. Overlay text at the top of each clip reports live run metrics; the **second-to-last** value is **realtime speedup** (simulated time per unit wall-clock). On this scene the NumPy path reaches up to about **100× realtime**; the cthreads path is shown for visual parity of the fluid, not as a matched FPS bake-off in the GIF encode.
+
+---
-Use `@Thread` on functions/methods and `@Threadable` on classes. The whitelist covers the usual scalars and containers, plus your own Threadable types. Code runs at native speed while you keep a Python-shaped control flow (jobs, pools, sync).
+Interactive CPU Mandelbrot (pool + Shared + TBuffer):
+[Open in Colab](https://colab.research.google.com/github/K-T0BIAS/CThreads/blob/main/colab/Mandelbrot.ipynb).
----
@@ -13,7 +28,7 @@ Use `@Thread` on functions/methods and `@Threadable` on classes. The whitelist c
- [Install](./docs/install.md) (includes GPU / Vulkan notes)
- [Guides](./docs/index.md)
-- [GPU guides (0.2.0)](./docs/guide/gpu/README.md)
+- [GPU guides (0.2.0+)](./docs/guide/gpu/README.md)
- [Release (GitHub / PyPI)](./docs/release.md)
- [Math & linalg](./docs/guide/math_and_linalg.md)
- [Compiler notes](./docs/COMPILER.md)
diff --git a/colab/Mandelbrot.ipynb b/colab/Mandelbrot.ipynb
new file mode 100644
index 0000000..22fd1e5
--- /dev/null
+++ b/colab/Mandelbrot.ipynb
@@ -0,0 +1,414 @@
+{
+ "cells": [
+ {
+ "cell_type": "markdown",
+ "id": "d5ae50cf",
+ "metadata": {},
+ "source": [
+ "# cthreads CPU demo: Mandelbrot (pool + Shared + TBuffer)\n",
+ "\n",
+ "Compares **pure Python** Mandelbrot with a **multithreaded cthreads** render that uses:\n",
+ "\n",
+ "| Piece | Role |\n",
+ "|---|---|\n",
+ "| `ThreadPool` | Fixed workers; many strip jobs without one OS thread per job |\n",
+ "| `Shared[list[int]]` | Cooperative host memory for small per-strip done flags |\n",
+ "| `TBufferI64` | Triple-buffered pixel store: **pointer shared**, no per-pixel list marshal |\n",
+ "\n",
+ "Why not one `@Thread` + `list[int]`? Packing/writeback of ~300k ints via per-element\n",
+ "ctypes can dominate runtime and make cthreads look slower than Python.\n",
+ "\n",
+ "## What you need\n",
+ "\n",
+ "1. `pip install cthreads` (**0.2.1+**)\n",
+ "2. A C++ compiler for the first kernel build (Colab: `g++` via `apt`)\n",
+ "\n",
+ "Warm up once (compile + link), then time steady-state runs."
+ ]
+ },
+ {
+ "cell_type": "markdown",
+ "id": "22bee8f3",
+ "metadata": {},
+ "source": [
+ "## 1. Setup"
+ ]
+ },
+ {
+ "cell_type": "code",
+ "execution_count": null,
+ "id": "d59fe3a2",
+ "metadata": {},
+ "outputs": [
+ {
+ "ename": "",
+ "evalue": "",
+ "output_type": "error",
+ "traceback": [
+ "\u001b[1;31mRunning cells with 'venv (Python 3.12.10)' requires the ipykernel package.\n",
+ "\u001b[1;31mInstall 'ipykernel' into the Python environment. \n",
+ "\u001b[1;31mCommand: 'c:/Users/tobik/Desktop/Better_Threads/venv/Scripts/python.exe -m pip install ipykernel -U --force-reinstall'"
+ ]
+ }
+ ],
+ "source": [
+ "%pip install -q -U \"cthreads>=0.2.1\" matplotlib\n",
+ "!apt-get -qq install -y g++"
+ ]
+ },
+ {
+ "cell_type": "markdown",
+ "id": "2823ab22",
+ "metadata": {},
+ "source": [
+ "## 2. Parameters\n",
+ "\n",
+ "Pixels live in a flat buffer of length `width * height`. Work is split into\n",
+ "**row strips** across the pool (Mandelbrot cost varies by region, thus more strips\n",
+ "than workers helps load-balance)."
+ ]
+ },
+ {
+ "cell_type": "code",
+ "execution_count": null,
+ "id": "cecb3a23",
+ "metadata": {},
+ "outputs": [],
+ "source": [
+ "import os\n",
+ "\n",
+ "WIDTH: int = 640 # image width\n",
+ "HEIGHT: int = 480 # image height\n",
+ "MAX_ITER: int = 120 # max iterations\n",
+ "X_MIN: float = -2.0 # x range (lower bound)\n",
+ "X_MAX: float = 1.0 # x range (upper bound)\n",
+ "Y_MIN: float = -1.2 # y range (lower bound)\n",
+ "Y_MAX: float = 1.2 # y range (upper bound)\n",
+ "\n",
+ "N: int = WIDTH * HEIGHT # pixel count\n",
+ "WORKERS: int = max(2, (os.cpu_count() or 2)) # OS threads\n",
+ "N_STRIPS: int = WORKERS * 4 # strips (4x workers for load-balance)\n",
+ "\n",
+ "print(f\"pixels={N} workers={WORKERS} strips={N_STRIPS}\")"
+ ]
+ },
+ {
+ "cell_type": "markdown",
+ "id": "6f2fd127",
+ "metadata": {},
+ "source": [
+ "## 3. Pure Python baseline\n",
+ "\n",
+ "Same escape-time math; **single-threaded**; writes a plain `list[int]`.\n",
+ "\n",
+ "### Why not `threading` for the Python side?\n",
+ "\n",
+ "This loop is **CPU-bound pure Python** (tight float/int work, almost no I/O).\n",
+ "CPython has one **GIL** per process: only one thread runs bytecode at a time.\n",
+ "\n",
+ "If you split rows across `threading.Thread`s:\n",
+ "\n",
+ "- you still serialize on the GIL -> little or no parallel speedup\n",
+ "- you **add** thread create/join, scheduling, and shared-list contention\n",
+ "- so multithreaded Python is often **slower** than one thread here\n",
+ "\n",
+ "`multiprocessing` can use cores (separate interpreters) but pays spawn + pickling\n",
+ "the image buffer, which is usually not worth it for this demo size.\n",
+ "\n",
+ "**cthreads** is different: `@Thread` kernels run as **compiled C++ off the GIL**,\n",
+ "and `ThreadPool` + `TBufferI64` share native memory without per-pixel Python marshal."
+ ]
+ },
+ {
+ "cell_type": "code",
+ "execution_count": null,
+ "id": "4b8c61a7",
+ "metadata": {},
+ "outputs": [],
+ "source": [
+ "def mandelbrot_python(\n",
+ " width: int,\n",
+ " height: int,\n",
+ " max_iter: int,\n",
+ " x_min: float,\n",
+ " x_max: float,\n",
+ " y_min: float,\n",
+ " y_max: float,\n",
+ " out: list[int],\n",
+ ") -> None:\n",
+ " for py in range(height):\n",
+ " for px in range(width):\n",
+ " x0: float = x_min + (x_max - x_min) * (1.0 * px) / (1.0 * width)\n",
+ " y0: float = y_min + (y_max - y_min) * (1.0 * py) / (1.0 * height)\n",
+ " x: float = 0.0\n",
+ " y: float = 0.0\n",
+ " i: int = 0\n",
+ " while i < max_iter:\n",
+ " if x * x + y * y > 4.0:\n",
+ " break\n",
+ " xt: float = x * x - y * y + x0\n",
+ " y = 2.0 * x * y + y0\n",
+ " x = xt\n",
+ " i = i + 1\n",
+ " out[py * width + px] = i"
+ ]
+ },
+ {
+ "cell_type": "markdown",
+ "id": "0d168a19",
+ "metadata": {},
+ "source": [
+ "## 4. cthreads kernel: strip worker\n",
+ "\n",
+ "- `pixels: TBufferI64`: shared native triple buffer (passed by handle; workers\n",
+ " write **disjoint** row ranges into the current write slot).\n",
+ "- `done: Shared[list[int]]`: small cooperative list on the **pool SharedHost**\n",
+ " (`done[strip_id] = 1` when a strip finishes).\n",
+ "- Launch many strips via `ThreadPool.group` (pins Shared safely for the wave).\n",
+ "- After all jobs finish, host calls `pixels.publish()` once, then `read_copy()`."
+ ]
+ },
+ {
+ "cell_type": "code",
+ "execution_count": null,
+ "id": "b1e5417a",
+ "metadata": {},
+ "outputs": [],
+ "source": [
+ "import cthreads\n",
+ "from cthreads import Thread, Shared, ThreadPool, prepare, load_kernels\n",
+ "from cthreads.sync import TBufferI64\n",
+ "\n",
+ "# safety check (usually not required, ensures that the native c++ backend is valid)\n",
+ "if TBufferI64 is None:\n",
+ " raise RuntimeError(\"cthreads.sync.TBufferI64 missing — need a CPU/GPU wheel with _ext\")\n",
+ "\n",
+ "\n",
+ "@Thread\n",
+ "def mandelbrot_strip(\n",
+ " pixels: TBufferI64,\n",
+ " done: Shared[list[int]],\n",
+ " strip_id: int,\n",
+ " y0: int,\n",
+ " y1: int,\n",
+ " width: int,\n",
+ " height: int,\n",
+ " max_iter: int,\n",
+ " x_min: float,\n",
+ " x_max: float,\n",
+ " y_min: float,\n",
+ " y_max: float,\n",
+ ") -> None:\n",
+ " \"\"\"\n",
+ " CThreaded mandelbrot function. Runs the same code as the python version,\n",
+ " execpt that it operates on a local slice of the full img. This way the work is\n",
+ " shared accross the OS Threads tha the cthreads threadpool launches.\n",
+ " \"\"\"\n",
+ " py: int = y0\n",
+ " while py < y1:\n",
+ " px: int = 0\n",
+ " while px < width:\n",
+ " x0: float = x_min + (x_max - x_min) * (1.0 * px) / (1.0 * width)\n",
+ " y0c: float = y_min + (y_max - y_min) * (1.0 * py) / (1.0 * height)\n",
+ " x: float = 0.0\n",
+ " y: float = 0.0\n",
+ " i: int = 0\n",
+ " while i < max_iter:\n",
+ " if x * x + y * y > 4.0:\n",
+ " break\n",
+ " xt: float = x * x - y * y + x0\n",
+ " y = 2.0 * x * y + y0c\n",
+ " x = xt\n",
+ " i = i + 1\n",
+ " pixels[py * width + px] = i\n",
+ " px = px + 1\n",
+ " py = py + 1\n",
+ " done[strip_id] = 1\n",
+ "\n",
+ "\n",
+ "def strip_ranges(height: int, n_strips: int) -> list[tuple[int, int]]:\n",
+ " \"\"\"\n",
+ " Helper function that splits the flat pixel buffer into N even ranges\n",
+ " \"\"\"\n",
+ " out: list[tuple[int, int]] = []\n",
+ " for s in range(n_strips):\n",
+ " y0 = (s * height) // n_strips\n",
+ " y1 = ((s + 1) * height) // n_strips\n",
+ " out.append((y0, y1))\n",
+ " return out\n",
+ "\n",
+ "\n",
+ "def render_cthreads(\n",
+ " pool: ThreadPool,\n",
+ " pixels: TBufferI64,\n",
+ " n_strips: int,\n",
+ ") -> list[int]:\n",
+ " \"\"\"\n",
+ " Run strip jobs on the pool; publish once; return a Python list snapshot.\n",
+ " \"\"\"\n",
+ " done: list[int] = [0] * n_strips # done flag for the ranges for work distribution\n",
+ " ranges = strip_ranges(HEIGHT, n_strips) # split the image into N strips\n",
+ " items = [ # prepare the range wise mandelbrot call arguments\n",
+ " (\n",
+ " pixels,\n",
+ " done,\n",
+ " s,\n",
+ " y0,\n",
+ " y1,\n",
+ " WIDTH,\n",
+ " HEIGHT,\n",
+ " MAX_ITER,\n",
+ " X_MIN,\n",
+ " X_MAX,\n",
+ " Y_MIN,\n",
+ " Y_MAX,\n",
+ " )\n",
+ " for s, (y0, y1) in enumerate(ranges)\n",
+ " ]\n",
+ " # submit the work to the cthreads.pool\n",
+ " group = pool.group(mandelbrot_strip, items)\n",
+ " # wait for all the strips to finish\n",
+ " group.results()\n",
+ " # check that all the strips finished\n",
+ " if sum(done) != n_strips:\n",
+ " raise RuntimeError(f\"Shared done flags incomplete: {sum(done)}/{n_strips}\")\n",
+ " # retrive the actual result from the TBufferI64\n",
+ " # (this is GIL free through the internal datastructure and therefore prefered for large data loads)\n",
+ " pixels.publish()\n",
+ " return list(pixels.read_copy()) # copy and return"
+ ]
+ },
+ {
+ "cell_type": "markdown",
+ "id": "35ba7ec0",
+ "metadata": {},
+ "source": [
+ "## 5. Warmup, then time both\n",
+ "\n",
+ "Warmup compiles the kernel and touches the pool path. Timed runs are steady-state."
+ ]
+ },
+ {
+ "cell_type": "code",
+ "execution_count": null,
+ "id": "e6c16d8b",
+ "metadata": {},
+ "outputs": [],
+ "source": [
+ "import time\n",
+ "\n",
+ "# Pool submit does not auto-prepare; compile + load kernels once.\n",
+ "binary = prepare()\n",
+ "load_kernels(str(binary))\n",
+ "\n",
+ "# create the threadpool\n",
+ "pool = ThreadPool(WORKERS).start()\n",
+ "# init the pixel buffer\n",
+ "pixels = TBufferI64(N)\n",
+ "\n",
+ "try:\n",
+ " # --- warmup (first Shared/TBuffer wave; do not time) ---\n",
+ " _ = render_cthreads(pool, pixels, N_STRIPS)\n",
+ "\n",
+ " # --- timed pure Python ---\n",
+ " out_py: list[int] = [0] * N\n",
+ " t0 = time.perf_counter()\n",
+ " mandelbrot_python(WIDTH, HEIGHT, MAX_ITER, X_MIN, X_MAX, Y_MIN, Y_MAX, out_py)\n",
+ " t_python = time.perf_counter() - t0\n",
+ "\n",
+ " # --- timed cthreads (pool + Shared + TBuffer) ---\n",
+ " t0 = time.perf_counter()\n",
+ " out_native = render_cthreads(pool, pixels, N_STRIPS)\n",
+ " t_native = time.perf_counter() - t0\n",
+ "\n",
+ " print(f\"pure Python: {t_python * 1000:.1f} ms\")\n",
+ " print(f\"cthreads: {t_native * 1000:.1f} ms\")\n",
+ " if t_native > 0.0:\n",
+ " print(f\"speedup: {t_python / t_native:.1f}x\")\n",
+ "\n",
+ " mismatches = sum(1 for a, b in zip(out_py, out_native) if a != b)\n",
+ " print(f\"pixel mismatches: {mismatches}\")\n",
+ " print(f\"Shared done flags: all {N_STRIPS} strips marked complete\")\n",
+ "finally:\n",
+ " pool.stop()"
+ ]
+ },
+ {
+ "cell_type": "markdown",
+ "id": "ff1291cd",
+ "metadata": {},
+ "source": [
+ "## 6. Visual check\n",
+ "\n",
+ "Left / right: Python vs cthreads. Right: timing bars."
+ ]
+ },
+ {
+ "cell_type": "code",
+ "execution_count": null,
+ "id": "747e6fbf",
+ "metadata": {},
+ "outputs": [],
+ "source": [
+ "import matplotlib.pyplot as plt\n",
+ "\n",
+ "img_py = [out_py[r * WIDTH:(r + 1) * WIDTH] for r in range(HEIGHT)]\n",
+ "img_native = [out_native[r * WIDTH:(r + 1) * WIDTH] for r in range(HEIGHT)]\n",
+ "\n",
+ "fig, axes = plt.subplots(1, 3, figsize=(14, 4))\n",
+ "\n",
+ "axes[0].imshow(img_py, cmap=\"magma\", origin=\"lower\")\n",
+ "axes[0].set_title(\"Pure Python\")\n",
+ "axes[0].axis(\"off\")\n",
+ "\n",
+ "axes[1].imshow(img_native, cmap=\"magma\", origin=\"lower\")\n",
+ "axes[1].set_title(\"cthreads pool+Shared+TBuffer\")\n",
+ "axes[1].axis(\"off\")\n",
+ "\n",
+ "axes[2].bar(\n",
+ " [\"Pure Python\", \"cthreads\"],\n",
+ " [t_python * 1000.0, t_native * 1000.0],\n",
+ " color=[\"#888888\", \"#2a9d8f\"],\n",
+ ")\n",
+ "axes[2].set_ylabel(\"ms\")\n",
+ "axes[2].set_title(\"Render time\")\n",
+ "\n",
+ "fig.tight_layout()\n",
+ "plt.show()"
+ ]
+ },
+ {
+ "cell_type": "markdown",
+ "id": "takeaways",
+ "metadata": {},
+ "source": [
+ "## Takeaways\n",
+ "\n",
+ "- **`TBufferI64`** — big pixel buffers stay native; pass one handle to every strip job.\n",
+ " Publish once after the wave, then `read_copy()` for Python/matplotlib.\n",
+ "- **`Shared[list[int]]`** — small cooperative state on the pool’s SharedHost; use\n",
+ " `pool.group` (or `submit_queue`) so the host stays pinned for the whole wave.\n",
+ "- **`ThreadPool`** — reuses workers; better than one OS thread per strip.\n",
+ "- Avoid timing a single job that packs a huge `list[int]` — marshal will dominate.\n",
+ "\n",
+ "Docs: [pools](https://github.com/K-T0BIAS/CThreads/blob/main/docs/guide/pools.md),\n",
+ "[sync / TBuffer](https://github.com/K-T0BIAS/CThreads/blob/main/docs/guide/sync.md),\n",
+ "[Shared marshal](https://github.com/K-T0BIAS/CThreads/blob/main/docs/guide/marshal_and_module.md)."
+ ]
+ }
+ ],
+ "metadata": {
+ "kernelspec": {
+ "display_name": "venv",
+ "language": "python",
+ "name": "python3"
+ },
+ "language_info": {
+ "name": "python",
+ "pygments_lexer": "ipython3",
+ "version": "3.12.10"
+ }
+ },
+ "nbformat": 4,
+ "nbformat_minor": 5
+}
diff --git a/docs/__ressources/cthreads_w6.gif b/docs/__ressources/cthreads_w6.gif
new file mode 100644
index 0000000..5f841be
Binary files /dev/null and b/docs/__ressources/cthreads_w6.gif differ
diff --git a/docs/__ressources/numpy_w6.gif b/docs/__ressources/numpy_w6.gif
new file mode 100644
index 0000000..c82eccb
Binary files /dev/null and b/docs/__ressources/numpy_w6.gif differ