diff --git a/.gitignore b/.gitignore index 06432db..7a4aba2 100644 --- a/.gitignore +++ b/.gitignore @@ -62,5 +62,6 @@ libcthreads_kernels.dylib # <<< cthreads (auto) todo.md -# demo artifacts -*.gif \ No newline at end of file +# demo artifacts (allow README showcase gifs) +*.gif +!docs/__ressources/*.gif diff --git a/ReadMe.md b/ReadMe.md index ec40238..b557388 100644 --- a/ReadMe.md +++ b/ReadMe.md @@ -1,11 +1,26 @@ ![LOGO](./docs/__ressources/CTHREADS_03_1.svg) +[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/K-T0BIAS/CThreads/blob/main/colab/Mandelbrot.ipynb) + ---- -**cthreads** compiles a typed Python subset into C++ so work can run on real OS threads **without the GIL** allowing true concurrency without `multiprocessing`'s process boundaries and pickling tax. +**cthreads** compiles a typed Python subset into C++ so work can run on real OS threads **without the GIL**, allowing true concurrency without `multiprocessing`'s process boundaries and pickling tax. Optional **Vulkan GPU** kernels (`@Gpu`) ship as a separate install ( **`cthreads-gpu`**) same `import cthreads`, do not install alongside the CPU-only `cthreads` wheel. + +Use `@Thread` / `@Threadable` on the CPU path and `@Gpu` when you installed `cthreads-gpu`. The whitelist covers the usual scalars and containers, plus your own Threadable types. Code runs at native speed while you keep a Python-shaped control flow (jobs, pools, sync; GPU: `gpu()` / `join`). + +| cthreads SPH | Python + NumPy SPH | +|:---:|:---:| +| realtime | linear speedup up to **100× realtime** on this scene | +| cthreads SPH simulation | Python + NumPy SPH simulation | + +**Fig. 1.** Side-by-side Smoothed Particle Hydrodynamics (SPH) fluid demo under identical scene setup (same particle set, forces, timestep, and camera). +**Left:** dynamics advanced with **cthreads** (`@Thread` kernels on a `ThreadPool`, shared native buffers). **Right:** the same solver expressed as **Python + NumPy** array kernels on the host. Overlay text at the top of each clip reports live run metrics; the **second-to-last** value is **realtime speedup** (simulated time per unit wall-clock). On this scene the NumPy path reaches up to about **100× realtime**; the cthreads path is shown for visual parity of the fluid, not as a matched FPS bake-off in the GIF encode. + +--- -Use `@Thread` on functions/methods and `@Threadable` on classes. The whitelist covers the usual scalars and containers, plus your own Threadable types. Code runs at native speed while you keep a Python-shaped control flow (jobs, pools, sync). +Interactive CPU Mandelbrot (pool + Shared + TBuffer): +[Open in Colab](https://colab.research.google.com/github/K-T0BIAS/CThreads/blob/main/colab/Mandelbrot.ipynb). ---- @@ -13,7 +28,7 @@ Use `@Thread` on functions/methods and `@Threadable` on classes. The whitelist c - [Install](./docs/install.md) (includes GPU / Vulkan notes) - [Guides](./docs/index.md) -- [GPU guides (0.2.0)](./docs/guide/gpu/README.md) +- [GPU guides (0.2.0+)](./docs/guide/gpu/README.md) - [Release (GitHub / PyPI)](./docs/release.md) - [Math & linalg](./docs/guide/math_and_linalg.md) - [Compiler notes](./docs/COMPILER.md) diff --git a/colab/Mandelbrot.ipynb b/colab/Mandelbrot.ipynb new file mode 100644 index 0000000..22fd1e5 --- /dev/null +++ b/colab/Mandelbrot.ipynb @@ -0,0 +1,414 @@ +{ + "cells": [ + { + "cell_type": "markdown", + "id": "d5ae50cf", + "metadata": {}, + "source": [ + "# cthreads CPU demo: Mandelbrot (pool + Shared + TBuffer)\n", + "\n", + "Compares **pure Python** Mandelbrot with a **multithreaded cthreads** render that uses:\n", + "\n", + "| Piece | Role |\n", + "|---|---|\n", + "| `ThreadPool` | Fixed workers; many strip jobs without one OS thread per job |\n", + "| `Shared[list[int]]` | Cooperative host memory for small per-strip done flags |\n", + "| `TBufferI64` | Triple-buffered pixel store: **pointer shared**, no per-pixel list marshal |\n", + "\n", + "Why not one `@Thread` + `list[int]`? Packing/writeback of ~300k ints via per-element\n", + "ctypes can dominate runtime and make cthreads look slower than Python.\n", + "\n", + "## What you need\n", + "\n", + "1. `pip install cthreads` (**0.2.1+**)\n", + "2. A C++ compiler for the first kernel build (Colab: `g++` via `apt`)\n", + "\n", + "Warm up once (compile + link), then time steady-state runs." + ] + }, + { + "cell_type": "markdown", + "id": "22bee8f3", + "metadata": {}, + "source": [ + "## 1. Setup" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "d59fe3a2", + "metadata": {}, + "outputs": [ + { + "ename": "", + "evalue": "", + "output_type": "error", + "traceback": [ + "\u001b[1;31mRunning cells with 'venv (Python 3.12.10)' requires the ipykernel package.\n", + "\u001b[1;31mInstall 'ipykernel' into the Python environment. \n", + "\u001b[1;31mCommand: 'c:/Users/tobik/Desktop/Better_Threads/venv/Scripts/python.exe -m pip install ipykernel -U --force-reinstall'" + ] + } + ], + "source": [ + "%pip install -q -U \"cthreads>=0.2.1\" matplotlib\n", + "!apt-get -qq install -y g++" + ] + }, + { + "cell_type": "markdown", + "id": "2823ab22", + "metadata": {}, + "source": [ + "## 2. Parameters\n", + "\n", + "Pixels live in a flat buffer of length `width * height`. Work is split into\n", + "**row strips** across the pool (Mandelbrot cost varies by region, thus more strips\n", + "than workers helps load-balance)." + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "cecb3a23", + "metadata": {}, + "outputs": [], + "source": [ + "import os\n", + "\n", + "WIDTH: int = 640 # image width\n", + "HEIGHT: int = 480 # image height\n", + "MAX_ITER: int = 120 # max iterations\n", + "X_MIN: float = -2.0 # x range (lower bound)\n", + "X_MAX: float = 1.0 # x range (upper bound)\n", + "Y_MIN: float = -1.2 # y range (lower bound)\n", + "Y_MAX: float = 1.2 # y range (upper bound)\n", + "\n", + "N: int = WIDTH * HEIGHT # pixel count\n", + "WORKERS: int = max(2, (os.cpu_count() or 2)) # OS threads\n", + "N_STRIPS: int = WORKERS * 4 # strips (4x workers for load-balance)\n", + "\n", + "print(f\"pixels={N} workers={WORKERS} strips={N_STRIPS}\")" + ] + }, + { + "cell_type": "markdown", + "id": "6f2fd127", + "metadata": {}, + "source": [ + "## 3. Pure Python baseline\n", + "\n", + "Same escape-time math; **single-threaded**; writes a plain `list[int]`.\n", + "\n", + "### Why not `threading` for the Python side?\n", + "\n", + "This loop is **CPU-bound pure Python** (tight float/int work, almost no I/O).\n", + "CPython has one **GIL** per process: only one thread runs bytecode at a time.\n", + "\n", + "If you split rows across `threading.Thread`s:\n", + "\n", + "- you still serialize on the GIL -> little or no parallel speedup\n", + "- you **add** thread create/join, scheduling, and shared-list contention\n", + "- so multithreaded Python is often **slower** than one thread here\n", + "\n", + "`multiprocessing` can use cores (separate interpreters) but pays spawn + pickling\n", + "the image buffer, which is usually not worth it for this demo size.\n", + "\n", + "**cthreads** is different: `@Thread` kernels run as **compiled C++ off the GIL**,\n", + "and `ThreadPool` + `TBufferI64` share native memory without per-pixel Python marshal." + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "4b8c61a7", + "metadata": {}, + "outputs": [], + "source": [ + "def mandelbrot_python(\n", + " width: int,\n", + " height: int,\n", + " max_iter: int,\n", + " x_min: float,\n", + " x_max: float,\n", + " y_min: float,\n", + " y_max: float,\n", + " out: list[int],\n", + ") -> None:\n", + " for py in range(height):\n", + " for px in range(width):\n", + " x0: float = x_min + (x_max - x_min) * (1.0 * px) / (1.0 * width)\n", + " y0: float = y_min + (y_max - y_min) * (1.0 * py) / (1.0 * height)\n", + " x: float = 0.0\n", + " y: float = 0.0\n", + " i: int = 0\n", + " while i < max_iter:\n", + " if x * x + y * y > 4.0:\n", + " break\n", + " xt: float = x * x - y * y + x0\n", + " y = 2.0 * x * y + y0\n", + " x = xt\n", + " i = i + 1\n", + " out[py * width + px] = i" + ] + }, + { + "cell_type": "markdown", + "id": "0d168a19", + "metadata": {}, + "source": [ + "## 4. cthreads kernel: strip worker\n", + "\n", + "- `pixels: TBufferI64`: shared native triple buffer (passed by handle; workers\n", + " write **disjoint** row ranges into the current write slot).\n", + "- `done: Shared[list[int]]`: small cooperative list on the **pool SharedHost**\n", + " (`done[strip_id] = 1` when a strip finishes).\n", + "- Launch many strips via `ThreadPool.group` (pins Shared safely for the wave).\n", + "- After all jobs finish, host calls `pixels.publish()` once, then `read_copy()`." + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "b1e5417a", + "metadata": {}, + "outputs": [], + "source": [ + "import cthreads\n", + "from cthreads import Thread, Shared, ThreadPool, prepare, load_kernels\n", + "from cthreads.sync import TBufferI64\n", + "\n", + "# safety check (usually not required, ensures that the native c++ backend is valid)\n", + "if TBufferI64 is None:\n", + " raise RuntimeError(\"cthreads.sync.TBufferI64 missing — need a CPU/GPU wheel with _ext\")\n", + "\n", + "\n", + "@Thread\n", + "def mandelbrot_strip(\n", + " pixels: TBufferI64,\n", + " done: Shared[list[int]],\n", + " strip_id: int,\n", + " y0: int,\n", + " y1: int,\n", + " width: int,\n", + " height: int,\n", + " max_iter: int,\n", + " x_min: float,\n", + " x_max: float,\n", + " y_min: float,\n", + " y_max: float,\n", + ") -> None:\n", + " \"\"\"\n", + " CThreaded mandelbrot function. Runs the same code as the python version,\n", + " execpt that it operates on a local slice of the full img. This way the work is\n", + " shared accross the OS Threads tha the cthreads threadpool launches.\n", + " \"\"\"\n", + " py: int = y0\n", + " while py < y1:\n", + " px: int = 0\n", + " while px < width:\n", + " x0: float = x_min + (x_max - x_min) * (1.0 * px) / (1.0 * width)\n", + " y0c: float = y_min + (y_max - y_min) * (1.0 * py) / (1.0 * height)\n", + " x: float = 0.0\n", + " y: float = 0.0\n", + " i: int = 0\n", + " while i < max_iter:\n", + " if x * x + y * y > 4.0:\n", + " break\n", + " xt: float = x * x - y * y + x0\n", + " y = 2.0 * x * y + y0c\n", + " x = xt\n", + " i = i + 1\n", + " pixels[py * width + px] = i\n", + " px = px + 1\n", + " py = py + 1\n", + " done[strip_id] = 1\n", + "\n", + "\n", + "def strip_ranges(height: int, n_strips: int) -> list[tuple[int, int]]:\n", + " \"\"\"\n", + " Helper function that splits the flat pixel buffer into N even ranges\n", + " \"\"\"\n", + " out: list[tuple[int, int]] = []\n", + " for s in range(n_strips):\n", + " y0 = (s * height) // n_strips\n", + " y1 = ((s + 1) * height) // n_strips\n", + " out.append((y0, y1))\n", + " return out\n", + "\n", + "\n", + "def render_cthreads(\n", + " pool: ThreadPool,\n", + " pixels: TBufferI64,\n", + " n_strips: int,\n", + ") -> list[int]:\n", + " \"\"\"\n", + " Run strip jobs on the pool; publish once; return a Python list snapshot.\n", + " \"\"\"\n", + " done: list[int] = [0] * n_strips # done flag for the ranges for work distribution\n", + " ranges = strip_ranges(HEIGHT, n_strips) # split the image into N strips\n", + " items = [ # prepare the range wise mandelbrot call arguments\n", + " (\n", + " pixels,\n", + " done,\n", + " s,\n", + " y0,\n", + " y1,\n", + " WIDTH,\n", + " HEIGHT,\n", + " MAX_ITER,\n", + " X_MIN,\n", + " X_MAX,\n", + " Y_MIN,\n", + " Y_MAX,\n", + " )\n", + " for s, (y0, y1) in enumerate(ranges)\n", + " ]\n", + " # submit the work to the cthreads.pool\n", + " group = pool.group(mandelbrot_strip, items)\n", + " # wait for all the strips to finish\n", + " group.results()\n", + " # check that all the strips finished\n", + " if sum(done) != n_strips:\n", + " raise RuntimeError(f\"Shared done flags incomplete: {sum(done)}/{n_strips}\")\n", + " # retrive the actual result from the TBufferI64\n", + " # (this is GIL free through the internal datastructure and therefore prefered for large data loads)\n", + " pixels.publish()\n", + " return list(pixels.read_copy()) # copy and return" + ] + }, + { + "cell_type": "markdown", + "id": "35ba7ec0", + "metadata": {}, + "source": [ + "## 5. Warmup, then time both\n", + "\n", + "Warmup compiles the kernel and touches the pool path. Timed runs are steady-state." + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "e6c16d8b", + "metadata": {}, + "outputs": [], + "source": [ + "import time\n", + "\n", + "# Pool submit does not auto-prepare; compile + load kernels once.\n", + "binary = prepare()\n", + "load_kernels(str(binary))\n", + "\n", + "# create the threadpool\n", + "pool = ThreadPool(WORKERS).start()\n", + "# init the pixel buffer\n", + "pixels = TBufferI64(N)\n", + "\n", + "try:\n", + " # --- warmup (first Shared/TBuffer wave; do not time) ---\n", + " _ = render_cthreads(pool, pixels, N_STRIPS)\n", + "\n", + " # --- timed pure Python ---\n", + " out_py: list[int] = [0] * N\n", + " t0 = time.perf_counter()\n", + " mandelbrot_python(WIDTH, HEIGHT, MAX_ITER, X_MIN, X_MAX, Y_MIN, Y_MAX, out_py)\n", + " t_python = time.perf_counter() - t0\n", + "\n", + " # --- timed cthreads (pool + Shared + TBuffer) ---\n", + " t0 = time.perf_counter()\n", + " out_native = render_cthreads(pool, pixels, N_STRIPS)\n", + " t_native = time.perf_counter() - t0\n", + "\n", + " print(f\"pure Python: {t_python * 1000:.1f} ms\")\n", + " print(f\"cthreads: {t_native * 1000:.1f} ms\")\n", + " if t_native > 0.0:\n", + " print(f\"speedup: {t_python / t_native:.1f}x\")\n", + "\n", + " mismatches = sum(1 for a, b in zip(out_py, out_native) if a != b)\n", + " print(f\"pixel mismatches: {mismatches}\")\n", + " print(f\"Shared done flags: all {N_STRIPS} strips marked complete\")\n", + "finally:\n", + " pool.stop()" + ] + }, + { + "cell_type": "markdown", + "id": "ff1291cd", + "metadata": {}, + "source": [ + "## 6. Visual check\n", + "\n", + "Left / right: Python vs cthreads. Right: timing bars." + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "747e6fbf", + "metadata": {}, + "outputs": [], + "source": [ + "import matplotlib.pyplot as plt\n", + "\n", + "img_py = [out_py[r * WIDTH:(r + 1) * WIDTH] for r in range(HEIGHT)]\n", + "img_native = [out_native[r * WIDTH:(r + 1) * WIDTH] for r in range(HEIGHT)]\n", + "\n", + "fig, axes = plt.subplots(1, 3, figsize=(14, 4))\n", + "\n", + "axes[0].imshow(img_py, cmap=\"magma\", origin=\"lower\")\n", + "axes[0].set_title(\"Pure Python\")\n", + "axes[0].axis(\"off\")\n", + "\n", + "axes[1].imshow(img_native, cmap=\"magma\", origin=\"lower\")\n", + "axes[1].set_title(\"cthreads pool+Shared+TBuffer\")\n", + "axes[1].axis(\"off\")\n", + "\n", + "axes[2].bar(\n", + " [\"Pure Python\", \"cthreads\"],\n", + " [t_python * 1000.0, t_native * 1000.0],\n", + " color=[\"#888888\", \"#2a9d8f\"],\n", + ")\n", + "axes[2].set_ylabel(\"ms\")\n", + "axes[2].set_title(\"Render time\")\n", + "\n", + "fig.tight_layout()\n", + "plt.show()" + ] + }, + { + "cell_type": "markdown", + "id": "takeaways", + "metadata": {}, + "source": [ + "## Takeaways\n", + "\n", + "- **`TBufferI64`** — big pixel buffers stay native; pass one handle to every strip job.\n", + " Publish once after the wave, then `read_copy()` for Python/matplotlib.\n", + "- **`Shared[list[int]]`** — small cooperative state on the pool’s SharedHost; use\n", + " `pool.group` (or `submit_queue`) so the host stays pinned for the whole wave.\n", + "- **`ThreadPool`** — reuses workers; better than one OS thread per strip.\n", + "- Avoid timing a single job that packs a huge `list[int]` — marshal will dominate.\n", + "\n", + "Docs: [pools](https://github.com/K-T0BIAS/CThreads/blob/main/docs/guide/pools.md),\n", + "[sync / TBuffer](https://github.com/K-T0BIAS/CThreads/blob/main/docs/guide/sync.md),\n", + "[Shared marshal](https://github.com/K-T0BIAS/CThreads/blob/main/docs/guide/marshal_and_module.md)." + ] + } + ], + "metadata": { + "kernelspec": { + "display_name": "venv", + "language": "python", + "name": "python3" + }, + "language_info": { + "name": "python", + "pygments_lexer": "ipython3", + "version": "3.12.10" + } + }, + "nbformat": 4, + "nbformat_minor": 5 +} diff --git a/docs/__ressources/cthreads_w6.gif b/docs/__ressources/cthreads_w6.gif new file mode 100644 index 0000000..5f841be Binary files /dev/null and b/docs/__ressources/cthreads_w6.gif differ diff --git a/docs/__ressources/numpy_w6.gif b/docs/__ressources/numpy_w6.gif new file mode 100644 index 0000000..c82eccb Binary files /dev/null and b/docs/__ressources/numpy_w6.gif differ