Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
55 changes: 55 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -271,6 +271,58 @@ jobs:
python ci/build_cuda_candidate.py
--output "$RUNNER_TEMP/tensorflow-cuda-e3"

cuda-e3-evidence-contract:
name: cuda-e3 / manual evidence contract / GPU-free
runs-on: ubuntu-24.04
timeout-minutes: 10
steps:
- name: Check out source
uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
with:
persist-credentials: false
- name: Set up CPython 3.11
uses: actions/setup-python@a26af69be951a213d495a4c3e4e4022e16d87065 # v5.6.0
with:
python-version: "3.11"
- name: Prove the contract lane remains TensorFlow-free
run: |
python - <<'PY'
import sys
assert not any(
name == "tensorflow" or name.startswith("tensorflow.")
for name in sys.modules
)
PY
- name: Run manual-harness and offline-verifier source tests
run: |
python -m pip install --upgrade pip==26.1 pytest==9.1.1
python -m pytest \
tests/test_cuda_e3_manual_harness.py \
tests/test_cuda_e3_evidence.py -q
- name: Exercise command help without TensorFlow, an extension, or CUDA
run: |
python scripts/certify_cuda_candidate.py --help
python scripts/verify_cuda_e3_evidence.py --help
- name: Prove both evidence scripts import without TensorFlow
run: |
python - <<'PY'
import importlib.util
import sys
from pathlib import Path

for name in ("certify_cuda_candidate", "verify_cuda_e3_evidence"):
path = Path("scripts") / f"{name}.py"
spec = importlib.util.spec_from_file_location(name, path)
assert spec is not None and spec.loader is not None
module = importlib.util.module_from_spec(spec)
sys.modules[name] = module
spec.loader.exec_module(module)
assert not any(
name == "tensorflow" or name.startswith("tensorflow.")
for name in sys.modules
)
PY

ci-gate:
name: ci-gate
if: ${{ always() }}
Expand All @@ -281,6 +333,7 @@ jobs:
- package
- core-compatibility
- cuda-e3-build-only
- cuda-e3-evidence-contract
runs-on: ubuntu-24.04
timeout-minutes: 5
steps:
Expand All @@ -292,6 +345,7 @@ jobs:
PACKAGE_RESULT: ${{ needs.package.result }}
CORE_COMPATIBILITY_RESULT: ${{ needs.core-compatibility.result }}
CUDA_E3_RESULT: ${{ needs.cuda-e3-build-only.result }}
CUDA_E3_EVIDENCE_RESULT: ${{ needs.cuda-e3-evidence-contract.result }}
run: |
python - <<'PY'
import os
Expand All @@ -303,6 +357,7 @@ jobs:
"package": os.environ["PACKAGE_RESULT"],
"core-compatibility": os.environ["CORE_COMPATIBILITY_RESULT"],
"cuda-e3-build-only": os.environ["CUDA_E3_RESULT"],
"cuda-e3-evidence-contract": os.environ["CUDA_E3_EVIDENCE_RESULT"],
}
failures = {name: result for name, result in results.items() if result != "success"}
assert not failures, f"public Alpha CI gates did not succeed: {failures}"
Expand Down
5 changes: 5 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,11 @@ Changelog and Semantic Versioning conventions.
CUDA-first fail-closed claim/lower routing, a machine-readable non-certifying
contract, and a synthetic-provider Linux compile/link-only CI gate that
never imports TensorFlow or loads/executes the candidate.
- Add an opt-in, manual real-NVIDIA first-stage evidence producer and offline
verifier. They require the frozen clean candidate checkout and record native
execution, parity, and lifetime evidence only; kernel activity and runtime
transfer profiling remain explicitly unverified, and
`support_claim=false` / `certification_ready=false` remain unchanged.
- Add rank-1 float32 CPU `tf.nn.softmax` with the final axis omitted or
supplied as literal positional/keyword `axis=0`. The lowering reuses the
owned same-wheel TFE `Softmax` unary path, preserves the existing explicit
Expand Down
4 changes: 4 additions & 0 deletions MANIFEST.in
Original file line number Diff line number Diff line change
Expand Up @@ -2,3 +2,7 @@ include ci/platform-contract.json
include ci/cuda-e3-build-only.json
include ci/build_cuda_candidate.py
include docs/cuda-build-only-0.1.2.md
include scripts/certify_cuda_candidate.py
include scripts/verify_cuda_e3_evidence.py
include tests/test_cuda_e3_manual_harness.py
include tests/test_cuda_e3_evidence.py
14 changes: 13 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,13 +24,25 @@ certification, release, or performance claim.
| Performance claim | **None** — no benchmark gate; Alpha does not claim speedups |
| Pure-Rust TensorFlow | **No** — native helpers call into the active wheel |
| Abandoned TF Rust crates | **Not used** as Cargo dependencies (`crate_dependencies() == ()`) |
| CUDA candidate | **Build-only**, `support_claim=false`, `certification_ready=false`; no real-GPU evidence |
| CUDA candidate | Build-only hosted CI plus opt-in first-stage real-NVIDIA evidence; `support_claim=false`, `certification_ready=false` |

The unreleased branch requires Core `rextio>=0.1.6,<0.2` and plugin API
**1.6**. It rejects boundary-free standalone Rust lowering. CUDA lowering also
requires exact authorization from
`rextio-device-cuda/cuda-tensorflow-tfe-linux-x86_64`.

The hosted CUDA job is compile/link-only: it never installs/imports TensorFlow,
loads the extension, or executes CUDA. A separately opt-in real-NVIDIA
first-stage harness may produce self-attested execution/parity/lifetime
evidence; its closed schema requires `kernel_activity_verified=false` and
`runtime_transfer_profiled=false`, and preserves `support_claim=false` and
`certification_ready=false`. The offline verifier establishes schema and
payload integrity only, not GPU execution, hardware certification, or CUDA
support. See
[the CUDA build-only and manual-evidence contract](docs/cuda-build-only-0.1.2.md)
for the exact Linux GNU/CPython 3.11/TF 2.21.0/Rust 1.93.1 pins, clean
candidate checkout, GPU:0/permitted-SM boundary, and commands.

Final release verification completed on 2026-07-18: GitHub Actions
[run `29597803215`](https://github.com/rextio/rextio-tensorflow/actions/runs/29597803215)
finished **13/13 jobs successfully**, and a no-cache CPython 3.11 install from
Expand Down
4 changes: 4 additions & 0 deletions ci/check_sdist_contract.py
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,10 @@
Path("ci/cuda-e3-build-only.json"),
Path("ci/build_cuda_candidate.py"),
Path("docs/cuda-build-only-0.1.2.md"),
Path("scripts/certify_cuda_candidate.py"),
Path("scripts/verify_cuda_e3_evidence.py"),
Path("tests/test_cuda_e3_manual_harness.py"),
Path("tests/test_cuda_e3_evidence.py"),
)


Expand Down
131 changes: 126 additions & 5 deletions docs/cuda-build-only-0.1.2.md
Original file line number Diff line number Diff line change
Expand Up @@ -78,13 +78,134 @@ owner.
The int32 axis handle for `reduce_mean(axis=1)` is the one bounded host control
input. It is not a user-tensor transfer.

## Hosted CI and manual testing
## Hosted CI

Hosted CI uses the real Core/provider orchestration with a deterministic
synthetic probe. It generates and links one cdylib. It never installs or
imports TensorFlow in that job and never loads or executes the extension.

Real-NVIDIA execution, numerical parity, kernel activity, lifetime, same-image
runtime identity, and absence of user-tensor transfers remain deferred manual
work. Any future evidence remains non-certifying until a separate review
explicitly changes the contract.
This build-only lane is deliberately not a substitute for a GPU test: it does
not load the extension, execute CUDA, establish numerical parity, or observe
the lifetime of borrowed TensorFlow objects.

## Opt-in manual real-NVIDIA first-stage evidence

`scripts/certify_cuda_candidate.py` is a manual, **first-stage evidence**
producer. It is not a hosted CI job and must be run only on a machine whose
operator has explicitly chosen to use a real NVIDIA GPU. Its output is checked
offline by `scripts/verify_cuda_e3_evidence.py`.

This is a frozen environment, not a portability recipe:

- Linux `x86_64-unknown-linux-gnu` with GNU/glibc; no macOS, Windows, musl, or
cross-compiled host is accepted.
- CPython 3.11, TensorFlow `2.21.0`, and Rust `1.93.1`.
- Core checkout exactly `7f47f0ce8cea0b6dbeb7fd3c733f65eeaa6bb5e0` and CUDA
provider checkout exactly `cf65733f06b91a801f9806367f09948ee7162540`.
- A clean TensorFlow-plugin checkout selected by `TF_REF`. After this PR is
integrated, set `TF_REF=0.1.2`; that is the default runnable path. Until
then, the current `0.1.2` target branch does not contain these two scripts,
so a reviewer/operator must set `TF_REF` to the current immutable full
PR-head SHA instead. Do not use a moving feature-branch name or record that
self-referential SHA in this document. After checkout, derive the full
lowercase SHA with `git rev-parse HEAD` and pass it explicitly via
`--expected-tensorflow-commit`; the harness verifies that it descends from
the frozen E3 base.
- Exactly one usable `GPU:0`, with a permitted architecture from this closed
set: `sm_60`, `sm_61`, `sm_70`, `sm_72`, `sm_75`, `sm_80`, `sm_86`,
`sm_87`, `sm_89`, or `sm_90`. Other ordinals and SM values are rejected
rather than generalized.
- GNU binutils, including `readelf`, on `PATH`. The harness records GNU build
IDs from the TensorFlow wheel images and fails closed when an expected image
has no build ID.

The harness deliberately has no `toolkit_root` setting or command-line option.
It reuses the active TensorFlow wheel and its already-loaded images; pointing
at an independent CUDA toolkit would violate the runtime-reuse contract.

Use independent checkout, output, and work directories. The output file and
the new exclusive work directory must be outside **all three** clean source
checkouts; the harness rejects paths inside any attested checkout. Do not
create the work directory itself: the harness requires it not to exist yet.

```bash
export E3_ROOT="$HOME/rextio-tf-e3-manual-$(date +%Y%m%d-%H%M%S)"
export E3_OUT="$E3_ROOT/evidence-output"
export E3_BUILD="$E3_ROOT/isolated-build"
mkdir -p "$E3_ROOT/checkouts" "$E3_OUT"

git clone https://github.com/rextio/rextio.git "$E3_ROOT/checkouts/rextio"
git -C "$E3_ROOT/checkouts/rextio" checkout --detach \
7f47f0ce8cea0b6dbeb7fd3c733f65eeaa6bb5e0
git clone https://github.com/rextio/rextio-device-cuda.git \
"$E3_ROOT/checkouts/rextio-device-cuda"
git -C "$E3_ROOT/checkouts/rextio-device-cuda" checkout --detach \
cf65733f06b91a801f9806367f09948ee7162540
export TF_ROOT="$E3_ROOT/checkouts/rextio-tensorflow"
export TF_REF=0.1.2
# Before this PR merges, replace 0.1.2 above with the current immutable full
# PR-head SHA. The 0.1.2 default becomes runnable only after integration.
git clone https://github.com/rextio/rextio-tensorflow.git "$TF_ROOT"
git -C "$TF_ROOT" fetch --tags origin "$TF_REF"
git -C "$TF_ROOT" checkout --detach "$TF_REF"
test -f "$TF_ROOT/scripts/certify_cuda_candidate.py"
test -f "$TF_ROOT/scripts/verify_cuda_e3_evidence.py"
export TF_COMMIT="$(git -C "$TF_ROOT" rev-parse HEAD)"
test "${#TF_COMMIT}" -eq 40

python3.11 -m venv "$E3_ROOT/venv"
"$E3_ROOT/venv/bin/python" -m pip install --upgrade pip
"$E3_ROOT/venv/bin/python" -m pip install tensorflow==2.21.0
"$E3_ROOT/venv/bin/python" -m pip install --no-deps \
"$E3_ROOT/checkouts/rextio" "$E3_ROOT/checkouts/rextio-device-cuda" \
"$E3_ROOT/checkouts/rextio-tensorflow"
rustup toolchain install 1.93.1 --profile minimal
command -v readelf
readelf --version
```

Import TensorFlow **in the same process that invokes the harness**. This is
required to establish the wheel-image reuse boundary, rather than an optional
smoke test. Set `TF_SM` to the actual permitted architecture of the sole usable
`GPU:0` (the example uses `sm_80` only as a placeholder):

```bash
cd "$TF_ROOT"
export TF_SM=sm_80
E3_OUTPUT="$E3_OUT/cuda-e3-first-stage.json" E3_WORK="$E3_BUILD" \
E3_CORE="$E3_ROOT/checkouts/rextio" \
E3_PROVIDER="$E3_ROOT/checkouts/rextio-device-cuda" \
"$E3_ROOT/venv/bin/python" - <<'PY'
import os
import sys

import tensorflow as tf

assert tf.__version__ == "2.21.0"
from scripts import certify_cuda_candidate

sys.argv = [
"certify_cuda_candidate.py",
"--output", os.environ["E3_OUTPUT"],
"--work-dir", os.environ["E3_WORK"],
"--core-root", os.environ["E3_CORE"],
"--provider-root", os.environ["E3_PROVIDER"],
"--expected-tensorflow-commit", os.environ["TF_COMMIT"],
"--sm", os.environ["TF_SM"],
]
raise SystemExit(certify_cuda_candidate.main())
PY
"$E3_ROOT/venv/bin/python" scripts/verify_cuda_e3_evidence.py \
"$E3_OUT/cuda-e3-first-stage.json"
```

The producer self-attests `native_extension_executed=true` only if the bounded
harness reaches that observation. It intentionally records
`kernel_activity_verified=false` and `runtime_transfer_profiled=false`.
The offline verifier checks canonical schema and payload integrity only; it
does not authenticate the producer, prove execution, recompute artifact
hashes, certify hardware, or confer CUDA support. Any evidence remains
self-attested execution, numerical-parity, and borrowed-object-lifetime
evidence, never a GPU-success claim, kernel-activity certification,
transfer/profile measurement, CUDA support, or a performance claim. It always
leaves `support_claim=false` and `certification_ready=false`.
Loading
Loading