Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
164 commits
Select commit Hold shift + click to select a range
05c2bd3
docs: add udma alltoall demo design
Jun 18, 2026
ecff707
docs: add udma alltoall implementation plan
Jun 18, 2026
cdf6b64
feat: add udma alltoall demo
Jun 19, 2026
e552736
Add UDMA demo checker collectives
Jun 22, 2026
6869b4c
Merge remote-tracking branch 'origin/main' into udma-demo-checker
Jun 22, 2026
8d15433
256p alltoall udma checker ok
Jun 24, 2026
8928ca3
udma alltoall ranksize core perf ok
Jun 25, 2026
8cfea67
4p udma alltoall perf normal
Jun 26, 2026
28a45d6
udma alltoall lmem register 64MB->128MB, pass count halved
Jun 26, 2026
cbd9666
8p udma alltoall perf on 156 (8M/32M/64M, ~272-283 GB/s)
Jun 27, 2026
2aca40e
add p2p/datacopy latency micro-kernels (testType 4/5) to isolate UDMA…
Jun 27, 2026
fd04a96
test: add UDMA alltoall bigdata layout plan
Jun 27, 2026
71f48e6
test: add UDMA alltoall bigdata source guards
Jun 27, 2026
cbf2cb3
feat(udma): add bigdata alltoall demo kernel
Jun 28, 2026
ead95b0
perf(udma): pipeline bigdata alltoall relay copies
Jun 28, 2026
dd8b1ef
perf(udma): split bigdata alltoall local copies
Jun 28, 2026
0eb4282
docs: design multinode udma bigdata alltoall
Jun 29, 2026
33beb61
perf(udma): add multinode bigdata slot layout
Jun 29, 2026
ae07f69
docs: design udma multi-eid route port
Jun 29, 2026
1b4c443
docs: plan udma multi-eid route port
Jun 29, 2026
859cc64
feat(udma): support weighted multi-eid routes
Jun 29, 2026
ef166e0
perf(udma): tune multinode bigdata alltoall sync
Jun 29, 2026
c3c1353
test(udma): align alltoall source guards after sync tuning
Jun 29, 2026
53ac4a9
perf(udma): checkpoint 64p remote put tail flag scheduling
Jun 30, 2026
cb3a7f5
perf(udma): restore validated remote put mteack path
Jun 30, 2026
6564bba
perf(udma): disable remote put tail check path
Jun 30, 2026
2288632
fix(udma): wait for remote put completion
Jun 30, 2026
4eea5bb
test(udma): add alltoall performance runner
Jun 30, 2026
a77406b
fix(udma): prefer aggregate routes for remote nodes
Jun 30, 2026
bbfc4f4
fix(udma): honor ranks per node in bigdata demo
Jun 30, 2026
b0d7df0
perf(udma): serialize remote put sends per link
Jun 30, 2026
148cab5
Revert "perf(udma): serialize remote put sends per link"
Jun 30, 2026
4e8b114
perf(udma): remove remote put ack wait
Jun 30, 2026
bfcacf7
perf(udma): prefer max-weight route for remote puts
Jun 30, 2026
0ca36a7
docs(udma): design fullmesh control route fix
Jul 17, 2026
6c1717c
docs(udma): plan fullmesh control route fix
Jul 17, 2026
b03c747
test(udma): require explicit fullmesh control routes
Jul 17, 2026
7c2ce28
fix(udma): pin fullmesh control traffic to valid routes
Jul 17, 2026
4c20ef1
docs(udma): design fullmesh GM tracing
Jul 17, 2026
b91b27e
docs(udma): plan fullmesh GM tracing
Jul 17, 2026
d2ab380
test(udma): define fullmesh trace capacity
Jul 17, 2026
daf376a
feat(udma): define fullmesh trace layout
Jul 17, 2026
08aafd2
feat(udma): convert fullmesh traces to Chrome JSON
Jul 17, 2026
11a5c19
test(udma): require fullmesh trace collection
Jul 17, 2026
9906318
feat(udma): collect fullmesh GM traces
Jul 17, 2026
8c6c7bc
test(udma): require per-core fullmesh trace spans
Jul 17, 2026
20e19ef
feat(udma): trace fullmesh core tasks
Jul 17, 2026
f6019d6
fix(udma): inline device trace offsets
Jul 17, 2026
23dd1e4
fix(udma): enable tracing for physical fullmesh
Jul 17, 2026
c02f3bf
fix(udma): separate fullmesh trace iterations
Jul 17, 2026
4d0beca
docs(udma): plan fullmesh 16-0 trace experiment
Jul 17, 2026
a1bd408
test(udma): require fullmesh 16-0 split
Jul 17, 2026
86a8f86
perf(udma): route fullmesh payload through core16
Jul 17, 2026
0e8d62b
fix(udma): handle terminal fullmesh shard boundary
Jul 20, 2026
f135bcf
docs: design HCCL two-host cluster validation
Jul 20, 2026
cd57ec1
fix(udma): stream fullmesh trace JSON output
Jul 20, 2026
e7ca40b
docs(udma): design isolated fullmesh task profiling
Jul 20, 2026
e89ebc5
test(udma): require wait-free isolated tasks
Jul 20, 2026
65a8bf2
feat(udma): add wait-free isolated task profiling
Jul 20, 2026
3f73a59
docs(udma): design grouped fullmesh alltoall
Jul 20, 2026
80fd6a1
docs(udma): plan grouped fullmesh alltoall
Jul 20, 2026
a5c1686
docs(udma): bound grouped alltoall to 128 ranks
Jul 20, 2026
c107930
test(udma): require grouped alltoall layout
Jul 20, 2026
b90d674
feat(udma): add grouped alltoall layout
Jul 20, 2026
dad5ee7
test(udma): require standalone grouped kernel
Jul 20, 2026
32be944
feat(udma): add grouped fullmesh alltoall kernel
Jul 20, 2026
1f50eef
fix(udma): build grouped kernel as separate library
Jul 20, 2026
3064705
test(udma): require grouped alltoall host mode
Jul 20, 2026
f53b1fb
feat(udma): add grouped alltoall host mode
Jul 20, 2026
a8da596
test(udma): require grouped alltoall trace conversion
Jul 20, 2026
7fd61e8
feat(udma): trace grouped alltoall pipeline
Jul 20, 2026
0b5ab40
fix(udma): inline grouped trace offsets on device
Jul 20, 2026
68e3b5f
fix(udma): isolate grouped trace cores by cache line
Jul 20, 2026
b254280
docs(udma): design grouped alltoall dual-route split
Jul 20, 2026
35ad253
docs(udma): plan grouped alltoall dual-route split
Jul 20, 2026
55bb7c7
test(udma): require balanced grouped dual-route policy
Jul 20, 2026
e148dcc
feat(udma): add grouped dual-route peer policy
Jul 20, 2026
7da9a5d
test(udma): require grouped device dual-route selection
Jul 20, 2026
fc38975
feat(udma): balance grouped peers across UDMA routes
Jul 20, 2026
ce9004a
docs(udma): design configurable grouped copyout workers
Jul 20, 2026
aeb8fcb
docs(udma): plan configurable grouped copyout workers
Jul 20, 2026
e198303
test(udma): require configurable grouped copyout workers
Jul 20, 2026
f3f3ae6
test(udma): require grouped copyout lane scheduling
Jul 20, 2026
1415be3
feat(udma): configure grouped copyout workers
Jul 20, 2026
fdeabc9
docs(udma): design grouped route-stage diagnostics
Jul 20, 2026
2667576
docs(udma): plan grouped route-stage diagnostics
Jul 20, 2026
6b0e943
feat(udma): classify grouped route stages
Jul 20, 2026
b1dc0da
feat(udma): filter grouped kernels by route stage
Jul 20, 2026
2e2c22d
feat(udma): stage grouped route diagnostics
Jul 20, 2026
00e3c61
test(udma): cover grouped stage trace filenames
Jul 20, 2026
dd29ca9
fix(udma): synchronize grouped stage startup
Jul 20, 2026
cf6ba17
fix(udma): batch grouped route stages
Jul 20, 2026
52f4475
fix(udma): synchronize staged ping-pong reuse
Jul 20, 2026
df8ee4b
docs(udma): correct staged self-copy trace count
Jul 20, 2026
abcf961
perf(udma): split local grouped stage profiling
Jul 21, 2026
5ec4926
docs(udma): add grouped alltoall experiment guide
Jul 21, 2026
56e5c6d
perf(udma): isolate combined remote sends
Jul 21, 2026
6e4a56b
perf(udma): add grouped contention diagnostics
Jul 21, 2026
ee703d3
perf(udma): trace combined grouped stage
Jul 21, 2026
c786578
perf(udma): add grouped no-copy diagnostic
Jul 21, 2026
3d2e822
perf(udma): throttle grouped signal polling
Jul 21, 2026
852dd34
Revert "perf(udma): throttle grouped signal polling"
Jul 21, 2026
1a8ad1f
perf(udma): split remote grouped copyout
Jul 21, 2026
0799297
perf(udma): add third grouped copyout worker
Jul 21, 2026
3e59d81
perf(udma): default grouped copyout to 48 workers
Jul 21, 2026
dddffad
Revert "perf(udma): default grouped copyout to 48 workers"
Jul 21, 2026
f072913
Revert "perf(udma): add third grouped copyout worker"
Jul 21, 2026
75bd08d
perf(udma): add third remote copyout worker
Jul 21, 2026
6f672d6
feat(udma): select grouped remote route mode
Jul 21, 2026
08e3271
docs(udma): design 1024-rank grouped alltoall
Jul 22, 2026
1274d80
docs(udma): size grouped trace for 1024 ranks
Jul 22, 2026
3d03557
docs(udma): plan 1024-rank grouped alltoall
Jul 22, 2026
77d79c9
feat(udma): expand communication capacity to 1024 ranks
Jul 22, 2026
2ddca3a
feat(udma): support 1024-rank grouped alltoall
Jul 22, 2026
4fed6b8
fix(udma): read expanded grouped traces
Jul 22, 2026
5960fdc
build(udma): support Ascend950 variants
Jul 22, 2026
35279b4
feat(udma): allow UDMA-only communicator init
Jul 22, 2026
8e0083a
test(udma): align 1024-rank trace capacity with signal protocol
Jul 22, 2026
748e15d
feat(udma): align grouped alltoall channel striping
Jul 22, 2026
9f9e70f
fix(udma): identify nodes independently of hostname
Jul 23, 2026
fd9124d
feat(udma): add grouped alltoall shared qp pool
Jul 23, 2026
e919e37
fix(comm): complete socket exchange transfers
Jul 23, 2026
a8eceae
fix(udma): exchange only required shared qp keys
Jul 23, 2026
001fe00
fix(comm): shard socket exchange across rank groups
Jul 24, 2026
102b5bf
perf(udma): batch grouped alltoall quiet operations
Jul 27, 2026
897f3f6
perf(udma): extend grouped quiet batches to 64
Jul 27, 2026
03cc964
fix(udma): restore direct quiet path for batch one
Aug 1, 2026
9319b4c
fix(udma): make rootinfo device offset order-independent
Aug 1, 2026
8aca7d0
feat(udma): add multi-region registered memory
Jul 30, 2026
d30ad52
feat(udma): support shared QPs across memory regions
Jul 30, 2026
0e7696d
feat(udma): reuse shared QPs across memory regions
Jul 30, 2026
4221476
feat(udma): allow 16 GiB grouped alltoall payloads
Jul 30, 2026
846193b
feat(udma): expand registered memory regions
Jul 30, 2026
0e76a37
perf(udma): restore single-region signal fast path
Aug 1, 2026
1e942b8
fix(udma): align grouped measured launches
Aug 1, 2026
35726b4
test(udma): use grouped layout assertion helper
Aug 1, 2026
079fbcc
fix(udma): extend CQ polling for large batches
Aug 2, 2026
9150bcd
Revert "fix(udma): extend CQ polling for large batches"
Aug 2, 2026
63c42e2
perf(udma): split grouped data and signal writes
Aug 2, 2026
62ba1f7
fix(trace): label grouped send workers correctly
Aug 2, 2026
aa0eae4
feat(udma): add grouped alltoall ingress credits
Aug 2, 2026
db811f6
docs(udma): record 64-rank ingress credit results
Aug 2, 2026
f61b36b
perf(udma): publish ingress credits through dedicated IPC
Aug 2, 2026
10f17d7
fix(udma): transfer grouped credits through MTE
Aug 2, 2026
869c75f
feat(udma): reduce grouped copyout cores with SDMA
Aug 2, 2026
b51dc2e
fix(udma): select grouped copyout resources by SDMA mode
Aug 4, 2026
2d4c65c
docs(udma): document grouped credit IPC MTE flow
Aug 4, 2026
748ca74
docs(udma): add reusable credit IPC implementation
Aug 4, 2026
2719f41
feat(udma): trace SDMA submit and wait phases
Aug 4, 2026
27025bf
fix(udma): use grouped error recorder in SDMA trace
Aug 4, 2026
bea490a
feat(udma): trace SDMA submit pipeline stages
Aug 4, 2026
24a95d0
perf(udma): close grouped ingress credit window
Aug 4, 2026
0b31864
fix(udma): avoid grouped terminal credit barrier
Aug 4, 2026
beeb7ab
Revert "fix(udma): avoid grouped terminal credit barrier"
Aug 4, 2026
7ab4b10
perf(udma): schedule ready grouped copyouts
Aug 4, 2026
a36e91d
perf(udma): prewarm grouped SDMA SQ pages
Aug 4, 2026
abdd99c
perf(udma): balance grouped 128p peer schedule
Aug 5, 2026
0dc483b
test(udma): add grouped single-stage isolation
Aug 5, 2026
e36784e
fix(udma): retain regions through grouped stage runs
Aug 5, 2026
425dcbe
fix(udma): isolate repeated demo barriers
Aug 5, 2026
cd712ca
fix(udma): wait for grouped terminal barrier
Aug 5, 2026
2a2ae01
fix(udma): extend grouped quiet polling
Aug 5, 2026
47dd5e7
fix(udma): align 128p grouped shared QP lanes
Aug 5, 2026
b4f03a5
test(udma): use balanced lane contract at 128p
Aug 5, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
626 changes: 626 additions & 0 deletions docs/alltoall-udma-success.md

Large diffs are not rendered by default.

1,048 changes: 1,048 additions & 0 deletions docs/grouped-alltoall-credit-ipc-mte.md

Large diffs are not rendered by default.

120 changes: 120 additions & 0 deletions docs/plans/2026-08-02-grouped-alltoall-ingress-credit.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,120 @@
# Grouped AllToAll Ingress Credit Implementation Plan

## Goal And Scope

Implement the default-disabled `window=1` ingress-credit protocol approved in
`docs/specs/2026-08-02-grouped-alltoall-ingress-credit-design.md`. The change
must bound payload admission to one source per destination lane while
preserving the existing path exactly when credits are disabled.

This plan does not change route weights, cabinet topology interpretation,
quiet batching, copyout assignment, or multi-pass credit semantics. Initial
hardware acceptance is single-pass only.

## Task 1: Pure Layout And Schedule Contracts

**Objective and role:** Define host-testable helpers for credit enablement,
token construction, receive ownership, and next-group peer mapping.

**Background and prerequisites:** The approved spec is authoritative. Existing
peer mapping and plan helpers live in
`tests/udma/demo/tilexr_udma_alltoall_group_layout.h`.

**Modification scope:** The grouped layout header and
`tests/udma/unit/test_tilexr_udma_alltoall_group_layout.cpp`.

**Constraints and non-goals:** Preserve existing peer mapping and default
layout behavior. Do not add device-only dependencies to the host-testable
header.

**Acceptance and verification:** Unit tests cover 256 ranks, group zero,
inactive final lanes, owner selection, shared-route credit identity, invocation
separation, and default-disabled behavior. Run the grouped layout unit binary.

**Artifacts and interfaces:** Stable helper contracts consumed by Host and
Kernel work.

## Task 2: Registered Credit Plane And Host Interface

**Objective and role:** Reserve ping-pong credit storage and pass a validated
ingress-window value and credit offsets into kernel launch.

**Background and prerequisites:** Task 1 token and layout contracts are fixed.
The grouped plan already owns payload, signal, and control regions.

**Modification scope:** Grouped layout plan, demo environment parsing and
allocation/launch wiring, grouped kernel declaration and launch wrapper, plus
focused source-contract tests.

**Constraints and non-goals:** `INGRESS_WINDOW=0` remains the default. Only 0
and 1 are accepted. Existing payload and signal offsets remain stable where
practical; registered-size overflow checks remain mandatory.

**Acceptance and verification:** Unit tests prove valid/invalid environment
and launch wiring contracts, credit planes do not overlap, and max registered
size accounts for both planes.

**Artifacts and interfaces:** Kernel receives two credit offsets and the
ingress-window value.

## Task 3: Kernel Credit Wait And Publication

**Objective and role:** Gate each destination send once per group, publish a
local request after both route-ready tokens arrive, and let the primary send
core issue the next lane credit.

**Background and prerequisites:** Tasks 1 and 2 provide mapping, storage, and
arguments. Existing MTE token wait and UDMA NBI primitives are reused.

**Modification scope:**
`tests/udma/demo/tilexr_udma_alltoall_group_kernel.cpp`, existing UDMA helper
interfaces only if required, and grouped source-contract tests.

**Constraints and non-goals:** Do not add per-credit quiet. Only cores 32..47
publish local requests; only primary send cores publish UDMA credits, preserving
the shared-SQ single-producer rule. Local request slots are 64-byte isolated so
independent receive-core cache maintenance cannot lose another lane's request.
Primary and secondary workers wait on the same destination credit. The
default-disabled branch must not execute a credit wait or write. If source
lifetime cannot be made safe without per-credit quiet, stop and revise the
design instead of weakening this constraint.

**Acceptance and verification:** Static/source tests prove placement,
single-producer ownership, request isolation, and unique shared-SQ completion
reclamation. Compile checks validate Ascend C overloads and launch signatures.
The exact final source passes a clean b131 Bisheng build and 4x8 hardware
validation; 256P admission-bound validation remains pending.

**Artifacts and interfaces:** Trace-visible or debug-visible credit wait and
publish phases where trace capacity permits.

## Task 4: Verification And Delivery Readiness

**Objective and role:** Establish local correctness and document remaining
hardware evidence.

**Background and prerequisites:** Tasks 1 through 3 are complete.

**Modification scope:** Tests and documentation required to keep validation
commands and limitations accurate.

**Constraints and non-goals:** Run only the explicitly authorized NPU matrix.
Do not modify or stage unrelated existing worktree files.

**Acceptance and verification:** Run focused unit tests, relevant Python
parser/source tests, formatting/build checks available on Windows, inspect the
final diff, and record the exact 2x8 then 4x8 then 256P hardware matrix. The
implemented path has completed 4x8 functional and warmup-5/repeat-50
performance validation; 256P admission-bound validation remains pending.

**Artifacts and interfaces:** A reviewable implementation, verification
summary, and scoped Git commit after the final checks pass.

## Key Risks

- Credit NBI source storage may be reused before hardware consumes it.
- Two receive copy slices may publish duplicate credits.
- Primary and secondary send workers may wait on different tokens by mistake.
- A final inactive lane may wait for or publish a nonexistent peer.
- Registered-memory growth may cross a region boundary or maximum-size check.
- Trace storage may need expansion if new phases are recorded.
177 changes: 177 additions & 0 deletions docs/specs/2026-08-02-grouped-alltoall-ingress-credit-design.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,177 @@
# Grouped AllToAll Ingress Credit Design

## Status

The original request-plus-UDMA-credit implementation was validated through 8x8
hardware on 2026-08-02. It has since been replaced by direct receive-core
publication through a dedicated credit IPC allocation and requires renewed
hardware validation. The experimental feature remains disabled by default.

## Goal

Bound each destination rank to at most 16 concurrent payload source ranks in
the grouped AllToAll data path. Preserve independent lane progress so one slow
source does not impose a full-rank or full-group barrier.

The first implementation is experimental and disabled by default. Existing
behavior and performance remain unchanged unless ingress credits are enabled.

## Non-goals

- Limiting the number of physical port operations to 16. A source may still
use primary and secondary routes after it receives one destination credit.
- Adding a global barrier between peer groups.
- Waiting for receive-copy completion before releasing network capacity.
- Changing topology-derived primary/secondary route weights.
- Enabling credit control for multi-pass payloads in the first hardware test.

## Existing Behavior

At `rankSize=256` and `groupWidth=16`, every rank has 16 logical peer groups.
Each group contains up to eight forward and eight backward peers. The 32 send
workers are split into 16 primary-route workers and 16 secondary-route workers.

Each worker advances to its next group immediately after its own quiet
completes. Workers and ranks do not share group progress. Consequently, a
destination may receive payloads from more than one logical group at the same
time even though each group contains no more than 16 source ranks.

## Credit Protocol

There are 16 independent destination lane chains. Group zero is initially
enabled. For group `g > 0`, a source must observe a credit from its destination
for `(invocation, g)` before either route posts payload data.

For destination rank `D`, lane `L`, and group `g`:

1. The designated receive owner waits until all payload routes expected from
`peer(D, g, L)` have published their data-ready tokens.
2. Before receive-copy, the owner directly writes one credit token through the
dedicated IPC mapping of `peer(D, g + 1, L)`, when that peer exists.
3. The next source observes the credit in its local dedicated credit buffer and
may post both primary and secondary payload routes to `D`.

At most one source per destination lane can therefore be admitted. With 16
lanes, the strict payload-source bound is 16 ranks per destination.

Credit dependencies are monotonic from group `g` to `g + 1`. Group zero has no
dependency, so the protocol does not introduce a cyclic startup wait.

## Credit Storage And Tokens

Each rank allocates a dedicated 1 MiB IPC buffer when
`TILEXR_ENABLE_CREDIT_IPC=1`. A credit received by source rank `S` is indexed by
destination rank `D`; each entry occupies 512 bytes and two fixed 512 KiB
ping-pong planes support 1024 ranks:

```text
credit[pingPongSlot][destinationRank * 512]
```

The expected value encodes at least the invocation and destination group. The
receiver accepts a value greater than or equal to the expected token, matching
the existing data-ready stale-token policy. Ping-pong slots prevent adjacent
invocations from reusing a live credit location.

Primary and secondary send workers for the same peer wait on the same credit.
Credits control source-rank admission, not route admission.

## Receive Ownership

With 32 copyout workers, workers 0 through 15 and 16 through 31 can wait on the
same peer signal while copying different payload slices. Only the first slice,
kernel cores 32 through 47, owns direct credit publication. Cores 48 through 63
never publish credits.

The owner publishes its credit after both expected primary and secondary data
tokens arrive and before MTE receive-copy begins. This avoids placing copyout
latency on the send admission path.

## Credit Submission

Credit publication is one 512-byte MTE copy by the receive owner through
`creditMems[nextSource]`. The token occupies the first 8 bytes of the slot;
copying the complete slot gives the remote IPC write an explicit MTE3
completion and keeps adjacent credits on separate transfer units. It does not
consume a UDMA WQE, QP, completion, or quiet, and it does not add another
producer to a shared UDMA SQ. Both send routes poll the same local dedicated
credit slot through a 512-byte MTE2 copy.

The dedicated allocation is independent of the existing optional communication
IPC buffer, so grouped ingress credit can run with `TILEXR_ENABLE_IPC=0`. It
adds one IPC mapping per peer and process, but only 1 MiB of device memory per
rank. Host code rejects ingress-credit execution if any mapping is missing.

## Configuration

Add an experimental ingress window setting:

```text
TILEXR_DEMO_ALLTOALL_GROUP_INGRESS_WINDOW=0|1
```

- `0`: current behavior and default.
- `1`: one admitted source per lane, at most 16 payload source ranks per
destination.

`window=1` additionally requires `TILEXR_ENABLE_CREDIT_IPC=1` before
communicator initialization.

A later `window=2` extension may trade a 32-source bound for more tolerance of
credit latency, but it is outside the first implementation.

## Validation

Host/unit tests must cover:

- 256-rank group and lane predecessor/successor mapping.
- No credit wait for group zero.
- One shared credit for primary and secondary routes.
- Single credit publisher when 32 copyout workers duplicate receive lanes.
- Correct inactive-lane behavior at the diameter and final partial group.
- Credit token separation across invocations and ping-pong slots.
- Default-disabled compatibility.

Hardware validation starts with 2x8 functional coverage, then 4x8, then 256P.
The 256P comparison uses 1 GiB/rank, multi channel, single pass, warmup 5 and
repeat 50. Evidence includes P50 and range plus per-phase `credit-wait`,
`send-quiet`, and `receive-wait` trace durations. A 1 KiB single-channel case
checks that the default-disabled path has no latency regression.

The first trace-enabled run is diagnostic only and is not compared directly
with trace-disabled performance.

The final 4x8, warmup-5/repeat-50, trace-disabled comparison passed data
validation on all 32 ranks:

```text
1 KiB/rank single: window0 P50 38.224 us, window1 P50 52.763 us
1 GiB/rank multi: window0 P50 5329.210 us, window1 P50 5317.925 us
```

This proved the previous request/UDMA-credit protocol progressed across repeated invocations
and does not regress 1 GiB throughput in the 4x8 environment. It does not prove
the 16-source global bound because 4x8 has only two peer groups.

The 8x8 A10 run used the fixed host order `226, 223, 220, 217, 198, 195,
192, 189`, single pass, shared QP, trace disabled, and warmup-5/repeat-50. All
64 ranks passed data validation:

```text
1 KiB/rank single: window0 P50 73.067 us, window1 P50 105.930 us
1 GiB/rank multi: window0 P50 6516.335 us, window1 P50 6336.370 us
```

Under the previous request/UDMA-credit implementation, ingress credit reduced
the 1 GiB P50 by 179.965 us (2.76%), from
approximately 153.46 GiB/s to 157.82 GiB/s. The 1 KiB case paid 32.863 us of
additional fixed control latency. This run exercises four peer groups and
therefore provides stronger repeated credit-chain evidence than 4x8, but it
still does not prove the 256P global admission bound.

## Residual Risk

Rank0-only trace cannot prove the global source-card bound. Hardware validation
needs either all-rank trace or lightweight per-destination admission counters.
The physical four-port capacity and cabinet-aware route weighting are separate
topology problems and are not solved by this credit protocol.
Loading