The tape says balanced. The split says one wallet and a crowd.
Net flow is a sum, and summing is lossy. Elephant Tracks splits the tape back apart β per side, how much of it is one wallet β from real swaps and their maker addresses.
Live, keyless, 2026-09-07: 800 swaps of AUSD in 29.4 s for 0 credits and no API key. One wallet was 62.0% of the sell side β $25,459,824 in 12 swaps β into 212 buying wallets, while net flow read +2.1%. Receipt β
No key. No signup. No install. One command:
python3 scripts/split_tape.py --address 0x00000000efe302beaa2b3e6e1b18d08d69a9012a --symbol AUSD --pages 8 --json ausd.jsonsplitting the tape β keyless, 8 page(s) x 100 swaps per token
token swaps avg buy $ avg sell $ ticket buy w sell w top buy top sell net flow note
--------------------------------------------------------------------------------------------------
AUSD 800 91,699.61 123,273.10 1.3x 212 111 11.7% 62.0% 2.1%
HERO β rule: max top-maker share of either side among |net flow| < 5.0% (1/1 qualified)
AUSD: one wallet is 62.0% of the sell side while net flow reads +2.1% β a dashboard calls this balanced
sell: $25,459,824 in 12 swaps from 0x9f681e397f51137215b8240b8bf4e523d898661b
buy: 467 swaps from 212 distinct wallets, largest 11.7%
supporting: 467 buys avg $91,700 Β· 333 sells avg $123,273 Β· ticket ratio 1.3x
net flow is one number. The split is four.
wrote ausd.json (29.4s wall clock, 0 credits β keyless)
That is a live call to CoinMarketCap's keyless
/public-apisurface. Nothing here is a fixture. These exact numbers are from 2026-09-07T07:00:52Z β run it yourself and they will differ, because they come from the market rather than from this file. The full receipt, including the 12 raw swap records behind the top seller, is committed atdocs/proof/ausd.jsonand walked through in DEMO.md.
Read the AUSD row. Net flow is +2.1% β every flow dashboard in existence renders that as "balanced". Split the same 800 swaps by maker address and the sell side is 62.0% one wallet β $25,459,824 across 12 swaps β sold into 212 distinct buying wallets, the largest of which is 11.7%. That is one seller distributing to a crowd, and the summed number cannot contain it.
The default run is the eight-token watchlist (python3 scripts/split_tape.py, one page each). It
applies the same published rule, and when no token reads balanced it says so on screen rather than
widening the band β on the 2026-09-07T07:11:15Z run, none did. Transcript in DEMO.md.
Every DEX flow tool reduces the tape to one signed number: buy volume minus sell volume.
That number cannot distinguish the two situations a trader most needs to tell apart:
| net flow | |
|---|---|
| One desk selling $27,000 clips to 4,000 people buying $180 clips | β 0 |
| 4,000 matched trades of the same size | β 0 |
Same number. Opposite market structures. One is distribution into retail; the other is churn.
Summing is lossy and cannot be inverted. You can always re-add a split back into a sum; nothing recovers the split from the sum. So we never sum β we pull the individual swaps and measure each side separately.
CoinMarketCap's /v1/dex/tokens/transactions returns, per swap: the side (tp), the USD
value (v), and the maker address (ma). That last field is what turns "one desk against a
crowd" from an inference into a measurement: per side, sum v per ma, and the largest
wallet's share of the side is the number.
The hero row is chosen by a published rule β highest top-maker share of either side among tokens whose net flow is within Β±5% and which clear both confidence floors β so it is a screen, not a cherry-pick. Taking the global maximum instead would pick a token a dashboard already flags as lopsided, which proves nothing.
Why maker share and not ticket size. The first version headlined average buy ticket Γ· average
sell ticket. At flat net flow buy volume β sell volume, so that ratio collapses to the count ratio
(n_sells / n_buys) β one quantity measured twice; across 54 flat tokens the two agreed to within a
median 5.8%, and the best real value was 5.2Γ. Maker share is a second, independent quantity. The
ticket ratio stays as a supporting column.
Four worked examples show that net flow can hide concentration. They do not show how often it
does, and four tokens chosen by a rule are still four tokens somebody chose. So we measured a
population β make base-rate, keyless, receipt in
docs/proof/base_rate.json:
| Tokens measured | 96 from the most liquid pairs on 7 (network, DEX) sources |
| Clearing the confidence floors | 92 |
| Reading balanced (|net flow| < 5%) | 22 β the denominator |
| With one wallet above half a side | 6 |
The median token whose net flow reads "balanced" has 25% of one side in a single wallet. A quarter of that side, one address, on a token a dashboard is calling balanced. At the 75th percentile it is 54% β more than half the side β and 6 of 22 (27%, 95% CI 13β48%) cross that line outright.
The interval is wide because 22 is a small denominator, and the receipt says so rather than rounding it away. The distribution is reported at every percentile precisely so the 50% threshold cannot be tuned after the fact to produce a friendlier headline β move the line and the whole shape is still there to check.
"Balanced" is not a statement about who is on each side. It never was.
pull the swaps β split by side β attribute each side to its makers β rank by concentration
No server, no database, no cache, no model. The product is one arithmetic operation applied to data only CoinMarketCap publishes, so everything that is not the fetch or the arithmetic was removed.
| Stage | Function | What it does |
|---|---|---|
| Fetch | get() |
One GET, keyless by default β a key exported as CMC_API_KEY is an optional escape hatch for a throttled IP. Backs off on 429/5xx; returns errors instead of swallowing them; explains an exhausted throttle instead of printing a body. |
| Paginate | pull_swaps() |
Cursor is data.lastId on the response envelope; de-duplicates on (tx, lgid); returns (swaps, meta). |
| Split | split() |
The product. Partition on tp; per side, Ξ£ v per ma β the top wallet's share, plus mean(v) and ` |
| Gate | _confidence() |
Is this ratio meaningful at all? Dust floor + minimum swaps per side. |
| Rank | main() |
Apply the published hero rule, print the table, emit JSON. |
| Layer | Technology |
|---|---|
| Language | Python 3.11, stdlib only on the judged path |
| Data | CoinMarketCap DEX API, keyless /public-api surface |
| Tests | pytest Β· hypothesis (property-based) Β· live contract tests |
| Quality | ruff Β· pip-audit Β· gitleaks Β· CodeQL Β· Dependabot |
Full derivation, including every failure mode and the deliberate non-architecture: ARCHITECTURE.md. The diagram as a page: elephant.edycu.dev/architecture.
| Endpoint | Used for | Key? | Called from |
|---|---|---|---|
/public-api/v1/dex/tokens/transactions |
per-swap side, USD value, maker address β the engine | π none | scripts/split_tape.py |
/public-api/v4/dex/spot-pairs/latest |
enumerating a token universe, to choose what to split | π none | scripts/sweep_universe.py |
/public-api/v4/dex/pairs/quotes/latest |
evaluated and rejected β its 24h_buy_volume / 24h_sell_volume are the aggregates this project exists to argue against |
π none | nothing β see below |
/v1/dex/holders/count |
how many addresses hold the token β context for the share | π none | scripts/split_tape.py |
The one Called from gap is deliberate and worth stating plainly, because a table of endpoint
names proves nothing on its own:
pairs/quotes/latestreturns volume and trade count per side, which is where this project started. FEEDBACK.md Β§7 derives why it was abandoned: at flat net flow the ticket ratio those fields produce reduces to the count ratio β the same quantity measured twice, exactly where a hidden asymmetry would be worth finding. It is listed because rejecting it is a finding, not because we call it.holders/countis called once per run, for the hero row only. It answers keyless β the parameter istokenAddressin camelCase, andaddress, which every neighbouring endpoint takes, returns error 4002 here. That is why it reads as key-only until you guess the casing. It is context, not an input: "one wallet is 62.0% of the sell side" reads differently for a token held by 1,150 addresses than for one held by 382,000. The published selection rule never sees it β adding a signal to the ranking after publishing the rule would be moving the target.
Everything else is runnable right now, with no key:
make sweep # /v4/dex/spot-pairs/latest -> a universe, ~9 s, 0 credits
make demo # /v1/dex/tokens/transactions -> the splitMost market-data APIs publish volume. Some publish trade count. CMC publishes the individual swaps with side, USD value and maker address together, keyless β and the maker address is the only reason the wallet count is available at all.
Remove CoinMarketCap and you would need a multi-chain swap indexer, a per-DEX pool registry, a balance indexer, a wallet-labelling pipeline and a symbolβcontract mapping service β five separate systems β to recompute what one endpoint returns in one call.
We also wrote up seven dated, evidenced findings for the CMC product team, including the one that changed this project's entire mechanism: FEEDBACK.md.
| Measurement | Value |
|---|---|
| Live run wall clock | 29.4 s β 800 swaps of one token, 8 calls, clean path Β· 10.0 s β the 8-token watchlist |
| Credits used | 0 β keyless, with every CMC env var explicitly unset |
| Tests | 115 (110 offline, 5 live) |
| Regression tests named for the defect they pin | 12 |
Property-based verification of split() |
2,000 generated tapes, 0 failing |
| Malformed-response boundary cases | 6 |
| Aggregation latency | p50 0.026 ms, p95 0.026 ms (n=200) |
| Live fetch latency | p50 1,403 ms, p95 17,474 ms (n=8, includes one throttle backoff) |
Coverage of scripts/ |
100% of 1,195 statements β make test fails under it |
The 2,000 is the number worth reading. Coverage says we ran the lines we wrote. The property
test says that across 2,000 generated tapes split() never violated six invariants: the ratio
never dropped below 1, wallet counts never exceeded trade counts, net flow never left Β±100%, the
elephant label never disagreed with the arithmetic, the top-maker share never escaped its own
denominator (0β1, attributed to a wallet on that side, never more volume than the side), and no
tape marked ok failed to clear both published floors. That last one is what stops another SHIB
reaching a headline.
pytest tests/test_high_signal.py -k invariants --hypothesis-show-statistics
# β 2000 passing, 0 failingticket_ratio divides two means. A mean over four trades is an anecdote; a mean over sub-cent dust
is spam. Both produce arithmetically correct ratios that mean nothing β and on 2026-09-07 one of
them reached the headline.
| Floor | Value | Pins |
|---|---|---|
DUST_USD |
$1.00 | a side made of spam rather than participants |
MIN_SIDE_SWAPS |
5 | a side too thin for "average" to mean anything |
Rows failing either are still printed, with the reason, and excluded from hero selection. Showing a disqualified row is more honest than hiding it. On 2026-09-07 an earlier build printed SHIB at 772.9Γ off 75 sub-cent sells; the floors exist because of that row.
- Python 3.11 or newer. That is the entire list.
- No API key, no account, no
pip install.split_tape.pyis stdlib-only.
git clone https://github.com/edycutjong/elephant.git
cd elephant
python3 scripts/split_tape.pyFor judges: there is no account to create and no credential to configure β the judged path is keyless by design. Start at JUDGE.md.
If CoinMarketCap's anonymous tier is throttling your IP, the script backs off (15 s, 30 s, 60 s) and, if the quota is exhausted, stops with a message that says so rather than printing a number. Optional escape hatch: a free key from coinmarketcap.com/api exported as
CMC_API_KEYmoves the identical call to the keyed endpoint. Never required β the default path is keyless and every receipt here was taken with no key set.
python3 scripts/split_tape.py --address 0x00000000efe302beaa2b3e6e1b18d08d69a9012a --symbol AUSD --pages 8 --json ausd.json # the headline, 800 swaps
python3 scripts/split_tape.py --address 0x6982508145454ce325ddbe47a25d4ec3d2311933 --symbol PEPE
python3 scripts/split_tape.py --json run.json # the watchlist, full result setmake install # dev deps (pytest, hypothesis, ruff, pip-audit) + the chromium the QA gates drive
make lint # ruff check + format check
make test # 110 offline tests, coverage gated at 100%, no internet
make test-live # 5 tests against the real CoinMarketCap contract
make demo # the judged capability, live, no key
make bench # deterministic p50/p95 over the captured tape
make bench-live # p50/p95 over the real keyless fetch
make site # re-render the landing page, the deck and JUDGE.md from docs/proof/*.json
python3 scripts/qa_site.py # 229 measured gates on the web surfaces (needs playwright + pillow)
make check # refuse to ship a placeholder, or a page that drifted from its receipts
make ci # lint + test + audit + check| Layer | Tool | Status |
|---|---|---|
| Code quality | ruff (check + format) | β |
| Unit testing | pytest, 110 offline tests | β |
| Property testing | hypothesis, 2,000 cases | β |
| Live contract testing | pytest -m live against real CMC |
β |
| Security (SAST) | CodeQL | β |
| Security (SCA) | Dependabot + pip-audit | β |
| Secret scanning | gitleaks, full history | β |
| Release automation | semver from Angular commits | β |
CI runs lint, tests, pip-audit, CodeQL, gitleaks, a placeholder gate, a deterministic benchmark β
and a live-api job that executes the real split on every push. It is keyless, so it runs on
forks and PRs too. If CMC changes the contract, it breaks in CI rather than in front of a judge.
The landing page, the deck and JUDGE.md are generated, never hand-edited.
scripts/render_site.py renders site/index.html, site/pitch/index.html and JUDGE.md from
the templates in scripts/site_templates/ and the receipts in docs/proof/. Every figure on any
of the three is a {{token}} filled from a committed JSON receipt; the render aborts if any slot
is left unfilled, and CI re-renders all three and fails on any diff. So no number a judge reads can
be typed in by hand, none can be a placeholder, and none can drift from the run that produced it β
the landing page and the judge guide cannot disagree, because they are the same render.
The web surfaces are measured, not eyeballed. scripts/qa_site.py drives both pages in headless
Chromium and runs 229 gates: no horizontal overflow at 375 / 768 / 1440, every link and image
resolves, a page height budget on phones, every text node β SVG text included β at or above WCAG AA contrast as actually painted, with
opacity flattened onto the background (minimum 5.65:1, gated against regression), the social card
measured from its own PNG header and description lengths counted, heading levels, prose-link
underlines and table headers checked, every interactive element proven to change on hover by
screenshot byte-diff, all eleven slides reachable by keyboard and inside the stage, the page
rendering with every external request blocked, prefers-reduced-motion leaving zero running
animations, and the animation itself sampled with the clock paused β the staged split must hold
the summed bar alone first, never reverse, never snap, and end in exactly the reduced-motion
frame; the one loop on the page must collapse and recover exactly once per cycle with zero
velocity across its seam. The suite drives the whole harness β tests/test_qa_site.py runs it
against the committed pages, and against a page built to fail so that every gate is shown able
to fail. Those tests skip where chromium is absent, which is why the browser gates are a local
gate and not a CI one. Installing a browser on the runner was tried and measured (#7): 228 of the
229 gates pass on ubuntu, and the mobile height budget fails there by 9 px β the landing page
shapes to 8,966 px on macOS and 9,009 px on ubuntu, against 34 px of headroom. That is a
text-shaping difference between platforms, and a budget is not something to widen until a runner
fits inside it.
elephant/
βββ scripts/
β βββ split_tape.py the product β fetch, split, rank, print
β βββ bench.py p50/p95, network and aggregation timed apart
β βββ seed.py capture a tape for replay (NOT the demo path)
β βββ render_site.py renders site/ and JUDGE.md from docs/proof/*.json β no hand-typed numbers
β βββ qa_site.py 229 measured gates on the two web surfaces (Playwright)
β βββ site_templates/ landing.html, deck.html, JUDGE.md β {{token}} slots, fail if unfilled
β βββ check_submission_readiness.py placeholder scanner
βββ tests/
β βββ test_split.py the maths + live contract tests
β βββ test_high_signal.py regressions Β· property Β· boundary
β βββ test_fetch_contract.py what get() and pull_swaps() promise their callers
β βββ test_cli.py every main(), through its own __main__ guard
β βββ test_render_site.py the generator, against the real receipts
β βββ test_qa_site.py the browser harness β and a page built to fail it
β βββ test_published_counts.py the test count on five surfaces vs the suite
βββ site/ generated: landing page (/) and deck (/pitch), dated snapshots
βββ data/seed_tape.json a recording. Nothing judged reads it.
βββ docs/proof/ ausd/shfl/bingo/gme.json receipts, live_run.json, benchmarks
βββ JUDGE.md Β· DEMO.md Β· ARCHITECTURE.md Β· FEEDBACK.md
βββ README.md you are here
- Per-swap aggregation from the keyless
/v1/dex/tokens/transactions - Confidence floors so a dust side can never produce a headline
- Backoff across both forms of the anonymous throttle; an exhausted quota explains itself
- Optional keyed escape hatch (
CMC_API_KEY) for a throttled IP β the default stays keyless - Live-run receipts, benchmarks, and property verification
- Cursor pagination that actually advances (
data.lastId),(tx, lgid)identity - Maker attribution per side β the top wallet's share, distribution and top ten
- Landing page and pitch deck (static, dated receipts) at elephant.edycu.dev
- Quadrant map on maker share: distribution Β· accumulation Β· wash-shaped Β· churn
- Hosted live tool with a token input β not built, and not just for lack of time.
CMC omits
Access-Control-Allow-Origin(FEEDBACK.md Β§6), so a browser cannot call it and a server-side hop is required. That hop then puts every visitor behind one IP against a per-IP anonymous rate limit β the same limit that turns CI red when several jobs start at once. Run from a clone, every reader gets their own quota; run from a shared proxy, the tenth reader in a minute gets a 429. A hosted tool would be less reliable than the CLI unless it carried our key, which the CLI deliberately does not need.
- A run measures a recent window, not 24 hours. Depth is
--pagesΓ 100 swaps; the published receipts use 8 pages. For a stablecoin that window spans days, for a meme coin minutes. (Until 2026-09-07 our paginator read the cursor off the last swap instead of the envelope'sdata.lastId, so every earlier number came from a single 100-swap window. Fixed, pinned by a test, and written up in FEEDBACK.md #4.) - A wallet is not an entity. One entity can spread across wallets, which makes the share a floor; a router or aggregator can pool many users into one maker, which inflates it. The number is "share of the side attributed to one address", exactly as the API reports it.
- The anonymous tier throttles, per IP, and reports it as HTTP 500 as often as 429. Run the
watchlist twice quickly and you hit the short throttle; the tool backs off (15 s, 30 s, 60 s)
and retries rather than failing the row, so that run is slow rather than broken. On 2026-09-07,
roughly 3,700 calls from one IP in a day exhausted the quota outright β three backoffs did not
recover it β and the run stopped with a message that says so and names the way through: wait,
or export a free key as
CMC_API_KEY. The key is an escape hatch, never a precondition; every receipt here was taken with every CMC variable unset. Filed as FEEDBACK.md #2. - Wallet counts are per-window β a maker trading in two windows counts once in each. These are distinct-maker counts within the measured tape, not lifetime holders.
- The 24h aggregate fields are not used, deliberately.
24h_buy_volume/24h_sell_volumereconcile withvolume_24hon only 6 of 200 pairs we sampled β 55 are off by more than 100Γ. This project computes from individual swaps instead. Full evidence in FEEDBACK.md #1. - The web surface is a snapshot, not a live tool. elephant.edycu.dev
and its
/pitchdeck carry dated receipts with the command that produced them. CoinMarketCap sends noAccess-Control-Allow-Origin, so a static page cannot call the API; the live capability is the CLI in this repository. /v1/dex/holders/trend/listis not available on the Startup plan, so the holder series is collected daily by us instead.
| Demo video | youtu.be/I1yOhXSutAk β 2:45, captioned. The run at 0:39 is a real keyless execution at real speed; nothing is sped up, and the terminal carries the label saying so. |
| Landing page | elephant.edycu.dev β the receipt beside its raw swaps. A dated snapshot, and it says so on every number. |
| Pitch deck | elephant.edycu.dev/pitch β 11 slides, arrow keys, P for notes. |
| For judges | JUDGE.md β the 30-second path. |
| The receipt | DEMO.md β the live run transcribed, with docs/proof/ausd.json behind it. |
| Submission | dorahacks.io/buidl/48344 β the BUIDL page for Build with CMC, Data and Visualisation track. |
The video's numbers and the landing page's differ slightly β 61.7% against 62.0% β because the 800-swap window moved between the two runs. Same wallet, same $25,459,824, same 12 swaps. The video says so on camera rather than hiding it: a number that moves with the market is the evidence it is live.
MIT β see LICENSE.
Built for the Build with CMC: API Hackathon, Data & Visualisation track. Thank you to the CoinMarketCap team for exposing per-swap maker addresses on a keyless endpoint β that single field is what made this measurable, and our feedback on the rest of the API is in FEEDBACK.md.
