What Feint deliberately does not do, and why. Stating this precisely matters more than the feature list: an emulator that lies about its coverage is worse than one that is small.
Scaleway Object Storage is S3-compatible, so emulating it is not the hard part (MinIO does it already). The obstacle is how clients reach it.
Every other Scaleway product can be redirected with one setting: SCW_API_URL
for the SDK and CLI, api_url for the Terraform provider. Object Storage cannot.
The Terraform provider builds the endpoint in code:
// internal/services/object/helpers_object.go
endpoint := "https://s3." + region + ".scw.cloud"and addresses buckets virtual-host style, https://<bucket>.s3.<region>.scw.cloud.
Redirecting that needs DNS interception plus a TLS certificate the provider will
accept. For a long time this document called that "a project of its own" — an
estimate nobody had made, which is the wrong shape for a refusal here. #76 asked
for the number; the section below is that number, measured. The short version:
the certificate half is cheap and safe, the DNS half is neither, and the whole
blocker reduces to this one product on this one client.
The SDK and the CLI are better off: they honour SCW_S3_ENDPOINT. So an S3
workflow driven by scw or by an SDK can already point at MinIO today; only the
Terraform path is blocked.
The consequence has since been observed live, from a stranger's stack rather
than a fixture (#262,
examples/stacks/surveyed.md): with
SCW_API_URL pointing at this emulator, the provider still sent its
CreateBucket to the real s3.fr-par.scw.cloud, which answered 403 on the
fake credentials. Nothing was created and nothing was billed — but the request
left the machine. A configuration carrying scaleway_object_bucket talks to
the real endpoint no matter where the rest of it is pointed, and any sentence
here promising that traffic never leaves your machine has to carve out this
one product on this one client. The escape-path section below (#280) carries
the full measured list, what warns before the run, and the egress cuts that
were tested rather than described.
The refusal above rested on an unmeasured cost. Measured against the real clients on a hardened Ubuntu workstation, without a byte leaving the machine (a loopback HTTPS server standing in for the cloud), it breaks into four numbers and one surprise: the halves are inverted. #76 wrote the cost as "DNS interception plus a certificate the provider will accept", as if both were the hard part. The certificate is the easy, safe half; the DNS redirect is the hard, dangerous one.
The blocker is "an endpoint built in code that no setting overrides". Swept across three providers' Terraform providers, CLIs and Go SDKs, that set is one:
| product | client | reachable by a setting? |
|---|---|---|
| Scaleway Object Storage | Terraform provider | no — newS3Client hardcodes https://s3.<region>.scw.cloud, virtual-host, no env var, no attribute |
| Scaleway Object Storage | SDK, scw CLI |
yes — SCW_S3_ENDPOINT |
| Exoscale SOS (object storage) | exo CLI, Terraform provider |
yes — sos_endpoint / --sos-endpoint, honoured by both |
| Outscale (all served) | octl, Terraform provider | yes — endpoints.api / OSC_ENDPOINT_API |
| every compute/network API, 3 providers | all clients | yes — SCW_API_URL, EXOSCALE_API_ENDPOINT, Outscale endpoint |
So the coverage cap #76 worried about is real but narrow: it is Object Storage
through Terraform, on Scaleway, and nothing else. Not a dozen scattered
endpoints — one product, one client. MinIO plus SCW_S3_ENDPOINT already covers
the SDK and CLI S3 paths; only this one corner needs DNS/TLS.
(The Exoscale Terraform provider's v2-client split is a different defect — a
missing ClientOptWithAPIEndpoint call, not a hardcoded host — and it is
DNS-independent: an endpoint option fixes it, filed upstream as
#573. It does not belong on this list.)
A local CA was minted with the standard library and an HTTPS listener stood up on loopback. Then each official client was pointed at it. The results, from the server's own handshake log:
| client | knob that works | proven by |
|---|---|---|
scw (Go) |
SSL_CERT_FILE |
scw instance server create completed end to end over local TLS |
exo (Go) |
SSL_CERT_FILE |
GET /v2/zone handshake accepted; refused as x509: unknown authority without it |
| terraform-provider-scaleway (Go plugin) | SSL_CERT_FILE, inherited from terraform's env |
terraform apply created 5 resources over local TLS; the plugin is a separate process and it saw the parent's SSL_CERT_FILE |
curl |
SSL_CERT_FILE, --cacert, CURL_CA_BUNDLE |
200 with, refused without |
oapi-cli (static binary, own trust store) |
none of the CA env vars; only --insecure |
ignored SSL_CERT_FILE and CURL_CA_BUNDLE; irrelevant here, Outscale has no hardcoded endpoint |
Two things settle the certificate question:
- The Terraform plugin inherits the environment. This was the open doubt in
#76 — a provider plugin is a child process go-plugin spawns — and it is
answered:
SSL_CERT_FILEset beforeterraform applyreached the Scaleway provider and it trusted the CA. So the durable, disqualifying option — a CA installed into the operator's system trust store — is not needed for any Go client. One process-scoped environment variable does it. SSL_CERT_FILEis scoped to the one command. It dies with the process, touches nothing else, and leaves no trace — exactly the property this tool's pitch (no account, no bill, no trace) requires.
No. A DNS server is the expensive answer and the measured case does not need it. What it needs is to make one hardcoded name resolve to loopback, and that is where the cost actually lives, because on a modern hardened Linux there is no per-process, disposable, unprivileged way to do it for the exact client that matters:
| mechanism | scope | verdict for the Scaleway S3 case (AWS SDK Go v2, static, pure-Go resolver) |
|---|---|---|
curl --resolve / --connect-to |
one command | works — proven landing s3.fr-par.scw.cloud and <bucket>.s3.fr-par.scw.cloud locally with a wildcard cert — but curl only; the SDK has no equivalent |
HOSTALIASES (glibc) |
one process | only the cgo resolver, and only single-label names — useless for a dotted FQDN |
LD_PRELOAD getaddrinfo shim |
one process | only cgo-resolver binaries. Measured: scw is dynamically linked with cgo getaddrinfo (interceptable); exo is static pure-Go (not). Terraform providers build CGO_ENABLED=0 — not interceptable |
network namespace + bind-mounted /etc/hosts |
disposable | works, with one named AppArmor profile. The station sets apparmor_restrict_unprivileged_userns=1, and the first measurement read unshare -r → EPERM as "namespaces are blocked". Re-measured: the namespace is created and root maps fine; what the restriction removes is capabilities inside it, so ip link add answers RTNETLINK: Operation not permitted. tools/install/apparmor/feint grants userns to one binary path and the interface is created — see below |
edit /etc/hosts |
whole machine, persistent | works for every client, but it is a durable change to the operator's machine — the thing the pitch forbids |
So for the one blocked client — the pure-Go, statically linked AWS SDK inside the
Scaleway Terraform provider, which offers no --resolve — the cheap scoped
mechanisms (curl --resolve, HOSTALIASES, an LD_PRELOAD shim) all miss it,
each for a different reason. What reaches it is a network namespace, and that
one is available.
The first measurement concluded that this station blocks unprivileged user
namespaces, from unshare -r answering EPERM. Re-measured with the exit codes
read properly — the earlier reading took $? through a pipe, which is the
false-verdict shape this repository keeps meeting — the three questions separate:
| asked | answer |
|---|---|
unshare --user --net true |
0 — the namespace is created |
| a uid map putting this user at root inside it | works — id -u answers 0 |
ip link add dummy0 type dummy inside it |
RTNETLINK answers: Operation not permitted |
So the restriction does not stop the namespace. Ubuntu transitions the process
into the stock unprivileged_userns profile, whose first line is
audit deny capability — the namespace exists and is powerless, which is a
better design than refusing outright and reads exactly the same from a script
that only checks whether unshare failed.
The fix is one named profile, not a sysctl.
tools/install/apparmor/feint grants userns to the feint binary path and
nothing else, in the same shape the distribution ships for lxc-unshare,
bwrap and a dozen others. Setting kernel.apparmor_restrict_unprivileged_userns=0
would lift the restriction for every program on the host; this lifts it for one
path.
Measured on a witness binary, both ways: with the profile loaded the interface is
created; unload it and RTNETLINK: Operation not permitted comes straight back.
sudo install -m 0644 tools/install/apparmor/feint /etc/apparmor.d/feint
sudo apparmor_parser -r /etc/apparmor.d/feintNothing in feint requires it. Without the profile, the paths that need a namespace refuse and say so; every other command works unchanged.
The certificate half is as cheap as #76 guessed. A CA, a leaf covering
s3.<region>.scw.cloud and *.s3.<region>.scw.cloud, and an HTTPS listener
are under 100 lines of pure crypto/x509 and crypto/tls, no dependency.
One caveat the wildcard exposes: *.s3.<region>.scw.cloud is single-level, so a
bucket name containing a dot (my.bucket.s3.<region>.scw.cloud) is not covered —
measured, curl refuses it. There is no DNS server in the standard library, but
per number 3 the measured case does not need one.
This is no longer a projection: internal/proxy/intercept.go mints exactly that
CA and leaf from the standard library alone, and feint proxy --intercept serves
the recorder over TLS with it. See number 6.
A process that trusts a feint-minted CA and resolves a real cloud hostname to
loopback is exactly how a real terraform apply silently hits a local emulator.
The measurement moves the danger: the certificate is safe because SSL_CERT_FILE
is process-scoped, but the name redirect is the hazard, and it is the half
that resists being scoped. /etc/hosts is machine-wide and persistent; an entry
left behind sends the operator's next real apply to a dead local port (loud) or
a stale emulator (silent, and the exact failure this project exists to avoid).
Whatever is ever retained here must scope the redirect as narrowly as the cert —
a devcontainer with its own hosts file, an explicit and temporary entry the
operator makes and removes — and must never install a CA into the system store or
edit /etc/hosts on the operator's behalf. Number 6 is how that scoping is done
in practice: the redirect lives in a namespace the client owns, never in the
operator's own /etc/hosts.
The two halves stopped being a projection. feint proxy --intercept <host>[,<host>] serves the recorder (see docs/proxy.md) over HTTPS with a
certificate internal/proxy/intercept.go mints from the standard library alone:
a short-lived, non-CA leaf covering the names a redirected client will address. It
writes the CA to a temporary file for SSL_CERT_FILE, prints the redirect recipe,
and removes the CA on exit. The scoping number 5 demands is structural here — the
command installs nothing into the system store and writes no /etc/hosts, because
it has no code that could.
That leaves one question, and number 3 answered it for the station: where the name resolves to loopback, disposably and without a durable trace. Measured across the places an operator actually has, rather than assumed:
| place | verdict | how, and what it costs the machine |
|---|---|---|
| a rootless container (podman, docker) | works, no host trace | the container has its own network namespace and its own /etc/hosts; --add-host=<name>:host-gateway redirects the name and SSL_CERT_FILE carries the CA. It reaches the namespace through the setuid newuidmap helper, a path the apparmor_restrict_unprivileged_userns sysctl does not touch, so it needs no profile — measured here, grep <name> /etc/hosts on the host stays empty afterwards |
| feint's own user + network namespace | works with the named profile | tools/install/apparmor/feint (number 3); feint itself opens the namespace, so the profile covers it. A separately-launched unshare or ip is not covered, by design, and that is the right architecture: a tool that asks the operator for sudo unshare has already lost no trace |
Incus container (--vm incus) |
works | a container is a namespace; already piloted by this repository |
a GitHub hosted runner (ubuntu-24.04) |
the container path works; feint's own namespace needs the profile | measured, not assumed (workflow run 31791022679, .github/workflows/userns-probe.yml, image 20260810.271, kernel 6.17-azure): the runner carries the same apparmor_restrict_unprivileged_userns=1. Bare namespace creates (exit 0), the root map is refused (exit 1) — an unconfined unshare is as powerless here as on the station. Rootless podman works out of the box and leaves no host trace (grep on the host stays empty), so the container row is the portable CI path. feint's own namespace would need tools/install/apparmor/feint loaded first, which a runner permits through passwordless sudo |
The through-line: in every permitted place the redirect is as disposable and as
narrowly scoped as SSL_CERT_FILE itself. It lives in a namespace the operator's
next real apply never enters, and it vanishes when that namespace does — which is
exactly what number 5 says any retained redirect must do.
Why this was built now: #92, not S3. The driver was the recorder, not Object
Storage. feint proxy records a real client by being its configured endpoint, but
a cloud that republishes its own address in a response body walks the client away:
Exoscale's GET /v2/zone hands back https://api-ch-gva-2.exoscale.com/v2, the
client follows it, and a session worth about ninety exchanges recorded eight
(#92). The plain proxy cannot hold a client it does not resolve for. Interception
can: with the republished name resolving to the proxy in a namespace of its own and
the CA trusted, the client follows the republished address straight back and the
whole session is kept. TestInterceptionRecordsThePostHandoffExchanges reproduces
it without an account — one session driven twice against the same recorder, the
republished name resolving to the proxy, then away — and records the whole thing
one way, only the pre-handoff exchange the other. Eight-of-ninety, and the fix, on
one run.
Everything above rests on the redirect being the hard half. feint proxy --forward (#336) changed one term of that: a Go client that installs no
Transport inherits http.DefaultTransport, which honours HTTPS_PROXY, so a
compiled-in name can be intercepted with no DNS trick and no /etc/hosts.
Whether the Scaleway Terraform provider's S3 client is such a client was
unknown, and #336 said so rather than predicting.
It is. Measured on Linux on 2026-08-21, terraform 1.15.4 with
scaleway/scaleway 2.81.0, against this repository's own emulator — no account,
no real endpoint, the public fake credentials of
tools/conformance/scaleway/fake-credentials.env. The recipe is the two
variables and nothing else, plus the mapping #357 added so the terminated host
lands on the emulator instead of the real cloud:
feint serve --addr 127.0.0.1:4760 &
feint proxy --record s3.jsonl --addr 127.0.0.1:4761 \
--forward 's3.fr-par.scw.cloud=http://127.0.0.1:4760,*.s3.fr-par.scw.cloud=http://127.0.0.1:4760'
# then, on a scaleway_object_bucket:
export HTTPS_PROXY=http://127.0.0.1:4761 SSL_CERT_FILE=/tmp/feint-intercept-ca-….pem
terraform apply -auto-approveOne tunnel terminated, one exchange recorded, and the User-Agent names the client beyond argument:
{"seq":1,"method":"PUT","path":"/","host":"feint-346-measurement.s3.fr-par.scw.cloud",
"status":404,"mounted":false,
"req":{"headers":{"Authorization":"REDACTED","X-Amz-Content-Sha256":"REDACTED",
"User-Agent":"aws-sdk-go-v2/1.43.4 … api/s3#1.107.0 terraform-provider-scaleway/2.81.0"}}}The PUT / is CreateBucket, virtual-host style, and the 404 is this
emulator answering: feint serves no object storage, so no pack claims that
route ("mounted": false). Nothing left the machine.
Read the two failures apart, because only one of them closes the door. A negative result would have to say which negative it was, and the control run shows exactly what the other one looks like. Driven again with the same environment against a proxy that does not name the S3 host:
recorded 0 exchange(s)
0 tunnel(s) terminated
12 connection(s) were refused because --forward does not name their host:
feint-346-measurement.s3.fr-par.scw.cloud
Error: operation error S3: CreateBucket, exceeded maximum number of attempts, 3,
request send failed, Put "https://feint-346-measurement.s3.fr-par.scw.cloud/": Forbidden
The client emitted CONNECT for the compiled-in name twelve times and the proxy
turned each one down. That is "arrived, and was refused" — the door opened. The
other negative, "did not honour the proxy", would have left the proxy with no
CONNECT at all and the client complaining about a certificate, which is what
macOS produces (number 2's warning in proxy.md) and why this was
measured on Linux.
So the hard half of #76 has a third door, and this one costs an operator
nothing: no namespace, no /etc/hosts, no privileged port, no AppArmor profile.
It does not by itself retain object storage — what it removes is the
redirect from the cost, not the S3 surface from the work — but the arbitration in
the verdict below now rests on the coverage argument alone, since the ceremony it
weighed has gone to zero. Reopening it is its own issue, with its own numbers.
Refused, now with numbers behind it — and the refusal changes shape twice.
Object Storage through Terraform stays out, but no longer for the reason first
written. The certificate was never the project it was called: it is feint proxy --intercept, under 100 lines of standard library, accepted by every Go client
including the Terraform plugin through one process-scoped environment variable. And
the name redirect, first read as undeliverable without touching the machine, is
deliverable after all — in a namespace the client owns (number 6): a rootless
container, or feint's own namespace under one named profile. And since #346
(number 7) it is not even that: the one blocked client honours HTTPS_PROXY, so
the redirect costs two environment variables. What is left is not a feasibility
wall and no longer an operator ceremony either — it is the coverage argument
alone: whether the S3-through-Terraform corner is worth emulating an S3 surface,
which is a product call and the only thing still holding the refusal.
The blocker is one product on one client, so the MinIO + SCW_S3_ENDPOINT page
remains the right answer for the S3 workflow, and the refusal caps coverage by
exactly one corner rather than quietly bounding the whole project. If
Object-Storage-through-Terraform is ever wanted, it is a roadmap item with a
named owner and a shape already measured, and number 7 shortened it again:
SSL_CERT_FILE and HTTPS_PROXY, both process-scoped, both proven on that exact
client — never a system trust-store install, never a hosts file this binary edits
itself, and now not even a namespace. What is left to cost is the S3 surface
itself; that is a product call, and it is no longer an unmeasured one.
Every sentence this project writes about locality has to be the one that is true: the APIs Feint serves run locally; a client can compose its own endpoint for a product outside that scope, and then its requests go where they always went. From the outside the two runs are indistinguishable — that is the whole problem — and the failure that would actually hurt is not the 403 the survey measured on fake credentials. It is somebody with real credentials in their environment — a developer's shell, a CI job that also deploys — running a stack they believe is sandboxed.
| path | measured | what warns today |
|---|---|---|
Scaleway Object Storage through Terraform: scaleway_object_* hardcodes https://s3.<region>.scw.cloud (top of this file) |
live on a surveyed stack (#262, flatcar-k3s): CreateBucket at the real endpoint, 403 on fake credentials. Reproduced for #280 with egress cut to a dead proxy: the same apply created its instance IP on this emulator and died on Put "https://<bucket>.s3.fr-par.scw.cloud/" — one run, half local, half not |
feint doctor and feint env scaleway, from the stack directory: the configuration's own text names the resource family |
| an S3 state backend or an aws provider pointed at real object storage | three surveyed stacks: kubic (state on Object Storage, endpoint in backend.conf), eu-data-platform (state on sos-ch-gva-2.exo.io, inline), platform (aws provider at sos-<zone>.exo.io, 403 at the real endpoint) |
the same scan, when the host is written in the .tf text — the kubic shape keeps its endpoint in a backend.conf the scan does not read, and stays invisible to it |
Outscale, OSC_PROFILE set: provider 1.1.x reads ~/.osc/config.json and ignores OSC_ENDPOINT_API |
#286, on 1.1.3: the plan left for api.<region>.outscale.com while the emulator received nothing |
feint doctor and feint env outscale, since #286 — this one lives in the shell, not in the stack |
| the Exoscale Terraform provider's split client, up to v0.70.0: egoscale v2 built with no endpoint option | #262/#284, section below: an apply splits between this emulator and the real cloud; fixed upstream in v0.71.0 and measured here (#644) | the emulator itself refuses a provider below v0.71.0 by user agent, naming the floor, and feint env exoscale names it too (#701) — the one escape that must pass through the front door to do damage, so the front door is where it is stopped |
The scan behind the first two rows reads the Terraform files around feint doctor and feint env (*.tf, *.tf.json, *.tofu, comment-stripped, dot
directories skipped) and matches the measured signatures each pack declares —
internal/providers/*/stackhazards.go. Warnings, never failures, and three
silence rules the tests hold: a directory whose own top level carries no
Terraform file produces nothing — Terraform only runs where root module files
sit, so a workspace that merely contains projects is not a stack
(TestADirectoryOfProjectsIsNotAStack, while a rooted stack's modules/ are
scanned); a commented-out resource produces nothing
(TestAStackHazardInACommentStaysSilent) — a warning that fires on dead text
is a warning people learn to ignore; and this repository's own fixtures scan
clean. The ok row states its own scope ("checks the measured list") because
that is all it checks.
Feint controls neither the client's process nor its DNS. feint proxy sees
every request that reaches it and, by construction, none that does not. The
scan above reads text, so it cannot see a value passed at runtime
(-backend-config, a variable), a module fetched at init, or the next product
whose client composes its endpoint upstream tomorrow — the nightly drift scan
reads SDK surfaces, not endpoint construction, so it will not see that one
either. In general, the escape is undetectable from inside the emulator, and
an approximate guard would be worse than none: it would license exactly the
belief it fails to protect. What exists is a measured list, said as such,
plus the one boundary that does not depend on the client cooperating — the
network the run executes in.
The tripwire costs one line and no privilege, and it is how the reproduction above was run — the escape became a loud failure naming its destination instead of a silent success:
HTTPS_PROXY=http://127.0.0.1:9 NO_PROXY=127.0.0.1,localhost terraform applyError: … CreateBucket, … Put "https://feint-escape-repro.s3.fr-par.scw.cloud/":
proxyconnect tcp: dial tcp 127.0.0.1:9: connect: connection refused
The emulator, on loopback, is reached directly through NO_PROXY; everything
else dies on a proxy that is not listening, and the error names the host that
was contacted. Scope, honestly: proxy variables are honoured, not enforced —
measured on Terraform with the Scaleway provider (above) and on the Exoscale
provider (#284, same technique); a client that ignores them walks straight
past. A tripwire, not a boundary.
The boundary is a network with no route out. Measured on rootless podman (4.9.3, netavark), no privilege and no host trace:
podman network create --internal feint-noegress
# feint on it (feint serve --addr 0.0.0.0:4599 --expose-to-network), the client beside it:
# GET /_feint/health → HTTP/1.0 200 OK
# connect s3.fr-par.scw.cloud:443 → "Network is unreachable", immediatelyOne caveat the measurement exposed: on an --internal network the container's
DNS still resolves real names (aardvark-dns forwards to the host's
resolver), so the cut is at connect, not at lookup — a name leaks as a query,
a connection goes nowhere. In CI the same property is whatever your runner's
egress policy provides; a hosted runner with open egress protects nothing, and
no flag here can change that. That sentence is the honest end of this section:
where the network permits the escape, Feint can at most name the measured
paths before the run — which is what feint doctor now does.
Two operations of baremetal/v1 are served, ListServers and ListOffers,
and the thirty-five others of the product are declined with their reason in
the pack. The listing exists because a fleet inventory has to enumerate the
product to describe a mixed fleet, and it was measured before it was written,
on 2026-09-06: scw baremetal server list (2.56.3) calls
GET /baremetal/v1/zones/{zone}/servers?order_by=created_at_asc&page=1, and
the Python SDK's list_servers_all (scaleway 2.12.0, what the
stephrobert.scaleway inventory plugin drives) sends organization_id on
every call and pages with page alone until a page comes back empty.
The catalogue is what the fleet declares. The issue sized the product at
one route; the CLI's source says two. Once ListServers answered, the same
command called GET /baremetal/v1/zones/{zone}/offers?page=1&subscription_period=unknown_subscription_period
and printed the catalogue's name for each server's offer_id in place of the
server's own offer_name (serverListBuilder, scaleway-cli), so a 501 there
failed the listing and an empty catalogue would erase every offer name the
operator seeded. ListOffers therefore answers one offer per distinct
offer_id among the zone's seeded servers, carrying the seed's name and
nothing invented around it: stock: empty, enable: false, no price, no
disk, no CPU. Any client reading GET /offers gets the same offers;
scw baremetal offer list is not one that gets through, because it then reads
GET /product-catalog/v2alpha1/public-catalog/products?product_types=elastic_metal,
a third product this emulator does not serve, and fails on that 501
(measured 2026-09-06). The offer catalogue is served for the join the server
listing makes, not for the offer command.
A server is seeded, never created. There is no CreateServer: ordering
Elastic Metal is a paid commitment on hardware the cloud delivers in minutes
to hours, and an offer catalogue with nothing in stock behind it is the
catalogue trap one product out. A bare-metal server enters through the state
door, serve --state <file> or PUT /_feint/state, as a resource of Kind
baremetal/server whose Attrs use the SDK's own field names:
{
"format": "feint-snapshot",
"version": 1,
"resources": [{
"ID": "6f1d2c3b-4a5e-4f60-8b7c-9d0e1f2a3b4c",
"Kind": "baremetal/server",
"Tenant": {"Provider": "scaleway", "Project": "11111111-1111-1111-1111-111111111111", "Zone": "fr-par-1"},
"State": "ready",
"Created": "2026-09-01T10:00:00Z",
"Updated": "2026-09-01T10:00:00Z",
"Attrs": {
"name": "db-1",
"offer_name": "EM-A210R-HDD",
"tags": ["db", "prod"],
"ips": [
{"address": "203.0.113.10", "version": "IPv4", "reverse": "db-1.example.net", "reverse_status": "active"},
{"address": "2001:db8::10"}
],
"boot_type": "normal",
"ping_status": "ping_status_up"
}
}]
}State is the server's status (the SDK's enum: ready, stopped,
delivering, ordered, ...), Tenant.Project its project_id, and every
field of the SDK's Server the seed omits is answered as the SDK's own "not
said" member (unknown_boot_type, ping_status_unknown, null for
install) rather than a plausible value. An address whose seed names no
version gets the family of the address itself. A restored resource is an
untrusted input: Attrs are decoded through the SDK's types, so a value of the
wrong type is dropped, never served under the SDK's field name.
What is not measured. No recording of this product exists in corpus/,
so three answers are this pack's convention rather than the cloud's: a zone
outside the six the product declares (fr-par-3, nl-ams-3, pl-waw-1,
it-mil-1) is refused with the invalid_arguments every other zoned list of
this pack answers, where the inventory plugin would read a 404 as "product
unavailable here" and a 400 as a discovery failure; name filters by
substring, instance/v1's documented reading; and a status value outside
the enum matches nothing rather than being refused. A recording that
contradicts one of these is the thing to change it, and GET on an empty
zone is the cheapest recording a real account can make.
Kapsule (Scaleway) and SKS (Exoscale) are the most-demanded unserved surface the survey measured (#262): SKS alone made two of the five Exoscale stacks not applicable, and the survey's own conclusion is that the public Scaleway ecosystem lives in Kapsule, RDB, LB and Object Storage. High demand does not say what "supporting it" means, so #283 asked the question that decides the cost: how far must the cluster work after it is created?
That is measurable, and it was measured twice, on 2026-08-19.
On the surveyed stacks: three of three continue past Terraform.
CentraleSupelec/kubic wires kubernetes and helm providers from
kubeconfig[0] and installs nine Helm releases (Argo CD, cert-manager,
Prometheus, Loki, Vault, Velero). datamindedbe/eu-data-platform feeds a kubernetes provider from
exoscale_sks_kubeconfig and creates namespaces and secrets in the same
layer. camptocamp/terraform-exoscale-sks does not even wait for a provider: a
null_resource polls the cluster endpoint's /healthz for five minutes and
fails the apply on timeout, then shells out to the exo CLI for a kubeconfig.
Not one observed stack treats the cluster as a record.
On the wild population: about half of Kapsule and two thirds of SKS. A
GitHub code search for resource "scaleway_k8s_cluster" and resource "exoscale_sks_cluster" in HCL returned 127 and 37 repositories. After
excluding the providers' own repositories and demos, verbatim copies, one
Scaleway mock, and collapsing classroom, interview-task and same-author
duplicates into one unit each:
| provider | distinct units | continue into kubernetes/helm/kubectl in the same configuration |
stop at cluster + pool + outputs |
|---|---|---|---|
| Scaleway Kapsule | 102 | 50 | 52 |
| Exoscale SKS | 20 | 13 | 7 |
The stop column is a floor, not a population that a CRUD emulation would
serve: it is dominated by classroom exercises, and the substantial stacks in
it output the kubeconfig with instructions to run kubectl next
(jpetazzo/container.training's lab harness consumes it seconds later). A
cluster is created to be talked to.
The verdict, and it is the same for both providers: refused at every level short of a real control plane.
- CRUD-only (create, read, update, delete, answer the stored attributes) is
refused. The clients themselves force the lie:
scaleway_k8s_poolwaits for the pool — nodes included — to reachreadyby default (wait_for_pool_ready, measured in the provider'spool.go), and the camptocamp module blocks on a live/healthz. Serving the API without a control plane means answeringreadyfor an API server that does not exist and issuing a kubeconfig that points nowhere — a lying 200 by construction, on exactly the field the majority of the measured demand consumes one resource later. A 501 naming the product is honest; a ready cluster with no API server is not, and that distinction is this project's founding rule. - Reproducing the managed service (versions, autoscaling, CNI, CCM, CSI, upgrades, maintenance windows) is refused permanently: feint would become a different product.
- The only admissible shape is a real local control plane behind
--vm, through the machine runtime, handing back a kubeconfig that answers — the same rule that makes a public address here "the provider's value, made to answer on the host". That shape is not scheduled: it puts somebody else's software lifecycle inside this project (the version the API claims versus the one the runtime ships, nodepools as real joined nodes, aLoadBalancerservice with no CCM behind it sittingpendingforever), and each of those is a place to start lying at one remove. If it is ever built, it is its own issue with those costs measured first.
Until then the refusal is served where a reader's client hits it: the Exoscale
pack declines every sks-* operation by name with this reason
(internal/providers/exoscale/pack.go), and /k8s/v1/ answers with a
Scaleway error envelope from the not-served prefix list
(internal/providers/scaleway/pack.go). The measured cost of the refusal is
the survey's: platform-shaped stacks stay not applicable, and serving the rest
of their products would not free them — eu-data-platform needs SKS and
DBaaS and SOS, so it stays blocked whatever
#284 decides.
Server types, prices and images are a small fixed table
(internal/providers/scaleway/catalog.go). The emulator has no fleet, no
inventory and no price list. It serves a table anyway because the clients read
it before creating anything: a 404 there makes scw instance server create
fail outright.
The 0.10.0 survey (#279) measured that the table is more than pre-flight
scenery: the Terraform provider validates a server's type against
/products/servers before it creates anything, so a type outside the table
fails a stack at plan — two of five surveyed stacks died exactly there.
Three consequences, each a decision this page records:
- The rows are measured, not invented. The served types are an excerpt of
the real answer to
GET https://api.scaleway.com/instance/v1/zones/fr-par-1/products/servers(public, no authentication), captured 2026-08-19 and embedded verbatim ascatalog_servers.json: every family the emulator carries, every size of it (PLAY2, DEV1, GP1, PRO2), plusSTARDUST1-S, which a surveyed stack named. The one deviation isper_volume_constraint, served empty and declared to the shapes gate, because a bound for local volumes this emulator never attaches would enter the client's size arithmetic with nothing behind it. - What the emulator will never enforce about a type stays unenforced. The RAM, CPU count, GPU count, bandwidth and prices of a row are answers, not behaviour: nothing meters a byte or bills a cent, and with a machine runtime every server boots the same class of container whatever its type claims. Treat any capacity, price or availability answer as decoration — it is now accurate decoration, which is strictly less misleading, and still decoration.
- One table serves every zone, and the real cloud varies by zone. The real
fr-par-1 lists 136 types where fr-par-3 lists 41, and
STARDUST1-Sexists in three zones of nine. A plan namingSTARDUST1-Sinfr-par-2passes here and fails against the real region. Zone-accurate inventory would mean carrying nine tables of a moving target; the divergence is accepted and stated instead.
The recording of that difference is committed, and it is a value rather than
a shape. corpus/scaleway/scw-cli.jsonl carries the real fr-par-1 answer,
all three pages of it, beside the emulator's own. The 118 types this table does
not stock are therefore measured rather than asserted — and the two gates that
read that recording agree on what they are: a key of that map is data, not a
field, which is transcript.DataKeyed, shared by feint shapes --check and
feint corpus --check so the same artefact cannot be graded two ways. Read as
fields they were 127 of the 136 findings the first corpus run produced, saying
one thing 127 times over everything else the file had to report (#355). What
is graded is the shape of an entry both sides carry, which is why the missing
per_volume_constraint.l_ssd bound is declined explicitly, in both spellings
the gates join on.
There is no ARM row, and that is a refusal with a reason, not a gap. The
one ARM type a surveyed stack asked for, COPARM1-2C-8G, is absent from all
nine zones of the real catalogue (measured 2026-08-19, every page, while
genuinely end-of-service families — START1, VC1, X64 — are still listed with
end_of_service: true): Scaleway withdrew the family, and resurrecting it
here would let a plan pass that production refuses.
TestTheRetiredArmFamilyStaysRetired keeps it out. The current ARM families
(BASIC2-A*, STANDARD2-A*) are real and could be carried the day a stack
asks — but an arm64 row costs more than a paste: the emulated marketplace is
x86_64 (its arch filter truthfully answers arm64 requests with an empty
list, #277), and a machine runtime boots containers on the host's
architecture, so an arm64 row must arrive with an arm64 image story or its
servers can never boot. The demand-driven rule above applies, with that bill
attached.
An unknown type is still accepted at create. The section below argues this
for image identifiers, and the reasoning transfers whole: a configuration that
names a type this table lacks — including a real one the excerpt has not
caught up with — must not die on the one thing that has nothing to do with
what it is testing. The clients that care already refuse client-side against
/products/servers (that refusal is the measured wall of #279, and growing
the table is its fix); an emulator-side refusal would add a second wall for
raw SDK users and catch nothing the first does not. The real cloud does
refuse an unknown type, so this is a divergence, recorded here on purpose.
The Outscale catalogue belongs to the emulated account, so AccountAliases: [Outscale] selects nothing (#700)
ReadImages and ReadSnapshots serve Filters.AccountAliases since #700, on
the owner's alias every image and snapshot carries. There is one owner in
this emulator, and the fixed catalogue is its by decision (catalog.go:
PermissionsToLaunch names the emulated account as the owner of every
catalogue image), so the alias the catalogue publishes is the account's own.
Two consequences a client should expect:
- the contract's own
ReadImagesexample,AccountAliases: [Outscale]withImageNames: [Ubuntu*, RockyLinux*], answers an empty list here: nothing this emulator serves is Outscale's, and it says so rather than relabel a catalogue whoseAccountIdis the account's; - selecting the account's own images apart from the catalogue by owner is not possible here, since both halves have the same owner. A residue check that wants the client's half alone lists without the filter and compares against the catalogue, which is fixed.
Giving the catalogue Outscale's identity would need Outscale's account id,
which the recorded corpus redacts along with every other AccountId; it is
not invented. That self is accepted as an alias by the real API was not
measured either, and it is not accepted here.
Measured on fr-par, 2026-09-02 (#394). Three refusals of the load balancer
product, and not one of them carries the scw error envelope:
| request | fr-par answers |
|---|---|
POST /lb/v1/zones/fr-par-1/routes, frontend that does not exist |
403 {"message": "Permission denied"} |
GET /lb/v1/zones/fr-par-1/frontends/{id}, absent |
404 {"message": "frontend not Found"} |
GET /lb/v1/zones/fr-par-1/lbs/{id}, absent |
404 {"message": "lbs not Found"} |
No type, no details, no resource. This emulator answers all three with the
envelope, and the reason is structural rather than chosen:
contracts/scaleway.json declares one errorSchema for the whole document,
scw.ResponseError with required: ["type"], extracted from the SDK — and
internal/probe validates every refusal against it. A body without type fails
the probe, which is the gate being right about the document it was given, and the
document being a document rather than a recording.
What is faithful is the status, which is the half a client branches on:
POST /routes answers 403 here now, where this emulator used to answer 404 and
where the two reads beside it answer 404 on both sides. Terraform retries a 403
and a 404 differently.
Lifting this means giving the contract an errorSchema per product rather than
per document, and a way to say "this product's refusals are not validatable
against the SDK's struct" that a recording rather than a document arbitrates.
RecordedFields already does the opposite direction — a field a recording
carries and the document does not declare — and the symmetric one does not exist
yet.
A create that names an image, a template or a machine type the emulator has never
heard of succeeds. Scaleway answers 201 for an image UUID that exists
nowhere, Outscale accepts ami-99999999, Exoscale accepts an invented template
id. The real clouds refuse all three.
This is deliberate, and it is the limitation on this page most likely to bite. The reason is the same one that makes the catalogue fiction: the emulator has no inventory, so the only ids it could recognise are the handful it invents. A configuration that hardcodes a production image UUID — the most common way a team first points an existing Terraform at the emulator — would then fail on the one thing that has nothing to do with what they are testing.
The cost is real and worth stating plainly: a typo in an image id is not caught here, and will be caught in production. Feint proves that a request is well-formed and that the response is shaped like the provider's, not that the resources it names exist.
If you need that check, declare your catalogue (#126):
feint serve --strict-catalog catalog.json{
"scaleway": {"images": ["debian_bookworm"], "types": ["DEV1-S", "PLAY2-PICO"]},
"outscale": {"images": ["ami-fe1a7001"], "types": ["tinav6.c2r4p2"]},
"exoscale": {"templates": ["11111111-1111-4111-8111-111111111111"], "types": ["21624abb-764e-4def-81d7-9fc54b5957fb"]}
}The file is the operator's own contract, versioned in their repository: per
provider, the identifiers each kind may name. A kind the file does not mention
is not checked, so the half you know can be asserted alone; a kind no pack
checks ("image" for "images") is refused before anything listens, because a
line nobody enforces reads exactly like one somebody does. Without the flag,
nothing above moves: the compatibility mode is byte-identical, and the
conformance suite runs in it.
With it, a create naming an identifier outside the declaration is refused in each cloud's own shape, and so is the lookup the client makes first, so the real client renders its ordinary not-found path:
| pack | what is checked | the refusal | measured? |
|---|---|---|---|
| Scaleway | images (label or marketplace UUID, either form covers the other), types |
404 not_found on image for GET /images/{id}, the marketplace label lookup and POST /servers; 400 invalid_arguments on commercial_type; /products/servers lists the declared types only |
the not_found shape on an identifier naming nothing, recorded 2026-08-21 (corpus/scaleway/scw-refusals.jsonl); the exact answer to a type fr-par does not sell, no |
| Outscale | images, types |
CreateVms: 400, code 5023, InvalidResource, "The ImageId '…' doesn't exist."; 400, code 4001 on a VmType; ReadImages and ReadVmTypes list the declared entries only |
the image refusal, recorded 2026-08-21 (corpus/outscale/oapi-cli-refusals.jsonl); the type refusal, no |
| Exoscale | templates, types |
create-instance and create-instance-pool: 404 {"message": …}; list-templates, get-template, list-instance-types and get-instance-type answer the declared entries only |
the 404 shape of a resource that does not exist, recorded; the cloud's answer to a create naming a template it does not offer, no |
An object the client registered itself (an image cut from a machine, a template it uploaded) is the client's own and is never refused by the declaration: the declaration is about what the fixed catalogue may answer for. Zones are not a kind yet.
A project identifier is the exception, since #391. It used to be the same
trade — GET /account/v3/projects/{id} echoed any identifier, so a stack
carrying a production project id kept working — and that stopped being tenable
the day the creates started refusing a project nobody holds, which is what the
cloud does:
| request naming a project that does not exist | fr-par answers |
|---|---|
instance/v1 create (server, IP, security group, placement group) |
403 permissions_denied, details: [{resource: project, action: read}] |
lb/v1 create |
403, details: [{resource: loadbalancer, action: write}] |
block/v1 CreateVolume |
403, details: [{resource: volume, action: write}] |
iam/v1alpha1 CreateSSHKey |
403, details: [{resource: ssh_key, action: create}] |
vpc/v2, vpcgw/v2, ipam/v1 creates |
404 not_found, resource: project |
GET /account/v3/projects/{id} |
404 not_found, resource: **project_id** |
DELETE /account/v3/projects/{id} |
404 not_found, resource: **project** |
Measured 2026-09-02. An echoing read beside a refusing create is worse than either on its own: the client resolves a project with 200 and is then refused every resource under it.
So the register decides, and it holds two kinds of project. What the
operator declared (feint serve --projects, cloud.projects in feint.yaml),
and what a client created — POST /account/v3/projects is served now, with
PATCH and DELETE beside it. A project that still holds a resource refuses
its delete with 412 resource_still_in_use, which is the cloud's own answer and
the reason the route is worth serving at all.
A stack carrying a production project id declares it, and that is the door the echo used to be:
feint serve --projects "platform-prod=8fad27e6-d8d8-45e5-8439-36ef7ff02fd6"An entry with no = keeps deriving its identifier from the name, so every
declaration written before #391 is unchanged. A request naming no project is
never refused: it is filed under the default, so a client configured from
feint env never meets any of this.
The list filters over that register: GET /account/v3/projects answers what
the operator declared plus what has been created, and a name or project_ids
filter that names something else answers an empty list. A filter is a question
about what exists, and emptiness is a truthful answer to it. A stack whose
project_name is not default needs that name declared — --projects platform-prod — or it fails on the provider's FindExact. The organization is
never compared at all, for the reason listSSHKeys records: scw names its own
configured organization on every list, and comparing told the CLI that the key it
had just created did not exist.
What this does not change is the rest of the section above. An image, a template or a machine type the emulator never minted is still accepted, and for the reason that has not moved: the emulator has no inventory of those. A project is different — it is the boundary every other resource is filed under, and the register is an inventory of exactly it.
What an unknown identifier can no longer do is boot a substitute. Measured
in #83, on all three packs: with a runtime configured (--vm incus, incus-vm,
incus-ovn), an image identifier no catalogue held was silently replaced at
boot — ask for Alpine, boot Ubuntu — while the API kept reporting the identifier
the client sent. Scaleway's resolution matched labels by substring, so centos,
rocky and ubuntu_focal all became Ubuntu 22.04 without a word.
Since then the create still succeeds and the boot refuses: the machine reaches
the provider's own failed state (stopped on Scaleway and Outscale, which
declare no error state for a machine; error on Exoscale) and the emulator's
log names the identifier. The state published is the one the effect produced,
not the one the intention aimed at. With --vm off, the default, nothing boots
and nothing changes — the control plane keeps accepting, exactly as this
section promises.
Two alternatives lost, and why:
- Refusing at the create, as the real clouds do, would turn the emulated catalogue into a whitelist and break the paragraph above: a configuration hardcoding a production image UUID must keep applying, because that is the first thing a team points at the emulator.
- Substituting out loud — a warning in the log, a mark on the resource —
keeps a machine whose
/etc/os-releasecontradicts the API for as long as it runs. A cloud-init that installs an Alpine package, or a playbook that branches on the OS family, still gets the wrong operating system with every signal saying success. A boot that fails with a stated reason is the only answer that cannot be misread.
An identifier resolves to nothing in two ways, and they end in the same
refusal without being the same case. An identifier nobody ever created is a
typo the control plane accepted, as above. An image the client registered —
CreateImage, served by Outscale and, since #131 (2026-08-13), by Scaleway
too (instance/v1/API.CreateImage; this sentence said "planned to follow"
for a fortnight after the route shipped, until 2026-08-27) — is the more
embarrassing one: ReadImages lists it, yet this
emulator keeps records, not disk contents, so there are no bytes to boot.
Booting the source's base image instead would silently drop whatever the client
baked into the image — and the golden-image workflow is precisely the one where
that difference is the point — so it is refused like the first case, and the
log says which of the two it was. If the emulator ever captures disk contents
(the runtime could: incus publish exists), that refusal is the line to
replace.
What that decision costs, measured 2026-08-24. The fifteen third-party stacks
of examples/stacks/surveyed.md were replayed under --vm incus-ovn for the
first time, and four machines started across all fifteen. Not because the
runtime failed — it started every machine it was asked for — but because
thirteen of the fifteen name an image the emulated catalogues do not hold:
ami-a3ca408c, ami-538af795, ami-47899c77, a talos image registered through
volume → snapshot → CreateImage, an RHCOS template registered by name. The two
that boot are the two that name a catalogue identifier rather than a
production one: kiwinet's debian_bookworm label, and
terraform-exoscale-vault's template_id read from the served template list.
So the rule of thumb, for anyone deciding whether --vm will do anything for
them: a configuration that hardcodes a real image id keeps applying and starts
nothing. That is the deliberate half of the decision above working exactly as
written — under --vm off the same configuration reports running and the
difference never shows. It is recorded here rather than filed as a defect,
because it is the documented behaviour meeting a population nobody had measured
it against.
Two moves close most of that gap without weakening the refusal, and neither guesses anything.
The image recipe derives from the family, so any version builds on the fly.
The build table used to enumerate five family/version rows, and the Scaleway
catalogue promised one it did not hold: debian_trixie maps on debian:13,
no debian/13 was ever built, and a client using the label the emulator itself
publishes got a machine without ssh. The irregular part of the recipe is per
family, never per version — the source is always
images:<family>/<version>/cloud, and only the ssh package name, the service
name and the package manager change — so machine.SpecFor now derives the
recipe from the ref: ubuntu, debian (apt), alpine (apk), almalinux, rockylinux,
centos and fedora (dnf). A known family at a version the station lacks is built
at first boot, announced when it starts and when it ends, through the same seam
feint images drives — one recipe, one per-image lock, and a builder container
named for its image and its process, because a Go lock cannot span two
processes and the pair that actually collided was a feint serve and a feint images in another terminal. feint images still exists for what it is
genuinely for: warming a station ahead of a run, so no terraform apply pays a
build. A family with no recipe (plan9:4, talos) still degrades exactly as
above, and a version the upstream image server has withdrawn — images: no
longer publishes ubuntu 18.04, debian 9 or debian 10, all named by surveyed
stacks — fails the boot with the ref and the source in the log instead of a raw
driver error.
An opaque identifier is the operator's to declare, and the refusal says how.
No table can name the OS behind ami-a3ca408c, and guessing one was refused in
#392: a stack that asked for an AlmaLinux and got an Ubuntu boots, then fails at
its first dnf. What replaces the guess is a declaration the operator signs:
FEINT_BOOT_IMAGES='ami-a3ca408c=ubuntu:22.04,ami-538af795=ubuntu:22.04' \
feint serve --vm incus-ovnEach entry maps one identifier onto family:version, with an optional @login
for the cloud where the login belongs to the template. The declaration is
consulted only when the pack's own catalogue resolved nothing, so no entry can
shadow a catalogue label; a malformed entry or an unknown family refuses to
start, naming the entry and the families a recipe exists for, because a typo
that surfaced hours later as a refused boot would be blamed on the stack. The
boot refusal itself names the identifier it got, the reason, and both gestures:
the declaration above, and the lookup below.
feint images resolve <id>... asks the providers' public listings what an
identifier is — network, but no account, and never on the boot path: a create
that waits on a third party's availability is the failure mode that cost the
Exoscale Terraform provider a fork. Measured on 2026-08-25, without
credentials: Outscale's Official OMIs
Reference
page is historical — it named all three withdrawn identifiers the surveyed
stacks hardcode (ami-a3ca408c = Ubuntu-22.04-2023.12.04-0, ami-538af795 =
Ubuntu-18.04-2021.02.04-0, ami-47899c77 = Debian-9-2019.11.29-0) while the
live API answers zero images for them to a valid account, and 401 without one.
Scaleway's marketplace (api.scaleway.com/marketplace/v2/…) and Exoscale's
template listing (api-<zone>.exoscale.com/v2/template) answer JSON with
identifier, name and — for Exoscale — family, version and default user. The
command prints what each listing calls the identifier and the exact
FEINT_BOOT_IMAGES line to paste; an identifier in no listing exits 2 and says
what remains (the stack's own docs, or whoever created the image), and a
listing that cannot be asked exits 1 and is never reported as an absence. What
it deliberately does not do is feed the emulator directly: the lookup fetches,
the operator declares, the emulator works offline afterwards — the same shape
as upstream:sync feeding the drift scan, and for the same reason a list baked
into the binary was refused: it ages into one more lie.
A derived image is a cache, and it is visible and individually removable.
Derivation splits two questions the enumerated table kept fused: what the
warm-up set requires (feint images --check grades that, exit 2 on a miss)
and what this emulator has put on the station — which now includes images no
list names. Both commands therefore print derived images as their own rows,
because an image an inventory cannot name is a silent residue on somebody
else's machine. feint clean deliberately removes no image, warm-up or
derived alike: clean removes what a killed run left half-alive, and an image
is the cache that spares the next run its minutes of build — the conformance
suite runs clean, and sweeping images there would rebuild the set on every
pass. Removal is explicit and targeted instead: feint images remove fedora/44 says the alias and fingerprint it is about to delete, refuses a
name the emulator never published, and can only ever spell aliases under the
feint/ prefix — an operator's own images cannot even be named through it.
A first boot that builds can outlive the answer to the call that caused it.
Measured on 2026-08-25: CreateVms on a declared fedora:44 built the image in
52 seconds and the machine reached running with its address published — while
the HTTP answer to that one call died on the server's 60-second write timeout
(curl: empty reply). The state converges and is honest; the response of the
call that paid the build is what may be lost, and a client that retries a
create it believes failed will make a second machine. The exposure predates
derivation — any boot longer than a minute had it — but a cold build makes it
ordinary, which is why the warm-up set exists and why debian/13 is in it:
feint images ahead of a run keeps every production path under the timeout.
Two details follow from the same decision. The Scaleway marketplace answers one
fixed UUID per label, so Terraform — which resolves a label into a UUID and
sends the UUID back — still names the distribution it chose; a single shared
UUID is how image = "debian_bookworm" used to boot an Ubuntu. And image and
login resolve together: whatever a pack resolves an identifier to carries
the login that image provisions — root on Scaleway, outscale on Outscale, the
template's own default-user on Exoscale — because the right distribution with
the wrong login is still a machine nobody can enter.
sbs_volume works since SW-3. tools/conformance/scaleway/terraform/
declares it, and the apply, the empty second plan and the destroy all pass. The
limit this section used to describe — no usable value at all — is over.
It is worth keeping how it read, because the fixture is the part that stings:
Omitting the block is the way through, and it is what
tools/conformance/scaleway/terraform/does — which is why the suite is green and shows none of this.
A fixture that avoids the one input that breaks is a test that cannot fail. The fixture now declares the block, and would go red if the fallback stopped working.
Measured by @vde-dis on #8, with OpenTofu 1.12.5 and scaleway/scaleway 2.80.0:
b_ssdwill not plan, and that is upstream's decision, not this emulator's. From provider 2.79 on it is refused before any request leaves: "b_ssd volumes are not supported anymore. Remove explicit b_ssd volume_type, migrate to sbs or downgrade terraform."sbs_volumeused to plan for ever, because the emulator overrode the type tob_ssdand the value read back never matched the value sent. It is now honoured, and since #365 it is also what a request naming no type gets: the disk is created inblock/v1, and the provider reads it back through the fallback it always used —instance.GetVolumefirst, thenblock.GetVolumeon a typed 404.- The local types are honoured too, since #393, and the override that used to
stand here is gone.
l_ssdwas answered asb_ssdwhatever the client asked, which is a typeinstance/v1has stopped minting at all: a create namingl_ssdonfr-paranswers201and the volume comes backl_ssd, measured 2026-09-02.scratchis refused, and on the argument the cloud names —image, not the type: "Cannot use an image with a scratch volume". The reason the override existed is unchanged and still holds the catalogue in place: the emulatedvolumes_constraint.min_sizeis 0 and the CLI sums local volumes against it, so a value above zero would make the CLI refuse the very creation it just asked for (TestCatalogueKeepsTheLocalVolumeTrapDisarmed). b_ssdis refused by this emulator now, the wayfr-parrefuses it, and the two routes do not share a wording:POST /volumespoints at the SBS migration page underargument_name: volume_type,POST /serverssays "Create volumes with volume_type=sbs_volume instead" undervolumes.<key>.volume_type. Both are transcribed ininternal/providers/scaleway/volumetypes.go.
What is still not emulated behind an SBS volume is the storage itself: the size,
the class and the IO/s are recorded and answered, and nothing is written
anywhere. perf_iops is a number a client reads back, not a rate anything
measures.
Volume encryption is refused rather than faked: kms_key_id names a key in
Scaleway's Key Manager, which this emulator does not serve, so a create carrying
one is rejected. Accepting it would let a client read its own key back from a
volume nothing encrypts.
CreateKeypair on the real Outscale, given a name and nothing else, generates
the pair and returns the private key in the response. This emulator refuses:
4001 PublicKey is required: this emulator does not generate keypairs,
because it would have to hand out a private key that only looks usable
The decision is deliberate and was reaffirmed on 2026-08-30. A generated private key here would be a real, usable key that unlocks nothing — the emulator boots machines whose authorised keys it controls, so the pair would be theatre. Worse, it is the kind of theatre somebody stores: a reader told to "save it, it is shown only once" saves a secret that is not one, and may reuse it.
What this costs, measured. The Outscale course at
blog.stephane-robert.info opens its first hands-on page with exactly this call:
octl iaas api CreateKeypair --KeypairName ma-cle-testand says, in as many words, "La commande retourne la clé privée dans la réponse JSON, à sauvegarder immédiatement, elle n'est affichée qu'une fois." Against this emulator that command fails, and it is the first thing a reader types.
The way through is one flag, and it is the shape the emulator wants anyway:
ssh-keygen -t ed25519 -f ./ma-cle-test -N ''
octl iaas api CreateKeypair --KeypairName ma-cle-test --PublicKey "$(cat ./ma-cle-test.pub)"The key is then one the reader really holds, and conformance:ssh proves a login
with exactly that arrangement on all three packs.
Not revisited unless somebody shows a use where a fabricated private key is better than an absent one. "The course says so" is not that argument: the course can carry the two lines above, and the emulator cannot un-hand a secret.
A server goes from stopped to running within the action call. Real hardware
takes a minute; reproducing that delay locally would only make every client wait
for information that does not exist here.
The states clients check are preserved: deleting a running server is refused
with transient_state, because Terraform depends on that error.
Since #637 the default has an opt-out, because the default has a cost that
only a client author could find. A reboot's target state is the state it started
from, so a waiter watching for running proves nothing: without task support the
only signal is watching the machine leave its initial state, and this emulator
answered running from the first read to the last. That made reboot the one
action whose waiting could not be exercised here at all — and it is the one where
waiting is subtle, which is what makes it worth being able to test.
feint serve --consistency eventualWith it, a Scaleway server walks the states fr-par walks, measured 2026-09-03
by polling a real DEV1-S through each action:
| action | states answered, in order |
|---|---|
poweron |
starting, then running |
reboot |
stopping, starting, then running |
poweroff |
stopping, then stopped |
States advance on reads, never on a clock, and that is the design rather than a shortcut. The emulator's clock is injected — tests freeze it — so a transient state timed on a wall clock would either never end under a frozen clock or turn every suite into a wait. A four-second suite stays four seconds.
Two consequences worth knowing. An action is not an observation: the read a
lifecycle action makes to change a resource does not advance it, or the action
would consume the first state it just pushed. And a failed action walks no chain
— a start that failed answers its failed state, because narrating a path towards
a running it never reached is exactly the plausible-wrong answer this project
exists to avoid.
The Outscale half walks under the same mode since #124. Against a real
Outscale account on 2026-08-08: CreateVolume answers State: "creating", a
CreateSnapshot issued before the volume settles is refused with
409 InvalidVolumeState (code 6007), and a snapshot is born in-queue with
Progress: 0 and only later completed. With the mode on, that is what this
emulator answers:
| action | states answered, in order |
|---|---|
CreateVolume |
creating on the create, then available |
CreateSnapshot during creating |
409 InvalidVolumeState, code 6007 |
CreateSnapshot once available |
in-queue with Progress: 0 on the create, then completed at 100 |
The refusal's Details wording was not recorded and is this emulator's; its
status, type and code are the measurement's, and osc.IsConflict classifies
the code the way it classifies the cloud's. With the mode off, a volume is
available and a snapshot completed at once, byte for byte what every suite
read before.
Serving that refusal had been tried and reverted once, and the reason is worth
keeping: a guard written for a state this emulator cannot reach is a control
that can never fire — the "a comment is not a control" defect, in code rather
than in prose. What made it servable is the mode, not a clock; and the guard
reads the volume without observing it (Store.Peek), because a read that
advanced the chain would consume the very state it checks and never fire, which
is the defect the first version had.
The line, for whoever adds the next resource: a state invariant is served when
the state it names is reachable here. LinkVolume on an already-linked volume,
DeleteVolume on a linked one, retyping a running machine, deleting a Net that
still holds a subnet — all reachable, all refused, all tested. A refusal that
would need an artificial delay to become reachable is not served, and belongs in
this list instead.
A recording of a real fr-par account (#427) put three more entries on it, and
all three are one decision seen from three products. Forty-seven of the
divergences that recording found are this paragraph, which is why they are
written here rather than patched one at a time.
- A block snapshot is
availablethe instant it is cut. Upstream it is born in a transient state, and aDeleteSnapshotissued while it settles is refused with412. Here the first delete succeeds, so the recorded second delete meets nothing (404against the cloud's204) and the read between them finds no snapshot. - A load balancer is gone the instant it is deleted. Upstream the read that
follows a
DeleteLBstill answers200, withstatus: to_delete, and only the read after that answers404. Here the first read already answers404. - A migration to the offer a balancer already carries is accepted. Upstream
it is refused with
400 invalid_arguments: the API accepts only migrations that change something (measured on a real account, 4 September 2026, #762). Here it is accepted and the read shows the same type it showed before. The reason is the oneCreateLBalready carries a few lines up: a refusal keyed on a VALUE turns a recorded200into a400whencorpus:checkreplays a recording whose values were replaced by synthetic ones of the same shape. The migration that changes the offer, which is the one a client actually makes, behaves as upstream does. - A public gateway is
runningthe instant it is created. Upstream it isallocatingfor a few seconds, and both aUpdateGatewayand aCreateGatewayNetworkissued in that window are refused with409and a body naming the state (current_state: allocating). Here both succeed.
The 409 is the interesting one, because it is the shape a client branches on
and this emulator can never answer it: the state that produces it is not
reachable here, which is the rule two paragraphs up. Serving the refusal would
mean inventing the window, and a guard for a state nothing can enter is a
control that can never fire.
Not a limit of the emulator, recorded here because it reads exactly like one.
In corpus/scaleway/scw-billed-shapes.jsonl the public gateway goes from
200 on its seventh read to 404 on its eighth with no DELETE anywhere in
the file, and its address follows: the recorded DeleteIP answers 404 and
the recorded GetIP answers 404 too. The destruction did not travel through
the proxy.
It is stated rather than assumed. The file holds three recording sessions whose
own sequence numbers run 1..14, 1..43 and 1..65 with no gap, so nothing was
dropped between the request that read the gateway and the request that could not
find it; and no DELETE on /vpc-gw/v2/…/gateways/ exists in that file at all.
So the eleven findings on those three exchanges measure the recording and not
the pack: this emulator was never asked to delete the gateway, and answering
404 for an object nobody destroyed would be the lie the whole project exists
to avoid. They go when the gateway is recorded again with its destruction in
the transcript.
This entry used to be the largest single divergence the 2026-08-24 recording
found, and it is over. CreateServer with no volumes in the body is answered
by fr-par with volumes: {"0": {"volume_type": "sbs_volume"}} — a block
volume — and the recording then reads that volume three times through
block/v1alpha1/API.GetVolume and deletes it there. This emulator gave such a
server a b_ssd volume in instance/v1, so all four of those calls answered
404: forty-three findings, one default. The eighteen acceptance entries that
carried them are deleted, and feint corpus --check compares those exchanges
for real now.
It is worth keeping why it took two steps, because the first was mistaken for a
price rather than a defect. The reason not to flip was measured: the whole
instance/v1 volume surface read a server's root disk out of the instance
store, so a block root was invisible to attach-volume, detach-volume, the
update's volume map, a create naming a volume and CreateSnapshot. But that was
already true for anybody who wrote root_volume { volume_type = "sbs_volume" },
which has worked since SW-3 — the defect existed, unmeasured, and the flip only
made it universal. The cost of a change and a defect it exposes are not the
same thing, and reading one as the other is what left this open for a month.
#571 fixed the five resolutions first; this became one line.
What the flip left standing, and it is a client behaviour rather than an
emulator one: scw instance snapshot create volume-id=<a server's root disk>
no longer works without unified=true. The CLI calls instance.GetVolume
itself before it sends anything (scaleway-cli 2.56.3,
internal/namespaces/instance/v1/custom_snapshot.go) and returns that error, so
the command stops one call before the emulator. The instance route does resolve
a block volume — unified=true reaches it and answers an sbs_snapshot — and
scw block snapshot create volume-id=<root> is the path the conformance suite
walks. Whether the real cloud answers instance.GetVolume for an SBS volume is
not measured here: no recording carries that call, and the SDK's own
getUnknownVolume only makes sense if it can 404.
And the snapshots have not crossed, which the flip put on the default path.
Measured on 2026-08-28, against this emulator, with scw 2.56.3:
| naming a block snapshot | naming an instance snapshot |
|---|---|
scw block snapshot create volume-id=<a server's root> — works |
n/a |
scw instance image create snapshot-id=… — 404 |
works |
scw instance snapshot get … — 404 (the fallback's own shape) |
works |
scw block volume create from-snapshot.snapshot-id=… — works |
404 |
A block snapshot is now the only kind a client can take of a server's root disk,
and scw instance image create cannot cut an image from one. So the golden-image
chain — snapshot a disk, cut an image, boot from it — is walkable from a volume
the client created and no longer from the server's own root. This is the shape
#571 fixed for volumes, one product over, and it is named rather than fixed:
the SDK says images built on block snapshots exist (Image.RootVolume.VolumeType
can be sbs_snapshot, and scw instance image list reads block.GetSnapshot
for exactly that value), so the gap is real and not a decision.
One consequence is already guarded, because it was created and measured inside
this change: an instance snapshot of a block volume is typed unified, never
sbs_snapshot. sbs_snapshot is a promise that the id resolves in
block/v1alpha1, and an image cut from a snapshot that broke that promise made
scw instance image list fail for the whole zone.
The store is memory: a dead process loses every emulated resource, and that is
the model, not a defect. The machines are different. A container the runtime
started does not die with the process that asked for it, so the policy below
exists, and it is deterministic — measured on 2026-08-13 by running every row
that can be triggered (tools/conformance/crash.sh triggers them on every run
where a runtime is configured).
| Event | The store | The machines, networks and rule sets |
|---|---|---|
Graceful exit (Ctrl-C, SIGTERM, feint stop) |
lost, and said out loud: feint stop names the count it is about to discard when no --state was recorded (saved first with --state) |
stay, labelled user.feint.provider |
Graceful exit with --cleanup |
lost (saved first with --state) |
swept before exit, counted out loud |
| SIGKILL, crash, power loss | lost | stay, labelled — nothing had a chance to run |
| Restart | starts empty, with the same notice stop prints, because restart goes through it |
named, never adopted: startup warns "labelled machines from a previous run exist; nothing was adopted, feint clean removes them", listing them by name |
feint clean |
untouched (it is a separate process) | everything labelled is removed; the runtime is queried, and the sweep reports what it could not remove instead of claiming success |
The store's own line was documentation only until #182. The model was right and
this page stated it, but the sentence was read after being bitten: an operator
reaching for restart mid-session paid with the whole fixture and learnt why
here, later. stop now says it at the moment it happens, on stderr, once, and
only when something is actually lost: with --state recorded it stays quiet,
because a warning on every healthy stop is the pattern people are trained to
ignore.
Why restart never adopts: the store that gave those machines meaning died with
the previous process. A machine resurrected from the runtime would be state
without an owner — trusted for the same bad reason a restored snapshot used to
be trusted, and snapshot.go documents where that leads. The startup notice
names the leftovers precisely so the operator decides, with
TestStartupNamesTheLeftoversItDidNotAdopt holding the line and
tools/conformance/crash.sh proving the whole sequence — kill, survive
labelled, warn, sweep to zero — against a real runtime, the runtime queried
directly rather than the store.
The notice keys on machines. Networks and rule sets alone stay silent: an empty emulated bridge is reused under its own name by the next run or refused as a block conflict out loud, the OVN uplink is deliberately kept across runs, and a warning that fired on every healthy restart would train everyone to ignore it — the exact way a gate dies, already measured on this repository.
Every row above sweeps objects — machines, networks, rule sets the runtime
can list. #316 measured, twice (2026-08-18 and 2026-08-19), a leftover no
object listing shows: a network's interface disappears while its dnsmasq
lives on, still bound to the gateway address. ip addr shows nothing, incus network list shows nothing; only ss -lnp disagrees, and the next run that
wants the block dies minutes in on dnsmasq: failed to create listening socket: Address already in use. Knowing that ss is the third place to look
cost three ten-minute runs the first time.
The attribution criterion is stricter than a label, because a process carries
none: a dnsmasq is the emulator's leftover only when its --interface
carries the fnt- prefix only this emulator derives and that interface no
longer exists. One whose interface is alive is somebody's working service,
whoever owns it — this station runs libvirt's and two other Incus projects'
dnsmasq beside feint's, and none of them is ours to name, let alone signal.
Four consumers share the one check (internal/core/machine/leftover.go):
feint doctor reports it and touches nothing; feint clean ends it, after
re-checking at the moment of the signal that the pid still is that leftover
(pids are reused); feint clean --check answers the same question without
ending anything, which is what the conformance suites ask before they start;
and the network-create error names the process when its listen address falls
inside the failing block. TestLeftoverDHCPRefusesAProcessItCannotAttribute
holds the refusal, and tools/falsify/specs/dhcp-leftover-ownership.json
proves the test bites.
Nothing here escalates, and the remedy is a command rather than a paragraph.
The runtime's dnsmasq runs as the incus user, so an ordinary sweep cannot
signal it: feint clean says so and exits 1 instead of claiming a clean host,
and every suite that takes an address block asks feint clean --check on its
doorstep (guard_leftovers, tools/conformance/guard.sh) rather than meeting
the state twelve steps into a run. That was measured — the runtime leg of
mise run evidence:update failed three times in a row, each time after every
client suite had already run, with the right remedy printed and nobody running
it (#375).
What none of them does is acquire a privilege it did not have. A conformance
suite that escalated to end a daemon it did not start would be a worse defect
than the one it works around: it is the question mustOwn asks of the driver,
one layer up, and a process nobody here created is not ours to end. So the
elevation is the operator's, in one line — sudo feint clean --vm <mode>, the
same sweep run by somebody who may signal it, re-asking every ownership question
at the moment of the signal. The permission probe behind --check is signal 0:
the kernel runs the check it would run for a real signal and delivers nothing,
which is the only acceptable shape for a question whose subject belongs to
somebody else.
One half of the remedy stays manual, and it is the #342 case above: when the bridge survived alongside its service, ending the service leaves the interface holding the same address, and nothing on the host proves this emulator created that bridge. Both commands are printed; only the second needs a human to decide the bridge is theirs.
A leftover is not always debris of an earlier run, and a doorstep cannot
prevent the other kind. #316 and #342 both measured leftovers surviving a run;
on 2026-08-21 the machines-on leg of mise run evidence:update produced one of
its own, in bridge mode, from a network it had created minutes earlier. The
runtime listed that network as unmanaged while its bridge and its dnsmasq
stayed up, and the emulator's log names the moment it broke:
could not isolate the subnet's network network=fnt-8488bc9e9e1
error="detach isolation from fnt-8488bc9e9e1: incus network:
open /var/lib/incus/networks/fnt-8488bc9e9e1/dnsmasq.raw:
no such file or directory"
The same run's log carries the load that produced it: NIC ACL writes failing on
Unknown or missing host side veth device, instance updates refused with
Instance is busy running a "delete" operation, a network delete refused with
The network is currently in use. So a network object can die under a bridge
this emulator created while the bridge and its service stay up, and the
doorstep then catches it at the next suite rather than at the start of the leg.
It is deterministic, and the lifecycle has a name. The leg was run twice
that evening and failed both times, in the same suite, on a subnet of
examples/stacks/outscale/main.tf — 10.50.1.0/24 the first time,
10.50.2.0/24 the second, both inside the 10.50.0.0/16 that stack declares.
Each failure is preceded by the pair above: two subnets torn down in the same
second, one answering Network not found (gone cleanly) and the other
answering open …/dnsmasq.raw: no such file or directory — a detach of the
isolation arriving at a network whose state directory the delete has already
removed. The one that answers the second way is the one left standing. Worth
noting because #316's original measurement holds 10.50.2.1: the three issues
of this family are downstream of the same teardown.
What #375 made cheap is the diagnosis and the remedy; the birth of that
leftover was a separate defect, and #386 is where it was closed. In OVN mode it
does not arise at all: an OVN network carries no dnsmasq, and
FEINT_VM=incus-ovn mise run conformance passed twice the same evening.
What closed it: a lock named after the network, and a question asked under
it. Two requests reach one network because a reconciliation lists the store
and a concurrent delete removes a member it listed — terraform destroy tears
down subnets in parallel, so this is the ordinary case rather than the unlucky
one. The driver now takes serialise.Lock("incus.network." + name) in
EnsureNetwork, IsolateNetwork and RemoveNetwork, so no config edit of a
network is in flight while its delete runs; and IsolateNetwork, holding that
lock, asks the daemon whether the network is still there before it edits
anything. Both halves are needed, and each has its own falsification in
tools/falsify/specs/teardown-race.json: the lock alone still lets a delete
that won it run first, and the question alone is a time-of-check a delete
crosses. Per network and never global — a global lock would queue every subnet
of a stack behind one delete, which is the mistake internal/core/machine/ serialise.go already records having made once.
A detach that could not happen is reported rather than counted as done:
machine.ErrNetworkGone is what the driver returns, and ReconcileIsolation
logs it at warn, naming the network. Warn and not error, because no rule set
was needed and none is missing — an error line here would fire on every
parallel destroy and teach a reader to skip it, which is how a log stops being
evidence. The rule set that isolated the network is dropped by the delete
itself, since the pass that used to drop it is now the one that refuses to run
against a network that is gone.
The section above reads The network is currently in use as one symptom among
the load of a parallel destroy. It is not. It was the whole of a second
producer, and it stayed invisible for four issues because the shape that hides
it is a delete that answers success.
The driver had Attach and no counterpart. So deleting a Scaleway private NIC
only forgot it in the store, and the device stayed on the container: the API
answered 204 while incus config device show still listed the interface.
DeletePrivateNetwork then called RemoveNetwork, Incus refused with The network is currently in use — correctly, a machine was still on it — the pack
logged that at error level and answered 204 anyway. The bridge, its rule set
and its dnsmasq outlived the run holding the block.
Measured on 2026-08-24 by inventorying the host around three consecutive runs of
tools/conformance/stacks.sh under --vm incus:
| run | exit | what it left |
|---|---|---|
| 1 | 0 | three bridges, three rule sets, three dnsmasq |
| 2 | 1 | nothing — it never got that far |
| 3 | 1 | nothing |
Runs 2 and 3 died on Address already in use for the blocks run 1 left. So the
failure is not intermittent: it is deterministic with one run of delay, and a
passing run is what arms the next one. Which of the three bridges survived
is Terraform's destroy scheduling, which is why the block named moves between
runs — the blocks themselves are fixed in examples/stacks/scaleway/main.tf and
are never chosen by the emulator.
Three things changed, and the third is the one that generalises:
- the driver's
Detach, required rather than optional, asking both ownership questions and removing only a device the instance itself carries. Both packs that attach now detach; the Exoscale handler had documented the gap as unclosable ("the driver deliberately has no hot-unplug"), which is how one defect lives in two packs. DeletePrivateNetworkrefuses withprecondition_failedinstead of logging. A network reported gone while its bridge holds the block is the same lie as a network created while nothing exists, and the create path already refused its half.- The sweep reads the host after itself.
feint cleansurveys, prunes, then surveys again, and anything present in both was asked to go, said nothing, and stayed. That is the only way this class is visible at all: no return value the remover produces can report it.--format jsonrecords one line per object with why it stayed, so the question is now "which mechanism produces the waste" answered byjqrather than by reading four issues.
feint clean --check --doorstep refuses a run whose host still holds a previous
run's machines or networks. The flag is separate from --check because the two
questions have different safe moments: guard_leftovers is asked before a run
starts and twelve steps into one, and mid-run those objects belong to the
emulator that is running. Asked at both, it failed a leg for owning what it had
just created — measured, and now held by
TestOnlyTheDoorstepAsksWhatAnEarlierRunLeft.
The two above leave objects a sweep can still remove. This one leaves objects nothing can remove, and the sweep was what made them permanent.
The upstream cause is three lines of Incus' own schema. networks_peers carries
three references, and only one of the three lacks a cascade:
network_id INTEGER NOT NULL,
FOREIGN KEY (network_id) REFERENCES "networks" (id) ON DELETE CASCADE
target_network_integration_id INTEGER ... REFERENCES networks_integrations (id) ON DELETE CASCADE
target_network_id INTEGER NULL, -- no foreign key, no cascadeThe schema above is read from sqlite_master on the station, Incus 7.2, not
quoted from anywhere: network_id cascades, target_network_integration_id
cascades, and target_network_id carries no foreign key at all.
Delete the source network and the peering row goes with it. Delete the
target and the row survives, holding an id that resolves to nothing, with
target_network_project and target_network_name both NULL. From that moment
the source network cannot be removed by anything, and the command that would
remove the row fails on the same missing target. Reproduced on 2026-08-25 by
planting one row against a network this emulator had just created, so that the
premise of this section is measured rather than inherited:
$ incus network peer delete fnt-387f1e8daa2 fnt-gonetarget
Error: Failed deleting peer: Failed loading target network: Network not found
$ incus network delete fnt-387f1e8daa2
Error: The network is currently in use
incus network peer show still lists the peering, with no target field left to
read, and incus network peer edit does not repair it: it returns 0 and
persists nothing, the target fields being immutable after creation. Incus 7.3
carries nothing that touches peer rows. The asymmetric cascade appears to be
unreported upstream.
Two further refusals belong to the same trap and are quoted from the
maintainer's station in #455 rather than from the reproduction above, because
they need a rule set to be attached and the two-subnet Net used here attaches
none: incus network unset <net> security.acls answers Failed applying router security policy: Failed loading target network: Network not found, and
incus network acl delete iso-<net> answers Cannot delete an ACL that is in use. That is the full cycle — the network cannot go because the rule set holds
it, the rule set cannot go because the network holds it, and the detach that
breaks the cycle is refused by the dangling row.
Three things this emulator was doing made it permanent, and all three are prevention rather than repair:
Prunetreatedfeint-uplinkas an ordinary network — it carries the sameuser.feint.providerlabel as everything else — and unset itsipv4.routesbefore a delete that then failed on the orphans still attached. From that instant every management path of those orphans failed validation withUplink network doesn't contain 10.2.4.0/24 in its routes, which is the detach they needed in order to go. The uplink now leaves the ordinary path, is deleted last, and its routes are never unset.- Neither
PrunenorRemoveNetworkdetachedsecurity.aclsbefore deleting a network, and the two hold each other: the rule set is "in use" by the network, the network is "in use" by the rule set. Both now detach first — and put the attachment back when the delete still refuses, because a network that survives with its rule set silently off is a firewall disarmed rather than a subnet swept. - Neither removed the half a peering leaves on its surviving target. Both now do, before the delete rather than after a refusal, since a delete that succeeds is exactly what leaves the neighbour holding a dead id.
The repair, for a station already carrying the state, is feint clean --force. It removes such a row through incus admin sql, which is Incus'
own supported mechanism for it rather than an edit of a file behind the
daemon's back, and it removes a row only when the network that row belongs to
carries the label this emulator wrote. That scope is the point and not a
detail: a dangling row of an operator's own is the same table, the same shape
and the same dead target, and a --force able to reach it would be a worse
defect than the one it repairs. The witness is
TestForceLeavesAThirdPartysDanglingPeerAlone, and
tools/falsify/specs/trapped-station.json proves it bites — fifteen mutations,
all red.
It was also driven once against the real runtime, on 2026-08-25: two dangling
rows planted, one on a network the emulator had labelled and one on a bridge
created for the purpose and deliberately named fnt-lab, so it carried the very
prefix a name check would accept. clean --check exited 1 naming only the first,
clean --force removed only the first and printed it whole, and the row on
fnt-lab was still there afterwards. The same reasoning governs the sweep one
level up: the half of a peering living on the far end is removed only when that
network carries the label too, since fnt- is a prefix anybody may type.
feint clean --check reports three states, and each is one no healthy run ever
produces: a peering row whose target no longer resolves, a network of the
emulator's whose block the uplink no longer carries, and a rule set attached to
a network already trapped by either. That last one is deliberately not the
bare form: IsolateNetwork ends by attaching a rule set to every OVN network
with a neighbour to keep out, so reporting "a rule set is attached" would refuse
every healthy run — the same way #426's doorstep once fired on hosts nothing was
going to fail on, which is how a check gets disarmed.
What remains a limit: a dangling row whose source network is not this
emulator's is left alone and named nowhere, deliberately. If an operator's own
tooling produces one, the repair is theirs to run, and incus admin sql global "DELETE FROM networks_peers WHERE id=<id>" is the statement — after reading the
row, because nothing here will read it for them.
The section above is about a row nothing can remove. This one is about the two rules that produce such rows in the first place, and both were read from the Incus source at v7.2.0 and then measured on the station, sequentially, with nothing concurrent.
A create is completed by any pending half aiming at the network, and by
exactly one. PeerCreate in internal/server/network/driver_ovn.go looks for
the half to consummate with GetNetworkPeers(TargetNetworkProject, TargetNetworkName) — the target network, and no clause on which network holds
the row — and answers More than one matching network peer was found when the
filter matches twice. So two pending halves aiming at one network make every
create on it fail, whichever pairs they belong to:
$ incus network peer create fnt-lab1 fnt-lab2 fnt-lab2
Network peer fnt-lab2 pending (please complete mutual peering on peer network)
$ incus network peer create fnt-lab3 fnt-lab2 fnt-lab2
Network peer fnt-lab2 pending (please complete mutual peering on peer network)
$ incus network peer create fnt-lab2 fnt-lab1 fnt-lab1
Error: Failed creating peer: More than one matching network peer was found
(lab1, lab2) and (lab3, lab2) are two different pairs, which is the whole
shape of #456: a Net of N subnets declares N(N-1) halves as its subnets appear,
and only a Net with three or more can have two of them aiming at one network at
once. That is why a stranger's three-subnet Net tripped it and the two-subnet
conformance fixtures never did — and why the emulator excludes both ends of a
pair rather than the pair, one lock per network taken in sorted order
(peerLock, internal/core/machine/incus_ovn.go).
A peer delete blanks every row aiming at the network it runs on. PeerDelete
clears target_network_id for every row whose target is that network, in the
same transaction, whatever pair the row belongs to. On a three-network mesh:
$ incus network peer delete fnt-lab2 fnt-lab3
Network peer fnt-lab3 deleted
$ incus query /1.0/networks/fnt-lab1/peers?recursion=1
{"name":"fnt-lab2","target_network":null,"status":"Errored"}
$ incus network peer create fnt-lab1 fnt-lab2 fnt-lab2
Error: Failed creating peer: A peer for that name already exists
The delete named lab2 and lab3, and it is lab1's half that came back
Errored. Redeclaring that half is then refused for the name, which this driver
used to tolerate as the peering already being there — so a peering was reported
applied and did not exist. A pair whose two halves the runtime does not both
call Created is now rebuilt from both ends instead, and every delete on that
path asks the label EnsureNetwork wrote before touching a network, never the
fnt- prefix.
What remains a limit. Because a delete invalidates every inbound half of a
network, repairing one pair of a mesh can damage its neighbours', so a mesh
already wrecked on the host is not guaranteed to come back complete in a single
reconciliation. Measured both ways on 2026-08-25, same shape both times: one run
of a three-subnet Net broken by hand and then given a fourth subnet came back
with twelve rows of which four were Created; the next, with the passes carried
on, came back 12 of 12, then 20 of 20, then 30 of 30 as subnets five and six
arrived, with no More than one matching network peer was found and no
could not peer in the log at all. A Net whose subnets are merely created is
the case that matters and it is exact: three subnets give three networks, six
rows, six Created. Nothing here is a reason to hand-repair a host — the next
subnet event reconciles, and feint clean --force is what frees a station that
carries the older, unremovable form.
No signature is checked, on any provider. Credentials must merely be well-formed, because the SDKs validate their shape client-side before sending anything.
This means Feint must never be exposed on a network you do not control. It is a development tool that grants everything to everyone, by design.
A refusal can now be produced on purpose, and that changes nothing here.
PUT /_feint/faults makes a named operation answer 401 or 403 (among others),
which is what lets a client's degradation path be observed at all. It is not
authentication: nothing inspects the credential, the rule fires on the operation
whatever the caller sent, and clearing it makes the same request succeed. An
injected refusal proves what the client does with a refusal, never that this
emulator — or the real cloud — would refuse that call.
Feint shipped a Docker driver and no longer does. Incus is the only machine runtime, and the reason is the network, not a preference between runtimes.
Emulating a cloud means emulating its addressing plan, and that needs four things from the runtime. Docker gives one and a half of them.
A network carrying a chosen block. Both can do this: docker network create --subnet=10.0.0.0/24 and incus network create n ipv4.address=10.0.0.1/24 are
equivalent. This is the half that works.
A fixed address on a machine. Docker takes --ip only for the first
network, and only on a user-defined one; every extra interface goes through
docker network connect, which means the machine comes up on one address and
acquires the others later. An emulated server with two private NICs, ordinary on
Scaleway, therefore cannot be started in one shot with both addresses known.
Incus takes them as device keys, at launch and afterwards, so the address the API
published is the address the machine carries from its first boot.
Enforceable rules. This is the decisive one. A security group has to be
something other than documentation, or the emulator repeats the flaw every local
AWS emulator has: MiniStack states that "security group rules are stored but
never filter traffic", and floci publishes host ports through socat sidecars
while its own documentation admits "the source CIDR value itself is not
enforced". Docker offers no rule layer: enforcing anything means writing iptables
or nftables rules into a container namespace by hand, then keeping them
consistent with the control plane. Incus has network acl natively, with ingress
and egress rules carrying action, protocol, ports and source, attachable to a
network or to a single NIC, enforced by the daemon. A security group maps onto it
almost term for term.
A real machine when the test needs one. A container shares the host kernel,
so it cannot carry a sysctl, a kernel module, or a systemd unit touching the boot
path. incus launch --vm gives a genuine KVM machine with its own kernel, which
Docker cannot do at all.
The cost of keeping Docker was not the 244 lines of driver: it was that every
network feature would have had to be written twice, once properly and once in a
degraded form, and that the degraded form would have set the ceiling for what the
emulator could honestly claim. Podman is not a substitute either; it shares the
same model. What is lost is real and accepted: Docker starts in under a second
where an Incus container takes a few, Docker is installed on more machines, and
CI images carry it more often. --vm off remains the default, so nothing in the
conformance suite depends on any runtime being present.
A security group here filters real packets, which is unusual enough to be worth stating precisely, along with what it does not do.
What is enforced. The group's default policies and its rules become an Incus network ACL attached to every interface of the machine. A port no rule opens refuses the connection, authorising one opens it without a restart, and revoking it closes it again. All four are checked end to end against a live daemon.
The second subnet no longer defeats the default-deny (#491). The
isolation rule set used to end with a catch-all allow in both directions,
which under OVN sat at rule priority 300 where a NIC's default action sits at
100/111 — so on any multi-subnet OVN run, a port no rule opens answered
anyway. The OVN isolation set now carries its foreign-block rejects (400) and
nothing else: unmatched traffic falls to the NIC default, which is the group's
default-deny, and the rejects outrank every allow a group can state (300), so
the two properties hold at once. Measured on the three example stacks under
feint up --runtime incus-ovn on 2026-08-26, listener proved from inside
before each probe: the forbidden port refuses from the station while an
allowed port opens in the same pass (Scaleway 443/9999, Outscale 443/9999 over
the public l2proxy path, Exoscale 443/9999), and a machine of another emulated
network is refused on a port whose group allows 0.0.0.0/0. A machine whose
groups enforce nothing wears the shared permissive posture set (opn-fnt,
catch-all allows at 300) on the NICs of an isolated network, so it stays open
to everything but the foreign subnets; the bridge-mode isolation set keeps its
catch-all, because there the network ACL filters at the bridge-host boundary
and would otherwise reject the station itself.
One leg of that pass did not hold, and it is the Scaleway one from the
station (#548). It holds since 2026-08-28. Measured on 2026-08-27, on the
same stack under the same runtime: the platform-web group opens 443 and
22-from-10.40.1.0/24 and names no port 80, a service was proved listening on
80 inside the machine, and the station reached 203.0.113.3:80 and
203.0.113.2:80. The negative control is in the same pass — the bastion, the
same shape of machine, refuses 80 because nothing listens there — so those two
were completed connections rather than a misread probe. The cause was #548's
routed NIC, which carried the published address and no rule set; the driver
now moves that address onto the filtered NIC as soon as the machine has one,
and the same probe refuses. The measurement of the remedy, in both driver
modes, is under "Two migrations were tried and refused, and a third one
works" below.
Measured: the three probes and the device dump in #548. Deduced, and worth
stating as a deduction: 9999 is a port nothing listens on, so its refusal
cannot be told from an absent listener, which is the failure this page warns
about under "The group covers a flexible IP" — the 08-26 pair's negative half
was therefore not in a position to see this. The Outscale and Exoscale legs of
that sentence were not re-measured here and are left standing.
What still holds from the 08-26 pass, and is re-measured on 2026-08-27 by
tools/conformance/functional.sh, is the same pair on the emulated
network, across two subnets: platform-web-0 reaches platform-app-worker-a
on 8080, which its group opens to the web block, and is refused on 9090, which
no rule names, with both ports proved listening inside the target first.
What this requires. security.acls on a bridged NIC exists from Incus
6.0.4 onwards. On 6.0.0, which Ubuntu 24.04 ships and will not move past, the
option is rejected outright and nothing is enforced. An ACL attached to a
network is accepted from 6.0, but only filters between the bridge and the
host, so it cannot separate two servers of the same subnet, which is precisely
what a security group is for. With an older runtime the emulator still serves
the whole product, and filters nothing: the log says so, once per group.
What has no equivalent here. Nothing in the Scaleway rule shape, as it
happens. A rule carries a protocol, a direction, an action, an ip_range and a
port range, and every one of them translates. There is no rule sourced by
another security group to worry about: that is the AWS model, where
UserIdGroupPair exists; instance/v1.SecurityGroupRule has only ip_range.
The stateful flag translates too, onto the runtime's allow and
allow-stateless.
The bound is that none of this applies to a server with no backing machine, which is the default configuration and what CI runs.
The bound a routed NIC adds (#337). A server that joins no private network
gets its public address on a routed NIC (#202) — and a routed NIC accepts
no security option at all. Measured on Incus 7.2, cold, one key at a time,
and the wording of the same refusal on the 7.3 CI runner:
| device | key | Incus answers |
|---|---|---|
nictype=routed |
security.acls |
Invalid device option "security.acls" |
nictype=routed |
security.acls.default.ingress.action |
Invalid device option … |
nictype=routed |
security.ipv4_filtering, security.mac_filtering |
Invalid device option … |
nic network=<managed> |
the same keys, together | Network ACL "…" does not exist |
The last row is the contrast that settles it: a NIC of a managed network
accepts the option names and complains only about the missing ACL; on a routed
NIC they are not options. There is no other per-NIC filtering mechanism to fall
back to, so a security group that restricts traffic is not enforced on a
server whose only interface is a routed NIC — a server with a public address
and no private network. The runtime declares it instead of pretending:
capabilities.firewall_public_only is false on /_feint/health, a rule set
bound to such a machine comes back as machine.ErrFirewallUnenforceable
rather than being half-sent, and the pack logs the declared limit as a warning
naming this section. The default security group — pure accept, filtering
nothing upstream — binds nothing anywhere, which is the faithful translation
of "filters nothing" and why an ordinary scw instance server create raises
no alarm at all.
A server with a private network keeps the claim on its emulated interfaces and on its public address, whichever order the two arrived in. That took two corrections, and both are worth keeping in view.
The sentence used to read "keeps the full claim … and a flexible IP, routed
through the filtered NIC, stays covered", and a measurement on 2026-08-27
contradicted it for the order the example stack produces (#548): a server
created with an ip_id boots carrying only its public address, so the driver
gives it a routed NIC, and the private NIC that arrives afterwards did not
take the address with it.
$ incus query /1.0/instances/feint-scw-936e816e-… | jq -c '.expanded_devices | with_entries(select(.value.type=="nic"))'
{"eth0":{"ipv4.address":"203.0.113.3","ipv4.host_address":"169.254.0.1","nictype":"routed","type":"nic"},
"eth1":{"ipv4.address":"10.30.1.10","network":"fnt-0fa112b2861","security.acls":"scw-ff5a3e3cbd3","type":"nic"}}
# the group opens 443 and 22-from-10.40.1.0/24, and names no port 80
platform-web-1 203.0.113.2:80 OPEN
platform-web-0 203.0.113.3:80 OPEN
platform-bastion 203.0.113.4:80 closed # nothing listens: the negative control| order | what the machine carries | the public address |
|---|---|---|
| private NIC first, address attached afterwards | one managed NIC, the address on its ipv4.routes |
covered — fifteen runs both ways, the paragraph below |
| address at creation, private NIC afterwards | eth0 nictype=routed addressless, eth1 on a managed network carrying the address |
covered since 2026-08-28: the driver moves the address onto the filtered NIC when the machine joins one (#548) |
| address at creation, no private network ever | eth0 nictype=routed carrying the address |
not covered: a routed NIC accepts no security option at all (#337), and there is no other interface to move to |
So the two capabilities say two different things, and a consumer needs both.
capabilities.firewall_public_only: false is about the interface: an address
living on a routed NIC is covered by nothing, and that is the third row.
capabilities.firewall_public_when_joined: true is the second row — a machine
that has an emulated interface ends up with its public address on it.
tools/conformance/functional.sh keyed a skip on the first of those two for as
long as the second was missing; it asserts the public pair now, on the address
the API publishes, and gates it on the second capability.
The group covers a flexible IP. The address is routed through the NIC
device's ipv4.routes (nic_bridged.go at v7.2.0 lists it among the device's
own fields and applies it host-side), so host traffic towards it crosses the
same bridge port the rule set filters. Measured with a deny-all group: fifteen
runs, public address dropping every time, and the counter-proof both ways —
authorise the port and the same probe connects, revoke it and it drops again.
An earlier version of this file reported the opposite: that in the conformance
suite the public address answered through a denying group. That was the suite
misreading its own probe, not the firewall: a dropped connection is killed by
timeout with an empty output, which the assertion then matched as a successful
answer. The verdict of every probe in network.sh is now an exit code.
A placement group is a scheduling constraint, and this emulator has no
scheduler: with --vm off nothing runs anywhere, and with a runtime every
machine is a container or VM on the single host that started feint. There is no
second hypervisor for max_availability to spread onto, and policy_mode: enforced refuses nothing here — a server in an enforced spread group boots
exactly like any other. A stack that needs machines to actually land apart is
not testing that property against feint.
The family is served anyway, for the reason security groups are served without
a runtime: everything a driven client does with a placement group is
create, read back and store. Measured on the Terraform provider at both pins
this repository drives — 2.43.0 (the surveyed terraform-talos stack) and
2.81.0 (the conformance fixture) — policy_respected is a Computed attribute
the provider only ever d.Sets; nothing waits on it, branches on it, or fails
an apply over it.
What keeps the record honest is that one field. policy_respected is computed
from the single-host reality at read time, never from the policy's wish:
low_latencyreads true — machines grouped on the same hardware is the one thing this emulator delivers by construction;max_availabilityreads true until two of the group's servers are running, and false from then on — two running members of a spread group share the only host there is, and saying otherwise is the lie the family was declined over before #285.
A server that is not running counts for nothing: it sits on no hypervisor, the
same doctrine that makes the server view answer location: null for it. On
the server endpoints (Server.placement_group.policy_respected) the value is
pinned false, because the SDK documents that the real API always answers
false there.
The same limit covers Exoscale anti-affinity groups, served since the
starter pack and used by the surveyed Exoscale stack: membership is recorded
and read back, and no machine is refused for exhausting hosts the way the real
platform refuses one when its anti-affinity group runs out of hypervisors.
Exoscale's API exposes no policy_respected equivalent, so there is no field
to keep honest — the record is the whole surface, and this section is where
its non-effect is stated.
Two defensible answers existed for which address a client reads on a server
(#116): the runtime address the machine got from its bridge, or the fictional
one the pack allocated from 203.0.113.0/24 (TEST-NET-3, RFC 5737). This
emulator publishes the fictional one, and routes it for real: the address
public_ips[0].address reports is the address the machine carries on its
interface, reachable from the host that runs the emulator, filtered by the
server's security group. ssh root@203.0.113.2 opens a shell.
The rejected option — publishing the runtime address — is rejected for three
measured reasons. It desynchronises the two views of one attachment, since
GET /ips/{id} must keep answering the address it allocated, and the real API
never lets server.public_ips[].address differ from the flexible IP it names.
It changes with the --vm mode and the operator's host, so the same test would
read different "public" addresses on two machines, which is the half-truth this
project exists to avoid. And it was never needed: the network conformance suite
had already proven the fictional address can genuinely answer through the
group; what #116 measured was two ordering holes, not a wrong address plane.
The two holes, and where they are held:
- An address attached before the boot was never routed.
attachAddressran while the server had no machine and silently did nothing, and nothing replayed it at poweron. Now the promised addresses ride the launch as device route keys — set while the instance is cold, because editing them on a live OVN NIC re-plugs the device and the guest loses its DHCP lease with nothing left to renew it — and poweron replays the guest half.TestPowerOnRoutesAnAddressAttachedBeforeBootandTestPublicAddressesAreRoutedBeforeTheFirstBoothold it. - A machine with no private NIC had no lawful interface for the route. It
used to boot on the operator's default profile bridge, which the driver
rightly refuses to route through (
mustOwn), and covering that NIC with a firewall meant overriding a profile device — a re-plug after boot that cost the guest its DHCP lease:incus listshowed RUNNING with no IPv4 at all. Machines with no attachment now boot on the emulator's own labelled network (fnt-default,10.209.84.0/24, deliberately obscure like the OVN uplink's block), created on first use and removed by the sweep.TestAMachineWithNoAttachmentBootsOnTheEmulatorsOwnNetworkholds it.
dynamic_ip_required follows the same mechanism (#117): poweron allocates an
ephemeral address from the same block — suppressed when a flexible IP is
already attached, which is upstream's own precedence — publishes it as a
dynamic: true entry of public_ips, and releases it on stop, standby,
terminate and delete alike. It never appears in /ips, because upstream never
lists it there. TestADynamicAddressFollowsThePowerCycle holds the cycle.
Bounds, stated rather than implied:
- The address answers from the host that runs the emulator (and from the emulated machines). It is a documentation address on purpose; nothing routes it beyond the host, and that is the point — a test that half-works against the real internet is worse than an address that visibly goes nowhere.
- A subnet-internal address — Outscale's
PrivateIp, Exoscale'spublic-ip, both of which this emulator fills with the machine's own address — answers the host in bridge mode and not in OVN mode, and the runtime declares which (capabilities.private_from_host). The cause is isolation's own machinery: the OVN router that separates two VPCs by construction also SNATs the host's connections on the way back, so the handshake never completes — measured, sshd up and answering its neighbours while the host read the port as closed. The routed public plane crosses that boundary in both modes, which is why every pack's ssh chain logs in through it: a Scaleway flexible IP, an OutscaleLinkPublicIp, an Exoscale elastic IP — each is genuinely routed to the machine, and each pack draws from its own RFC 5737 block (TEST-NET-3, -2 and -1 respectively) so two emulated clouds on one host can never route the same /32 to two machines. - On a virtual machine (
--vm incus-vm), the host half of the route is in place from the first boot, and the guest half — the address on the guest's own interface — lands on the first read after the agent answers, the same read that publishes a VM's address. - An address attached to a running server in OVN mode still bounces the NIC for an instant (the route keys are not live-updatable there; the driver restores what it can, and a DHCP-owned lease is the runtime's to re-issue). Attach before boot when the order is yours to choose; it always is in Terraform, where the IP and the server share a plan.
- A stored address — a flexible IP's, a dynamic one — is revalidated before it
reaches the driver: one outside the emulated block is refused and logged,
never routed, because a restored snapshot carries these values verbatim.
TestAPoisonedStoredAddressIsNeverRoutedholds the refusal.
Exoscale's Elastic IP is designed to be held by several instances at once: an
Elastic IP carries a healthcheck, and the platform sends the address to whichever
instance is passing it. Measured on ch-gva-2 rather than assumed — exo compute instance elastic-ip attach was accepted twice for one address, and both instances
then reported holding 185.19.28.243.
The emulator keeps that in its control plane, because it is what the real one does: both instances go on listing the Elastic IP, and a client reading either of them sees what Exoscale would show.
The runtime cannot follow. Two containers answering ARP for one /32 make the
host pick arbitrarily, and an emulator that publishes an address its machine may
or may not carry has given up the one thing it exists to promise. So the address
is routed to the most recently attached instance, and taken back from the
previous holder before it is handed over.
That rule is feint's, not Exoscale's. There is no healthcheck here, and inventing an election would put a winner in a client's hands that nothing measured. If you are testing a failover, the emulator will not perform it: attach the address to the instance you want to reach.
The same bookkeeping serves all three packs (machine.Binding.RouteAddress), so
the runtime carries one address on one machine whatever the pack's API allows —
Scaleway moving a flexible IP on request, Outscale refusing without AllowRelink,
Exoscale accepting both holders.
tools/conformance/parity.sh drives one equivalent request through the three
providers' own surfaces and counts what the host carries, read from the
runtime rather than from the API under test. Run against main on 2026-08-24
with FEINT_VM=incus, it reports four divergences with a single root cause.
| row | scaleway | outscale | exoscale |
|---|---|---|---|
| a machine on a private network, no public address asked for | 1 iface / 1 addr | 1 / 1 | 2 / 2 |
| the same, one public address explicitly requested | 1 / 2 | 1 / 2 | 2 / 3 |
The extra address is 192.0.2.1 on eth0. The instance's own read answers
public_ip: null, and no other field of the API names it, so it is carried and
published nowhere. It is present from boot: the first run aborted before any
Elastic IP existed and already measured 2/2, which rules out a leftover from
an earlier attachment.
Two things follow, and the second is the reason this is written here rather than fixed in passing:
- The parity claim itself is sound. Remove that one address and all four findings go: both rows read the same interface and address counts on the three clouds, and the orphan check reads zero. The equality this suite asserts is not too strong for Exoscale.
- What the fix should be is a product decision, not a patch. Either the
machine must not carry the address, or the API must publish it —
machines.gostates the intent ("Exoscale's eth0 is the public interface, the address this pack publishes as public-ip is the primary interface's"), and the measurement says the published half is missing. Choosing costs a reading of what a real Exoscale instance answers when no Elastic IP is attached.
Until then mise run conformance:parity runs on demand and is deliberately not
in the conformance aggregate: a red suite inside the gate every other change is
judged by teaches people to skip the gate.
Measured on a real account on 2026-09-04 (fr-par-1), on a flexible IP
created and released in the same minute, while the Day-2 leg's finding that
{"reverse": null} left the reverse in place was being settled:
- a fresh address reads
"reverse": null; PATCH {"reverse": null}answers 200 and readsnull, and so doesPATCH {"reverse": ""}: the SDK'sNullableStringValuesends the first, and the account normalises the second to it. This emulator now does both;PATCH {"reverse": "feint-676.invalid"}answers 400invalid_arguments,argument_name: reverse,reason: constraint, "Your reverse must resolve. Ensure the command 'dig +short feint-676.invalid' matches your IP address".
The last line is the limit. This emulator accepts any reverse: its addresses
are documentation blocks nothing resolves, a stack that sets a reverse on one
would fail against a copy of that refusal, and checking DNS from an emulator
that runs offline would refuse everything or nothing. What it cannot do is
what follows from that: prove that null clears a reverse that was set,
because no reverse can be set on the account without a resolving name. The
three answers above are the whole of the measurement; the clearing of a set
reverse is what TestANullReverseClearsIt asserts of the emulator on the
strength of the SDK's own type, and it is marked here as unmeasured upstream.
The driver reserves the address on the NIC (ipv4.address) and, until
2026-08-29, left the guest to ask for it. That worked on two of the three images
the Outscale catalogue serves and never on the third, which is what a
verdict that depends on the host looks like when the host is not the variable.
One Subnet, one OVN network, one machine per image, twice each. The number is
how long after CreateVms answered the machine carried the address the API had
published, read from the runtime (incus query /1.0/instances/<name>/state),
not from the emulator:
| image | AMI | address carried after the create answered |
|---|---|---|
ubuntu:24.04 |
ami-fe1a7001 |
27 ms, 32 ms |
debian:12 |
ami-fe1a7002 |
46 ms, 35 ms |
alpine:3.21 |
ami-fe1a7003 |
never |
(The three were ami-00000001..3 when this was measured; #395 moved them out
of the corpus sanitiser's minting space, and nothing about the boot changed.)
never is measured, not inferred: six alpine machines, ceilings of 45 s, 90 s
and 180 s, none of them carried it. Station of the maintainer, 2026-08-29,
--vm incus-ovn, Incus 7.2, OVN 24.03.6, CreateVms itself 22.0–22.5 s
throughout (#562's constant, unchanged).
The alpine cloud image ships dhcpcd 10.1.0, and ifupdown-ng prefers it to
busybox's udhcpc whenever it is installed. dhcpcd ARP-probes the address it
was offered before accepting it. Inside the guest:
eth0: soliciting a DHCP lease
eth0: offered 10.182.9.4 from 10.182.9.1
eth0: probing address 10.182.9.4/24
eth0: 10:66:6a:88:5e:54(10:66:6a:88:5e:54) claims 10.182.9.4
eth0: DAD detected 10.182.9.4
once a second, for ever. The MAC that "claims" the address is not the machine's
own (10:66:6a:4a:aa:69 in that run) and not another machine's: it is the OVN
gateway router's external port, 10.209.83.128 on feint-uplink. The probe
is flooded to the network's localnet port, reaches the uplink, and the gateway
answers it — answering ARP for the block it fronts is exactly that router's job,
and the block is on the uplink because ipv4.routes puts it there so the host
can reach the network at all. The guest reads its own address as taken, declines
its own lease, and asks again.
Ubuntu and Debian configure the interface from the lease without an ACD probe, so they never see the answer.
tools/conformance/outscale/network.sh waits wait_until 24 for that address
and expired at 24.816 s on this station while CI passed the same suite. The
temptation is to raise 24. Three measurements say it would have been the wrong
fix:
- at the end of a 90 s wait the guest's DHCP client is gone — nothing is asking any more, so no budget reaches it;
- in that same guest, at that same moment, one
udhcpcexchange is answered in 97–143 ms: the path was open the whole time; - waiting longer makes the DAD answer more certain, not less, because the gateway is programmed by then. CI's green on this suite is its probe winning a race, not the address arriving.
Incus.Start now calls settleFirstBoot on the launch branch: for every NIC of
ours that reserves an address, the address is configured inside the guest and —
under OVN — the private aggregates are laid through the network's router. That
is what Attach already does for a hot-added NIC and what restoreGuestNetwork
puts back after a reboot (#548, #549); the first boot was the last door still
trusting the lease. A device that reserves nothing is left to DHCP, which is
correct and is the case #202 exists for.
After it, on the same station and the same images: alpine:3.21 carries its
address 31 ms and 36 ms after the create answers, like the other two.
The sentence above is right about the address and was wrong about the wait. A
device that reserves nothing must not be given an invented address; it still has
to be waited for, because PowerOn returning is what the verification of
#670 takes as its cue to read the machine.
Measured on the maintainer's station on 2026-09-25, --vm incus-ovn, on the one
machine in this repository that attaches with no address — an Outscale Vm with
no SubnetId, created by tools/conformance/outscale/ssh.sh:
| what | when |
|---|---|
| the verdict was published | 0.6 s after the container started |
| the lease actually landed | 0.54 s |
| ssh into that same machine | succeeded moments later, in the same suite |
So the emulator reported "a machine does not carry what its plan claims" about a
machine that carried it, and runtime-proof.yml went red four nights running.
The suite passing and the verdict failing in the same run is the signature: the
suite asks at one moment, the verdict is a cumulative counter on
/_feint/health that remembers every reading, including the one taken too
early.
settleFirstBoot now waits for the lease on that branch too. Waiting invents
nothing — the guest is read, never written — so #202 stays closed, and one wait
mends both broken claims because the lease carries the private aggregates with
it as DHCP option 121 (proto dhcp on all three). The budget is 10 s and not
guestRouteWait's 90 s: a lease lands in about half a second, or the client has
declined it for good, which the measurement above this section established at
45 s, 90 s and 180 s alike.
TestAFirstBootWaitsForTheLeaseOfANICThatPinsNone fails without it, and
tools/falsify/specs/first-boot-takes-its-address.json mutates it.
Why leg.sh runtime said the opposite. It runs the four network suites; the
CI job runs those plus the three ssh ones, and the machine above is created by
outscale/ssh.sh. A local run measured broken=0 and the gate was armed on it.
Replaying the CI population — same seven suites, one emulator, in that order —
reproduced it outside the runner, and answers held=67 broken=0 unreadable=0 repaired=0 with the wait in place.
- It is not a statement about dhcpcd being wrong. A real cloud's DHCP server does not sit behind a gateway that proxy-ARPs the block, so the probe is answered here and not there.
- Nothing was measured about
--vm incus-vm. The same door is used, and a virtual machine'sincus execneeds its agent; no VM was booted for this. - The station also carried 394 stale
ovn-bridge-mappingsentries for two live OVS bridges, one added per OVN run, which madeovn-controllerlog ~4 MB ofBridge 'incusovnNNNN' not found for network 'feint-uplink'per machine creation. Pruning them changed nothing about the measurement above — the alpine machines still never took their address — so it is a separate matter, recorded here because it is real and nobody had looked.
Whether an emulated machine can reach the internet is a property of its
primary interface's shape, not of the host it runs on. Measured on
2026-08-26, one emulator (incus-ovn), one station, one minute apart, the same
#cloud-config declaring package_update: true and packages: [nginx] on
both:
| shape | outbound | DNS | cloud-init | port 80 |
|---|---|---|---|---|
| routed NIC (Scaleway server, public IP, no private network at create) | none | none | status: error on package_update_upgrade_install |
never listens |
NAT-backed network (Outscale public-Cloud Vm on fnt-default) |
yes | yes | status: done |
listening — nginx installed |
The port-22 positive control was OPEN on both, so the error measures the
machine, not the reader.
Why the routed shape has no route out is #202's decision: a machine carries
exactly the addresses its provider's API publishes, on a routed NIC with no
network underneath — and therefore no NAT and no resolver, because the images
feint images builds carry their ssh daemon already (#203) and nothing else
about the machine needs the internet.
An OVN network names its own gateway as the lease's resolver — never the
uplink's address, never a public one, never nothing (#660, #684, #697).
The machines that boot on it get the name server they use another way:
feint serve --resolver (default 1.1.1.1), written where systemd-networkd
reads it and set with resolvectl (#694, #696). What the lease names decides
one thing, the on-link /32 a guest's RoutesToDNS= lays towards it, and
that route is dead the moment the address is off the segment. Three values
were measured, all under incus-ovn, and the invariant they share is what
TestAnOVNNetworkLaysNoRouteTowardsItsResolver holds: the uplink's own
address, which is also the one the station dials a public address from —
2026-09-04, ip route get 10.209.83.1 from the guest answered dev eth1 and
the platform's published address answered nothing; a public resolver,
2026-09-04 — 1.1.1.1 dev eth0 proto dhcp scope link, ping 1.1.1.1
unreachable, and the machine could not resolve; and none at all, 2026-09-07 —
Incus names the uplink's address when nobody else does, and the first shape
came back for four scheduled nights on platform-web-0, a public address and
a private network, while the same stack at v0.12.1 passed each of them
(#697). With the gateway named, the guest holds no route towards the uplink,
no dead /32 anywhere, the reply leaves via <gateway>, and the station
reaches the published address. Name resolution rides the drop-in, not the
lease: a machine with a way out resolves through it (measured on the shapes
of #694 and #696, dns: ok after a cold attach and after a reboot), and one
holding a public address has no way out under OVN to resolve through, which
is #695's outbound half. The default route inside the guest is
via 169.254.0.1, the routed NIC's link-local next hop, in bridge mode,
where it wins even when a private NIC with a NAT-backed network is attached
later, the ordinary Terraform order on Scaleway and Exoscale (server first,
NIC after); under OVN the routed device carries no address and the guest has
no default route at all ("Which door a reply leaves by", below).
Outscale's Subnets are NAT-less on purpose and faithfully: upstream, a
Subnet reaches out only through an internet or NAT service, and this emulator
records those without routing them (see "Outscale's gateways and NAT move
records, not packets").
This is not a per-environment fact. The runtime-proof runner showed the same
boundary from both sides: machines whose outbound is routed and NATed through
the host crossed the Docker FORWARD policy (the comment in
runtime-proof.yml), and the routed shape without NAT kept the ssh suite red
for five consecutive nights until the images stopped needing a repository
(#335). Two shapes, one boundary, both environments.
Consequences, stated rather than implied:
- A stack's
packages:clause cannot complete on a routed-shape machine. cloud-init runs the user data — the config reaches/var/lib/cloud/instance/user-data.txt— and ends instatus: error, in a journal inside the guest that nobody opens. The example stacks therefore declare none (#507); their user data does what an offline machine can hold (write_files,runcmd, a python listener — python is guaranteed wherever cloud-init runs). - The emulator says it at boot: a client cloud-config that declares a
package step on a machine booting with no emulated network under it is
warned about in the emulator's log, naming this section.
TestAPackageStepWithNoRouteOutIsSaidOutLoudholds the warning, and its absence on the shapes that can install. - No NAT is added for the routed shape here. A routed NIC has no network
object, so there is no
ipv4.natto switch on: feint would have to write masquerade rules on the operator's host outside anything Incus manages, and provide a resolver besides — for an interface thatfirewall_public_onlyalready documents as unfiltered. The real cloud does give such a machine outbound, so this is a divergence, and it is recorded here rather than papered over; widening the machine layer's contract is the boundary work of the architecture audit (#514), not a patch.
A machine holding a public address has to answer at that address. #660 named the failure in one sentence — "its reply was leaving by a door the request had not come in through" — and #672 asked for the door to be measured per machine shape rather than deduced, because a verifier that invented its expected values would be a second reconciler, wrong with the same confidence as the first.
It is measured now, and the answer is that the door is not what decides.
One emulator per mode, a Scaleway server holding a flexible IP on a private
network, both orders of construction. The service is proved listening from
inside the guest by reading /proc/net/tcp for state 0A before any
reachability verdict, and the reader itself is checked against a planted witness
first (fnl_listen_reader_control), because an instrument that cannot prove
itself turns every "nothing is listening" into a measurement of the harness.
| mode | order | the door: ip route get <peer> from <public> |
neighbour | station reaches it |
|---|---|---|---|---|
incus-ovn |
address first, NIC hot | dev eth1, on-link, no via |
INCOMPLETE |
no |
incus-ovn |
NIC cold, then address | dev eth0, on-link, no via |
INCOMPLETE |
no |
incus (bridge) |
address first, NIC hot | dev eth1, on-link, no via |
resolved (DELAY) |
yes |
incus (bridge) |
NIC cold, then address | dev eth0, on-link, no via |
REACHABLE |
yes |
The door reads the same in all four. A door claim comparing via and
dev — which is what machine.door.Check compares — would report held on the
two machines that answer nobody. That is why no door claim is derived from this
table: the silence in claims.go stays, and it is now silence for a measured
reason rather than for an unmeasured one.
A published address is reached through the uplink, so the address the station answers from is not its LAN address. Read on the station while the machine is alive:
--vm incus-ovn 203.0.113.2 dev feint-uplink src 10.209.83.1
--vm incus 203.0.113.2 via 10.211.0.2 dev fnt-943cfdf1d2c src 10.211.0.1
So the guest must reply to 10.209.83.1 under OVN, and to 10.211.0.1 under
the bridge. The first is the uplink's own address and is not on the guest's
subnet; the second is the bridge's gateway and is. The guest's route to
either is on-link, so it resolves the peer itself, and only one of the two can
be resolved:
--vm incus-ovn 10.209.83.1 dev eth1 INCOMPLETE <- nobody answers the ARP
--vm incus 10.211.0.1 dev eth0 REACHABLE
The request lands either way — one connection in SYN_RECV inside the guest
while the station dials, under OVN — so the emulator is not dropping anything
inbound. It is the reply that never reaches the wire.
The on-link route itself is DHCP's: 10.209.83.1 dev eth1 proto dhcp scope link. systemd-networkd lays an on-link /32 toward every DNS server a lease
announces (RoutesToDNS=), and the OVN network named none — #693 had taken
the public resolver out of the lease for the same reason (#684) — so Incus
named the uplink for it. A /32 beats any default route by longest prefix,
which is why #672's two variants of this shape read "same gateway, same
device, opposite results": the gateway and device were never the deciding
part. Since #697 the lease names the network's own gateway, the one address
on the segment, and the /32 is live.
docs/limits.md has described the routed shape as keeping default via 169.254.0.1 even when a private NIC arrives later. That holds in bridge
mode, where the same reading shows both routes side by side and the routed
interface carrying the way out:
--vm incus default via 169.254.0.1 dev eth0
default via 10.211.0.1 dev eth1 proto dhcp src 10.211.0.2
1.1.1.1 from 203.0.113.2 via 169.254.0.1 dev eth0
Under OVN it does not. The routed device exists — nictype: routed,
ipv4.host_address: 169.254.0.1 — and carries no address at all in the
guest, so it lays no default route, and the address rides the managed NIC
instead as ipv4.routes.external. The guest then has no default route of any
kind, only the RFC1918 aggregates, and ip route get 1.1.1.1 answers Network is unreachable. That is the outbound half #695 is about, and this section is
where its cause is recorded.
Whether the real cloud's reply leaves by the interface carrying the public address, and with which next hop. No recording carries it. What is measured is this emulator's behaviour and the difference between its two modes, which is what a fix has to preserve on one side and change on the other.
Upstream, two private networks of two different VPCs do not reach each other.
Whether the emulator delivers that depends on which --vm mode backs it.
With --vm incus (bridges), they reach each other. The cause is the
runtime's, and it is documented: "traffic between managed bridge networks on
the same server isn't NATed as it's routed directly between the bridges". Two
attempts were measured. A rule set attached to the network holds when nothing
else is attached, and stops holding once the NICs carry a security group of
their own. Adding the foreign blocks to that NIC rule set did not hold either.
Both are in the tree, because they cost nothing and separate the simple case;
neither is trusted, and the conformance suite reports the state rather than
asserting a separation that is not there.
With --vm incus-ovn, they do not. An OVN network is a logical network
with its own router: another OVN network is simply not on it, so the
separation is the topology's, not a rule's. Measured before any code was
written: two OVN networks, one instance on each, ten probes per direction,
zero connections, with the control probe confirming the listener answered.
Two networks of one routing VPC are joined the runtime's own way, with
network peer — five probes, five connections once peered. The conformance
suite asserts the cross-VPC separation as a hard failure in this mode, where
the bridge mode keeps it a documented skip.
Measured again on 2026-07-29, both modes end to end: 16 checks green on
incus with the isolation one skipped, 17 green on incus-ovn with nothing
skipped for want of a capability. Exactly one assertion of the whole suite
changes verdict with the mode, and the second skip that remains in both — a
machine carrying an address the API does not publish — is not mode-dependent.
Since that run the suite no longer compares a mode name: it reads
capabilities.isolation from /_feint/health, so a runtime that declares
isolation and fails to deliver it is a hard failure, and one that declares none
is skipped rather than silently passed. What a mode can prove is declared by the
driver (machine.Capabilities) instead of inferred from its name.
The mode has prerequisites the bridge mode does not — ovn-central,
ovn-host, and an Open vSwitch pointed at the local northbound socket — so it
is asked for by name — but --vm auto tries it first and falls back to the
bridge, which is the reverse of what this paragraph said until the ordering was
fixed. The old order was backwards for the reason that matters here: an operator
who had installed all three got the one mode that cannot isolate two VPCs, and
nothing told them. What auto chose is printed at startup, with its isolation
capability beside it.
Three behavioural differences are accepted and worth knowing. A security
group's drop default answers as a reject on an OVN NIC (the NIC's own
default for unmatched traffic; the port is closed either way, but an RST is
visible where upstream is silent), because the NIC-level default-action keys
cannot be changed on a live OVN NIC without re-plugging it —
UpdatableFields in the runtime's nic_ovn.go at v7.2.0 lists no
security.* key but security.acls itself (the rest are limits.* and
connected), and the re-plug was measured to cost the guest every
address it carried. A flexible IP still rides the NIC's own route keys as in
bridge mode (ipv4.routes.external, whose l2proxy ingress mode answers ARP
for the address on the uplink and delivers packets carrying the public
destination), but since those keys are not live-updatable either, attaching
or detaching one bounces the interface for an instant while the driver
restores the guest's addresses. And between two servers of one subnet whose
groups both restrict traffic, the sender's group wins: the runtime evaluates
every NIC rule set in a single pipeline where an egress allow (priority 300,
or 111 for a default) is tested before the receiver's ingress default
(priority 100) — the constants are in the runtime's acl_ovn.go at v7.2.0,
and the bypass was measured before being believed. The emulator narrows the
gap the only faithful way available: a group that enforces nothing, such as
the default security group, attaches nothing, so the common case — one
restrictive group probed from an unrestricted neighbour — filters exactly as
upstream does. Across two subnets the receiver's policy always holds, because
traffic enters through the router port and no sender rule matches there.
--vm incus-vm gives each server its own kernel. It does not give it a private
NIC after boot: Incus refuses to add the device to a running virtual machine,
and says why.
Failed to start device "eth1": Failed adding NIC device:
GenericError: PCI: slot 0 function 0 not available for virtio-net-pci,
in use by virtio-balloon-pci,id=qemu_balloon
Measured on Incus 7.x with the emulator's own fixture: a server created, started and then attached to a private network. The container modes attach it without trouble, because a veth pair needs no PCI slot.
Two things follow, and both are in the code rather than only here.
The NIC says so. When the runtime refuses the attachment, the private NIC is
answered as syncing_error, which is a state PrivateNICState declares. It used
to stay available while the failure lived in a log line, so the API published
an address through IPAM that the guest never carried — the one failure this
project exists to avoid. Polling the guest for three minutes confirmed it never
took the address.
The address is applied on the guest's own interface, once its agent answers.
Two separate defects were fixed on the way to measuring this one: the driver
passed Incus's device name (eth1) into a guest that names the interface
enp6s0, and it configured the interface before the virtual machine's agent had
started, which answers VM agent isn't currently running rather than "not
running". Both are fixed and tested through the injectable runner; neither is
enough to work around the PCI limit above.
The order that does work with incus-vm is to attach before the first boot,
which is what terraform apply does when the NIC is part of the same plan as
the server.
The lifting condition is upstream's, not this emulator's: the refusal is the
runtime's own (PCI: slot … not available), so this section is to re-measure
when the station's Incus moves past the version above — nothing here can widen
what the daemon refuses.
Outscale declares DryRun on every action, and the real API uses it to say
whether the request would have been accepted: it validates the arguments and
answers without changing anything.
This emulator honours the first half. The flag is answered at the mount point, before the handler runs, so nothing is created, started or deleted — which is the property that matters for a host, and the one a client relies on to probe safely.
It does not honour the second half. A dry run of a malformed request answers 200
here and 400 upstream, because no handler ever sees it. A client using DryRun
as a validator therefore learns nothing from this emulator, and a client using it
to avoid a side effect is served correctly.
The reason it is not implemented per handler is measured rather than aesthetic:
the first attempt honoured the flag inside six handlers, and Outscale declares it
on all twenty served actions. DeleteVms --DryRun true destroyed the machine —
a control implemented per handler was missing from exactly the destructive ones.
Answering at the mount point cannot have that gap, and gives up validation to
get there.
TestDryRunReachesNoHandler holds the half that is served.
terraform apply works against Scaleway here, through the provider's api_url
attribute. It does not work against Exoscale, and the reason is not something
this emulator can fix.
exoscale/exoscale 0.70.0 builds two clients in pkg/provider/provider.go:
// egoscale v3
if ep := os.Getenv("EXOSCALE_API_ENDPOINT"); ep != "" {
opts = append(opts, exov3.ClientOptWithEndpoint(exov3.Endpoint(ep)), ...)
}and an egoscale v2 client, created with no endpoint option at all. The variable is therefore honoured for one and ignored for the other.
An apply does not fail cleanly and does not work: it splits. Some resources answer from this emulator, and the rest are created on the real cloud, in the same run, with whatever credentials the environment holds.
Measured on 0.70.0 without a byte leaving the machine — outbound traffic routed to a proxy that was not listening, so the attempt is visible and cannot succeed:
Error: Post "https://api-ch-gva-2.exoscale.com/v2/ssh-key"
proxyconnect tcp: dial tcp 127.0.0.1:4740: connect: connection refused
with EXOSCALE_API_ENDPOINT=http://127.0.0.1:4733/v2 set. The emulator saw
nothing.
There is no endpoint setting to reach for. tofu providers schema -json
lists five provider attributes — key, secret, timeout, environment,
sos_endpoint — and their own documentation lists the same five. sos_endpoint
is Object Storage only; environment composes a %s-%s.exoscale.com domain.
Neither points anywhere local.
Nor is there one deeper in the client. egoscale/v2 reads no environment
variable of its own, and the .exoscale.com suffix is compiled into
v2/api/request.go. The option that would do it, ClientOptWithAPIEndpoint,
exists in egoscale/v2 and is never called by the provider — a grep over
exoscale/ and pkg/ returns nothing. Three sites build a v2 client without
it: CreateClient, getClient, and the plugin-framework provider's
Configure.
So the emulator refuses that client, by the user agent it sets itself
(Exoscale-Terraform-Provider/…), with a message saying what is happening. Half
serving it is the worst of the three outcomes: a half-success is
indistinguishable from working until the invoice arrives.
That was true of every published provider up to and including v0.70.0. Upstream fixed it in #576, released as v0.71.0 on 2026-08-31, and the refusal came off on 2026-09-05 — on a measurement, not on the release note.
Between 2026-08-26 and that date the refusal started one layer earlier still:
no Terraform for Exoscale at all, fork included. #525 had measured why the
user-agent guard alone was not enough — a feint down on the example stack,
run without the fork's dev_overrides, resolved the published 0.70.0 and its
refresh sent five signed requests to api-ch-gva-2 and api-ch-dk-2. Traffic
that leaves for the real cloud never reaches an emulator-side guard, and what
stood between those requests and a real account was one unasserted property:
feint up appends the pack's deliberately public credential pair after the
caller's environment, so the fake pair wins and Exoscale refuses the signature.
That property is now held by
TestThePacksOwnCredentialsOutrankTheCallersShell.
The same instrument that closed the door opened it. feint proxy --forward '*.exoscale.com=<emulator>' accepts the CONNECT, terminates the TLS with a
certificate minted for the run, records every host asked for, and sends the
request to the emulator instead of the real host — so a provider that ignores
its endpoint is caught rather than obeyed, and no measurement of this can
reach a paying account whatever its answer. Driven against
examples/stacks/exoscale, the published provider, no dev_overrides and no
fork, 2026-09-05:
| provider | apply | second plan | destroy | hosts on the wire |
|---|---|---|---|---|
| v0.70.0 | 15 created | no resource changes | 15 destroyed | 57 to api-ch-dk-2 and api-ch-gva-2 |
| v0.71.0 | 15 created | no resource changes | 15 destroyed | none |
The v0.70.0 row is the control, and it is what makes the v0.71.0 row mean anything: an empty transcript proves nothing until the instrument has been shown able to report a full one.
A configuration pinned below v0.71.0 resolves a provider that still splits, so
the refusal did not disappear — it acquired a version. The emulator reads the
version out of the user agent the provider sets itself
(Exoscale-Terraform-Provider/0.70.0 (c4e8499d) …, measured off the wire) and
refuses anything older, naming the floor. A client it cannot parse is served:
the refusal is for a version measured splitting, not for everything unfamiliar.
examples/stacks/exoscale states version = ">= 0.71.0", and
tools/conformance/exoscale/terraform.sh asserts the resolved version against
that floor before it applies anything.
Three things went with the decision, rather than staying as names:
FEINT_EXOSCALE_ALLOW_TERRAFORM, whose only subject was verifying a candidate fix by hand. The fix is released; driving a provider older than the floor is a different feature with a different name, not a variable that turns this off. The declaration schema's refusal of that name went with it.- The
feint up/feint downdoorstep veto (packEngineVeto), which no pack implements any more. A mechanism with no implementer and guards with no subject are what this repository calls an intention written in the past tense. - The
up.go VetoEngineproof kind in the capability matrix, for the same reason.
The exo CLI is unaffected and is driven by the conformance suite: it reads
EXOSCALE_API_ENDPOINT for everything.
Closing this properly needed an endpoint option on the provider's v2 client,
which was upstream work. It was filed as
exoscale/terraform-provider-exoscale#573, with the mechanism, the
three construction sites and a reproduction, and fixed in v0.71.0 on
2026-08-31 — the condition written into every refusal above, and met. While
it was open, feint env exoscale printed the warning on stderr, where eval
cannot swallow it; the warning outlived the refusal by a day and told a
Terraform user not to do what the pack had just started supporting (#701), so
the note now names the floor, read from the constant the refusal reads.
Obsolete since 2026-09-05: upstream released the fix (v0.71.0) and the
published provider drives this pack, so there is nothing left for a fork to
do. It had been retired from use on 2026-08-26 (#525). This section stays as
the
dated record of what the fork was for and what it proved — this repository
does not rewrite its past, it dates it. Do not follow the recipe below to
drive a stack; feint up refuses the engine, and
tools/conformance/exoscale/terraform.sh refuses with the same reasons. The
recipe's remaining audience is whoever verifies a candidate upstream fix.
The fix is four lines per site, so it is also carried on a fork, pinned:
stephrobert/terraform-provider-exoscale@fix/v2-client-honours-api-endpoint, commit2e78b42, branched fromde9d60c2(0.70.0 plus six commits).
This recipe is a snapshot, and nothing re-checks it. It was last verified on
2026-08-11: the branch tip was still 2e78b42 and the build below succeeded.
No gate clones a third-party repository — deliberately, that would put someone
else's availability in this project's CI — so past that date the honest claim
is "it worked then", not "it works". The recipe checks out the measured commit
rather than the branch tip for the same reason: a tip can move under a reader,
a commit cannot. If the build breaks or the fork disappears, check
exoscale/terraform-provider-exoscale#573 first — upstream landing an
endpoint option is the outcome that makes this whole section obsolete, and the
released provider is then the thing to use.
It passes ClientOptWithAPIEndpoint at the three sites, and nothing else.
Terraform resolves it without a registry, through dev_overrides:
git clone -b fix/v2-client-honours-api-endpoint \
https://github.com/stephrobert/terraform-provider-exoscale
cd terraform-provider-exoscale && git checkout 2e78b42
go build -o /tmp/tfp/terraform-provider-exoscale .
cat > /tmp/dev.tfrc <<'RC'
provider_installation {
dev_overrides { "exoscale/exoscale" = "/tmp/tfp" }
direct {}
}
RC
eval "$(feint env exoscale)"
export TF_CLI_CONFIG_FILE=/tmp/dev.tfrc
terraform apply # no `init` for an overridden providerMeasured against this emulator, with a security group (v3 client) and an SSH key (v2 client) in one configuration:
exoscale_security_group.v3_side: Creation complete after 0s
exoscale_ssh_key.v2_side: Creation complete after 0s
Apply complete! Resources: 2 added, 0 changed, 0 destroyed.
Both calls arrived — POST /v2/security-group and POST /v2/ssh-key, 200 each
on /_feint/trace — the second plan was empty, and destroy removed both. The
same configuration on the published 0.70.0 creates the security group here and
sends the SSH key to api-ch-gva-2.exoscale.com.
FEINT_EXOSCALE_ALLOW_TERRAFORM=1 is still required: the emulator refuses by
user agent, and the fork does not change the user agent it sets.
One limit the fork does not lift. setEndpointFromContext in egoscale/v2
rewrites the request host from the zone context unless the configured host is
an IP literal. The fork is therefore honoured end to end for
http://127.0.0.1:4599/v2 — which is what feint env exoscale prints — and a
hostname endpoint such as http://gateway.internal:8080 would still be
rewritten back to *.exoscale.com. Closing that half is a change in egoscale,
not in the provider.
The v3 resources are not out of reach, and a measurement said otherwise for
a day. tools/conformance/exoscale/terraform.sh first read the whole stack as
blocked at block_storage on list zones: ListZones: Not Found: no route matches this path, but it is served under /v2/, and #448 recorded the cause as "the fork
corrects the v2 client, and the stack now uses v3". That was wrong. The script
was exporting EXOSCALE_API_ENDPOINT as a bare host, and the v3 client resolves
its zone endpoint against exactly what it is given — the path belongs in the
value, which is why feint env exoscale prints one and why
examples/stacks/surveyed.md records one. With the path restored, the same
stack applies sixteen resources through the same pinned commit, block storage
and private networking included, plans empty and destroys clean (2026-08-24).
The error is worth keeping written down: the message named the fix in as many
words, and it was read as a statement about the fork.
It does not count towards conformance, and must not. The north star of this
project is that the official client cannot tell the difference; a client this
project patched is no longer the official client. What the fork proves is real
and worth having — that the rest of the emulated Exoscale surface holds under
Terraform, and that #573 is the only thing in the way — but it is a weaker claim
than a route driven by a published client, and adding the two together would
repeat the error probed exists to avoid. Exoscale's preview label came off
on what exo proves, at EXO-2, and not on this.
Deleting an Outscale Vm here leaves the record readable, reporting
State: terminated, and it stays that way until the emulator restarts.
The visibility is not a convenience, it is required. The Terraform provider
answers DeleteVms by polling ReadVms until the Vm reports terminated; a
record that vanished makes it read an empty list, and the plugin crashes outright
— "Plugin did not respond", on every destroy. The real API keeps a terminated Vm
readable for the same reason: a state a client waits for has to be observable.
What differs is how long. Upstream, a terminated Vm disappears after a few minutes. Here it never does, because this emulator has no clock of its own to expire anything on and inventing one would mean a background timer whose only purpose is to make a resource vanish while a test is looking at it.
The practical consequences, none of them silent:
ReadVmswith no filter lists terminated machines. A client counting machines must filter onVmStates, which this pack serves — that is what the filter is for, and it is why refusing an unserved filter matters.- A terminated Vm holds nothing: it is ignored when a Subnet is deleted and when
addresses are counted. Otherwise
terraform destroyfailed on the Subnet, naming the machine it had just terminated. feint restartclears them, like everything else: the store is in memory.
Since #378 and #389, CreateVms cuts every machine a root BSU volume from the
snapshot its image names, ReadVolumes answers for it, ReadVms publishes it
under /dev/sda1, and DeleteVms destroys it — DeleteOnVmDeletion is true
on it, which is what the recorded account answers on its own machine in every
read of its life.
The chain behind it exists so that no identifier a response publishes is
decorative: the image names a snapshot ReadSnapshots answers for, the snapshot
sizes the volume, the volume names the machine. A fictional root VolumeId was
tried once and the Terraform provider resolved it — volume vol-rooti149 not found ended a whole conformance run.
What that volume is not is a disk. It carries a size, a type, a state and a provenance; it holds no bytes, exactly like every other volume here. Three consequences a user can meet:
- Growing it changes a number and nothing else. There is no filesystem to extend, and nothing inside the machine sees the new size.
CreateImagefrom aVmIdanswers an emptyBlockDeviceMappings. An image is a copy of a disk's bytes, and there are none to copy, so no snapshot of the machine's device is cut. Cutting an image from a snapshot works and is whattools/conformance/outscale/terraform/storage.tfdrives. A machine created from such an image still gets a root volume, with a size and noSnapshotId: naming one would be a relation that resolves and is false.- "Nothing left behind" now includes a volume per machine. A suite that counts what a teardown leaves has one more object per machine to account for, and it goes with its machine rather than surviving it.
The catalogue's own three snapshots are the emulator's, not the client's:
DeleteSnapshot refuses one with a ResourceConflict naming the catalogue,
because deleting it would leave every catalogue image pointing at a snapshot no
read answers for. Their VolumeId names the volume each image was cut from, and
that volume is gone — which is the state every OMI of a real account is in, and
one this emulator reaches by ordinary means (CreateVolume, CreateSnapshot,
DeleteVolume). It is the one identifier of the chain that names nothing, and
the one no client follows.
The thirteen block-storage operations (EXO-4) are a control plane. A volume is a
record carrying a size, a state and the instance holding it; nothing is
allocated, nothing is written, and no machine gains a disk when one is attached.
A --vm run does not change that — the machine driver knows nothing about this
product, and a volume attached to a running container is a fact of the API and
of nothing else.
What follows from it, stated rather than left for a reader to find out:
- A snapshot is instantaneous and costs nothing, because it copies nothing.
Its
sizeandvolume-sizeare both the volume's declared size — there is no compression to model, and a ratio invented here would be fiction a client could compute against. blocksizeis 4096 on every volume. One storage class, one number; a value varying per volume would be arithmetic nobody measured.encryptedis always true, which is Exoscale's own default rather than a claim about this emulator. No byte is encrypted because no byte exists.- A resize is a number changing. The refusal to shrink is real and enforced, because that one has a consequence a client can hit; growing is a field write.
What is enforced is every relation a client's plan depends on: one volume
attaches to one instance and refuses a second, an attached volume refuses its
delete, a volume with snapshots refuses its delete, and a detach of nothing is an
error rather than a success. Those are the rules a destroy walks in order, and
they are driven by exo compute block-storage on every pull request.
One refusal's wording is load-bearing, and only Terraform could show it. The
provider's destroy calls detach unconditionally, and tells a tolerable refusal
from a real failure by reading the message:
strings.HasSuffix(err.Error(), "Volume not attached"). So this emulator answers
exactly that sentence. It first answered "the volume is not attached to an
instance" — the same fact, refused the same way, and terraform destroy died on
unable to detach volume for every volume that had never been attached. exo
cannot show this: it does not detach before deleting. Measured against the
patched provider pinned below,
which is what that fork is for — a check by hand that does not count towards
conformance, and did find something.
Both packs exist to prove the core stays protocol-neutral: three genuinely
different dialects, one store, one port. The generated tables in the README and
the artefacts under coverage/ carry the current counts; this paragraph
deliberately names none, and its own history is the reason. From the first
commit (2026-07-30) to 2026-08-27 it called the two packs "starters" and
enumerated their coverage by hand — the machine lifecycle, inventory and
addressing for Outscale; instances, catalogue and SSH keys for Exoscale — and
that sentence then survived, unretouched, Exoscale's block storage (#12), its
NLB (#345) and its instance pools, and Outscale's load balancer with its OVN
dataplane. Ten days of understatement in the file whose job is to be exact
about what is not served — and an understatement misleads exactly like an
overstatement: somebody reads it and does not try what works.
What remains true, and is the limit: the two packs cover less of their
upstream surface than the Scaleway pack covers of its own (measured by the
same drift scan for all three, coverage/*-coverage.json, regenerated by
mise run drift:update), and every unserved operation is triaged in each
pack's Declined() with its reason. Adding an operation is ordinary work (see
the provider-pack-author skill), not a redesign. The lifting condition is
the tables', never this paragraph's: the day the counts converge, the
generated artefacts say so without anyone rewriting a sentence here.
Every response the emulator writes is checked against the provider's own API
description. The three descriptions are not equally strong, and the artefact
under contracts/ records which case it is rather than leaving a reader to
assume they are alike.
The table below is generated from those artefacts by feint docs, and
feint docs --check fails when it drifts. That is not decoration: this section
previously read "Exoscale — assumed. 7 of 299 schemas carry it", a denominator
that had gone stale and a numerator nobody could recompute. Rewriting it by hand
produced a worse number still — 464 of 468 — which measured the extractor's own
--assume-closed flag rather than anything Exoscale declares. A figure that
counts your own assumption and reads like evidence is the exact failure this
project exists to avoid.
The schema check is one-directional: it catches a field a response invents,
and it can only catch an omitted field where the provider declared it
required — which Scaleway does on 9% of its schemas, Outscale on 27%,
Exoscale on 35%. On the rest, an omission never violates anything.
Since #88 the other direction has its own control, and its precision was
measured before its semantics were chosen. Every answer a conformance run
provokes is compared with the full property list the provider's document
declares; an absence fails the run (fields.missing on /_feint/conformance,
gated by tools/conformance/score.sh) only when a recorded real-cloud answer
carries the field too. The corroboration is not caution, it is arithmetic:
on the first instrumented run, of the 106 declared-but-absent fields a
recording could arbitrate, 83 were absent from the real cloud's answer as well
— pagination tokens that only exist when a further page does, client tokens
echoed only when sent. A gate that failed on all of that would be red on
purpose, and a gate that is red on purpose is a gate somebody turns off.
What no source can arbitrate is published rather than failed on:
fields.unconfirmed lists, per operation, every field only the document
vouches for — 317 fields across 97 operations at the time of writing. Each
entry is one recording away (feint shapes --record) from becoming either a
failure or nothing. That list is this page's kind of sentence: it names
exactly what nobody has proven.
A missing response schema in contracts/*.json has three causes, and only one
of them can be checked. The artefact distinguishes them because folding them
together is what left thirty-one served Scaleway operations reading unchecked
on the contract axis, and thirty-two reading none on probed, for months
(#429).
-
The document declares a body.
responsenames its schema, and the answer is validated against it. 306 of Scaleway's 370 documented operations, all 236 of Outscale's, 371 of Exoscale's 374. -
The document declares a success with no content at all. Scaleway writes
204: {description: ''}on 64 operations, and it is the only provider here that does.noContentcarries the status, and the answer is held to it in both directions — a body where none is declared, and a status the document does not name. This is a validation, not a silence, and reading it as one is what put those thirty-one at zero.The 64 are not simply "the DELETEs", and that is what makes the field worth reading rather than the method: 52 of Scaleway's 56 DELETEs are here, and the other four —
vpcgw/v2.DeleteGateway,vpcgw/v2.DeleteGatewayNetwork,lb/v1.RemoveBackendServers,lb/v1.UnsubscribeFromLb— declare a body and answer one. Twelve operations that are not DELETEs are here too, theSet*user-data and cloud-init family among them. -
The document declares a body this extraction cannot name. A top-level array, a free-form object, a media type that is not JSON. Three Exoscale operations are in this case —
list-events,get-sks-cluster-inspection,list-sks-cluster-deprecated-resources— and they stayunchecked, correctly: nothing about the emulator's answer is known.
One thing the probe cannot reach, and it is a property of the probe rather than
of any pack: instance/v1/API.{Get,Set,Delete}ServerUserData address a key by
name, and no call in a probe run produces one. Measured — a server the probe
creates answers {"user_data":[]}, because the only operation that could put a
key there is the one that needs the key. A client invents the name
(scw instance user-data set key=cloud-init does), and the probe may not invent
anything. Those three earn the contract axis from client traffic and stay at
zero on probed.
start backgrounds the emulator itself, with no & and no container runtime,
and the README makes a point of it: a JVM or a CPython process cannot cleanly
daemonise, a static Go binary can.
Measured on 2026-07-30, on this repository's first public CI run: on both macOS
runners (macos-15, Apple Silicon, and macos-15-intel) the detached process
exits immediately and feint start reports it. feint serve in the foreground
serves the three control planes there normally, which is what the cross-platform
job asserts.
So the claim holds where it was measured, Linux, and nowhere else yet. On macOS,
run feint serve and background it yourself. The cause is not diagnosed and no
fix is promised here; what is promised is that the page says which platform the
sentence was measured on, which is the whole difference between a claim and a
proof.
| Provider | API description | Schemas | Unknown fields |
|---|---|---|---|
| Exoscale | 2.0.0 |
473 | assumed by this emulator |
| Outscale | 1.42.0 |
655 | declared by the provider |
| Scaleway | instance/v1, instance/v2alpha1, vpc/v2, ipam/v1, iam/v1alpha1, marketplace/v2, block/v1, block/v1alpha1, lb/v1, vpcgw/v2, account/v3, baremetal/v1 |
708 | declared by the provider |
Declared means the provider wrote additionalProperties: false themselves:
refusing an unknown field enforces their rule, and a violation is theirs to
answer for. Assumed means their document is silent and this emulator closes
the schemas anyway — a deliberate choice, because a field their document does not
describe is a field this project has no way to emulate faithfully, but one that
is the emulator's own and could be wrong. The distinction is recorded so nobody
reads the second as the first.
Scaleway sits apart again: its descriptions are generated from protobuf, where
the wrapper types (google.protobuf.StringValue and its family) carry no
additionalProperties at all because the concept does not exist upstream. The
extractor unwraps them, which is why the policy reads declared without a count
that would mean anything.
The page under /_feint/ui is held by five things. Four read the asset as text
and one runs it, and the difference matters when reading a green check.
Read as text, in Go, on every go test:
TestTheEmbeddedPageNamesNoProvidergreps the shipped files for the names the packs declare, so a provider name written into the page fails the build.TestThePageNeverBuildsMarkupFromAStringrefusesinnerHTMLand its family, which is the cloudinit lesson applied to HTML.TestEveryNodeTheScriptWritesToExiststies every identifier the script looks up to an element the document carries.TestThePageAddsOnlyGETRoutesenumerates the mux — about the server rather than the page, and the only one a rewrite of the script could not defeat.
Run, in a real browser, by tools/ui/screenshots.sh (mise run docs:ui), on
every pull request in CI and on demand locally:
- the page is loaded against a live emulator, and the harness waits until it has rendered that emulator's data;
- eighteen assertions compare what the document displays with what the endpoints answered: the routes mounted, driven, probed and never proven; the driver name and the resource count; a created resource with its identifier, its kind and one of its attributes; a refusal reason, in full, in the product it belongs to; a call in the log with its path, the field no handler read, and the one that found no route mounted; the number of rows in the drill-down behind the headline figure; and that the page threw no exception on the way.
So a renamed node, a region that fails to render, a number written into the wrong element, or a script that throws halfway now fail CI. That is the hole this section used to describe, and it is closed for the values above.
Where the guarantee stops, said plainly because a reader will take this as a promise:
- It is a smoke test of one state of the page, not a test of its behaviour. The harness clicks three things — one legend button, one product, one resource — and asserts what appears. Pausing the log, the "problems only" filter, the search box, the theme toggle, the reconnect after a dropped stream and the no-flicker refresh are all exercised by nobody.
- It asserts values, never appearance. A stylesheet that renders every card invisible, illegible or overlapping passes: the nodes still carry the right text. Only a human looking at the images catches that.
- It runs one browser engine. Chromium is what CI has; Firefox and Safari are
unmeasured, and this page uses
color-mix(),<details>styling andEventSource, all of which are supported everywhere and none of which anyone here has checked outside Chromium. - It needs a browser. Without one the script exits 3 and says so — loudly, never as a silent pass — and on that machine the page's DOM is simply unchecked.
docs/assets/ui/*.png are regenerated by mise run docs:ui and committed. The
freshness gate is feint docs --check, which the pre-commit hook, mise run docs:check and tools/release/preflight.sh already run — one rail, not a second
one somebody has to remember.
What it compares is a digest of the three files the page is made of, recorded in
docs/assets/ui/manifest.json when the images were written, against the page the
binary serves. Change the stylesheet without regenerating and the gate fails.
What it does not compare is the images themselves, and that is a decision rather than an omission. This page renders wall-clock values by design — the time of each call, the age of each resource — so two captures a second apart differ; and the same capture taken on a workstation and on a runner differs again, because font rendering does. A gate demanding byte equality would be red permanently, and a permanently red gate gets disarmed, which is worse than no gate because it still looks like a control.
The consequence to hold on to: the pictures are guaranteed to be of this page, and nothing guarantees they are good pictures of it. That is a human's job at review time, and the images are in the pull request diff for exactly that.
feinttest.Start pulls the published image and runs it through the local
container runtime. That works on this station, on amd64 CI runners, and under
docker run --platform linux/arm64 with qemu — and it does not work on
ubuntu-24.04-arm, where the container starts and dies at once:
feinttest: the emulator never answered on http://127.0.0.1:33437: context deadline exceeded
container exited (code 255), and it wrote nothing
Both tests of the package, twice, on 17 August 2026. The exit code and the empty
log are the whole of what is known: the image is a multi-arch index carrying a
real arm64 binary — docker manifest inspect lists linux/arm64, and that
variant serves its routes here under emulation — so the runner's runtime refuses
something the binary is not doing. Nothing available here reproduces it, and a
guess written down would be the kind of sentence this document exists to avoid.
What follows from it. The arm64 leg of the test matrix does not set
FEINT_TESTCONTAINER, so the package skips there. Every other leg sets it, and
go test ./... on a developer's machine skips it too unless asked — the package
reaches a registry, and mise run check is offline by contract.
The diagnosis above is itself a fix: the first version of this package passed
--rm, so a container that died was gone before anything could read it and the
CI printed No such container where the reason should have been.
TestAFailedStartSaysWhatTheContainerDid is the control, and
tools/falsify/specs/container-diagnosis.json puts the flag back and requires
it to fail.
What is not affected. The image itself. feint start, the OCI image as a CI
service, and the composite action all run on arm64; this is about a Go test
driving a container runtime, on one class of runner.
The CI gate watches the products the emulator has begun serving. Asking it to account for all 1700 upstream operations would fail forever and train everyone to ignore it. Widening the scope is a decision to make when a product is started, not before.
The Outscale suite drives octl since 2026-08-25 (#460), and four of the things
it sends cannot be expressed as octl flags. None of the four is a limit of the
emulator; all four are limits of the client, and they are written here because
somebody reading tools/conformance/outscale/octl.sh will otherwise read
--payload as sloppiness.
Three are a flag that is missing or miswired, and they are unchanged since the pin moved to v0.0.32 — re-measured 2026-09-14, each case gives the same answer on both binaries. The fourth arrived with v0.0.32 and is a deliberate removal rather than a defect: see below.
octl generates one flag per field of the SDK's own request struct. Its flag
builder (pkg/builder/build.go) has a case for bool, int, int32, int64,
string and map, and none for float, so a field typed as an array of
numbers gets no flag at all:
| call | field | what octl v0.0.31 and v0.0.32 both do |
|---|---|---|
ReadVmTypes |
Filters.MemorySizes ([]float32) |
no flag is generated |
ReadVolumes |
Filters.CreationDates (a slice of iso8601.Time) |
the flag exists and is unusable |
ReadNetPeerings |
Filters.ExpirationDates (same) |
the flag exists and is unusable |
The second and third are a different defect and worth naming precisely, because
the error message points at the value and the value is not the problem. The flag
is registered as a scalar Var — buildFlags's reflect.String case takes
the FlagValue branch without ever looking at f.Slice — while the setter asks
pflag for a string slice. Every value is refused, well-formed or not:
$ octl iaas api ReadVolumes --Filters.CreationDates 2026-01-01T00:00:00.000Z
an error occurred: invalid Filters.CreationDates value: trying to get
stringSlice value of flag of type osctime
The fourth is not a defect at all. v0.0.32 withdrew --ResultsPerPage from
the 43 reads that carried it and --NextPageToken from 36, and paginates
itself behind one --max-pages (measured 2026-09-14 by reading --help across
the 234 iaas actions it serves, and the only flag it added in return is
--style). The client stopped letting its caller choose a page size; the API
did not stop declaring one, and this pack refuses a size outside the published
bound. So the suite sends the parameter as a payload rather than dropping the
assertion, because letting a client's convenience decide what the emulator is
held to is how coverage disappears without anybody choosing it.
All four go through --payload instead, which is still octl composing and
signing iaas api <Call> — not a hand-rolled curl, which would stop measuring
what a real client does. And the substitution cannot pass silently: octl
decodes a payload with DisallowUnknownFields and drops it without failing
when the decode errors, so a mistyped payload would send the request with no
filter, the emulator would answer 200 with the whole inventory, and the suite's
refuse_call fails on "accepted what it must refuse". The assertion is what
proves the field arrived.
The same version changed three flags from string to a new base64File kind —
CreateKeypair --PublicKey, CreateVms --UserData, UpdateVm --UserData, and
those three only. They take a file path now and the client encodes its bytes,
which is what Outscale documents those three fields to be. The suite passes
paths; both suites carry a version floor so an older binary fails on a sentence
rather than on an opaque 400.
One more cost of the same client, and v0.0.32 is where most of it went away.
Measured 2026-08-25 on v0.0.31 with the return code and a slice of output beside
every timing: octl spent about 700 ms starting up on every invocation with no
network at all — --version 678 ms, --help 689 ms — against roughly 30 ms for
the HTTP request. Re-measured 2026-09-14 on v0.0.32: --version 102 ms, and the
octl conformance leg end to end 38 s against 184 s, both green, same
machine, one run each.
That is why the Outscale suite fills its public-IP block through one --waitfor
process rather than spawning one per address: 255 calls in 51 s that way,
against 186 s as separate processes, measured on v0.0.31. Those two numbers have
not been re-measured, and the block is written that way regardless — a fixed
cost of 100 ms times 255 calls is still the dominant term.
The comparison with oapi-cli that justified the swap stays as recorded: the
whole suite ran in 369 s against oapi-cli's 177 s on v0.0.31, and the trade was
deliberate, the old client's 409 back-off costing 12 s a refusal on eleven
refusals against a startup that was cheap. v0.0.32 removes most of what was
being traded away.
The API description enumerates eight zone names and this emulator publishes one per process. It is not an oversight and it is not laziness — it is what the official CLI forced.
Serving all eight was the obvious fix and it was worse. The CLI queries every
zone it is told about and merges the answers, so eight zone entries pointing at
one emulator turned a single instance into eight identical rows in
exo compute instance list. A resource duplicated per zone is a defect a user
sees on their first command. At Exoscale a zone is a property of the endpoint
(api-<zone>.exoscale.com), so one endpoint honestly serves one zone.
Which zone is the operator's choice since #278: FEINT_EXOSCALE_ZONE
selects any zone their document publishes, and unset keeps ch-dk-2, the
CLI's own default — serving a zone the CLI does not default to makes every
unflagged command fail before it calls anything, on
find zone: not found in ListZonesResponse. Three of the five surveyed
Exoscale stacks target another zone (#262), which is what made the choice
worth a knob.
Two consequences, stated rather than hidden:
- a client asking for a zone the process does not serve is refused, where the
real cloud would serve it. That is a visible, honest difference; the
alternative — one endpoint pretending to be eight zones — is a silent, wrong
one. Where the refusal lands depends on the client, because the two
families do opposite things with the zone list. The exo CLI merges every row
into every listing, so it keeps the single row and a mismatch dies inside it
as
find zone: not found in ListZonesResponse. The Terraform provider (behindFEINT_EXOSCALE_ALLOW_TERRAFORM=1) resolves endpoints by name and never merges, so its zone list also carries a signpost row for each of the seven other published zones, pointing at/v2/unserved-zone/<zone>— and the next call is refused naming the deployment's zone, the resolved zone and theFEINT_EXOSCALE_ZONEremedy. Measured onexoscale_domain, which the provider resolves through a hardcodedch-gva-2: before the signpost its apply died client-side asfind zone: "ch-gva-2" not found in ListZonesResponse, a message that sends the reader after their zone configuration when DNS is simply not served (#262, #284); after it, the same apply is refused with the mismatch named, and a wrong-zone create is refused instead of silently served as the deployment's zone (TestAnUnservedZoneSignpostNamesTheMismatch). - on any zone but
ch-dk-2,exo compute instance-type listandshowfail withfind zone: "ch-dk-2" not found, and that is the client, measured at its source: exo 1.95.1 passes its compiled default to the v3 zone switch (cmd/compute/instance_type/instance_type_list.go:85, withcmd/common.go:15fixingDefaultZone = "ch-dk-2"); no flag, config or env reaches it. Invisible against the real cloud, whose zone list always namesch-dk-2. Every other command of the driven flows takes--zoneor honours the account default;tools/conformance/exoscale/zones.shpins the exact failure so the day exo heals it, the suite says so.
GET /v2/template?visibility=private used to answer the public catalogue,
each entry declaring "visibility": "public" inside a response filtered to
private (#271). The cause was a handler that discarded its request, so no
query parameter could have an effect — and the same signature sat on four more
Exoscale operations whose contract declares filters. All five now read what
their operation declares: visibility and family on list-templates,
visibility on list-security-groups, manager-id, manager-type and
ip-address on list-instances, instance-id on list-block-storage-volumes,
and list-events validates its from/to window (over an audit trail that is
empty by design, see above). A gate holds the rule from here on:
TestDeclaredQueryParametersAreRead fails any route whose contract declares
query parameters while its handler never reads the query.
Two answers differ from the real cloud, and both are the honest half of a choice rather than an accident:
labelson list-instances is refused with a 400, not implemented. Their document types it as a bare string with no format and no description, egoscale v3 exposes no option that sends it, so any wire encoding this emulator picked would be an invented format — the thing rule 4 exists to forbid. A client that sends it learns so at the moment it happens; the real cloud would filter.TestInstanceListRefusesTheLabelsFilterholds the refusal.?visibility=publicon list-security-groups answers an empty list. The real cloud publishes public security groups of its own; this emulator publishes none, and listing the private groups under a public label would be the same lie #271 names, pointed the other way.
The SDK that Terraform provider 2.82.0 and later embed asks the gateway for
its metadata, and uses the domain it answers to compute a srn://… client-side
for every product that gained one. This emulator mounted no such route until
2026-09-15, and answered 404: measured through feint proxy --record, one
apply plus destroy of the conformance fixture sent 148 of them, which
was 148 of that run's 169 refusals.
Nothing failed, and that was luck rather than a decision: the callers discard the error, so the SRN stayed empty rather than wrong.
What is served now, and what it costs. Three constants:
GET /metadata -> 200
{"platform": "external", "partition": "scw", "domain": "scw.eu"}
They are the real cloud's own answer, read on a read-only shot at an fr-par
account on 2026-09-14 and again on 2026-09-15, and committed as
corpus/scaleway/scw-gateway.jsonl. They are constants, and that is the
limit: the real gateway describes the platform a caller reached, so an
account on a different platform or partition would be told this one. Every
account this project can reach answers these three, and an emulator serving one
platform cannot honestly vary them — but a reader building a multi-platform
expectation on this route would be building it on a fixture.
The corpus grades the shape rather than the values, because a committed corpus
is sanitised: it keeps the status, the field tree and the types, and replaces
every value with a synthetic one. The values are held by
tools/contract/scaleway-gateway.yaml and by a test that writes them out.
#271's gate catches a handler that never reads its query at all. Its comment
names what it cannot see — a handler that reads some declared parameters and
drops the rest — and #277 measured that residual class on this pack: every list
read its page, so the gate was green, while ?order_by=created_at_desc
answered ascending. The per-operation gate
(TestNoDeclaredQueryParameterIsDroppedByItsHandler) now requires every
declared parameter to be named in its own handler's call graph, and everything
it found is served, with these decisions where the emulator's model and the
real cloud part ways:
- Every
order_by(and instance/v1'sorder) is served with the SDK's documented default. A bare list answerscreated_at_ascwhere block, vpc and iam declare it,created_at_descwhere instance and ipam do — including a barescw instance server list, which now answers newest first like the real API. A value outside the operation's enum is a 400 naming the parameter, never a silently different order. order_by=attached_at_*on ipam ListIPs is refused with a 400. This emulator records no attachment time, and sorting by a stand-in would answer an order nobody asked for.include_deletedon block lists is read and never widens the answer. Deletion here is immediate — the store retains nodeletedvolume or snapshot for the flag to reveal. The state filter is real code on the default path; the difference from the real cloud, which retains deleted volumes for a while, is this line.object_storage_private_access_enabled=truematches nothing, and so doess3_integration_enabled=true, which is the name Scaleway retired on 2026-08-25 and this filter still accepts (#570): a client that has not been rebuilt keeps sending it. No VPC or Private Network here integrates with Object Storage, which is not emulated (see above), so true truthfully answers an empty list and false answers everything. The body answers the new name alone — both upstream sources declare only that one, and inventing a deprecation alias in an answer is not the same decision as tolerating a name a client still sends in a request.archandtypeon marketplace ListLocalImages are equalities against the one published image.arch=arm64ortype=instance_localanswers an empty list — the catalogue is x86_64 andinstance_sbs— where dropping the filter answered an image of the wrong architecture with a 200.without_ip=falseon ListServers filters nothing. The SDK documents only the true direction ("list Instances that are not attached to a public IP"); a complementary meaning for false would be invented.disabledon iam ListSSHKeys is an equality on the field when present. The SDK comment ("defines whether to include disabled SSH keys") could also read as a widener over a default exclusion, but nothing upstream states that default, and a bare list here has always answered every key.organization/organization_idfilters scope to the whole account. One organization lives here and identifiers are unchecked by decision (see above), so a named organization resolves to that account —scopeOf's long-standing rule, now applied to every list that declares the parameter. The equality reading was tried first and the CLI refuted it within the hour:scw iam ssh-key listnames its configured organization on every call, and nothing obliges a client's configuration to spell the emulator's constant, so comparing answered "no keys" about keys the same client had just created.project/project_idstay real filters: a project is an isolation boundary this emulator does honour.
TestServersHonourTheDeclaredOrder, TestServersFilterByLinks,
TestBlockListsHonourTheDeclaredFilters, TestVPCListsHonourTheDeclaredFilters,
TestIPAMListHonoursOrderAndResourceFilters,
TestInstanceListsHonourTheirRemainingFilters,
TestSSHKeysHonourTheDeclaredParameters and
TestLocalImagesHonourTheirDeclaredParameters ask with non-default values —
the whole class survived every suite that only asked for defaults — and
tools/falsify/specs/scaleway-list-parameters.json proves each family's test
red without its fix.
Outscale's Tag carries a ResourceType, and their OpenAPI declares it as a
bare string. The values come from the SDK instead, where they are a deliberate
patch: TagResourceType in osc-sdk-go/pkg/osc/client.gen.go, twenty of them,
listed in the enum block of patch.yaml.
An internet service is not among the twenty — and their InternetService schema
declares Tags, which the Terraform provider sets. Both are true at once, and
they cannot both be honoured in the flat view.
So the emulator splits the two questions rather than picking one:
CreateTagsandDeleteTagsaccept it. This is what #99 was: the pack answeredthe resource igw-… does not existabout a resource it was serving, and anoutscale_internet_servicewith atagsblock failed its apply.ReadTagsleaves it out. Every row of that view carries aResourceType, and there is no value upstream declares for this one. Inventinginternet-gatewaybecause AWS spells it that way is the invented format rule 4 forbids, and it would be indistinguishable from a measured value to anyone reading the answer.
The tag is not lost: ReadInternetServices returns it on the resource, which is
where the provider reads it back. TestReadTagsOmitsAKindUpstreamDoesNotName
pins both halves.
This is one row of taggable in internal/providers/outscale/tags.go, the only
one with an empty type. If Outscale adds a value, that field is where it goes.
InternetService, LinkInternetService, NatService and the routes that
name them are served, and none of them makes a packet flow.
LinkPublicIp is no longer on that list: a linked address is routed to the
Vm's machine, answers from the host that runs the emulator, and
ssh outscale@<PublicIp> opens a shell — the outscale ssh conformance suite
drives exactly that. The limit was real while the machines sat on the
operator's default bridge, which the driver rightly refuses to route through;
they boot on emulator-owned networks now.
The rest is structural rather than unbuilt, and the difference matters because
everything around it was buildable and got built. The emulator has no data
plane beyond that host: a NAT service is a managed appliance in a facility
this machine is not in. A public address allocated here comes from
198.51.100.0/28 — TEST-NET-2, reserved by RFC 5737 and routed nowhere on
purpose, so beyond this host an address goes visibly nowhere rather than
quietly somewhere.
Fourteen addresses, and that is a measurement rather than a limit of the
model. The block was a /24 until 2026-08-25. What the emulated pool has to
hold is that allocation refuses past the last address with a typed 9029 and
that a released address returns to it — neither claim depends on how many
addresses there are. Exhausting a /24 cost the Outscale conformance suite 254
DeletePublicIp calls, and at roughly 700 ms of client process startup each
that was the larger half of an eleven-minute workflow. Measured across every
suite, fixture and example stack in this repository, the peak of addresses held
at once is two. ReadPublicIpRanges publishes the block the allocator serves,
derived from the same constant rather than written twice, and
TestTheAllocatorStopsWhereTheCatalogueSaysItDoes fails if either side is
edited alone. A project that needs more can widen the prefix in
publicips.go; nothing else has to move with it.
What is real is the resource algebra, and it is what a plan actually depends
on: an address a NAT service holds refuses to be released, a gateway refuses to
be deleted while linked, a Net refuses to go while a gateway is attached to it,
a route through a gateway that is not linked to the Net is refused, and a
subnet refuses to vanish under a NAT service placed in it. terraform apply
of the provider's own examples/net_vm, its second plan and its destroy all
pass against this, which is exactly what those refusals are for.
So: a plan that builds a routable topology applies, reads back and destroys correctly. A machine inside it still cannot reach the internet. Use the emulator to test the shape of your infrastructure, never its connectivity.
The LBU family is served as far as the surveyed stacks exercise it (#281): create, read, update (health check, security groups, secured cookies), register/link and unlink backend Vms, delete. Three of the five surveyed Outscale stacks (#262) stand on exactly that lifecycle, and all three apply, re-plan empty and destroy against it.
Since #344 the listeners can also be moved after the create
(CreateLoadBalancerListeners, DeleteLoadBalancerListeners), which is a
second-apply operation and never a first one: CreateLoadBalancer carries its
listeners inline. Providers 1.1.3, 1.7.0 and 1.8.0 all call the pair from their
Update path and from nowhere else, and all three delete the departing front port
before creating the arriving one, so a single-listener port change really does
pass through a balancer holding no listener at all. That transient state is
allowed rather than refused, and the runtime follows it: a balancer with no
listener left is withdrawn from the host instead of going on distributing on a
port the API has stopped listing.
What a 200 from CreateLoadBalancer means here, stated rather than implied:
-
The configuration is recorded and round-trips. Listeners, backends, health-check settings, tags, the security groups and the subnet come back field for field, and the delete guards hold the same algebra as the rest of the pack: a subnet or a security group under a balancer refuses to go.
-
Its own private address distributes, under
--vm incus-ovnand there only. A balancer'sPrivateIpis an address of the Subnet it sits in, and it is handed toincus network load-balanceron the Subnet's own network: connections from inside that network are spread over the registered Vms, an unlinked Vm stops receiving them, and deleting the balancer takes it off the host. Measured on 2026-08-20 with three machines on one OVN network — two backends answering their own name on:80, one client — 6/6 answered at t0 and 6/6 again a minute later, over both backends each time, and 6/6 to the survivor after an unlink.tools/conformance/outscale/balancer.shis that measurement, replayed on demand.The claim is declared, never deduced, and it has two halves (#481):
capabilities.balancingsays the runtime can distribute — the OVN mode alone sets it, startup verification clears it on a host with no OVN wiring — andenforced.balancingsays which packs hand their balancers to it, exactly asenforced.firewalldoes one capability over. A build that does not know a key answers nothing, which reads as absent. A suite that wants to assert distribution gates on the conjunction and never on a mode name:$ curl -s localhost:4599/_feint/health | jq '{capabilities: .capabilities.balancing, enforced: .enforced.balancing}' { "capabilities": true, "enforced": ["outscale"] }
The two halves exist because each has been measured true while the other was false: under
incus-ovnthe capability was true on a process whose Scaleway stack held onelb/lbin the API and zero balancers on the host (#481, 2026-08-25), andtools/conformance/outscale/balancer.sh, keyed on the capability alone, would have been correct for Outscale and wrong word for word if copied for Scaleway. -
The public face distributes nothing, and that is a measurement too. The
DnsNamefollows the measured format (<name>-<digits>.<region>.lbu.outscale.com,internal-prefix included) and resolves nowhere; the public address of an internet-facing balancer comes from203.0.113.0/24— TEST-NET-3, RFC 5737, routed nowhere on purpose, and deliberately not the block ReadPublicIps allocates from, because the real service associates an address the account does not own.The reason it stays that way is on the record (#315, measured 2026-08-19). A VIP outside the network, delegated through the uplink's
ipv4.routes, answered 6/6 probes at t0, 6/6 at t+60s, and 0/6 from t+180s onwards, permanently,ip neigh flushincluded: the runtime announces such an address with a burst of gratuitous ARPs at creation time and never again — the same defectinternal/core/machine/incus_ovn.gorecords for network forwards. The driver'sipv4.routes.externalpath holds indefinitely and is strictly one-machine, so it cannot carry a multi-backend VIP either.So
EnsureBalancerrefuses a listen address outside the network's own block rather than configuring one. A balancer that passes a test and fails three minutes later is worse than a balancer that was never claimed. -
A backend on another subnet is withheld; the ones on the balancer's own subnet are distributed (#457 measured the boundary on Incus 7.2 with OVN, 2026-08-25, re-measured 2026-08-26; #483 measured what refusing the whole spec over it cost). The two-tier architecture — a load balancer on the public subnet, machines spread over others — is the ordinary shape, and the runtime cannot take it whole: the driver distributes to the backends inside the balancer's own block, withholds the others by name before the write, and the pack records both halves on the balancer's
Runtime(balancer-distributed,balancer-undistributed, readable through/_feint/state) beside one WARN. The200stands andReadLoadBalancersgoes on describing every registeredBackendVmId, because that is the configuration the client asked for and the real cloud would honour it whole.Why partial rather than none, decided rather than implied: the first fix refused the spec entire, and #483 measured the result — the host held a registered balancer with no backend and no port while the API described two healthy ones, apply exit 0, zero ERROR lines. A balancer that distributes to the machines it can reach, with the withheld ones named in the record and the log, leaves a witness on the host and a trace a gate can read; a balancer that distributes to nobody is indistinguishable from one never handed over.
The four measurements, because the shape is common enough that "it should be possible" deserves an answer rather than an opinion. On two OVN networks of the emulator's own making,
10.181.0.0/24and10.181.4.0/24:- a balancer on the first with a backend in the second is refused outright,
Target address is not within the network subnet for backend "b1-80"; - peering the two networks — both halves
CREATED— does not relax it: the same refusal, word for word; - putting the balancer on the backends' network instead, which is the
placement that would serve the shape, is refused on the other end,
Load balancer listen address "10.181.0.5/32" overlaps with another network or NIC; - and there is no key to declare the address with: an OVN network answers
Invalid option for network "fnt-mxb" option "ipv4.routes", since only a NIC carriesipv4.routes.externaland a VIP has no NIC.
The one placement the runtime accepts is a listen address belonging to no emulated block at all, delegated through the uplink — which is exactly the address class of the paragraph above, the one that goes dark in minutes, and not the address the API published either.
What full delivery would take, named so nobody re-derives it: either Incus lifting the same-subnet restriction on
network load-balancertargets — the OVN load balancer underneath can reach any address its logical router routes to, and the peered route exists; the boundary is the daemon's own validation,Target address is not within the network subnet— or this emulator mapping every subnet of a Net onto one runtime network, which Incus's one-ipv4.address-per-network model does not offer. Until one of those moves, the split above is the whole truth the host can carry.The withheld half is a limit, not an incident, so the pack says it at WARN rather than ERROR: it holds for the life of the stack, and an error that is permanent on a working configuration teaches people to skip errors. The level is bearable because the state carries the fact — the record is what a gate reads, the WARN is for the human watching the boot.
capabilities.balancingcovers the balancer whose backends are on its own network and says nothing about this one; what withdraws the claim entirely is the host refusing a write the driver had accepted. - a balancer on the first with a backend in the second is refused outright,
-
No backend health exists, and none is invented.
ReadVmsHealthstays declined, and it stays declined after the dataplane landed:incus network load-balancer infoanswers "No load-balancer health information available", so nothing probes a backend even under OVN, and a backend reportedUPthat nothing checked is the exact answer this project exists to refuse. The health-check settings round-trip because Terraform plans on them; the health states do not exist until something measures them. -
The stored health-check defaults are the vendor's own (interval 30, timeout 5, unhealthy 2, healthy 10, TCP on the first listener's backend port — the defaults Outscale's user guide documents), so a stack that never touches
outscale_load_balancer_attributesreads back what a fresh real balancer would say.
What is not served, by name and on purpose: the public-Cloud form
(SubregionNames without a Net — no surveyed stack takes it, and nothing is
measured about what it answers), access-log enablement (there is no OOS
bucket here to publish into), listener policies, listener rules, LBU tag
CRUD, and server certificates. Each answers a refusal naming the line, never
a silent 200.
The next wall a day-2 edit meets is the LBU tag CRUD, and it is named rather
than left to be found. Measured on 2026-08-21 with provider 1.8.0: changing a
tags block on an existing outscale_load_balancer answers Error: Unable to update Load Balancer carrying feint does not serve DeleteLoadBalancerTags,
and the plan then stays at 0 to add, 1 to change, 0 to destroy. It is out of
#344's scope because that issue served the path carrying traffic and a tag
reaches no runtime; the demand for it is now written down rather than guessed.
DeregisterVmsInLoadBalancer is refused on reachability, not on demand.
Provider 1.1.3 is the only version whose code contains the call, on the update
path of the load balancer's own backend_vm_ids, and that path panics before a
request is built: the attribute is declared schema.TypeList
(resource_outscale_load_balancer.go:150) and the update casts it to
*schema.Set (:726). Measured 2026-08-21 against this emulator —
interface conversion: interface {} is []interface {}, not *schema.Set, an
upstream defect rather than an emulator one. Providers 1.7.0 and 1.8.0 removed
the call; detaching a backend goes through UnlinkLoadBalancerBackendMachines,
which is served.
One choice here is not a measurement, and says so. Naming a front port that
carries no listener in DeleteLoadBalancerListeners is accepted rather than
refused: nothing here has watched a real account answer that request, and what
the caller asks for — that these ports carry no listener afterwards — is already
true of a port that carried none. A refusal would have been just as much of a
guess, and a riskier one, since it would break a client that retried a
half-applied update.
The lb/v1 ZonedAPI and vpc-gw/v2 families are served as far as the measured
clients exercise them (#282): the surveyed kubic and terraform-talos stacks,
Scaleway's own LB and VPC modules, scw lb and scw vpc-gw. The dataplane
#315 built for the Outscale LBU has not been wired here: the mechanism is the
same and the pack is not, so the honest statement is that a Scaleway balancer
still records its configuration and forwards nothing. capabilities.balancing
says what the runtime can do, never what a given pack asked it for, and this
pack asks for nothing — which is why it does not appear in
enforced.balancing (#481). That absence is the declaration: a suite gating on
the conjunction of the two halves skips here instead of asserting a
distribution this pack never promised, and the measurement behind the sentence
is on the record — under incus-ovn on 2026-08-25, a Scaleway stack held one
lb/lb, one frontend, one backend and one route in the API, and a sweep of
every managed network on the host found zero load balancers.
What a 200 means here, stated rather than implied:
- The configuration is recorded and round-trips. The balancer, its backends (pools, millisecond timeouts, the one-of-seven health-check config), frontends, inline ACLs and routes come back field for field; so do the gateway, its IP and the GatewayNetwork. The wrong destroy order gets a refusal — a backend under a frontend, a gateway under a connection — never a silent success.
- No traffic is forwarded. A balancer's IPv4 comes from
198.51.100.0/24(TEST-NET-2), a gateway's from192.0.2.0/24(TEST-NET-1) — RFC 5737, routed nowhere on purpose, distinct from the instance flexible block so no two products ever publish the same address. The gateway NATs nothing, pushes no route into a machine, and its bastion accepts no connection, which is why the bastion allow-list operations are declined rather than recorded. - No backend health exists, and none is invented.
GetLBStatsandListBackendStatsstay declined: nothing probes a backend, and a backend reportedUPthat nothing checked is the exact answer this project exists to refuse. The health-check settings round-trip because Terraform plans on them; the health states do not exist until something measures them. - Both attachment spellings are served. SDK generations up to
v1.0.0-beta.29 attach a Private Network at
/lbs/{id}/private-networks/{pnID}/attach; the current one says/lbs/{id}/attach-private-network. terraform-provider-scaleway v2.43 — the pin of a surveyed stack — sends the old one, production still accepts it, and the emulator does too (Route.Legacycarries the measurement). - The attachment's address is a first-class IPAM citizen. An attach
without
ipam_idsbooks an address from the Private Network's own pool; a GatewayNetwork does the same, or holds theipam_ip_idthe client booked first. Both read back through/ipam/v1filtered byresource_type=lb_serverorvpc_gateway_network, which is exactly how the Terraform provider resolves them.
Only vpc-gw v2 is served. The portal publishes no v1 document any more
(measured 2026-08-19) and every mounted route here is checked against the
portal's document, so v1 is declined wholesale, by name: a provider pinned
below 2.52 — terraform-talos's ~> 2.43.0 among them — meets a named 501 on
/vpc-gw/v1/..., and the recorded fix is the provider bump to ≥ 2.52, the
release that moved the product onto v2. Also declined by name: MigrateLB and
UpgradeGateway (a capacity move nothing performs), certificates (nothing
terminates TLS), subscribers (no event to deliver), the PAT rules (a rule
recorded and never applied is indistinguishable from protection), and both
type catalogues (ListLBTypes, ListGatewayTypes — unmeasured inventory; the
gateway create still refuses an offer outside VPC-GW-S/M/L/XL, the #279
lesson).
An Exoscale network load balancer records its configuration, names its backends, and grades none of them
The NLB family is served whole since #345 — the balancer, its services, and the two per-field resets — after a year declined by #14. What #14 refused is what this section exists to keep refused, and it is worth stating in the same breath as what is now served.
The configuration round-trips, and a real client converges on it. Name,
description, labels, and per service the protocol, the ports, the strategy, the
instance pool and the whole healthcheck block come back field for field. The
example stack under examples/stacks/exoscale/ applies with an exoscale_nlb
and an exoscale_nlb_service, re-plans empty, and destroys clean (15 resources,
measured 2026-08-21 with the patched provider this document pins).
The health of a backend is not measured here, and none is invented. A
service publishes healthcheck-status, and every entry it publishes carries the
backend's public-ip and no status. That is upstream's own shape rather
than a compromise: the element schema load-balancer-server-status declares no
required property, so an entry naming a server with no verdict on it is
well-formed. The official CLI prints it as
{"instance_ip":"192.0.2.2","status":""}, which is the honest sentence — these
are the servers behind the service, and nothing graded them.
The two alternatives were both worse, and both were considered rather than
skipped. Publishing success is the fabrication #14 declined the family over.
Publishing an empty array reads as this service has no backend, which is a
claim about the pool, false for every pool this emulator holds, and one a client
could plan against.
What is not measured about it, said plainly. No recording of a live NLB
exists here: shapes/exoscale.json carries GET /v2/load-balancer from an
account that held none, so it pins the envelope key and nothing inside it.
What is measured is that their published document allows an entry with no
status, and that the official CLI and the Terraform provider both accept one.
Whether the real API ever omits the field on a service it is actually probing is
a question a recording would settle and nothing here answers.
Nothing forwards a packet, and the reason is an address rather than a
missing feature. The Outscale LBU distributes real connections under
--vm incus-ovn because its PrivateIp is an address of the Subnet it sits in,
and machine.EnsureBalancer accepts exactly that. An Exoscale NLB publishes one
address, ip, and their load-balancer schema declares no other — no subnet,
no private network, no counterpart to PrivateIp. That single address comes
from 192.0.2.0/24 (TEST-NET-1, RFC 5737), which is outside every emulated
network's own block.
Measured on 2026-08-21 against a live incus-ovn host, on an OVN network of
this emulator's own making (10.63.7.0/24):
EnsureBalancerwith192.0.2.1— the address this pack gives an NLB — answers "balancer … listens on 192.0.2.1, which is outside … 's own block 10.63.7.0/24: an address the runtime has to announce goes dark within minutes (#315)";- the same call with
10.63.7.240and one backend answers<nil>, so the refusal is about the address and not about the call; - and the daemon itself refuses the public address before any guard of ours is
consulted: Failed creating load balancer: Uplink network doesn't contain
"192.0.2.1/32"in its routes.
Those three lines are quoted as they were printed that day. The refusal has
since gained the sentinel #457 added (machine.ErrBalancerNotDistributed), so
its wording differs; the verdict does not.
So capabilities.balancing is irrelevant to this family: the pack never asks
the runtime at all, because the only call it could make is one whose refusal is
guaranteed. The runtime's balancing half needed no provider-shaped concession to reach
that answer, and internal/core gained no Exoscale knowledge — what is missing
is an address upstream does not publish, and no field of an interface can supply
one.
Two client facts that are the pack's and not the API's, recorded because they surprised.
- A service mutation's operation refers to the balancer, never to the
service. egoscale v2 passes that reference straight to
GetNetworkLoadBalancerand finds the new service by diffing the balancer's list (v2/network_load_balancer_service.go:121at v0.102.4), so referring to the service makesterraform applyfail withGet …/v2/load-balancer/<service id>: resource not found. Measured; the exo CLI cannot arbitrate it, because it resolves every object by listing and never reads a reference. - This CLI clears no field of this family. Every other family here records
that the CLI clears a field by sending the update with an empty value; on the
NLB it does not.
exo compute load-balancer update --description ""sendsPUT {}, and the service form sends only the healthcheck block it re-sends on every call. The per-field DELETEs are served because their document declares them, and no published client issues one.
What a service's backends are. The members of the instance pool it targets,
which is where upstream takes them from too — a service names a pool, never a
list of machines. Pool members carry a public address since #345 (their pool's
public-ip-assignment decides, inet4 by default); before that they carried
none, and every service in front of a pool answered an empty backend list.
BootMode, Performance and VmInitiatedShutdownBehavior used to be
accepted at create with a 200 while every read answered a constant of the
pack — the client asked medium/restart/legacy, the same create's answer
said high/stop/uefi, and a Terraform stack setting any of them
re-planned the same in-place change for ever (#276, the #268 pattern on
per-machine scalars). They are stored and restituted now, on the create and
on UpdateVm where upstream declares them, values validated against their
enums, and Performance honours upstream's own precedence: a performance
flag inside the VmType (tinavW.cXrYpZ) wins over the parameter. The same
sweep covers the neighbours with the same symptom: TpmEnabled,
ActionsOnNextBoot.SecureBoot, ShutdownBehaviorConfiguration (whose
defaults are now the SDK's own "By default" lines — GuestAction stop,
HostAction restart; the old constant said stop/stop) and UpdateVm's
IsSourceDestChecked.
What is served is the datum. The behavioural half of these fields has nothing to act on in this emulator, and saying so is the difference between an echo and a lie:
VmInitiatedShutdownBehaviorandShutdownBehaviorConfigurationdescribe what the platform does when the guest shuts itself down. No path here watches a guest-initiated shutdown —StopVmsstops the machine whatever the field says, which matches upstream, where the API stop is not a VM-initiated one. Aterminatebehaviour will therefore never terminate a machine here, because the event that would trigger it is never observed.TpmEnabledandActionsOnNextBoot.SecureBootround-trip; no vTPM and no secure-boot state is presented to any guest the runtime boots.IsSourceDestCheckedround-trips; nothing enforces the check on traffic.BsuOptimizedstays the constantfalseon every read, and that one is upstream's own behaviour, not this pack's shortcut: "This parameter is not available. It is present in our API for the sake of historical compatibility with AWS" (osc-sdk-go client.gen.go:3029).
The NetPeering family is served — create, accept, reject, delete, read, with
the SDK's own states (pending-acceptance, active, rejected, failed,
deleted) and its refusals (accepting or rejecting anything but a pending
one, deleting a rejected or failed one). Three of the upstream behaviours
cannot exist here, and each is stated rather than approximated:
- One account. Upstream, the owner of the accepter Net accepts, and a
pending peering is deletable only by the requester. The emulator's single
account owns both ends of every peering, so the identity rules are satisfied
by construction and only the state machine is measurable. An
AccepterOwnerIdnaming any other account is answered as an unknown Net, because in a one-account world that is what it is. expiredis unreachable. Upstream it is what seven days of silence produce; no clock here advances a state on its own.failedhas one reachable door. Upstream it is what overlapping IP ranges produce, but this emulator refuses to create two overlapping Nets in the first place — every Net backs a real block on the host — so the only overlap left is a Net peered with itself.
What an accepted peering does depends on the runtime mode, same rule as "Subnet isolation depends on the runtime mode" above:
- Under
--vm incus-ovn, accepting the peering peers the backing networks of the two Nets the runtime's own way (network peer), and deleting it separates them again. The outscale network suite asserts the whole cycle — unreachable before, unreachable while pending, reachable once active, unreachable after delete — gated on the declaredcapabilities.isolation, never on a mode name. - Under
--vm incus(bridges), two Nets already reach each other, so an accepted peering grants nothing a measurement could see, and the suite skips and says so. - With
--vm off, the whole lifecycle is control plane, proven byoctland the Terraform provider.
One deliberate simplification either way: upstream, traffic flows only once both Nets' route tables carry a route through the peering. Here the acceptance alone grants reachability, and the route stays a record — the same limit as "Outscale's gateways and NAT move records, not packets" one section up. Do not use the emulator to prove a peering's routing configuration; use it to prove the plan's shape and the state machine.
Scaleway's custom routes (scaleway_vpc_route, vpc/v2/API.CreateRoute and
its family) are served since SW-4, and they are records: create, read, update
and delete round-trip for the Terraform provider, the nexthop is validated to
exist, and no runtime mode programs it. A route sending 192.168.42.0/24
through an instance NIC is a row the client reads back, not a path packets
follow.
What is real is what a routing-enabled VPC delivers between its own Private
Networks, and it does not come from these records: EnableRouting reconciles
the machine driver's isolation the moment it flips — under OVN the VPC's
networks are peered, in bridge mode their mutual reject rules are lifted — so
two networks of a routing VPC reach each other and two networks of two VPCs
still do not, mode permitting (see "Subnet isolation depends on the runtime
mode"). TestEnableRoutingReconcilesThePeering in
internal/providers/scaleway pins that link to the driver.
Same rule as Outscale's gateways, one paragraph up: use the routes to test the shape of a plan, never to test where traffic goes. The day a nexthop is worth programming for real, it will arrive as a declared driver capability, measured under OVN, not as a silent upgrade of these records.
vpc/v2/API.GetACL and vpc/v2/API.SetACL are served since #343, and they are
records in the exact sense the section above gives a custom route: the whole
rule set round-trips for scw vpc rule get/set and for scaleway_vpc_acl, the
protocols and actions are held to the SDK's own enums, the sources and
destinations are parsed as CIDRs, and no runtime mode programs a filter at the
VPC edge. A rule dropping 0.0.0.0/0 here closes nothing.
They were declined until #343, under a reason worth repeating because half of it still stands: "a filter recorded but never applied is indistinguishable from protection". What changed is not the enforcement — it is who is told. A 501 stopped the client that was only ever going to read its rules back, and said nothing to anyone about enforcement; this page says it, in the place a reader looks for it, and the pack's own file repeats it. The refusal protected nobody and cost every stack that declares the resource.
What decided the split was a measurement rather than the SDK's shape
(2026-08-21, recorded through feint proxy --record and ranked with
feint coverage --observed):
| operation | what called it | verdict |
|---|---|---|
vpc/v2/API.GetACL, SetACL |
scw vpc rule get/set, and the official provider's scaleway_vpc_acl |
served, as records |
the five *IngressRule |
nothing: scw has no ingress-rule subcommand, no surveyed stack names scaleway_vpc_ingress_rule |
still declined |
the five *VPCConnector |
scw vpc vpc-connector list/create, both recorded taking a 501 |
still declined, and the demand is real: peering two VPCs is the one property the bridge mode cannot deliver |
The last row is the one to read twice. A recorded call is not on its own a reason to serve: the connectors are declined despite the demand, because answering them would report done the very thing this project measures the absence of. Demand decides what is worth serving; it never decides what can be served honestly.
Use the ACL to test the shape of a plan and the round-trip of a rule set. Never use it to test whether a packet is dropped. The day a rule is worth programming for real it arrives as a declared driver capability, measured under OVN, not as a silent upgrade of these records.
Until #475 was fixed, internal/providers/scaleway/firewall.go was the only
file in any pack that referenced machine.Firewaller: an Outscale or
Exoscale security group was served as a control plane, echoed back, and
reconciled onto nothing — with a runtime configured, every port of the machine
stayed open whatever the group said, and the API answered success on every
rule. All three packs now hand their rules over through the shared layer
(machine.Binding), and the witness is observable on the host: incus network acl list carries scw-*, osc-* and exo-* sets marked feint security group, attached to the machines that wear the groups — on the interfaces each
provider says its groups cover, which is not every interface (#574, below;
measured 2026-08-26 on
examples/stacks/scaleway — three scw-* sets, six interfaces between them —
examples/stacks/outscale and examples/stacks/exoscale, each under
feint up --runtime incus-ovn).
The history matters because the claim drifted once before. That gap was
published as the opposite until #180: (*Incus).Capabilities() declares
firewall: true in every mode, one set for the whole process, and this
repository tells a consumer to key on the capability rather than on a mode
name. Following that advice, a user probed a port a deny-default group should
have closed and found it answering.
/_feint/health carries both halves, and the honest check is their
conjunction:
$ curl -s localhost:4599/_feint/health | jq '{capabilities: .capabilities.firewall, enforced: .enforced.firewall}'
{
"capabilities": true,
"enforced": ["exoscale", "outscale", "scaleway"]
}capabilities.firewall is what the runtime can do. enforced.firewall is who
asks it to. A pack absent from that list either does not wire it or has not
said, and a consumer cannot tell those apart — which is intended, because both
mean the same thing to whoever is about to open a socket.
The two bounds, both measured. First, an interface the runtime declares
unenforceable stays unenforceable: a routed NIC accepts no security option
(#337, capabilities.firewall_public_only: false), and that is the primary
interface of every Exoscale instance and of every Scaleway server whose only
address is public. The pack hands the set over, the driver refuses with the
typed error, and the log names the declaring capability instead of crying
wolf. The capability is about the interface, not about a machine with only
one — and machine.Capabilities.FirewallPublicOnly says so in its own words
since #548, because a declaration whose subject is wrong reads like proof.
Until 2026-08-28 that list had a third member, and it is the one #548 was
filed on: a Scaleway server created with its address, whose private NIC
arrives afterwards and did not take the address with it (measured 2026-08-27
on examples/stacks/scaleway: eth0 routed and bare beside eth1 on a
managed network carrying the rule set). That machine is covered now — the
driver moves the address onto the filtered NIC — and the paragraph after the
next one carries the before-and-after in both driver modes.
Reproduced from the API alone on 2026-08-27, without the stack, under
--vm incus-ovn: a group whose inbound default is drop with one rule
allowing 443, a server created with its flexible IP, its private NIC attached
afterwards. The escape and its negative control are in the same probe.
$ incus query /1.0/instances/feint-scw-d3eaa40c-… | jq -c '.expanded_devices | …'
{"eth0":{"ipv4.address":"203.0.113.2","nictype":"routed","type":"nic"},
"eth1":{"ipv4.address":"10.181.7.2","network":"fnt-c9b63dbeff0","security.acls":"scw-3bab95b997b","type":"nic"}}
203.0.113.2:22 connect_ex=0 OPEN # no rule opens 22: the group is not on eth0
203.0.113.2:80 connect_ex=111 refused # reached the machine, nothing listening
10.181.7.2:80 connect_ex=113 no route # the private side, where the policy holdsconnect_ex=111 on a port the group never opened is the decisive line: the
packet reached the guest and was refused by it, where a covered interface
would have dropped it.
Two migrations were tried and refused, and a third one works — and ships.
#548 left one thing untried — whether the driver could move the address onto
the managed NIC once that one arrives, which is the shape the other creation
order already produces and the one the rule set covers. Three attempts, and
the third is what RouteAddress does today.
The first two, by hand on 2026-08-27:
$ incus network set feint-uplink ipv4.routes "…,203.0.113.2/32"
Error: Failed to add route {… Dst: 203.0.113.2/32 …}: file exists
$ incus config device remove feint-scw-d3eaa40c-… eth0
$ incus query /1.0/instances/feint-scw-d3eaa40c-… | jq -c '…'
{"eth0":{"network":"incusbr0","type":"nic"}, "eth1":{…}}The uplink cannot be given the /32 while the routed NIC still owns the host
route for it — the collision #498 documents, met from the other side — and
removing the routed device unmasks the profile's eth0 on incusbr0, the
operator's own default bridge, which this emulator refuses to put anything on.
Neither is a remedy, and neither should be proposed again.
Both come from one place: the host route. The third attempt takes the address off the device without taking the device off the instance, which is neither of the two, and Incus 7.2 accepts it on a running instance:
$ incus config device set <machine> eth0 ipv4.address= # 1, the host route goes
$ incus network set feint-uplink ipv4.routes "…,203.0.113.2/32" # 2, no longer refused
$ incus config device set <machine> eth1 \
ipv4.routes.external=203.0.113.2/32 # 3, routeAddressOVN's own gestureStep 1 releases the address, step 2 is the first refusal now unblocked, step 3
is the address arriving on the interface that wears the rule sets. The device
stays, so incusbr0 is never unmasked. In the bridge mode the third step is
ipv4.routes on the managed device instead, which is that mode's own way of
carrying a public address, and the first two are unchanged.
Before and after, in both modes, read on the NIC and from the station. The
shape is #548's, reproduced from the API alone: a group whose inbound default
is drop with one rule allowing 443, a server created with its flexible IP,
its private NIC attached afterwards, a listener on 443 and one on 80, and 8080
bare as the negative control. Measured 2026-08-28.
--vm incus-ovn, before --vm incus-ovn, after
eth0 routed ipv4.address=203.0.113.2 eth0 routed (no address)
no security.acls no security.acls, and nothing to cover
eth1 network=fnt-… eth1 network=fnt-…
ipv4.address=10.199.0.2 ipv4.address=10.199.0.2
security.acls=scw-… ipv4.routes.external=203.0.113.2/32
security.acls=scw-…
203.0.113.2:443 OPEN 203.0.113.2:443 OPEN a rule opens it
203.0.113.2:80 OPEN ← the escape 203.0.113.2:80 refused a listener is there, no rule is
203.0.113.2:8080 refused 203.0.113.2:8080 refused no listener: the negative control
--vm incus, before --vm incus, after
eth0 routed ipv4.address=203.0.113.2 eth0 routed (no address)
eth1 network=fnt-… security.acls=scw-… eth1 network=fnt-… ipv4.routes=203.0.113.2/32
security.acls=scw-…
203.0.113.2:443 OPEN 203.0.113.2:443 OPEN
203.0.113.2:80 OPEN ← the escape 203.0.113.2:80 timed out (the bridge default is drop)
203.0.113.2:8080 refused 203.0.113.2:8080 timed outTwo readings of that table are worth spelling out. The verdict is taken on the
NIC and not only on the connection: security.acls is on the interface
that carries the address in the "after" column, which is what tells this apart
from a connection that happens to fail. And the bridge column's refusals are
timeouts rather than resets, because a bridged NIC's default action drops where
OVN's isolation set rejects — so on that mode the negative control cannot tell
"no listener" from "dropped", and the pair that carries the verdict is 443 open
beside 80 closed with both proved listening inside the machine.
What a restart costs, and what had to be fixed for the move to survive one.
The address now lives on the interface a hot attach created, and nothing inside
a guest remembers what this driver configured on it: measured on 2026-08-28,
a machine rebooted through the API came back with no address at all on eth1,
the driver's own wait ran its ninety seconds and gave up with it carries no IPv4 address, and the published address stopped answering — 443 timed out
where it had been open, in both modes. The cause was named by repairing it by
hand, in three steps, until the symptom went: restoring the interface's private
address alone did not bring 443 back, and a route towards the station's block
did. So the restart path restores the address a NIC device reserves (both
modes) before it waits for a lease nobody offers, and the routes towards the
peered subnets follow as before (#549). Re-measured after that: 443 open, 80
refused, 8080 refused, identical before and after the restart.
What is still not delivered on this shape. The machine has no way out: a
routed NIC has no NAT, its default route points at the link-local host address,
and ping 1.1.1.1 from inside answers nothing — before the migration and after
it, with the station as the control in both passes, so the move neither gave
nor took that away. A guest that needs outbound access wants a default route
through its emulated network, which this driver deliberately does not invent
(see repairGuestInterface: inventing one would route a machine the control
plane declared isolated). The #507 bound therefore stands unchanged.
What #548 delivered before the remedy, and keeps. The refusal a routed NIC
still earns names every escaping interface and the addresses it delivers
(eth0 (203.0.113.2)), read from both ipv4.address and ipv4.routes so an
address attached after the boot is named too, and the warning does not call the
machine public-only. That is what an operator reads on the one shape the
remedy cannot reach — a machine with no emulated network to move the address
onto. TestTheUnenforceableRefusalNamesTheAddressThatEscapes,
TestTheUnenforceableRefusalNamesAnAddressRoutedAfterTheLaunch and
TestTheUnenforceableWarningDoesNotCallTheMachinePublicOnly fail without it,
and tools/falsify/specs/uncovered-interface.json replays all three. A routed
NIC that carries nothing is not named at all, and that half is
TestARoutedNICThatCarriesNothingIsNotAnEscape: after the move the device
stays on the instance with no address, and reporting it would be describing an
escape that had been closed.
Second, between two machines of one subnet the sender's permissive
egress still wins over the receiver's ingress default (the single-pipeline
divergence "Subnet isolation depends on the runtime mode" documents, measured
again on the Exoscale stack's two pool members on 2026-08-26: a member reaches
its neighbour's unopened port, the station does not). The bound this paragraph
used to carry — the isolation set's catch-all allow at 300 defeating every
NIC-level default deny on multi-subnet OVN runs — was #491, and it is gone:
the OVN isolation set carries only rejects now, so a group's default-deny
holds with the isolation attached, measured on the three example stacks under
feint up --runtime incus-ovn (the section "The firewall enforces, within
stated bounds" carries the runs).
Group-sourced rules (the tiering statement, "tier 2 accepts tier 1") are
expanded into the member machines' addresses, since the runtime has no group
selector, and re-expanded whenever a member boots or gains an interface.
Which interfaces a group covers is the provider's answer, not this
emulator's (#574). Exoscale documents the difference in one sentence —
"Security group rules do not apply to traffic inside private networks"
(Private Network Overview)
— and from a344f8d (#494/#475) to 2026-08-27 this emulator applied them
there anyway. The emulated default group carries no ingress rule, so its
rule set translates to a drop default, and the driver wrote it onto the
membership NIC. Measured 2026-08-27, one private network and two instances
created with --public-ip none, whose only interface is therefore the
membership one:
# --vm incus, before # --vm incus, after
eth0: eth0:
ipv4.address: 10.186.0.20 ipv4.address: 10.186.0.20
network: fnt-49f6b5af641 network: fnt-378ef075b46
security.acls: exo-11e594f4819 type: nic
security.acls.default.ingress.action: drop
type: nic
0/10 probes connected 10/10 probes connectedUnder --vm incus-ovn the same wrong rule set was written and did not
bite: the sender's catch-all egress allow at priority 300 outranks the
receiver NIC's default deny at 100/111, the ordering the paragraph above
records as #491. Both runs are in the record because they say different
things — every green this segment produced under OVN rested on that accident,
so the after-state is read off the NIC (incus config device get) and never
off a connection succeeding.
The scope is declared per interface by the pack
(machine.Attachment.Unfiltered) and reaches the driver on
FirewallBinding.Unfiltered; Scaleway and Outscale declare nothing, keep
filtering their private NICs, and their own network suites assert it. Two
consequences for whoever reads the witness at the top of this section. An
Exoscale instance with no public address and only private memberships now
carries no rule set on any interface — it has no interface a security
group covers, which is what upstream describes. And under OVN, a membership
NIC on a network carrying the emulator's isolation set wears the permissive
posture set (opn-fnt…) rather than nothing, because a network-level ACL
forces the reject default onto every NIC attached to it: unfiltered has to
mean open inside the segment, not closed by the network.
Three instruments had to agree to hide this for as long as it lasted, and all
three are repaired in the same change: runtime-proof.yml ran two of the
three network suites and not the Exoscale one, evidence:update's runtime leg
overrode an exported FEINT_VM in silence, and this suite's own header
claimed the pack "does not yet sync its security groups onto the machines" —
true before a344f8d, read as a live fact for as long after.
The station reaches an OVN private address only via the network's router, and the posted uplink routes do not go there (#496)
The runtime posts a scope-link route on the uplink for each OVN subnet, and the
station cannot reach a private address through it. Measured on 2026-08-26
(--vm incus-ovn, Incus 7.2), with zero ACL on the path — network detached from
its isolation set, NIC rule set empty — and the listener proven alive inside
the machine:
$ ip route | grep 10.30.1
10.30.1.0/24 dev feint-uplink proto static scope link
$ python3 probe.py 10.30.1.11 443
10.30.1.11:443 connect_ex=113 CLOSED
$ sudo tcpdump -ni feint-uplink host 10.30.1.11
ARP, Request who-has 10.30.1.11 tell 10.209.83.1 (x5, no answer)One variable changed, everything opens:
$ sudo ip route replace 10.30.1.0/24 via 10.209.83.128 dev feint-uplink
$ python3 probe.py 10.30.1.11 443
10.30.1.11:443 connect_ex=0 OPEN10.209.83.128 is the network's volatile.network.ipv4.address — the OVN
router's own address on the uplink. Reproduced identically on the Exoscale
stack's networks (10.90.1.0/24, 10.90.2.0/24). Outscale is not concerned:
its station probes go through public addresses, whose /32s l2proxy answers
ARP for.
Measured: the routes, the capture, the two flips above. Deduced, not
instrumented: the OVN router answers ARP only for its own addresses and the
l2proxy /32s, which is consistent with the capture — the scope-link route makes
the station ask who-has for the internal address, and nothing on the uplink
answers. One counter-witness is on record and stays on record: #491's probe of
2026-08-26 00:00 read connect_ex=0 from the station on 10.99.1.4, a
station→private measurement that worked at least once in a configuration this
finding does not explain. Both observations are written here as they were made.
Re-measured on 2026-08-27 on fresh networks (10.181.7.0/24, one machine,
zero ACL relevance to the result), with the capture this time rather than
beside it:
$ ip route show 10.181.7.0/24
10.181.7.0/24 dev feint-uplink proto static scope link
$ sudo tcpdump -ni feint-uplink 'host 10.181.7.2 or arp'
ARP, Request who-has 10.181.7.2 tell 10.209.83.1 (x3, no answer)
10.181.7.2:22 connect_ex=11 10.181.7.2:80 connect_ex=113
$ sudo ip route replace 10.181.7.0/24 via 10.209.83.128 dev feint-uplink
10.181.7.2:22 connect_ex=111 10.181.7.2:80 connect_ex=111The flip's new fact is the value it flips to. With the via route the two
ports stop being unreachable (11, 113) and become refused (111) — the
receiving NIC's own rule set answering, since the group under test opened
neither. So the router does forward for the subnet when it is addressed
directly; what fails is the scope-link form's ARP, exactly as the capture
shows.
Why it will not lift here, measured. The earlier wording said the fix was
to post the route via the network's router instead of dev-only, as if the
emulator were choosing the form. It is not: the emulator never writes a host
route. Its only host binary is incus — every ip in internal/core/machine
runs as incus exec <machine> -- ip …, inside a guest — and the host route
above is Incus's own materialisation of the uplink bridge's ipv4.routes,
which an OVN network's subnet must sit inside or the network is refused
("Uplink network doesn't contain … in its routes"). That key takes CIDRs and
nothing else, measured on a scratch bridge on 2026-08-27:
$ incus network set probe496br ipv4.routes=10.223.0.0/24
$ ip route show 10.223.0.0/24
10.223.0.0/24 dev probe496br proto static scope link
$ incus network set probe496br ipv4.routes="10.224.0.0/24 via 10.222.222.2"
Error: Invalid value for network "probe496br" option "ipv4.routes":
Item "10.224.0.0/24 via 10.222.222.2": invalid CIDR addressSo lifting this means the emulator running ip route as root on the operator's
host, which is a different program from the one this repository ships: the
machine driver's whole blast radius is what incus will do for the user who
started it. The remedy stays where it belongs, with whoever wants the path —
one ip route replace … via <volatile.network.ipv4.address> per subnet, which
the block above is the recipe for.
Who this touches: any harness that probes a private address from the station,
firewall proofs first. Machines between themselves are not concerned. The
capability already says it — capabilities.private_from_host is false under
OVN — so a consumer asks /_feint/health instead of probing blind (see "A
public address is the provider's value, made to answer on the host").
A VPC created without enable_routing answered routing_enabled=false where the real cloud answers true (#497, lifted 2026-08-27)
The premise was verified on the real cloud before anything else (real account, 2026-08-26, the test VPC deleted afterwards):
$ scw vpc vpc create name=feint-premise-routing # no enable_routing
RoutingEnabled trueThis emulator stored the request field as-is
(internal/providers/scaleway/vpc.go,
res.Attrs["routing_enabled"] = req.EnableRouting), so the Go zero value
became the default — the inverse of upstream. The Scaleway example stack
created its two VPCs without the field, and both read back false (its
workload VPC has written enable_routing = true out since #503, for this
reason; its management VPC still says nothing):
$ curl -s …/vpc/v2/regions/fr-par/vpcs | jq '.vpcs[] | {name, routing_enabled}'
platform-workload routing_enabled=false
platform-management routing_enabled=falseConsequence measured on the host: the web and app networks of the same
workload VPC are not peered (incus network peer list empty), their isolation
sets reject each other, and app→web:443 reads connect_ex=111 CLOSED while
the web group's rule accepts 0.0.0.0/0. On the real cloud, two Private
Networks of one routed VPC reach each other. Measured: the three blocks above.
Deduced: nothing.
The limit is gone; this section stays as its dated record. The create no
longer stores the Go zero of an absent field: createVPCRequest.EnableRouting
is a pointer, and only a value the client actually sent overwrites the true
newVPC already carried for the lazily provisioned default VPC. An explicit
false is still stored as false, in either direction — the real-cloud
measurement above covers the absent field and nothing else, and inventing an
answer for a field that was sent is the guess this repository refuses.
Re-measured after the fix, same runtime, 2026-08-27:
$ curl -sH 'Content-Type: application/json' \
…/vpc/v2/regions/fr-par/vpcs -d '{"name":"audit-497"}' | jq .routing_enabled
true
$ incus network peer list fnt-e5ba51ab1fa # two Private Networks of that VPC
| fnt-aac94982924 | default/fnt-aac94982924 | local | CREATED |The peering is the half that matters and the one the flag was never only about:
reachableFrom reads it, so a default corrected in the view alone would have
left the host exactly as it was — which is why
TestAVPCCreatedWithoutEnableRoutingRoutes asserts on the peer list and not on
the field, in both directions, and why
tools/falsify/specs/vpc-routing-default.json replays it with the Go zero put
back.
An API reboot used to log Failed to add route: file exists for its own public /32 (#498, lifted 2026-08-27)
The limit is gone; this section stays as its dated record, because this
file dates its past rather than erasing it. As measured on 2026-08-26
(--vm incus-ovn, examples/stacks/scaleway), a reboot — and, as the fix
established, an ordinary poweron — of a server carrying a routed public
address logged one ERROR per replay:
01:06:50 ERROR could not route the public address to the machine address=203.0.113.3
server=e2297f8a-… error="set routes of uplink feint-uplink: incus network:
Error: Failed to add route {… Dst: 203.0.113.3/32 …}: file exists"
over a host that ended up correct — the /32 answered — which made it exactly the noise that teaches ignoring the next true ERROR.
The record above left the fault undecided between two candidates: the stop
half not removing the route, or the re-install not tolerating file exists.
The measurement that closed it (2026-08-27, #498, delivered with the
#547/#549/#498 lifecycle batch, PR #561) answered neither. A
Scaleway server is created before its private NIC exists, so its public
address rides a routed NIC (#202) and the runtime installs the host route with
the device; by the time the private NIC is there, Plan.RouteVia names the
private network, so the replay a poweron or reboot runs asked for the very
same /32 through the OVN path — a collision with a route already delivering,
not a lost route needing re-install. RouteAddress now answers "already
delivered" for an address a routed NIC of the machine carries, and touches
nothing; tolerating file exists instead would have hidden the collision by
writing an uplink route for an address the uplink does not carry.
Re-measured after the fix, same stack and runtime: a reboot and a poweron of
the same server, zero file exists lines, zero ERROR.
TestReRoutingAnAddressARoutedNICAlreadyCarriesTouchesNothing,
TestAnAddressNoRoutedNICCarriesStillTravelsTheOVNPath and
TestAMachineWithNoRoutedNICIsRoutedAsBefore fail without the guard —
replayed green for this record on 2026-08-27 — and
tools/falsify/specs/lifecycle-tells-the-truth.json replays them with the
guard neutralised in both directions, the refusal and the acceptance.
Fixed on 2026-08-28. The refusal below is logged at WARN, and the rule the
measurement suggested is now written where the refusal is
(Binding.refuseUnknownImage) and held by
TestADocumentedRefusalIsAWarningAndAFailureStaysAnError:
An ERROR is something this emulator did not do that it was built to do. A WARN is something it deliberately declines and documents, where the API answer stays honest.
The half the issue had not measured is measured now: of the 48 ERROR sites
under internal/ on 2026-08-28, that one call was the only one on the wrong
side of the line. Its own neighbours are failures and stay ERROR — a start the
driver refused, an image build that could not fetch its source, a pack that
declares no interface plan (plan.go says why that one is not a decline). The
test asserts both directions for exactly that reason: a change that made every
refusal a warning would pass its first half and quiet the lines the log exists
for. Nothing else moved — the boot is still refused, the machine still reads
back its FailedState, and the refusal still names the identifier, the reason,
the consequence and both gestures.
What follows is the measurement that established it, kept because the reasoning is the reusable part.
Replaying the fifteen surveyed stacks under a machine runtime (main@72d861d,
--vm incus-ovn, logs of 2026-08-25), five runs printed level=ERROR —
fourteen lines between them (O1-rke 1, O2-ztiac 5, O4-k3s 1, O5-kasten 5,
S1-talos 2) — and every one is the same deliberate, documented refusal to boot
an image identifier the catalogue does not hold (the call site built around
DeclaredImageSyntax in internal/core/machine/binding.go). The refusal text
itself is right: reason, consequence, and both gestures (feint images resolve, FEINT_BOOT_IMAGES). The level is what is wrong, and the same log
says so 200 ms later:
19:54:23.185 level=ERROR msg="refusing to boot: nothing says which operating system this image identifier names" …
19:54:23.385 level=WARN msg="this runtime does not distribute a load balancer of this shape; the API goes on describing it and its backends" …
Two documented degradations, same run, two levels — the second is #457's, and it is the one that follows the rule. The run that printed five of these ERRORs was a success: ztiac applied 54 of 54, matched its reference exactly, and destroyed 54 cleanly.
The lesson that outlives the fix: do not grade an emulator log by grep ERROR
alone, and do not level a limit like an incident. Under those replays the grep
found fourteen lines about a documented behaviour and nothing about the run
that really failed, which is precisely how a team learns to skip a log's
errors.
Lifted 2026-08-29 by #605. What follows is the measurement that produced the
limit, kept because the fact underneath it has not moved: images: still
publishes neither ubuntu/18.04 nor debian/9, and two of the three identifiers
the surveyed stacks hardcode still resolve to them.
What changed is that the command no longer hands you a line that cannot boot. It reads the image server's own simplestreams index, leaves an unbuildable version out of the declaration, and says why:
$ ./feint images resolve ami-a3ca408c ami-538af795 ami-47899c77 ; echo "rc=$?"
ami-538af795
Outscale official OMIs reference: Ubuntu-18.04-2021.02.04-0
→ ubuntu:18.04 — the image server no longer publishes ubuntu/18.04/cloud, so this declaration
would refuse to boot; it is left out of the line below.
ubuntu versions published today: 22.04, 24.04, 26.04, jammy, noble, resolute
naming one of them is your call, not this command's
declare and restart:
FEINT_BOOT_IMAGES='ami-a3ca408c=ubuntu:22.04' feint serve ...
rc=0A run whose every identifier resolved to a withdrawn version exits 2, the way an identifier in no listing already does.
What it deliberately does not do is substitute. The published versions are printed as a fact; choosing among them stays the operator's act, because naming the nearest buildable version for them is the silent replacement #83 measured and closed on all three packs.
The cost the limit used to carry, for the record. Pasting the old line bought
nothing: replayed on michaelcourcy/kasten-on-outscale @f1fcc87, 29 applied, 0
machines started, five Vms tainted — the same figures as with no declaration at
all. The absence was measured with a control so that "absent" is not "looked
nowhere": the images: listing read as JSON held 726 aliases, ubuntu/18.04 and
debian/9 absent, ubuntu/22.04 and debian/12 present as controls. Read as
JSON deliberately — incus image list -c l prints one alias and (7 more), and
grepping that column reports present aliases as absent too.
scw 2.56.3 prints a recovered panic on every successful lb acl delete, and the defect is upstream (#505)
Every conformance pass of the Scaleway load-balancer chain carries one orphan
line, glued before the step's ok::
runtime error: invalid memory address or nil pointer dereference ok: the balancer chain round-tripped, …
Bisected command by command on 2026-08-26, stderr separated per step
(emulator --vm off, 24 + 11 commands replayed): exactly one command emits
it, scw lb acl delete <id> — which succeeds (rc=0, the ACL is deleted,
the suite goes on). The recovered panic's stack points into the CLI itself
(main.cleanup, cmd/scw/main.go:45; interceptACL.func21,
internal/namespaces/lb/v1/custom_acl.go:83): the interceptor guards its
pre-fetch with a type assertion on *ZonedAPIDeleteCertificateRequest where
the argument is a *ZonedAPIDeleteACLRequest, so getACL stays nil and the
delete's success path dereferences getACL.Frontend.LB.Tags. The panic is
recovered, printed, and the process exits 0.
Measured: the exact command, rc=0, the ACL deleted, the stack, the
reproduction under --vm off as under --vm incus-ovn, and no ERROR on the
emulator's side.
Could this emulator answer something else, without lying, that does not trigger the fault? That is the only version of #505 that would be a fix rather than a note, so it was answered by reading the two functions and then by experiment, on 2026-08-28.
Read (scaleway-cli/v2@v2.56.3 and scaleway-sdk-go in the module cache):
ZonedAPI.DeleteACLends ons.client.Do(scwReq, nil, opts...)— the response is decoded intonil. No body this emulator sends reaches the faulty path, so the shape of the answer is not a lever at all.lbACLDelete'sRunreturns&core.SuccessResult{Resource: "acl", Verb: "delete"}unconditionally whenever that call returns a nil error. So any truthful success produces the value the interceptor then dereferences.- The interceptor is installed on all four ACL verbs, and its pre-fetch is
guarded by
argsI.(*lb.ZonedAPIDeleteCertificateRequest)— never true here, sogetACLis nil on every path.
Measured, with the emulator's own fault injection, which is what makes this an experiment rather than a second reading:
| what the emulator answers | rc | panic on stderr | the ACL |
|---|---|---|---|
204, as it does (DeleteACL succeeds) |
0 | yes | deleted |
500, via PUT /_feint/faults |
1 | no | survives |
And the three sibling verbs — acl get, acl update, acl create — carry the
same interceptor, exit 0 with empty stderr, and never panic: they answer an
*lb.ACL rather than a *core.SuccessResult, so they never reach the branch.
That locates the fault exactly at "the runner returned a success", and nowhere
near the emulator.
So the emulator's only lever is to fail a delete that worked, which loses the resource and lies about it — the one thing this project exists not to do. Option (1) is disproved, and #505 closes as documentation rather than as a fix.
What to do with it: nothing, here. This is not a divergence; the line is noise
in every conformance log, tolerated because delete stderr passes through and rc
is 0. Not filtered either, which would be a workaround hiding a real upstream
defect: it is tolerated and the suite now asserts what makes tolerating it
honest — scw-cli.sh reads the ACL list back after the delete, so "rc=0 with a
panic on stderr" is measured rather than trusted. Whoever meets the line in a
log: #505 is the reference to point at. What would lift it: the upstream fix in
scaleway-cli — one type in the assertion, to be reported through
scw feedback bug; the line disappears when a fixed scw ships.
Still not measured, and it does not change the verdict: what a real account answers. Nothing in the faulty path reads a response, so the same panic is expected there — but that sentence is a deduction from the source above, not a measurement, and it is written here as one.
The Exoscale stack's second plan is not empty: two per-id outputs read back null at apply time (#520)
feint up --runtime incus-ovn applies examples/stacks/exoscale (the pinned
maintainer fork, dev_overrides) green, and the apply's own Outputs: block
lists five outputs — front_by_id and ingress_address_by_id are missing
from it. The second plan is then not empty:
$ terraform plan -detailed-exitcode # exit 2
Changes to Outputs:
+ front_by_id = "platform-front"
+ ingress_address_by_id = "192.0.2.2"No resource changes: the whole diff is those two outputs appearing, and
applying that plan changes 0 resources and converges — the third plan exits 0
(#506, the earlier filing of the same defect). Reproduced on three
independent trees on 2026-08-26 — main@46230cc,
fix/473-493-ovn-under-concurrency, fix/481-483-…@f11d384 — with the same
two lines, so it is deterministic and belongs to no branch. The two outputs
are exactly the per-id doors #478 added: data sources that read the NLB and
the elastic IP back by id (data.exoscale_nlb.front.name, the ingress
address by its id); every output that reads the resources or the list-shaped
data sources is present from the first apply. The clean second plan
docs/clients.md records for 2026-08-18 was measured on the 13-resource
stack, before those doors existed.
The reading above was deduced rather than instrumented: Terraform omits an
output whose value is null, so the first apply's by-id read — issued
immediately after the create — must have answered a document whose name (and
the elastic IP's address) was empty. That pointed at the pack's by-id GET
racing its own async create-completion.
Measured on 2026-08-28, and that reading is wrong. The by-id read was
driven against the pack directly, which is the surface this repository has
decided to drive, with --vm off because the question is the document and not
the runtime:
| what was asked | reads | empty answers |
|---|---|---|
GET /v2/load-balancer/{id} on one balancer, back to back |
100 | 0 |
| create a balancer, read it by id with no delay | 100 | 0 name, 0 ip |
| create an elastic IP, read it by id with no delay | 1 | 0 |
| create an instance pool, read its size by id with no delay | 50 | 0 |
There is no window to race. createLoadBalancer puts the resource with name
and ip in its attributes before it writes the operation, and the
operation it writes is already state: success — the async shape is in the
envelope, never in the document. The same holds for the elastic IP. The exo
CLI reads the same balancer back with its name and ip_address filled.
Two consequences, and the second is why this stays here rather than becoming a fix:
- The emulator does not discriminate between the three per-id doors. The stack has three by-id data sources — the NLB, the elastic IP and the instance pool — and only the first two came back null. The emulator serves all three from one store and answers all three completely, so whatever produced the nulls tells them apart and this pack does not.
- What tells them apart is on the client side, and it is the client this
repository refuses. The provider builds two clients and only one honours
EXOSCALE_API_ENDPOINT(the section above, and #573);setEndpointFromContextinegoscale/v2does not rewrite a literal address, so which door answers depends on which client the provider built for it. That is a reading of the provider, not a measurement of it — the fork is not checked out here — and it is written as one.
What to do with it: nothing drives this stack, by decision (#525), so this is
not a defect anybody can meet through a supported path. The replay that found
it was the pinned maintainer fork, and -detailed-exitcode does not
distinguish an outputs-only diff from a resource diff either, so an honest
replay reports "exit 2, outputs-only, resources at zero" rather than "red".
What would lift it is upstream #573; what would not lift it is
any change to this pack, and the table above is why.