Repository navigation
Route every reputation penalty through one scale, and require 3 trust samples before acting - #34
Merged
Merged
Conversation
A real node on the test droplet (ndebvzgeoij5t2l) was blocked permanently off ONE trust sample. TrustScoreState.score with a single out-of-threshold sample is exactly 0.0; evaluate_reputations feeds that straight to NodeReputation.evaluate_trust, which takes 0.15 off per pass by design, and the evaluator runs every 60 s — from 1.0 that crosses the 0.2 block threshold in six minutes. Once blocked, frames are dropped and apply_reward is a no-op, so nothing recovers. The operator's stance is that for now trust must never lower a node's reputation. That is one switch, not edited literals: apply_penalty now multiplies by NodeReputation.penalty_scale, and every downrating source already funnels through apply_penalty (trust 0.15/0.05, heartbeat 0.1, detection rate 0.05, neighbour consistency 0.08, and the retina-server backend's direct 0.1 for ADS-B cross-validation). At scale 0 the penalty is not applied AND not recorded — "no new penalties" is how an operator checks the switch is on — while rewards keep working. summary() publishes the scale so /api/radar/analytics shows whether downrating is live. penalty_scale is a ClassVar, not a dataclass field: the backend snapshots reputations with dataclasses.asdict() and restores them with NodeReputation(**saved), so a field would freeze an operational stance into saved state and make older snapshots fail to construct. The library default stays 1.0 — escalation to a block is still the design when downrating is on. The backend sets the scale from REPUTATION_PENALTY_SCALE and currently defaults it to 0. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
evaluate_reputations called rep.evaluate_trust(ts.score) whenever a node had any samples at all. TrustScoreState.score is a good/total ratio, so at one sample it is either 0.0 or 1.0: a single out-of-threshold residual reads as "trust critically low", and the 60 s evaluator then takes 0.15 off per pass until the node is blocked. That is what happened to ndebvzgeoij5t2l on the test droplet. The two readers of this score had drifted apart. retina-server's node_bias.get_node_trust grew a 3-sample bar with a neutral 0.5 prior below it when trust started weighting the solver — after this evaluator was written — but the evaluator kept reading TrustScoreState.score raw. So one claim residual from the identity-first lane could start a block that the solver weighting itself would never have acted on. TRUST_MIN_SAMPLES now lives in trust.py, next to the score it qualifies, and the backend imports it instead of keeping its own copy. The evaluator skips evaluate_trust below the bar (the 0.5 prior sits between the warn and reward thresholds, so acting on it would be a no-op anyway) and substitutes the prior for neighbour_trust, so a one-sample neighbour can no longer count as a "trusted neighbour" whose disagreement condemns anyone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
On the test droplet a real node (node_ref
ndebvzgeoij5t2l) was permanently blocked off one trust sample.evaluate_reputations()calledevaluate_trust(ts.score)for any node with at least one sample; one out-of-threshold claim residual gives a score of exactly 0.0;evaluate_trustapplies 0.15 per pass (deliberately per pass), the evaluator runs every 60 s, so 1.0 crossed the 0.2 block threshold in six minutes. Blocked nodes have every frame dropped,apply_rewardis a no-op while blocked, and the backend persists the block across restarts. The backend's own trust reader (node_bias.get_node_trust) already applied a 0.5 neutral prior below 3 samples; this evaluator bypassed it, so the two paths disagreed.What
Two commits, independent of each other:
NodeReputation.penalty_scale(a ClassVar, so it never round-trips throughasdict()snapshots) andset_penalty_scale().apply_penaltymultiplies by it and, at an effective amount of 0, records nothing and changes nothing. Every downrating source routes throughapply_penaltyalready (trust 0.15/0.05, heartbeat 0.1, detection rate 0.05, neighbour consistency 0.08, and the backend's direct ADS-B cross-validation call), so a scale of 0 means no node can be blocked by any of them. Rewards are untouched.summary()publishes the scale. Library default stays 1.0; the backend sets it fromREPUTATION_PENALTY_SCALE, default 0 — a temporary stance while the trust input is a single claim residual (retina-server PR to follow).TRUST_MIN_SAMPLES = 3intrust.py;evaluate_reputationsonly acts on a score at or above it, and a neighbour below it counts as the 0.5 prior rather than a "trusted neighbour".node_biaswill import the same constant so the paths cannot drift apart again.Tests
test_bad_actor_gets_blockedkept and made explicit about the scale (1.0). New: zero scale never blocks or records across all sources; 0.5 halves; validation; summary/asdict; one bad sample + 20 evaluator passes leaves reputation 1.0, three samples penalise. 518 passed; ruff clean.🤖 Generated with Claude Code