Shadow Replay
Any verification scheme has to answer one awkward question: what stops a miner from doing careful, honest work on the requests it knows are being checked, and cheap garbage on everything else? Shadow replay is AETRON's answer. The requests picked for verification are ordinary user requests, so there is no "being checked" state a miner can detect and no moment it can single out for good behaviour.
The Attack It Closes
Call it selective answer fraud. A miner registers with an expensive model, passes every check the network sends it, and serves users with something much cheaper: a smaller model, a quantized copy, or a cache of stale answers. The user gets a worse answer, the network sees clean verification results, and the miner keeps the difference in running costs.
This attack works only if the miner can tell the two kinds of traffic apart. That is the whole hinge. A verification system built around a dedicated endpoint, a fixed test set, or a recognisable prompt format hands the attacker exactly that ability.
Test Prompts Are Easy to Spot
This is not a theoretical worry. In an internal test, realistic user queries were compared against formal validator-style queries drawn from a standard probe template set. The two groups separated on trivial surface features: capitalisation, formal against informal vocabulary, punctuation, whether the text contains numbers.
A classifier of a few lines built from those features labelled every query in both samples correctly. No machine learning, no effort, no cost to the attacker.
The conclusion carried into the protocol design: any scheme where a verifier sends specially prepared test prompts is bypassable, because prepared prompts look prepared.
How Shadow Replay Works
There is no separate verification channel. Verification rides on the real traffic.
- A user request arrives at the routing layer.
- With probability
rthe routing layer marks it as a shadow replay candidate. The AETRON default forris 7%. - The request goes to the miner as an ordinary request. Nothing in it says it was marked.
- The miner generates the answer and the validation artifact, and returns both.
- The routing layer passes the answer to the user, exactly as it would otherwise.
- If the request was marked, another miner in the same task independently re-executes it in canonical mode and compares its result with the miner's committed artifact hash.
A mismatch feeds the same fraud path described in Inference Verification: statistical scoring, arbitration by a third miner, and slashing when the verdict is confirmed.
The marking decision is made after the request already exists, in a component the miner has no view into. The miner sees the same request either way.
Why It Is Indistinguishable
Indistinguishability here is a property of construction rather than a claim about clever disguise. Shadow replay candidates are drawn from real user traffic, so they follow the traffic distribution by definition. There is nothing to imitate and nothing to give away.
This was tested adversarially. An attacker model was given the request stream and allowed to separate replayed requests from ordinary ones using length, character entropy, word length, and punctuation rate. No feature separated the two groups by more than ordinary variation inside one traffic source, and the attacker's classifier ended up below the accuracy of guessing.
In the same simulation, a miner substituting a degraded model was caught on every replayed request. A control run using a naive dedicated verification endpoint, where the attacker's classifier identified every verification query, gave a detection rate of zero.
The Arithmetic of Getting Caught
The miner is forced into a blind bet. Every request it degrades might be the one that gets replayed, and nothing in the request tells it which. Cheating once or twice is survivable. Cheating as a business is not: the chance of escaping review collapses as fraudulent responses accumulate, so a miner that cheats systematically is caught quickly rather than eventually.
That gap between "sometimes" and "as a business" is why shadow replay is paired with stake. Slashing is designed so that the loss on a single confirmed fraud event outweighs what a miner could earn over the queries it takes to get caught, which removes the profit from short bursts of cheating as well as sustained cheating. The pallet applies an immediate slash on one confirmed fraud verdict, and also acts on repeated flags inside a sliding window of 100 requests.
Cost
The miner pays nothing for shadow replay, because it does no extra work and gets no extra signal. The cost falls on the re-checking side: at a 7% rate the network runs 7% extra verification passes, and for language models those passes are prefill-only, a small fraction of the cost of generating a full answer. Inference Verification has the per model class breakdown, and Heterogeneous Mining covers how the comparisons stay valid across different hardware.
Privacy
Shadow replay works on real user requests, so it is bound by the same rules as the rest of the request path:
- The routing layer is the only component that knows which request was marked and which miner received it. That mapping is not exposed to the network.
- The re-checking miner is selected from other miners in the same task, and the request reaches it through the same routing layer rather than through a third party.
- Requests that were not selected are dropped after they are served. Only shadow replay candidates are held, and only until verification finishes.
See Privacy for what the request path does and does not reveal.
Small Pools: One or Two Miners
Peer shadow replay needs someone else to do the replay. In a task with one or two miners, that does not exist in a meaningful form.
With a single miner there is simply no other miner to select. With two, the selection is formally valid but the pair only ever checks each other, and one operator holding both hotkeys turns it into self-verification wearing a costume. Peer checking becomes genuinely collusion-resistant only once a pool is large, around ten miners or more.
The protocol handles this by switching verification mode based on pool size rather than pretending the small case works:
| Miners in task | Mode | Verification |
|---|---|---|
| 1 or 2 | SmallPool | The Neuronet Owner runs mandatory reference probes on a fixed interval (about a day). Without a fresh Owner verdict the task earns no emission. Shadow replay is Owner-initiated. |
| 3 or more | PeerVerified | Full peer shadow replay with random verifier selection and cross-miner statistical scoring, automatically. |
A single-miner task is allowed on purpose. It keeps the barrier to entry at zero and lets a task grow from nothing, and the switch to peer verification happens on its own once the third miner joins.
The Owner is a reasonable verifier here: it computed the reference probes when it registered the task, and it holds a share of the task's emission plus whatever product depends on the output, so it has a direct reason to want honest results. The degenerate case, an Owner running the only miner and checking itself, is not really verification. It is also not profitable. A task's share of emission is its verified useful work divided by the network total, so a fake task dilutes its own share rather than taking anything from anyone else. The Owner only cheats the Owner.