Skip to main content

Inference Verification

This page explains how AETRON verifies inference work: the everyday case where a miner serves a model and returns an answer. Verification is done by miners for each other, enforced by the protocol. There is no separate validator role and no subjective scoring.

The Determinism It Relies On​

Running a model on a given input produces a deterministic distribution of next-token probabilities (the logits), even though the sampled token can vary. Two honest runs of the same model, with the same settings, produce the same distribution. A different or tampered model produces a different one. This is what makes a re-check meaningful: the result either matches or it does not.

The practical consequence is a wide separation between honest noise and fraud. Measured on the same model across different GPUs, the L2 distance between logit vectors sits around 0.01 to 0.03. Substituting a different model pushes it to roughly 0.3 to 0.8. A threshold in the 0.05 to 0.10 range separates the two cases with room on both sides.

The Validation Artifact​

When a miner serves a request, it records a compact validation artifact describing the key states of the computation. The artifact itself stays with the miner. Only a hash reaches the chain, and the artifact is handed over on request during a re-check.

What goes into the artifact depends on the model class.

Language models (autoregressive). The artifact holds the top-k probabilities for the first generated token, the sampling seed, the model hash, a request id, and a timestamp. The first token is the expensive one to fake, because producing it requires processing the entire input through the real weights. Size is around 200 bytes.

Diffusion models (image, video, audio). Sampling here is iterative, so the artifact pins the trajectory rather than a single distribution: hashes of the intermediate latent at fixed steps (for example step 5 and step 25), a hash of the final output bytes, the seed, the step count, and the model hash. Size is around 200 bytes.

Encoder models (BERT, CLIP, ViT). These have no sampling stage and are fully deterministic, so the artifact is minimal: a hash of the produced embedding or class output plus the model hash. Size is around 100 bytes.

The Verification Protocol​

Step 1: Execution​

A consumer request reaches the miner through the Neuronet owner's API. The miner runs the inference, saves the artifact locally, and returns the answer with a signed receipt. Nothing has touched the chain yet, and the user waits only for the inference itself.

Step 2: Batching the Commitments​

Artifacts are committed in two layers so that on-chain cost does not scale with traffic.

A miner collects the artifacts it produced over a window (on the order of 5,000 requests in 30 minutes) and builds a Merkle tree over them, with one leaf per artifact. It sends only the resulting batch root to the Neuronet owner, off chain. The owner collects batch roots from all its miners (200 miners in the reference case) and builds a second tree over those roots.

One transaction then submits the aggregated root for the whole Neuronet. That transaction is around 200 bytes and can cover a million requests, which works out to roughly 0.0002 bytes of chain state per request. Every individual artifact is still provable against it by a Merkle path.

Step 3: Selecting What Gets Checked​

The protocol derives the selection from a verifiable random function over the block hash and the request id. The VRF decides two things: which requests are challenged (about 7% of traffic) and which miner from the same task is assigned to re-check each one.

Because the selection is a function of chain state rather than a choice made by any participant, no one can steer a check toward a friendly verifier or away from a fabricated answer.

Step 4: Replay​

The assigned miner asks the executor for the artifact, checks the Merkle path to confirm that this exact artifact was committed on chain, then reproduces the computation. Replay is deliberately partial, and the shape differs by model class.

For a language model the verifier runs prefill only: a single forward pass over the input, which yields the logits of the first generated token, compared against the artifact by L2 distance. For a 500 token answer that is one pass out of 500, about 0.2% of the original work.

For a diffusion model the verifier reruns the same prompt and seed but stops early, for example after 5 of 50 steps, and compares the hash of its own latent against the committed one. That costs about 10% of a full generation. Video follows the same idea on a single frame or step, at 3% to 10%.

Encoder models are replayed in full and the output hashes are compared directly. The relative cost is 100%, but these models are small enough that this is cheap in absolute terms.

The aggregate overhead is the challenge rate multiplied by the replay cost. For language models that is 7% of 0.2%, or about 0.014% of network GPU time. For diffusion it is 7% of 10%, about 0.7%.

Step 5: Recording the Verdict​

The verifier submits its verdict on chain, roughly 100 bytes, and only for the challenged share of requests.

Verification is not optional. At the end of an epoch the runtime checks each miner against the assignments the VRF gave it. If a miner was assigned 23 checks and submitted 23 verdicts, its reward is paid in full. Missing verdicts reduce the reward in proportion. A miner that ignores its verification duty does not get paid for its own work.

Shadow Replay​

A miner that could tell a verification request apart from a real user request might behave honestly only when watched. AETRON prevents this with shadow replay: the requests chosen for re-check are real user requests, indistinguishable from normal traffic. The miner cannot know which of its answers will be verified, so the only safe strategy is to be honest on every request.

The Verdict Quorum​

Single checks are combined into a quorum so that one unlucky comparison does not decide a miner's fate, and so a small group cannot quietly wave each other through.

A challenged job is finalized against K = 20 verifiers, with a fraud threshold of T = 10. Verifiers that abstain (for example because they could not obtain the artifact) are excluded from the denominator entirely, so an abstention is neither a vote of confidence nor an accusation. Finalization requires 20 non-abstaining votes.

The verdict follows directly from the fraud count among those votes:

Fraud votes out of 20VerdictEffect
10 or moreFraudPenalty applied to the executor
1 to 9SuspiciousNo payout on the job, the miner stays under scrutiny
0HonestWork is credited and earns emission

The Suspicious band matters. A single dissenting verifier is not enough to punish a miner, which protects honest miners from an isolated bad comparison or a griefing verifier, but it does not let the job pass clean either.

Finalization is permissionless: anyone can trigger it once the votes are in, so a miner cannot delay an unfavorable verdict by refusing to act.

Replication and Arbitration​

Two separate mechanisms handle the case where verifiers do not agree with each other or with the executor.

Replay quorum requires that at least M = 3 independently committed replay hashes converge on the same value before a job's verdict is treated as replicated. One verifier reporting its own result to itself proves nothing, so the design demands several blind commitments landing on the same answer. If they do not converge before the replay deadline, the job finalizes without credit for useful work rather than being scored on a single unconfirmed opinion.

Arbitration handles a direct conflict. If a verifier returns a mismatch, the protocol does not yet know who is lying: the executor may have substituted a model, or the verifier may be making a false accusation. A third miner replays the same request. If its result lines up with the executor, the accuser is the one at fault. If it lines up with the accuser, the executor is. The majority decides, and the party out of line is penalized, including the case where that party is the accuser. False accusations carry a cost, which is what keeps arbitration from becoming a free weapon against competitors.

Cross-Hardware Tolerance​

Because miners may run different hardware, comparison uses the tier set for the task: a bit-exact match within one hardware class (Tier A), or a calibrated tolerance across architectures (Tier B). Tier B thresholds are derived from measured differences between honest machines, which is why the gap between honest variation and model substitution described above is the load-bearing assumption of the whole design.

See Heterogeneous Mining for how the tiers work, and Proof of Intelligence for the wider mechanism.