Attack Resistance
A verification scheme is only as good as the attacks it survives. This page lists the fraud strategies that were considered against Proof of Intelligence and the mechanism that catches each one. The measurements quoted here come from the attack testing rounds run against the reference implementation, not from live main network traffic. See Project Status for where the protocol actually stands today.
The Attacker's Position
A dishonest miner wants emission without paying for the compute that earns it. Every practical fraud reduces to one of three moves: run a cheaper model than declared, run less of the declared model than claimed, or behave honestly only when it suspects it is being watched. Each attack below is a variant of one of those three.
Faking the Entrance Check
A task publishes reference probes, a small set of fixed inputs whose correct output distribution is known. A miner must reproduce them to join.
The obvious attack is to train a cheap surrogate model that imitates the target's outputs on exactly those probes, pass the check, then serve real queries with the cheap model. This was tested by distilling a same-architecture surrogate until it was aggressively overfitted on the probe set.
The surrogate matched none of the probe hashes. Distribution matching does not produce bit-identical logits: bf16 rounding shifts the low bits, top-k ordering changes, and the hash diverges. Reproducing a hash requires the same weights in the same numerical representation, which is the point of the check.
Batching Real Traffic
Serving many prompts in one batch is the standard efficiency trick, and it changes the arithmetic. Padding tokens alter the accumulation order inside attention, and floating point addition is not associative, so the same prompt produces different logits at different batch sizes.
Tested across a range of batch sizes, every one produced a different hash. This is why the canonical execution spec fixes batch size at 1 for verified work. It is a correctness requirement rather than a performance preference. A miner is free to batch for its own throughput, but the values it commits for verification must come from a batch-of-one run.
Answering Well Only When Watched
If verification traffic is distinguishable from user traffic, a miner can serve honest results to the verifier and cheap results to everyone else. This is not a theoretical concern: a five line classifier separated validator queries from user queries with full accuracy in testing. A dedicated verification endpoint is insecure by construction.
The answer is shadow replay. Requests selected for re-execution are real user requests, drawn from ordinary traffic and carrying no marker. At a 7% replay rate the detection rate against a selectively fraudulent miner was 100%, and an adversary trying to guess which requests were replays scored 40%, worse than a coin flip. Details in Inference Verification.
Sampling Games
A miner might declare greedy decoding and quietly sample, or claim a seed it did not use. Generated text differs run to run, so text comparison would be fooled.
Verification does not compare generated text. It compares the logit distribution from the forward pass, which is computed before any sampling decision is made. Across greedy decoding and a range of temperatures and seeds, the generated tokens differed every time and the forward pass hash was identical every time. Sampling spoofing has no surface to attack here.
Half-Finished Image Generation
Diffusion gives an attacker a natural place to cut corners: run the first few denoise steps honestly, then substitute cheap or random latents for the rest. The final image still looks plausible while the compute bill drops.
Each denoise step is a deterministic function of the previous latent and the seed, so committing a hash per step turns the run into a chain. In testing, fraud injected part way through a run was detected at every step where it appeared. Partial substitution has nowhere to hide once the chain is committed.
Mixing Models Together
For multi-modal work, an attacker could try to run the image encoder from a cheap model and the text encoder from the declared one, then concatenate the embeddings.
In most cases this fails before verification even runs, because different model families produce different embedding widths, and concatenation is simply not defined across them. When two versions of the same family happen to share dimensions, the logit hash still separates them. The structural incompatibility is the primary barrier and the hash is the fallback.
Faking the Expensive Half
Prefill, the pass over the whole prompt before generation starts, is the compute-heavy phase. A miner could try to run it on a small model and continue generation on the declared one.
Three independent checks catch this. The forward pass hashes differ because the models differ. The vocabularies differ, and the tested pair had no overlap at all in top token IDs. And the context hash over the full prompt does not match a hash over a partial prompt, which flags prompt manipulation before any replay happens. An attacker has to defeat all three at once.
Long Contexts
Verification of a long request is only meaningful if a long forward pass is reproducible. Tested across context lengths from a few thousand tokens up to 32K, repeated runs produced identical hashes. Cross-architecture agreement on the leading token held over the same range.
The published ceiling so far is a 128K run. Going further stopped at the model's own context window rather than at anything in the protocol, which is the distinction that matters here. Verification has no context limit of its own, so million-token contexts are a question of model and memory, and those runs are being measured now with results to follow.
One caveat matters here. Replaying only a chunk of a long prompt in isolation does not reproduce the logits from the full-context run, because attention over a chunk sees less than attention over the whole prompt, and the mismatch grows the further into the prompt the chunk begins. The verifier therefore replays the full context and compares at the position of interest, rather than replaying fragments. Chunked shortcuts are only sound for models with sliding window attention.
Full-context replay is what makes long prompts expensive to check, and that cost was measured rather than assumed. Re-prefilling a single position on a long prompt costs close to what the executor's own run cost, so the "verification is a fraction of a percent" figure holds for short answers and not for long ones. A KV cache in the runner was also a precondition for long contexts being practical at all, rather than an optimization.
Agent Workflows and Tool Output
In a multi-turn agent conversation, the tampering target is not the model output but the tool output fed back into it. A miner could replace a price returned by an API and continue generating honestly from the altered context.
Two layers cover this. The next turn's logits depend on the whole conversation through attention, so altered tool output changes the hash: in the tested case, swapping a single number moved the next-token distribution far outside anything honest variation produces. Separately, the hash over the full conversation text changes as soon as any turn changes, which is a cheap runtime check that fires before a replay is needed.
The same conversation hash makes it safe to move a session between miners. In a multi-turn conversation handed from one miner to another mid-way, the final hash matched the single-miner run, so load can be rebalanced without breaking reproducibility.
Backdoors in Training
For training work, the concern is a poisoned update that survives verification. Every individual step is a genuine forward, backward and optimizer update, so a step-level replay has nothing to catch. The mechanism that covers this reads a stage's update history rather than any single step, and it is described in Proof of Training.
Trust Levels for Code and Environment
Reproducible hashes prove that two parties ran the same computation. They do not, by themselves, prove which code was on disk. AETRON handles this in layers, each stronger and more expensive than the last.
Level 1, inferred verification. If the results match bit for bit, the code was effectively the same. Any modification that changes numerical output is caught by replay, and that covers model substitution, altered forward passes and environment mismatches. It does not catch modifications with no numerical effect, such as added telemetry, and it does not catch an executor and a verifier who modified their code identically. The defense against that second case is that verifiers are assigned by chain randomness rather than chosen by the executor. With a 20% colluding fraction and 5 independent replays, the chance that every replay lands on a colluder is 0.2^5, about 0.00032. This level assumes an honest majority, the same assumption every proof-of-work and proof-of-stake system makes.
Level 2, environment fingerprinting. The task publishes hashes of the reference inference runner and of the full dependency list, and optionally a container image digest. Miners report their own hashes at registration and on periodic re-checks, and a mismatch is rejected or penalized. This closes code changes with no numerical effect, library monkey-patching, dependency substitution, and running one code path for probes and another for user traffic. Its limit is honest: the report is self-declared. A miner can submit a correct hash and run something else, and runtime patching through injected libraries or direct memory writes stays invisible.
Level 3, hardware attestation. With confidential computing on supported GPUs, the hardware root of trust signs a statement about which code and which weights are loaded inside a sealed enclave. That converts self-reporting into a cryptographic guarantee, and it also hides prompts and weights from the machine operator. The cost is a documented 4% to 8% throughput overhead plus a fleet restricted to attestation-capable hardware, and the trust shifts to the chip vendor's key infrastructure. Side channels are reduced but not eliminated. This level is reserved for confidential tasks, not general work.
The chain's role is the same at every level. It stores the canonical spec and the reference hashes, supplies the randomness that assigns verifiers, holds the committed roots, and enforces penalties. It does not execute inference, does not hold prompts or answers, and cannot check code directly.
What Happens After a Fraud Verdict
Penalties are graded by class rather than applied uniformly. Repeated offenses climb a ladder.
Serious classes are treated as permanent marks on a miner's record: confirmed fraud, a false accusation against an honest miner, signing off on work that was in fact fraudulent, submitting a verdict that could not be reproduced, missing a witness deadline after prior warnings, and staying silent as part of a colluding group. The first such offense is a temporary ban, and a second one is permanent. These are also the classes that put the miner's bond at risk.
Softer classes decay over time and are handled with a gentler ladder, because a single failure is usually an operational problem rather than fraud. A missed witness duty costs half the reward first, then all of it, then a temporary ban. Downtime starts at a 10% reduction, then 50%, then zero reward. A dispute that turns out to be unfounded starts at a 20% reduction and escalates the same way.