Skip to main content

Canonical Execution Spec

"The same model" does not mean "the same result". Two honest miners can load identical weights, answer the same prompt correctly, and still produce outputs that differ in the last few bits. Verification has to tell that apart from a miner who quietly swapped in a smaller model. The canonical execution spec is how a task removes every source of variation the protocol can control, so that whatever variation is left is understood and bounded.

Why Logits Are the Right Thing to Compare​

At each generation step a language model produces a probability distribution over the whole vocabulary. The sampled token is random, so the text two miners return will often differ. The distribution behind it is not random: for the same input and the same settings it is the same distribution.

That is the property verification rests on. Two honest miners running a 405B model return the same top probabilities even when one samples "4" and the other samples "four". A miner that substitutes an 8B model returns visibly different probabilities, and the distance between the two vectors makes the substitution obvious. Comparing text would prove nothing, comparing distributions proves which model actually ran.

What Varies Even With Correct Weights​

Quantization, inference engine, continuous batching, speculative decoding and MoE routing all change the logits produced from one set of weights. So does the system prompt. Some of these differences are small, and small is exactly the problem: measured against each other, the common numeric formats disagree by about as much as a fine-tuned substitute model would.

A tolerance wide enough to accept an unfixed dtype is therefore also wide enough to accept a swapped model. The answer is not a wider tolerance, it is to fix the environment.

What the Spec Fixes​

Each task carries a canonical spec that every miner serving it must reproduce during verification:

FieldWhat it pins down
model_hashThe exact weights, by hash
dtypeNumeric format of weights and activations (bf16, fp16, fp32, int8, fp8, nvfp4, mxfp4)
attention_implAttention implementation (SDPA, FlashAttention 2, cuDNN)
gpu_archGPU architecture family, mandatory for Tier A and optional for Tier B
compile_modeTorch compilation mode (eager, reduce-overhead, max-autotune)
seedThe seed used for any randomness in the run

The task configuration around it adds the rest of the canonical conditions: batch size of 1 during validation, speculative decoding off, deterministic routing for MoE models, and a hash of the canonical system prompt where one exists. Miners are free to serve real users any way they like, with continuous batching, speculative decoding and a full KV cache. The canonical mode applies to the replay pass, not to production traffic.

Attention implementation deserves a specific warning, because it looks cosmetic and is not. On one GPU with one dtype, SDPA and eager attention produced completely different hashes, with no probe matching between the two backends. Same GPU, same weights, same precision, no match. A validator and the miner it is checking have to use bit-identical strategies, or the comparison fails for a reason that has nothing to do with honesty.

Cross-Architecture Behaviour​

Within one architecture, results are reproducible bit for bit. Across architectures they are not, and no amount of configuration fixes that.

Measured between two NVIDIA generations on identical prompts, in bf16, fp16 and fp32 alike, not a single hash matched. The results are numerically equivalent, the per-logit differences being far below anything a user would notice, but tensor cores on the two generations use a different reduction order and accumulation schedule, so the bits differ and the hashes diverge.

The same signature appears crossing a further generation boundary, and there it went beyond bits: neither the hashes nor the leading token ids matched, meaning the divergence was large enough to reorder which tokens win rather than only to nudge their probabilities. Those runs also used different PyTorch and CUDA versions, so the gap mixes hardware and software effects and should not be read as a pure hardware number.

Two conclusions follow. Crossing an architecture boundary breaks bit-exact comparison as a general rule rather than as a quirk of one vendor generation, and within an architecture determinism holds, which is why a single-architecture task can be verified with an exact hash. This is what the two verification tiers are built on, and Heterogeneous Mining covers how Tier A and Tier B differ.

Enforcing dtype With Reference Probes​

The spec declares a dtype, but a declaration alone does not stop a miner from claiming fp16 and running bf16 to save memory. Reference probes close that gap.

When a task is created, the Owner computes reference logits on a canonical environment: a fixed set of 10 prompts, the task's canonical spec, batch size 1. Only the hashes go on chain, which is 10 hashes of 32 bytes each.

  1. A miner registering for the task receives the same prompts and computes logits on its own hardware in canonical mode.
  2. It submits a hash per prompt.
  3. The runtime compares against the reference. A match admits the miner, a mismatch rejects it.
  4. Periodically the protocol picks 3 of the 10 prompts and requires the miner to reproduce them again, which costs 3 forward passes and 96 bytes on chain.

The check is a hash comparison rather than a distance, deliberately. A hash gives a binary answer with no threshold to tune and no grey zone: the right dtype and the right weights match, anything else does not. Running the correct dtype on two different GPUs of the same architecture reproduced the probes exactly, so the exact match is a realistic requirement and not a trap for honest miners.

The periodic recheck also covers the case of a miner that registers honestly and changes its dtype or its model afterwards.

Quantization Is Part of the Spec​

Quantization is not an implementation detail a miner chooses. It is part of the task fingerprint, because each format produces its own deterministic output.

A range of formats has been validated as canonical-ready, meaning each one is bit-identical across repeated runs on the same GPU. The set covers the standard floating point types, the common integer and 4-bit quantization schemes, and the newer microscaled formats. Several of the newest need Blackwell-class hardware and a recent PyTorch build, so choosing them narrows the pool of miners that can serve the task. The runtime is the authority on which formats a task may declare.

A high L2 distance between a quantized format and the bf16 baseline reflects discretization loss, not instability. A format can sit far from bf16 and still be perfectly deterministic within itself. Since reference probes are computed for the declared format, a miner running a different format than the one declared fails the probe check.

One family is excluded: AWQ cannot be used for verification. Repeated AWQ runs on the same GPU produce different hashes, and the non-determinism lives in its dequantization kernel rather than in any hardware, so no architecture rescues it. A miner may keep AWQ for serving users, but a replay pass has to run a canonical-ready format.

Coverage Across Model Types​

The spec is not specific to text generation. What gets pinned down depends on the kind of model:

Model typePinned by the spec
LLM (text generation)Quantization, engine, system prompt hash, batch size 1
DiffusionSeed, number of steps, CFG scale, scheduler
Encoder / embeddingQuantization (models are small, so overhead is minimal)
Agent (multi-turn)System prompt hash, tool configuration hash
MoEDeterministic routing, quantization, batch size 1
Audio (Whisper class)Quantization, language, task (transcribe or translate)

The validation artifact carries a context hash covering the system prompt, conversation history and user message, alongside the model hash and the generation parameters. A verifier checks that hash before replaying anything: if the context does not match, the comparison fails immediately and no compute is spent.

Where This Fits​

The canonical spec is what makes a re-check meaningful. Without it, a mismatch could mean fraud or could just mean a different attention backend, and the network would have no way to tell. See Inference Verification for how a re-check runs, Heterogeneous Mining for the two tiers, and Task Configuration for how an Owner sets these values.