Skip to main content

Heterogeneous Mining

A defining feature of AETRON is that miners do not need one specific kind of hardware. The same verification works across very different machines, so a single Neuronet can combine many hardware classes at once.

Supported Hardware​

A Neuronet can be served by a mix of:

  • NVIDIA server and consumer GPUs across several generations (Ampere SM 8.6, Ada SM 8.9, Hopper SM 9.0, Blackwell SM 12.0)
  • AMD data-center GPUs (CDNA, MI250 and MI300X)
  • Apple Silicon (M1 through M4, via MPS)
  • CPU configurations (Intel, AMD, Apple Silicon), running in fp32

This means the pool of possible miners is far larger than on networks tied to a single GPU vendor. Data centers, workstations, and consumer machines can all contribute to the same network, and no single hardware vendor is a bottleneck.

The Cross-Hardware Problem​

Different architectures produce slightly different numerical results from the same input. Kernel libraries differ (rocBLAS on AMD against cuBLAS on NVIDIA), reduction orders differ, and floating point rounding follows different patterns. A naive re-check on a different kind of machine would flag honest miners as fraudulent, because their outputs never match bit-for-bit. This is the reason most decentralized AI networks support only homogeneous hardware.

AETRON solves it at the protocol level with two verification tiers. The tier is part of the task's canonical spec, so the Neuronet owner chooses it up front rather than having it negotiated per job. See Execution Spec for the rest of the canonical parameters.

Two Verification Tiers​

Tier A: Bit-Exact​

Within a single hardware class running the same dtype, results are reproducible bit-for-bit, so verification is a plain hash comparison with zero tolerance. Any deviation is treated as fraud.

Tier A holds inside a cluster, not across vendors. Consumer NVIDIA against consumer NVIDIA works, Hopper against Hopper works, and Apple Silicon against Apple Silicon works, because MPS was measured to be bit-deterministic within its own architecture. AMD against NVIDIA cannot work in Tier A for a physical reason: different kernel libraries emit different bytes for the same math.

Tier A is intended for cases where the strongest guarantee is worth restricting the miner pool, such as training step replay, Witnessed Checkpoint trajectories in Proof of Training, and high-stakes inference.

Tier B: Calibrated Tolerance​

Across different architectures, AETRON compares results using calibrated limits rather than an exact match. The comparison does not look at the generated text. It looks at the probability distribution the model produced before sampling, which is deterministic for a given input and configuration.

Tier B is the universal tier. All measured architectures fall inside one threshold set, so a Neuronet can mix NVIDIA, AMD, Apple Silicon, and CPU without any per-pair configuration.

What Tier B Compares​

Tier B is a composite check rather than a single number. Each signal is sensitive to a different kind of manipulation, so passing one does not help an attacker against the next. The values below are the ones currently set in the runtime.

SignalThresholdWhat it catches
KL divergence, p99 over probesbelow 0.05A substituted smaller model, a swapped top-1 token, a skewed sampling temperature
L2 distance over full logitsbelow 25Noise injected directly into logit space, which barely moves KL
Top-20 token overlap (Jaccard)above 0.9Perturbation of the distribution tail that leaves the top-K intact
Probability mass deviationbelow 0.12Redistribution of mass that keeps ranking order
Top-1 rank displacementat most 5The honest token falling out of the leading positions

The verdict trips if any signal fails. In the runtime these are stored as scaled integers (a KL threshold of 50000 means 0.05, an L2 threshold of 25000000 means 25), because the chain does not use floating point. Not every signal carries the same weight for every workload: L2 is a hard check for diffusion and advisory for language models, and top-20 overlap is advisory, since measured honest overlap can drop well below the nominal limit on short outputs.

Where the Thresholds Come From​

The limits are set from measurement rather than from a guess. Every supported architecture was run against every other one on the same model and the same probe set, and each threshold was placed well clear of the widest disagreement two honest machines produced. The measurement data behind them is not published.

That margin is what allows a single threshold set to cover the whole matrix instead of a separate configuration for each architecture pair, which is the practical difference between a network that can mix vendors and one that cannot.

Forming a K-Quorum from Mixed Hardware​

A single cross-architecture comparison is a noisy signal, and Tier B is tolerant by construction, so one tripped job is not evidence of fraud. Verdicts are decided over a quorum instead.

The protocol challenges K = 20 jobs per miner per batch and declares fraud at T = 10, meaning at least half of the sampled jobs must trip. Checkers are drawn by verifiable randomness from chain state, so neither the executing miner nor its checkers can steer the assignment. Nothing in the draw restricts the quorum to one architecture: an AMD miner can be checked by Apple Silicon and NVIDIA machines in the same batch, since Tier B thresholds hold for every pair.

That aggregation is what makes the tolerance safe in both directions. An honest miner sitting near the boundary would have to trip on half the batch rather than on one unlucky comparison, which keeps accidental penalties rare. A miner cheating consistently is caught reliably. Partial cheating is where any single signal is weakest, which is why the verdict rests on the composite rather than on KL by itself.

Bit-exact Tier A does not need this smoothing for false positives, but the quorum still applies, because it is what makes collusion expensive. See Proof of Intelligence for how quorum fits the wider verification flow.

INT8 and Quantization​

The obvious way to cheat on a Neuronet is to declare fp16 or bf16 and quietly run INT8, which is much cheaper. Quantization is not a subtle effect at the distribution level. It moves the output distribution by far more than honest cross-architecture noise does, so the two cases separate with a wide margin rather than a fine one, and the composite signals above carry the verdict.

Declared dtype is checked before any compute is spent as well, because an fp32 substitution shifts the distribution too little to be caught by distance alone.

Quantization is not treated as fraud in itself. Several quantized formats are valid declared configurations, including MXFP4 and native ternary weights. What matters is that the format is declared in the canonical spec and matches what actually ran.

Why It Matters​

  • Lower barrier to entry. A miner can join with the hardware it already has, including consumer GPUs and Apple Silicon.
  • No vendor lock-in. The network is not tied to one chip supplier, which makes it resilient to shortages and restrictions.
  • Larger, more decentralized network. More kinds of machines can participate, which spreads the work more widely.
  • Collusion is harder. A colluding group would need to control the hardware mix a random draw produces, not just one architecture.