Skip to main content

Troubleshooting

This page covers the problems miner operators hit most often, what each one looks like from the outside, and which command or log line tells you what is actually wrong. The reference client is split into a background daemon and a CLI that talks to it over a local socket, so almost every diagnosis starts the same way: ask the daemon what it thinks its own state is, then read the log. For installing and configuring the client first, see Installation and Configuration.

First Three Commands​

Before digging into any specific symptom, run these:

aetron-miner status
aetron-miner daemon status
aetron-miner engine status

status prints config path, daemon PID and uptime, the configured network, the chain backend actually in use, mining state, GPU, memory, Pulse tier, inference backend, and cache size. daemon status reports the pidfile, the socket path, and whether an RPC ping succeeds. engine status reports jobs_completed, jobs_failed, avg_latency_ms, last_error, and the Pulse counters. Between them they separate "the process is dead" from "the process is alive but doing nothing".

The daemon writes its log to ~/.aetron-miner/daemon.log. Stream it with aetron-miner logs. If that file does not exist, the daemon is logging to stderr instead, and you need to redirect it yourself or run it under a supervisor such as systemd.

The Daemon Will Not Start​

aetron-miner daemon start spawns the daemon detached, then waits for the pidfile and the socket to appear. If neither shows up in time it fails with daemon failed to start within <timeout>. Check ~/.aetron-miner/daemon.log. That message means the process was launched and then died, so the answer is in the log, not in the CLI output.

Things to check, in order:

  • Binary not found. aetron-miner-daemon binary not found means the CLI could not locate the daemon next to itself or anywhere on PATH. Install both binaries into the same directory.
  • Missing config. If status prints config: missing (...) - run \aetron-miner init``, the daemon has nothing to load. Run the init wizard.
  • A stale process. daemon already running with a PID means a live pidfile exists. Use aetron-miner daemon restart rather than starting a second one.
  • Release runner misconfigured. If a signed release runner is configured but the path does not exist, the daemon aborts rather than silently dropping to the development path. The log says the binary is configured but not found and names the path from AETRON_RUNNER_BINARY or runner_binary in the config. This is deliberate fail-closed behavior: a silent downgrade would let an attacker demote a tampered binary past signature checking.

If daemon status says RPC: socket exists but daemon not responding (probably dying), the socket file outlived the process. Run aetron-miner daemon stop (which removes it) and start again.

The GPU Is Not Detected​

The symptom is that status shows a CPU inference backend, or a Pulse tier far below what the hardware should reach, while the machine clearly has a GPU. Two causes account for most cases.

Sandbox versus CUDA. The daemon's sandbox blocks network egress at the syscall level on Linux. An overly broad filter that denies every socket() call also stops the CUDA driver from enumerating GPUs, which surfaces in the runner log as Error 304 (cudaErrorOSCall) together with torch.cuda.is_available()=False. The work then silently proceeds on CPU. This was fixed by narrowing the filter to internet address families only, so if you see those two lines together you are on an old build. The runner log prints the selected device as device=, which is the fastest way to confirm what it actually picked.

Driver against build. A runner bundle built for a newer CUDA major version falls back to CPU on a host with an older driver. The tier label comes from nvidia-smi and will still look correct, so trust device= in the runner log over the tier label.

Running on CPU when the task expects a GPU class is not just slow. Results computed on an unexpected architecture are more likely to fail verification, so treat a wrong device= as a hard problem rather than a performance issue.

The Miner Does Not Reach the Node​

The client is honest about this rather than pretending. The chain field in aetron-miner status shows one of three states: a real connection, a by-design simulation, or a fallback with the reason attached. If it shows the fallback, the reason string tells you which step failed:

  • AETRON_WALLET_PASSWORD not set: the daemon cannot unlock the keystore, so it has no key to sign with. Set the variable in the daemon's environment.
  • keystore unlock failed: ...: wrong password, or a keystore file that is missing, corrupt, or written by an unsupported version. aetron-miner wallet info confirms which keystore is being read.
  • connect failed: ...: the node URL is unreachable or rejected the connection. Check the URL in the config and that the node is accepting WebSocket RPC.
  • register failed: ...: the connection succeeded but registration did not, so submissions would fail against the pallet. Check the configured Neuronet and task IDs and the account balance.

A simulated backend appears only when the config asks for it. Both public networks run the PoI runtime, so testnet and mainnet connect for real, and a failure to reach them leaves the miner offline rather than quietly simulating. If you see a simulated backend without having configured one, you are on an old build. See Project Status for what is deployed.

No Peers Are Found​

Peer discovery uses a DHT that needs at least one entry point. The startup log warns explicitly when there is nothing to bootstrap from: no DHT seed and an empty peer cache means the miner never lands in anyone else's routing table, and a gateway cannot resolve its address. Set AETRON_BOOTSTRAP_ADDRS to one or more multiaddresses.

Common mistakes here:

  • A seed address without a peer ID. The log warns that a seed address lacking /p2p/<peer_id> was skipped. Multiaddresses for bootstrap must carry the peer ID.
  • Behind NAT with no relay. A separate warning fires when AETRON_RELAY_ADDRS is empty: a miner behind NAT is unreachable without a relay reservation, even though its own outbound connections work.
  • Peer cache disabled or unwritable. The cache lets a restarted miner rejoin without seeds. Its path comes from AETRON_PEER_CACHE_PATH; setting it to an empty value turns the cache off.

Once a miner has connected successfully at least once, the cached peers usually make later starts independent of the seed list. See Peer-to-Peer for how the layer is meant to behave.

A Model Will Not Download or Verify​

Downloads run through the daemon, not the CLI, so aetron-miner models download --from <url> fails immediately with a message telling you to start the daemon if it is not running. The daemon fetches manifest.json from the base URL and verifies every file against it before publishing anything into the cache.

The download is fail-closed at several points, and the error names which check tripped:

  • A file whose sha256 does not match the manifest aborts the whole download.
  • A file whose size does not match the manifest is rejected the same way.
  • A non-2xx response reports the HTTP status together with the URL that returned it.
  • A manifest entry with a suspicious path or artifact hash is rejected outright as a path traversal attempt.
  • An oversized manifest.json is refused rather than parsed, to avoid an out-of-memory condition.
  • When you pass --model-hash, the Merkle root computed over the manifest files must equal the expected on-chain model hash. A mismatch rejects the download, which is what stops a fake mirror from serving substituted weights.
  • Re-downloading an artifact that is already cached is refused. Delete it first.

For inspecting what is already on disk, aetron-miner models list shows each artifact's status. Corrupt carries a reason, Partial means an interrupted download with the byte count so far, ManifestOnly means metadata without the bytes, and MissingManifest means files with no manifest to check them against. aetron-miner models verify <prefix> does a metadata check by default and --full re-hashes every file in the manifest, reporting each mismatching path. Hash prefixes must be at least 6 hex characters, and an ambiguous prefix is reported with the list of artifacts it matched.

Jobs Are Not Arriving​

If the daemon is up and engine status shows jobs_completed stuck at zero, work through this order.

  1. Is the engine even running? aetron-miner daemon start starts only the daemon. aetron-miner start starts the daemon and then the engine. Check the mode in engine status, and that the engine is not paused.
  2. Read last_error. The engine snapshot carries the last failure it saw, which usually names the stage that broke.
  3. Check the chain backend. A simulated backend means registration never reached a real network, so no real task will ever be dispatched to you.
  4. Check discovery. If nothing can resolve your address, requests cannot arrive no matter how healthy the local process is. See the peer section above.
  5. Check jobs_failed. Jobs arriving and failing is a different problem from jobs never arriving, and it usually points at the runner rather than the network.

A related failure mode is the Pulse loop rather than the job loop. Fast Pulse stages have a 15 second timeout and the proof stage has 60 seconds. If the memory-hard Pulse table cannot be built inside the fast timeout on a high-memory tier, the daemon concludes the runner has hung, respawns it, and the engine ends up in an error state repeating that cycle. Repeated Pulse timeouts in the log with the engine never reaching steady state is the signature.

Verdicts Disagree​

A verification verdict is a binary Match or Mismatch that a checker commits and then reveals. Disagreement is expected occasionally and is designed for: a miner's outcome is decided over a quorum of independently drawn checkers rather than by any single comparison, and the runtime is currently set to a quorum of 20 with a threshold of 10, so a fraud verdict needs at least half the quorum to report a mismatch. One stray mismatch against you does not decide anything.

If your own checks keep disagreeing with the rest of the quorum, the cause is usually local rather than protocol level:

  • Wrong device. Work that fell back to CPU produces results from a different architecture than the one you declared, which pushes comparisons outside the calibrated tolerance. Confirm device= in the runner log.
  • Wrong model or precision. Verification compares against the task's declared execution parameters. Serving a different quantization or precision than the task fixes will not verify. See Execution Spec.
  • A stale or partial model cache. Run aetron-miner models verify <prefix> --full on the model the task uses.
  • Environment drift in the runner. A runner missing a module or running a mismatched library version fails in ways that look like disagreement rather than a crash. Run aetron-miner env verify first, then read the runner log; a self-test that no longer passes explains a whole class of mismatches at once.

For what the comparison actually measures and why cross-architecture results are compared with tolerance instead of bit equality, see Inference Verification and Heterogeneous Mining.