Troubleshooting¶
Symptom → root cause, using the E_* classes. The full inventory with
sources is the error inventory; the wire
envelope is specified in the HTTP API.
Daemon won't start / CLI can't connect¶
$ shardr models
shardr: E_DAEMON_UNREACHABLE: shardhive daemon not reachable — is it running? (…)
- Is
shardhive serverunning? Check the socket:ls -l ${SHARDR_SOCKET:-${XDG_RUNTIME_DIR:-${TMPDIR:-/tmp}/shardhive-$(id -u)}/shardhive.sock} shardhiverefusing the socket path (refusing to remove regular file/directory/symlink at socket path …): a wrongSHARDR_SOCKETpoints at something that is not a socket. The daemon never deletes it — unset or fix the variable.socket … is in use by another shardhive: a second daemon on the same socket; pick another--socket.
Config errors at startup (loud, before the daemon serves)¶
$ shardhive serve
shardhive: config ~/.config/shardr/config.toml: unknown [swarm] key "bogus" (known: enabled, seed, upload_limit, dht, no_seed_verify, webseed_addr)
Unknown [swarm]/[catalog] keys, wrong value types, and malformed
lines are startup failures by design — a typo must not quietly
disable seeding. See the
configuration reference for the
valid keys.
Refused imports¶
E_SOURCE_NOT_REGULAR — a local import path contained a symlink,
FIFO, or device (executed):
$ curl -s --unix-socket "$S" -X POST http://localhost/v1/import/local \
-d '{"paths":["/tmp/evil-link.gguf"],"as":"gold/toy"}'
{"error":{"code":"E_SOURCE_NOT_REGULAR","message":"importer: import root is a hard boundary — source is not a regular file (symlink/fifo/device); refusing fail-open (E_SOURCE_NOT_REGULAR): /tmp/evil-link.gguf"}}
The import root is a hard boundary — resolve the symlink target or import the real file.
E_NOT_ANCHORED (catalog) — the listing's Hugging Face source is
gone (rescued). Without --trust-catalog / trustCatalog: true the
pull refuses; with it, digests verify against the catalog provider's
own checksums (a trust shift, printed as a warning — see
Catalog pulls).
E_SOURCE_FORBIDDEN (HF) — gated repo without a token, or a
nonexistent repo id when anonymous (HF answers 401/403 for both). Set
SHARDR_HF_TOKEN in the daemon's environment and re-check the repo
id.
Partial fills / unreachable sources¶
E_SOURCE_UNAVAILABLE with "swarm client disabled" — fills,
/import/bt, and catalog pulls need [swarm] enabled = true.
E_SOURCE_UNAVAILABLE with "manifest … is not local" — the name
resolves, but no import ever brought the manifest blob in. Ensure by
importing (local/HF/catalog/bt) first.
Catalog pull fails after two minutes with "no torrent metadata within 2 minutes" (executed against a live listing):
{"error":{"code":"E_INTERNAL",
"message":"swarm: catalog: no torrent metadata within 2 minutes — the swarm is unreachable (seeders online? tracker up?); a magnet alone never carries the file tree"}}
Torrent metadata comes from peers, not from the webseed. A listing with zero seeders online cannot deliver its file tree — the webseed only serves file bytes after metadata. Retry later or pick another listing; byte integrity is never at stake (the anchor gates every file).
Known gap (tracked on issue #65, post-merge follow-up): after the
metadata phase, a catalog job has no stall detection — a peer
delivering metadata but no bytes leaves the job fetching
indefinitely (context.Background() in the catalog import handler).
If a catalog job sits in fetching far beyond any plausible transfer
time, restart the daemon: jobs live in memory only, the CAS and state
are unaffected, and already-sealed files stay verified.
Wrong quant / selector errors¶
E_NO_MEMBER— the repo's index has no such member. List what exists:shardr models(executed:gold/smollm2-135m-instruct index present q4_k_m,q4km,raw).E_AMBIGUOUS_SELECTOR— the prefix matches several members;candidateslists them; spell the full quant.- Model imported as
rawalthough the filename saysQ4_K_M— the quant vocabulary is lowercase-only; an uppercase token classifies asraw(see the table in the URI grammar page). Rename the file to a lowercase quant form and re-import; identical bytes converge into the existing index.
Integrity problems¶
E_CORRUPTION (a present blob is not valid), E_NO_INDEX,
E_INVALID_INDEX — enumerate with the re-hash (executed):
$ shardhive cas verify --all
verify --all: 0 mismatched, 0 missing, 0 state errors
all blobs clean
Exit codes: 0 clean, 1 mismatch, 2 missing. Corrupt or missing
blobs are re-fetched by an ensure/import of the affected artifact —
nothing is ever silently rebuilt.
Runner lifecycle failures¶
shardr stop enforces the supervisor duty (002 §4): SIGTERM, clean
exit within 30 s, then SIGKILL — idempotent, with pid identity
verification before signalling (executed):
$ shardr stop docs-e2e
stopping docs-e2e (pid 85034)…
stopped docs-e2e
E_STATE: refusing to stop …— the pid fails identity verification (start token or served-ref mismatch): a stale registry entry or pid reuse. The message names the state file; verify and clean it manually — the runner refuses to signal a process it cannot prove is its own.- Foreground
shardr runhangs after Ctrl-C — SIGTERM is sent; the supervisor waits up to 30 s before SIGKILL. A second Ctrl-C does not skip the window. - llama-server not found (
E_BINARY) — the runner resolvesllama-servernext to theshardrexecutable (bin/llama-server, a symlink into the pinned extract dir); the message names all remedies. Runmake llama(part ofmake all), or point$SHARDR_LLAMA_SERVERat a binary, or install llama-server on$PATH.