Quickstart¶
From a fresh checkout to a chat completion. Every command below was executed for real on macOS (Apple Silicon); outputs are the real ones, paths generalized. Linux/x86-64 differs only in the fetched runtime platform.
1. Build everything¶
make all builds both binaries and fetches the pinned llama.cpp
runtime (digest-verified, from runtime/llama.lock — see
Runtime & releases):
$ make all
>> building shardr + shardhive into bin
go build -o bin/shardr ./cmd/shardr
go build -o bin/shardhive ./cmd/shardhive
>> fetching prebuilt llama.cpp b10684 (darwin_arm64)
fetched darwin_arm64 -> bin/llama-b10684 (sha256 8310138372444cbe…)
ln -sf llama-b10684/llama-server bin/llama-server
>> bin/llama-server ready (pin: b10684)
(On an unauthenticated machine the GitHub release API may answer 403 —
export a token: GITHUB_TOKEN=$(gh auth token) make all.)
2. Start the daemon¶
$ ./bin/shardhive serve
shardhive: swarm: seeding 1 artifact(s)
shardhive 0.0.1-dev listening on /var/folders/…/T/shardhive-501/shardhive.sock
In a second shell the CLI speaks to it over the Unix socket (the socket permission is the access boundary — mode 0600). A fresh store is honest about being empty (executed against a fresh CAS):
$ ./bin/shardr models
no models imported yet — shardr import local <paths> --as ns/name
3. Import a model¶
Local import (namespace is mandatory — files must be regular files;
filenames drive quant classification, so keep quant tokens lowercase;
executed with the pinned SmolLM2-135M Q4_K_M weights, digest
ed5fa30c487b282e…):
$ ./bin/shardr import local /tmp/toy-model --as gold/smollm2-135m-instruct
import started: job 4bce521b40d9ea6e
fetching 0/1 fetching 1/3 done 1/1
quant q4_k_m
done
The same flow works over the API (POST /v1/import/local), from
Hugging Face (shardr import hf <repo> — revision-pinned to the
resolved commit), or as an anchored
catalog pull (shardr catalog search <terms> →
shardr pull <owner/repo>). Note the operational reality of catalog
pulls: torrent metadata comes from peers, so a listing with zero
seeders online fails loudly after two minutes — see
Troubleshooting.
4. Check the inventory¶
$ ./bin/shardr models
gold/smollm2-135m-instruct index present q4_k_m
(Re-importing the same bytes under other names converges into the same index and adds the derived quants as members — see the table in the URI grammar.)
5. Run the model¶
$ ./bin/shardr serve shardr:///gold/smollm2-135m-instruct:q4_k_m --id docs-e2e
resolve shardr:///gold/smollm2-135m-instruct:q4_k_m
spawn <worktree>/bin/llama-server
weights <cas>/blobs/sha256/ed/5fa30c48… (zero-copy mmap)
serving shardr:///gold/smollm2-135m-instruct:q4_k_m
id docs-e2e
endpoint http://127.0.0.1:63441
model shardr:///gold/smollm2-135m-instruct:q4_k_m
The weights path is the CAS blob itself — llama-server mmaps it
directly, no copy (zero-copy serving). shardr run <ref> is the same
flow in the foreground (endpoint from llama-server's own startup log);
shardr stop docs-e2e stops a background instance (SIGTERM, SIGKILL
after 30 s).
6. Chat completion (OpenAI-compatible)¶
$ curl -s http://127.0.0.1:63441/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"smollm2","messages":[{"role":"user","content":"Name the largest planet in our solar system. One word."}],"max_tokens":10}'
{
"choices": [
{
"finish_reason": "length",
"index": 0,
"message": {"role": "assistant",
"content": "The largest planet in our solar system is Jupiter,"}
}
],
"model": "shardr:///gold/smollm2-135m-instruct:q4_k_m",
"system_fingerprint": "b10684-cc83d7b48",
"usage": {"completion_tokens": 10, "prompt_tokens": 42, "total_tokens": 52}
}
The response names the canonical reference that was served and the runtime fingerprint (llama ref + upstream commit from the pin) — what answered is traceable to what was stored.
7. Verify the store (optional, explicit)¶
$ ./bin/shardhive cas verify --all
verify --all: 0 mismatched, 0 missing, 0 state errors
all blobs clean
Next: the concept tour for how the pieces fit together, or the references for the full API.