Skip to content

Deploy the training service ​

This guide is for operators deploying a JevK5 training service or maintaining its approved miner pool. Customers using hosted Zils can follow the training workflow without creating a Supabase project or running these services.

The coordinator handles customer authorization and miner assignments. Supabase provides authentication, job state, and private object storage; dataset and checkpoint files transfer directly to storage using signed URLs. A separate processor validates data and evaluates candidates. Approved miners claim jobs over outbound HTTPS and train candidate adapters.

These instructions use JevK5 4B. For older jobs or recorded pilot reproduction, use the legacy Kev queue guide. Training and evaluation do not by themselves make a prediction endpoint ready; configure automatic assignment and activation and the decision API deployment for accepted-model serving.

Prerequisites and isolated project resources ​

Start with Git, uv, and Python 3.13 on Linux or WSL 2. Complete Install and create the reference from a fresh clone. The processor and miners require a BF16-capable NVIDIA GPU, a compatible CUDA driver, and the verified models/jevk5-reference; CPU and MPS are unsupported for this runtime. Allow at least 25 GB free disk for the model and packages, plus space for datasets and candidates. The coordinator API needs no GPU or model checkpoint.

Use a Supabase project with Auth and Storage enabled, and configure Auth for the intended customer accounts and redirect URLs. For an existing deployment, reuse its configured resources rather than applying first-install steps again. Choose a Storage plan/global file limit compatible with the configured limits: 128 MiB per uploaded dataset split and 512 MiB per artifact, with a 512 MiB aggregate checkpoint limit. Project-wide limits may be lower than bucket limits.

Review and apply supabase/migrations/202609300001_training_jobs.sql with the Supabase SQL editor or a privileged PostgreSQL connection:

Bash
psql "$SUPABASE_DB_URL" -v ON_ERROR_STOP=1 \
  -f supabase/migrations/202609300001_training_jobs.sql

SUPABASE_DB_URL is a server-side connection string supplied by the operator; never put it in the web app or a miner configuration. Apply this migration once. It creates only the fez_training_* tables/functions, private fez-training-data and fez-training-models buckets, and policies protecting those resources. It neither resets the project nor changes unrelated data. A restrictive Storage policy prevents existing broad client policies from exposing the training buckets. Service-role access remains privileged.

Customers can select only their own job records and cannot directly change job state. Miner registrations, assignments, replay nonces, and queue RPCs are server-only. The coordinator verifies each customer session against Supabase Auth. Miners instead sign requests with Bittensor hotkeys; they receive no Supabase project credentials.

Run the API and processor ​

From the repository root, create the protected environment file if it does not already exist:

Bash
mkdir -p .private
cp .env.example .private/training.env
chmod 600 .private/training.env

Edit the copied file with the project's URL and server service-role key. Set ZILS_TRAINING_API_URL to the coordinator's reachable URL and ZILS_WEB_ORIGIN to the exact browser origin. Set ZILS_TRAINING_MODEL=jevk5-4b-v0.3 in both the API and processor environments. The example uses loopback for local development; a self-hosted service needs its own public URL. Approved participants in hosted Zils use https://training.zils.ai and operator-provided access. Load the configured file in each coordinator/processor terminal:

Bash
set -a
. .private/training.env
set +a
.venv-kev/bin/python -m zils.coordinator serve

In another terminal with those same environment variables:

Bash
.venv-kev/bin/python -m zils.coordinator process \
  --state .private/queue-processor --reference models/jevk5-reference --device cuda

--once processes at most one pending stage and exits. The API itself needs no GPU or checkpoint; the processor does. They can run on separate machines with the same Supabase project and API URL. Processors can recover abandoned database leases; each processor directory has an exclusive process lock.

The API listens on 127.0.0.1:8910 by default. For remote operation, run it behind an HTTPS reverse proxy with request/concurrency limits and configure the exact public URL on both coordinator and miners. The built-in threaded HTTP server is a development server; it is not an internet edge server. A persistent host, such as a DigitalOcean host managed using doctl, can run the API under a service supervisor. The processor belongs on the configured CUDA model host. No DigitalOcean resources are required or created by this repository.

Connect the web app ​

In the independent Zils website, configure:

Environment
NEXT_PUBLIC_SUPABASE_URL=https://your-project.supabase.co
NEXT_PUBLIC_SUPABASE_PUBLISHABLE_KEY=replace-with-public-publishable-key
NEXT_PUBLIC_ZILS_TRAINING_API_URL=https://your-training-api.example

The web app also supports NEXT_PUBLIC_SUPABASE_ANON_KEY as a legacy fallback. Only public keys belong in NEXT_PUBLIC_*. Configure the Supabase Auth redirect allowlist to include the actual /train URL, and set the coordinator's ZILS_WEB_ORIGIN to that site's origin. Redeploy/rebuild the web app after setting public environment variables. The /train dashboard is implemented in the web app repository, not this repository's standalone benchmark preview.

The customer training workflow covers data preparation, export consent, uploads, evaluation, and API readiness. Customers sign in through the dashboard; prediction API keys do not replace their training-service session.

Approve miners and run a queued miner ​

A job enters awaiting_approval after validation. It does not expose data to all registered miners. With the coordinator environment loaded in an operator terminal, register an approved miner's public hotkey and an unused stable UID. Replace the placeholders below; 1 is an example local UID:

Bash
export MINER_HOTKEY='REPLACE_WITH_PUBLIC_SS58_ADDRESS'
export JOB_ID='REPLACE_WITH_JOB_UUID'
.venv-kev/bin/python -m zils.coordinator worker --hotkey "$MINER_HOTKEY" --uid 1
.venv-kev/bin/python -m zils.coordinator approve \
  --job "$JOB_ID" --hotkeys "$MINER_HOTKEY"

MINER_HOTKEY is the miner's public SS58 address; JOB_ID is the UUID displayed in the dashboard. Supply up to sixteen distinct approved hotkeys. UIDs are local queue identities, not a claim of chain registration. IDs are frozen in each assignment. worker --disable disables future authenticated miner requests; previously issued download URLs remain valid until expiry and downloaded data cannot be recalled.

On the miner, complete the JevK5 reference setup and provision its signing wallet. Follow the wallet creation steps in Bittensor registration. Queue-only miners do not need funding or chain registration. Install the wallet SDK in the model environment and create the private configuration directory:

Bash
uv pip install --python .venv-kev/bin/python -r requirements/testnet.txt
mkdir -p .private

Create .private/queue-miner.json with the miner's existing wallet name, hotkey name, public SS58 address, and the deployment's coordinator URL. For hosted Zils, use https://training.zils.ai; replace the example URL for your own deployment:

JSON
{
  "coordinator": "https://your-training-api.example",
  "hotkey": "REPLACE_WITH_PUBLIC_SS58_ADDRESS",
  "wallet": {"name": "REPLACE_WITH_WALLET_NAME", "hotkey": "REPLACE_WITH_HOTKEY_NAME"}
}

Save it as .private/queue-miner.json and run:

Bash
chmod 600 .private/queue-miner.json
.venv-kev/bin/python -m miner.queue \
  --config .private/queue-miner.json --state .private/queue-miner \
  --reference models/jevk5-reference --device cuda

For a disposable local test identity, a configuration may use seed instead of wallet and hotkey; generate it with Keypair.create_from_seed from a securely generated 32-byte seed, keep it private, and register the resulting public address. Never give miners the Supabase service-role key. Each miner downloads and hashes its assigned training export, reuses the cached base model, trains a candidate, and uploads only the three checkpoint files. Saved candidates survive miner restarts. If the miner and processor share a GPU, use separate service identities and the shared compute lock. The same running miner can subsequently claim another customer's job; no new fleet bundle is required.

Completion, failures, and limits ​

Jobs transition through uploading → validating → awaiting_approval → queued → running → evaluating → completed, or failed. There are at most five active jobs per customer. Miner leases last twenty minutes and renew every minute; requests are signed for the exact API URL, route, body, nonce, and timestamp. Expired claims cannot submit with an old token. Each assignment permits up to three attempts. A job is evaluated when all assignments finish/fail, or after its twenty-four-hour deadline. There is no additional score reward for accepting more jobs or signing repeated requests.

The processor compares candidates against the calibrated starting checkpoint, using the JevK5 acceptance gate. For an explicit version upgrade, the comparison instead uses the pinned previous customer model and its existing serving temperature on the new test data. No qualifying candidate leaves the current API version intact. It uploads accepted artifacts and a release manifest to private Storage. The customer receives aggregate results and short-lived download URLs, never other customers' records or raw validator predictions. The release still requires the pinned base model to run. completed can mean no_qualifying_model; this is a valid experimental result. Job weight vectors are diagnostic within-job scores and are not sent to Bittensor by the queued workflow.

Storage/network interruptions can be retried after lease recovery. Invalid data or model/runtime failures produce a failed job for operator inspection. Local private processor state contains the evaluation reports/logs. Cancellation stops new claims and fences processing completion; it cannot erase data already downloaded or immediately terminate remote compute. For a transient failed job, inspect the cause before creating a new one. Do not reset database states by hand without understanding lease ownership.

Training data is readable by approved miners, who may retain it. Signed URLs limit access, not the use of downloaded bytes. Files and raw local runs are retained until operator cleanup; automatic retention/deletion, billing, resumable multipart uploads, and confidential compute are not implemented. The automatic workflow can activate accepted artifacts through the separate prediction runtime after verifying their identity. The separate Zils decision API provides shared-model inference and bulk processing with its own credentials, queue, and retention controls. The initial upload path uses direct PUT; retry restarts a failed file transfer.

Verification ​

Bash
make check
make check-queue-db

The second command requires PostgreSQL 16+ binaries (initdb, pg_ctl, psql). Set ZILS_PG_BIN to their directory if they are not on PATH. It creates and deletes a separate temporary database cluster, never connects to your existing Supabase database, and checks migration execution, tenant/storage isolation, service-only functions, replay rejection, lease recovery, and bounded attempts. CI runs both suites. HTTP tests exercise real signatures and file transfers with fixture miners; they do not contact Supabase or send chain transactions.

Implementation references: Supabase database functions, row-level security, and signed uploads.