TEHAS

Commands

On the host as the tehas user (in the example, boss on tehas-tehas). Interactive shells get the environment from ~/tehas/env.sh; a non-interactive ssh boss@tehas-tehas '<command>' must source it first. The scripts live in ~/tehas/bin, installed from bin/ in the platform repository.

Every command loads the instance settings first and stops with exit 78 and one line if a setting it needs is missing or malformed. See The host.

Instance settings

Set in instance.env in the instance directory. instances/example/instance.env in the platform repository has every one with a comment.

Setting Default Meaning
TEHAS_REPO required owner/name of the repository whose issues are the queue
TEHAS_CONTROL_ISSUE required Issue that carries tehas:pause
TEHAS_ALLOWED_LABELERS required Logins whose tehas:ready counts, space separated
PI_ENSEMBLE_SUBAGENT_MODEL required Model for every subagent role
PI_ENSEMBLE_SUBAGENT_PROVIDER required Its provider, as named in models.json
TEHAS_HOST_ADDR 127.0.0.1; required with previews Address the docs server binds to, and with previews the edge
TEHAS_WORK_REPO ~/tehas/repos/<name> The product checkout cycles run in. Must be a clone of TEHAS_REPO
TEHAS_BRANCH the checkout's origin/HEAD Branch cycles start from and pull requests target
TEHAS_WORK_CPUS the host's CPUs minus one, at least 1 docker --cpus of a cycle's container
TEHAS_WORK_MEMORY half the host's memory docker --memory of a cycle's container
TEHAS_CONCURRENCY 1 Workstreams side by side inside one cycle
TEHAS_MAX_AWAITING 3 Open tehas:awaiting-merge issues allowed before holding
TEHAS_WORK_TIMEOUT 21600 Seconds before a cycle is given up
PI_ENSEMBLE_HOST_ALIASES empty name:address pairs the sandbox must resolve, comma separated
PI_ENSEMBLE_IMAGE tehas/pi-rukas:uid<uid> The sandbox image
TEHAS_PREVIEWS 0 1 installs and runs the staging edge, stacks and previews
TEHAS_CLONES 0 1 installs the database clone and refresh commands
TEHAS_MAX_PREVIEWS 3 Pull request previews allowed at once
TEHAS_STACK_BOOT_TIMEOUT 1800 Seconds a stack may take to become healthy
TEHAS_STACK_DOMAIN localhost Domain the dashboard links stacks under when the route table cannot be read

The example instance sets the five required ones, TEHAS_HOST_ADDR and PI_ENSEMBLE_HOST_ALIASES, and leaves the rest at the defaults.

tehas-dispatch

One poll of the issue queue. Deterministic, no model.

tehas-dispatch --dry-run       # report the decision, change nothing
tehas-dispatch                 # run at most one cycle, then return
tehas-dispatch --init-labels   # create the tehas labels in the repository

Order of rules: pause label on the control issue, the cap on issues awaiting merge, who applied tehas:ready, then urgent first and lowest number. Log: ~/tehas/logs/dispatch.log when run by systemd, the terminal otherwise. Each cycle's runner output: ~/tehas/logs/dispatch-work-<issue>.log.

Before a cycle it puts the product checkout on the default branch and fast-forwards it. With previews on it also builds a preview for each pull request a cycle opens and keeps the standing stack current.

tehas-work

One pi-rukas /work cycle, headless, in the sandbox image.

tehas-work <issue> [--restart]

It starts pi --mode rpc --approve in a container, sends /work <issue> and waits until .pi/work-state/<issue>.json in the checkout leaves running.

Exit Meaning
0 The cycle ended as merged or handoff (parked for a human)
1 The cycle ended as aborted
2 Pi exited, or the time limit passed, while the cycle was running
3 The sandbox cannot push to the repository; no cycle was started
64 No bot token or no git identity; no cycle was started
78 The instance settings are invalid; no cycle was started

The container gets the bot token and the PI_ENSEMBLE_* settings, a CPU and memory limit, and no Docker socket or SSH keys. git inside it is given the bot's identity and a credential helper that reads the token, through environment-level configuration; before a cycle starts, a dry-run push proves that works. The RPC stream is kept in ~/tehas/logs/work-<issue>-<time>.jsonl. The limits are TEHAS_WORK_CPUS, TEHAS_WORK_MEMORY and TEHAS_WORK_TIMEOUT.

The bot's git identity is the tehas user's global user.name and user.email on the host.

tehas-status and tehas-console.ts

Feed the factory floor and the plain dashboard. tehas-status writes the queue snapshot to ~/tehas/live/status.json; the dispatcher runs it at every poll and every two minutes during a cycle. tehas-console.ts is started by tehas-work and follows the cycle's agents into ~/tehas/live/console-<issue>.txt, masking tokens and passwords. Neither needs running by hand.

tehas-health

tehas-health        # prints nothing; refreshes ~/tehas/live/health.json
curl -i http://tehas-tehas:8080/health

Writes the factory's health from what is already on disk: the queue snapshot and the current or last cycle. tehas-status runs it on exit, so it is refreshed at every poll and every two minutes during a cycle. The docs server serves the file at /health and /healthz as application/health+json.

status Meaning HTTP
pass Taking work, idle or in a cycle 200
warn Paused (tehas:pause on the control issue) 200
fail The last cycle ended as anything but merged or handoff, or there is no queue snapshot 503

output says the same in a sentence. checks.queue has the snapshot time and the counts in line, building, in review and stopped; checks.cycle has the issue, status, step and times of the current or last cycle.

A stopped dispatcher cannot report itself: the file simply stops changing and the endpoint keeps answering with the last state. A monitor should also treat a time older than about ten minutes as down.

tehas-energy

tehas-energy            # one sample; prints nothing; writes ~/tehas/live/energy.json
tehas-energy --every    # one sample at every 15-second mark; what the service runs
systemctl --user enable --now tehas-energy.service

Asks the model mesh for its cluster snapshot and keeps what the Energy use section shows: per machine the draw, kWh and RAM; per model the slots and how many are busy; and the last 30 minutes of draw. Only machines that are alive and serve a model the factory uses are kept, and only those models. The factory's models are the ones named by PI_ENSEMBLE_SUBAGENT_MODEL and PI_ENSEMBLE_MODEL_*; with none set, the file has no machines. It works only against a viiwork mesh, which it recognises by the answer.

The mesh answers The file
A cluster snapshot Rewritten
Nothing, a timeout, a server error or a busy answer (408, 429) Left as it is; the page calls it stale
Anything else Removed, and the section with it

The mesh is the one the factory's model provider points at. Set TEHAS_ENERGY_URL in ~/.config/tehas/env to ask another cluster URL, and TEHAS_ENERGY=off to remove the file. The file is readable by anyone who can open the dashboard and names the machines, their draw and their models.

The service reads its settings when it starts: run systemctl --user restart tehas-energy.service after changing either one. Stopping the service does not hide the section; the last file stays and the page calls it stale. To hide it, set TEHAS_ENERGY=off and restart, or stop the service and delete ~/tehas/live/energy.json.

Against a backend that is not viiwork the service still asks <provider>/v1/cluster every 15 seconds and writes nothing. On such a host leave the service disabled, or set TEHAS_ENERGY=off.

The gates (run by the cycle, usable by hand)

A product names its gates in .pi/verify-cmd, .pi/verify-cmd-full and .pi/smoke-cmd: the first line that is not empty or a comment is the command. Run them by hand from the root of the product checkout. For the example product:

scripts/verify.sh          # fast: the lint, shellcheck, the shell tests
scripts/verify.sh --full   # plus deno lint over the TypeScript

install.sh

Installs tehas on the host from the platform checkout, for the instance the host points at. See The host.

cd ~/tehas/platform
./install.sh --check     # change nothing; Instance, Dependencies, Installed, Tests
./install.sh

Exit 0 when done or nothing to report, 1 when --check found something or a step failed, 78 when the instance settings are invalid (nothing was changed).

bootstrap.sh

Prepares a fresh host with everything tehas and pi-rukas need, at the versions in host/versions.env. See Bootstrap a host.

sudo ./bootstrap.sh root boss   # system packages, Docker, docker group, lingering
./bootstrap.sh user             # the rest, under the user's home
./bootstrap.sh check            # change nothing; found and expected version per tool

tehas-test

The checks of the platform itself. Run from the platform checkout.

bin/tehas-test

It runs shellcheck over the shell scripts, then every test/*_test.sh. Those need bash, git, jq and shellcheck only: no sandbox image, no product checkout and no network. On an instance with previews it also runs the Deno tests of the staging edge, which import the product's route table. Exit 0 when every check passed, 1 when one failed, 2 when nothing could be checked.

scripts/check-generic.sh is the lint that keeps one product's names out of platform files. scripts/verify.sh runs it and bin/tehas-test together.

Only with previews or clones

These commands are installed only when the instance turns the feature on. They belong to one product's stack. Staging describes them in use.

tehas-stack

A private copy of the product for the default branch or one pull request.

tehas-stack up dev | pr-123 [<ref>]
tehas-stack smoke <name>
tehas-stack urls <name>
tehas-stack down <name>
tehas-stack list

dev is the name of the standing stack; it follows the product's default branch. up builds the images first and only then replaces a running stack, so a broken build leaves the previous one serving. A build that fails is retried once on its own, because some tests are timing-sensitive while several builds compete. Logs: ~/tehas/stacks/<name>/build-<service>.log, docker logs tehas-<name>-api. The smoke table is kept in ~/tehas/stacks/<name>/smoke.md.

tehas-probe

Load a page in headless Chromium, print a JSON verdict, optionally screenshot. It runs the probe script from the product checkout inside the sandbox image.

tehas-probe <url> --shot page.png
tehas-probe <url> --shot page.png --mobile --full-page
tehas-probe <url> --marker "text that must be in the page"

Screenshots land in ~/tehas/shots. The verdict has status, finalUrl, title, canonical, h1, console errors and failed requests. Exit 0 when the page is ok, 1 when it loaded but failed a check, 2 when it did not load.

gen-routes.ts

The staging edge's route table for one stack, generated from the product's tenant registry. Run from the product checkout.

deno run -A --no-config ~/tehas/edge/gen-routes.ts --api HOST:PORT --fe3 HOST:PORT \
  --fe4 HOST:PORT [--suffix pr-123] [--frontend prod|v3|v4]
deno run -A --no-config ~/tehas/edge/gen-routes.ts --format json [--suffix pr-123]

check-no-tehas-hosts.sh

A script in the product repository, not in this one. It fails if any file outside the tehas tooling names a .tehas host, and is part of that product's verify gate and release path.

pg-clone

Throwaway copies of the nightly snapshot database.

pg-clone create pr_123     # seconds for 100 GB; prints the DATABASE_URL
pg-clone url pr_123
pg-clone list              # name, size, age in hours, connections
pg-clone drop pr_123
pg-clone gc 72             # drop clones older than 72 hours
pg-clone init              # once per host: the role and its password file

A clone is a reflink copy, so it shares disk with the snapshot until written to. The role tehas_test owns its clones and can read nothing else; it is the only database credential agents get. The script touches only databases with its own clone prefix. Creating a clone fails while the snapshot has open connections or the nightly refresh is swapping it in. The server is reached at TEHAS_HOST_ADDR, port 5433, unless TEHAS_PG_HOSTPORT says otherwise.

pg-refresh

Restores the newest production dump into the snapshot database. A systemd user timer runs it at 03:30 UTC; it takes about 80 minutes. Status: ~/tehas/state/pg-refresh.json. By hand: pg-refresh or pg-refresh --force.