What pi-tehas does
pi-tehas turns GitHub issues into reviewed pull requests, using pi-rukas agents on self-hosted models. A person writes the issue, applies one label, and later reviews and merges. Everything in between is tehas. An instance can also give each pull request a running preview; that part belongs to one product so far.
It is a set of shell scripts, two small Deno programs and a few containers on
one host. There is no server process of its own and no database of its own:
state lives in GitHub labels, in pi-rukas's state files and in plain files under
~/tehas.
Each feature below says how far it has been proven. "Verified" means it was run on a tehas host and the result checked. Unless a row says otherwise, that was the first instance, before settings moved into an instance directory. Nothing here has run for longer than a day.
Queue and dispatch
| Feature | Detail | Status |
|---|---|---|
| GitHub issues as the only control plane | Labels are the state machine: tehas:ready, tehas:running, tehas:awaiting-merge, needs-human-attention |
Verified |
| Authorization by label | A cycle starts only if tehas:ready was applied by an allowlisted account. A label applied by the bot or anyone else is ignored and logged |
Verified |
| Polling | A systemd user timer runs one poll two minutes after the previous one ended. A poll is a few GitHub calls and no model call | Verified |
| One cycle at a time | Urgent first (tehas:urgent), then lowest issue number |
Verified |
| Pause switch | tehas:pause on the control issue stops new cycles; a running one finishes |
Verified |
| Review backpressure | No new cycle while 3 issues await merge | Built, not yet reached |
| Mesh check | The queue is held, not burned, while the model endpoint does not list the configured model | Built after a real failure; the hold path itself is untested |
| Retry | Removing needs-human-attention and re-applying tehas:ready starts a fresh cycle |
Verified |
| Reporting | Start and outcome comments on the issue; token usage per agent role on the pull request | Issue comments verified; the pull request comment awaits the first cycle that opens one |
The work cycle
| Feature | Detail | Status |
|---|---|---|
| Headless pi-rukas | /work <issue> is sent to pi --mode rpc inside the sandbox container; no terminal needed |
Verified |
| Sandbox | Agents run in a container with CPU and memory limits, the bot token and nothing else: no Docker socket, no SSH keys, no host environment | Verified |
| Project gates | The cycle runs the project's own .pi/verify-cmd, .pi/verify-cmd-full and .pi/smoke-cmd |
Gates verified in the sandbox; first full cycle in progress at the time of writing |
| Protected gates | pi-rukas halts a cycle whose patch touches .pi/, .github/ or AGENTS.md |
pi-rukas feature, not exercised here |
| Human merge | Nothing merges on its own | By configuration (no merge grant in AGENTS.md) |
| Self-hosted models only | One model for every role, set in the instance's instance.env |
Verified |
| Concurrency knob | TEHAS_CONCURRENCY sets how many workstreams a cycle runs side by side; 1 by default, raise it when the model backends have capacity |
Setting reaches the sandbox; its effect on a cycle is not yet observed |
Staging
Only on an instance with TEHAS_PREVIEWS=1 and TEHAS_CLONES=1. Off by
default. See Staging.
| Feature | Detail | Status |
|---|---|---|
| Fake domains on the tailnet | Every tenant at its production domain plus .tehas; DNS and an HTTP edge proxy answer only on the tailnet address |
Verified on the host; split DNS for other devices is not configured |
| Standing stack | The product built from the main branch, redeployed when the branch moves | Verified by hand; the automatic redeploy is untested |
| Preview per pull request | Built when a cycle opens a pull request; links and a smoke table are posted on it; removed when it closes; at most 3 | Stack verified by hand; the dispatcher-driven path is untested |
| Database per stack | A reflink clone of the nightly production snapshot: seconds to create, no extra disk until written | Verified |
| Unprivileged database role | The only credential agents and stacks get. It owns its clones and can read nothing else | Verified |
| Inert stacks | No outbound sync, purge or submission jobs, analytics dropped, ads in test mode, no LLM backend | By configuration; checked by reading the product's switches, not by observing traffic |
| Smoke run | Four pages per stack: status, title, canonical on the production domain, no staging host in the HTML | Verified |
| Safe replacement | A stack is rebuilt before the running one is touched; a failed build leaves it serving; a failed build is retried once alone | Verified |
Keeping staging out of production
Only with previews.
| Feature | Detail | Status |
|---|---|---|
| No product awareness | The product is never told about staging hosts; all links name production | Verified on running stacks |
| Static leak gate | A script in the product repository fails if any file names a .tehas host; run in every cycle and before a release tag |
Verified |
| Runtime leak check | The smoke run fails a page that contains .tehas |
Verified |
Visibility
| Feature | Detail | Status |
|---|---|---|
| Plain dashboard | At /dashboard: in line, being worked on, waiting for review, stopped, completed (10 each, then a link to GitHub), staging stacks |
Verified |
| Live console | The running cycle's agents, the commands they run and the start of each result, streamed to the dashboard; tokens and passwords masked | Verified on a real cycle |
| Factory floor (game view) | The front page: the stages as a 16-bit factory floor with the live queue, a work order monitor, an issue board and a breakroom | Tested with fake files |
| Production console | The running cycle's console in the game view, read from the same /live/ files as the dashboard |
Tested with fake files |
| Energy use | In the game view when the model backend is a viiwork mesh: draw now, kWh for 24 hours and 30 days, a 30-minute chart, the machines serving the factory and its models | Tested with fake files |
| Protolab | Tickets at the prototype stage: open issues labelled tehas:prototype, listed in the game view. Not in the queue snapshot yet, so it needs a loaded export. Hosting their images, videos and static HTML prototypes is planned, not built |
Sample or export only |
| Health endpoint | /health and /healthz on port 8080: pass, warn (paused) or fail (last cycle failed, answered with 503) |
Tested with fake files |
| Operating manual | This site, Markdown rendered by the docs server with no build step | Verified |
| Page probe and screenshots | Headless Chromium from the sandbox image; a JSON verdict and an optional screenshot. Previews only | Verified |
Host management
| Feature | Detail | Status |
|---|---|---|
| Configuration from git | ./install.sh copies this repository's files into place and keeps the previous version; --check reports drift |
Verified |
| Instance settings | Everything that differs between instances is in one directory; no script names a product or a host | Covered by the shell tests |
| Install by feature | Level 1 always; previews and clones only when switched on; a feature switched off is stopped before it is removed | Covered by the shell tests |
| Host bootstrap | bootstrap.sh installs every tool at a pinned version and checks the result |
Check and a second run verified; a run from scratch is not |
| Read-only source | The host clones this repository with a read-only deploy key; the bot has no access to it | Verified |
| Pinned sandbox image | Base by digest, tools by version; each build is also tagged by date for rollback | Verified |
| Nightly snapshot refresh | Restores the newest production dump, checks it, swaps it in. Clones only | One manual run verified; the timer has not fired yet |
A second instance
tehas was built for one product. The check that it serves another is the
instance on the host tehas-tehas, which works on this repository. Each row
is what was seen there or in the tests, with its date. The record is
docs/bringup/2026-10-08-tehas-tehas-log.md.
| Check | Result | Status |
|---|---|---|
| No product or host names in platform files | scripts/check-generic.sh passes. It is a lint against regressions, not proof |
Passing, 2026-10-08 |
| No script edits | Installed from instances/tehas-tehas/ with no file changed outside instances/. One local line (TEHAS_BRANCH) was added on the host, see below |
Verified, 2026-10-08 |
| Fresh level 1 install | 39 files; the docs server and the dispatcher timer run with no preview or clone file, container, network or directory on the host | Verified, 2026-10-08 |
| Invalid settings change nothing | install.sh and tehas-dispatch exit 78 with an empty or incomplete instance; ~/tehas identical before and after |
Verified, 2026-10-08 |
| Feature on, then off | The feature's services are stopped before its files go; a failed stop leaves everything in place | Shell tests only |
| Gates in the sandbox image | The repository's own .pi/verify-cmd-full and .pi/smoke-cmd pass inside the image as the unprivileged user |
Verified, 2026-10-08 |
| What a cycle's container can reach | The bot token and the listed mounts; no Docker socket, no SSH keys, no ~/.config/tehas, not the platform checkout |
Verified, 2026-10-08 |
| Dispatch | The timer picked up a labelled issue unattended, checked who labelled it, relabelled it, commented, passed the push check and started the cycle | Verified, 2026-10-08 |
| A cycle to a pull request | Two attempts on one issue, neither reached a pull request. Both were stopped by the model endpoint refusing calls (no free slot), the first in explore, the second in plan after a completed 22-minute explore. See below |
Not proven |
| Host stays responsive during a cycle | Load 0.2 to 0.5 on 8 CPUs; the cycle's container at 5% of a CPU and 212 MB. Seen over 40 minutes of cycle time, not over a whole cycle | Seen, 2026-10-08 |
| Bootstrap | bootstrap.sh check passes and a second bootstrap.sh user changes nothing. The host itself was installed by hand before the script existed |
A run from scratch is not proven |
| The first instance on the new scripts | Not migrated. Its instance directory is written in this repository and reproduces the old built-in values | Not done |
Why no pull request yet: the configured model has one slot on the model endpoint, a request that waits more than 20 seconds for it is refused, and reading a 70,000-token prompt alone takes about 30 seconds. As soon as two conversations want the slot in the same minute, one is refused; pi-rukas retries the step three times and then stops the cycle. In the first attempt the other conversation was another job on the model host. In the second, that job was off, and the other conversation was the cycle's own: pi-rukas's driving session takes model turns of its own while a step's agent is working, and the two starved each other. So a cycle needs a model that serves at least two conversations at once, whatever else is running. Nothing in this was a fault of the instance's setup, and nothing in tehas can fix it: the instance needs a model with two or more slots and enough context (the driving session's prompt reached 81,000 tokens), or a much longer wait at the endpoint. Until a cycle completes, the claim that a second instance works end to end rests on everything above this row.
The local line: a cycle starts from the product's default branch, and this
repository's contract files were on a work branch when the instance was set up.
TEHAS_BRANCH named that branch on the host until it was merged.
Not built
- Monitoring: availability probes, incident issues, weekly report.
- Electricity price scheduling.
tehas:urgentonly changes ordering today. - Model evaluation; every role uses the same model by decision, not by measurement.
- Watching tehas itself: a stuck cycle, a failed refresh, a full disk.
- More than one project per host. See A new project.
- Previews and database clones for any product. What exists is one product's implementation behind two switches.
Known limits
- One host, one project, one cycle at a time.
- The cycle's container holds the bot's token, so agent code can do what that token can. Give each instance its own bot account. See The host.
- Previews are plain HTTP, so features that need a secure context cannot be tested on them.
- Stacks and clones hold real production rows.
- A pull request closed without merging leaves
tehas:awaiting-mergeon its issue until someone removes it.