Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Training a site

A business describes itself in plain words: what it does, who comes to it, what it keeps track of and through which stages, and who handles what. The trainer, portal plan and portal build, turns that into a working content repo: a question series for every role, the state machines behind them, and a desk per handling role. No software is written for the business; the only thing produced is content, and every line of it is held to the lint (portal lint), which is the same code as the site, at the same version.

The trainer was its own repo, uhhm/portal-trainer, until portal 0.6; its runs (runs/, the lab notebook of every site trained so far) stay there.

Build: structure from the plan, words from the model

portal build --plan plans/contractor.yaml --out out/nordvik \
    --lint "portal lint" --llm "MODEL=sonnet engines/claude.sh"

A plan.yaml is a needs spec that also says what each form collects and, per need, max_fields: how many required fields that person will tolerate, taken from what the business said about them, never from the form. build refuses a form over its budget, and hands the budget to the verifier, so the check still means something after handover. Everything structural is derived from the plan with no model involved: the state machines (read off stages: walk the line, be turned down at any point, succeed only from the end, pass over skippable stages), the front door, a followup page per need, a gated directory per group with its desk and the pages only that group may open, and every button the machine allows. The verifier accepts the result by construction; what it still checks is the budget.

The model writes words only, one page at a time, into a JSON object whose shape is fixed. Its strings are merged over built-in fallbacks key by key, so a bad reply costs wording, never a working site. --llm-for <path prefix> points the model at some pages only (aggregates.yaml for the mails alone); a reply in the wrong language for the plan is refused, and a mail that lost the submitter’s link gets it back.

Onboarding from the plan

invites_by: [board] on a plan, or invited_by: [...] on a group, puts an invite form on the inviting group’s desk for each group it may fill, and a list of the invites sent. Portal does the rest (an account in Kanidm, the group, a one-time credential link) through invite: on the alternative. Every group with a desk invites into itself by default.

Rebuilding a live site

--words-from <repo> lifts the words a repo already has back into the words map by what each card is (its bucket, the page it leads to, the group it invites into), across the whole repo, so a rebuild keeps hand-edited wording even when a card moves page (a door taken away, a need brought onto the front door) and adds only the new slots. Cards are matched by what they are (the bucket they record into, the page they lead to, a field’s name, the group they invite into), not by position, so a rebuild that reorders or reshapes a page still hands each card its own words; on uhhm.no the live front door’s three headings follow their buckets onto the new page. --words <json> overrides slots by page and may repeat, each file layered over the last (the words a build used, then a round of hand changes); words.json in the output is the words as used. See runs/tomtervel-invites-2026-09-23 and runs/uhhm-2026-09-23.

Same contractor, same local model (nemotron 30B-A3B on ergo)Free-form generatebuild
A site the verifier acceptsnever, in 6 rounds and 16 minfirst try, 130 s
Held-out personas, with descriptions-8 of 8
Held-out personas, headings only-7 of 8

The wording numbers are from the re-score of 2026-09-22, one task per call (runs/contractor-build-2026-09-21/rescore-2026-09-22/). The first score, all tasks in one prompt, said 8 of 8 on headings only; the apprentice picking the building survey over “Join our carpentry team” only showed once no other visitor was in the prompt.

For comparison, a Sonnet-class model writing everything free-form scored 8 of 10 on headings only and silently dropped a stage the owner had described. build cannot drop a stage: it never decides one. The denominators differ: the build site gives the two crew personas a one-option page, so they generate no task, and 8 of 8 is eight public personas where the generate number is ten. Held out means held out from the model: the plan’s who and wants lines and the personas were written by the same person about the same needs, and the fallback heading is the need’s wants, so the two are cousins.

The front door opens modestly

A front door with more buttons than text leaves a visitor paralysed. So the built front door opens with what the business does, in one or two features and no button at all (does: on the plan, at most two, each for: a need so it speaks to one role and is headed by what that role gets); then a look-around card, things a visitor may inspect before asking for anything (shows:, each a portal resource block passed through verbatim, plus every public view the plan opens); and only then the offers, as more alternatives on the same page. Nothing is layered: doors still exist for a plan that wants them, but a second tier is what this shape is there to avoid, and on uhhm.no a flat front door with an opener routed better than the doored one.

The house and its players: a role the opener invites in has to see its win, and the machine has to give it something to do after it sits down. build warns when a featured need gives the submitter neither a move (moves_by: {<stage>: [submitter, ...]}) nor a public view; a player with no move is a spectator.

Doors: a first question before the forms

Five or six needs on one page is more than a visitor sorts through (the full contractor plan had six). A plan can put public needs behind doors: each door is one alternative on the front door leading to a page of its own, where the needs that name it sit.

doors:
  - id: patients
    who: someone with pain or an injury, or a patient who already has a time here
needs:
  - id: new-patient
    door: patients
    ...

The door’s page is a nav page, public, holding the forms. The spec gives such a need max_steps: 2 and pins via to the door’s heading, so the verifier walks both pages. A door nobody is behind, a door on an internal need, or a door named like a group is refused, and a front door over five alternatives is warned about. The visitor is scored on both pages: a persona who picks the wrong door is a misroute on /, one who picks the wrong form is a misroute on the door’s page.

Cases with doors: plans/physio.yaml (a physiotherapy clinic, patients and referrers), plans/fotballklubb.yaml (a football club, in Norwegian, parents and helpers), plans/hire.yaml (a machine hire yard, hiring and firms), and plans/contractor-full.yaml (three doors over six public needs). Each has an intake under intakes/ in the business’s own words and held-out personas.

What a plan can say about a form and its readers

  • multiple: true on a select: pick any number (portal renders toggles). What a volunteer can do, which events, which committees.
  • groups: at the top, a list of { id, label }: the groups with a name the public knows, such as a vel’s committees. A select with options_from: groups offers exactly those labels, so the application form and the desks cannot drift apart.
  • watched_by: [public]: an ungated read-only page of the bucket at /status/<bucket>, linked from the followup page. Everyone sees what was sent in and where it stands. email and tel fields, and any field marked private: true, are projected away before it is shown.
  • events: { open: applied, active: onboarded } on a need: the event name a stage is entered by, where the builder’s default (the stage’s own name, received for the first) is not what a live site already has. Portal replays every record by event name, so a rebuild of a site that holds records keeps the names its aggregates.yaml has, or every record vanishes from its desk. plans/uhhm.yaml rebuilds uhhm.no this way.
  • mails: { <stage>: submitter | <address> } on a need: who hears when a case enters that stage, sent by portal the moment it moves (portal 0.3.43, delivered by uhhm/gdo on the host). submitter goes to the email the form collected, so the form must collect a required email; an address is the treasurer, the board. The plan then needs mail: { from, reply_to } at the top, the sender every mail carries into site.yaml. The words model writes each mail’s subject and body (aggregates.yaml in the words map); the fallback says what happened and, where the submitter has a move from that stage, the link to make it.
  • stages_text: { <stage>: <words> } on a need: each stage in the plan’s language, where the ASCII id will not do (ikke_lost reads as English). Used by the fallback mails, the “what happens next” line, the desk buttons’ fallback labels, and written into needs.yaml as stage_words so the desk score describes an outcome in the site’s own words. With them, the vel’s desks score 26 of 26.
  • turned_down_from: [<stage>, ...] on a need: the unfinished stages a case can still be turned down from; omitted, every one. An applicant is declined before they start and departs after, which is what uhhm.no’s live machine says and what two near-synonymous desk buttons had been standing in for.
  • reached_from: { <stage>: [<stage>, ...] } on a need: where a stage is entered from, when the line a case walks does not say it. Two shapes need it and nothing else does - a case SENT BACK to a stage it was already past (a building application returned for the drawing it was missing, after the hearing turned something up), and a GOOD ENDING PARTWAY ALONG (an authority that changes its own decision while it is looking at the appeal, rather than passing it up to the board). A stage named there is entered from exactly those stages and nowhere else, so saying where a refusal may come from also says where it may not: an applicant may withdraw from anywhere, and be refused only once the papers are in, which turned_down_from could not express because it was one list for every bad ending at once. Found by replaying a building authority’s real cases through portal’s engine, which refused three legal paths and accepted two illegal ones until the plan could say this; see runs/permits-2026-09-25/README.md.
  • attended_by: on a need: who moves cases that no desk button moves, an automation listening on the bucket’s events, written into aggregates.yaml as portal’s own annotation for the reader.

Plan: the intake, read by a model, checked by a person

portal plan --intake intakes/contractor.yaml --out out/nordvik \
    --personas sectors/construction/needs.yaml \
    --llm "MODEL=nemotron-3.5-lightning:30b-a3b engines/ollama.sh"

This is the one step where a model decides structure: which needs the intake implies, the stages, who handles what, the form, the budget. Its YAML is parsed and validated exactly as build will, the errors go back to it, and what passes is written verbatim for a person to read before building from it. Personas are never asked of this model; they are attached from a file, and where their need: names an id the plan does not have, plan says so and leaves the mapping to the reader.

generate (below) is the older path, a model writing pages free-form against the lint. It is kept for the record of what it measured.

How it works

 intake (plain words)
    │
    ▼  phase 1                       frozen after this phase: the generator
 needs.yaml ───────────────────────  may not weaken the spec to pass
    │
    ▼  phase 2   ◄── iris findings, until it exits 0
 aggregates.yaml + questions/** + desks
    │
    ▼  phase 3   ◄── misroutes from a blind visitor model
 wording
  1. Needs. The intake becomes needs.yaml: one need per kind of person and thing they want, including the roles inside the business. Each names the bucket the answer lands in, the group that handles it, the states that count as finished, and the effort the visitor will tolerate.
  2. Structure. Pages, state machines and desks are generated, then revised against the lint, iris check, until every need is met: a path within budget, a desk held by the right group, every finished state reachable, no record stranding. The page-writing model sees the needs and never the personas.
  3. Wording. A separate model plays each persona, blind, first with headings only and then with descriptions. A misroute is a wording defect and goes back to the generator. A rewrite that breaks structure is reverted.

Model output is only ever written to aggregates.yaml, site.yaml and questions/**.yaml inside the output directory.

portal generate --intake intakes/contractor.yaml --out out/nordvik \
    --lint "portal lint" --llm engines/ollama.sh
portal answer --tasks tasks.jsonl --llm engines/ollama.sh > answers.jsonl

answer puts one task per call to the visitor model, with the options shuffled per task and the call seeded from the task id: no visitor sees another visitor’s situation, a heading’s place on the page cannot carry the answer, and a rerun on the same site is the same experiment. The earlier runs under runs/ answered all tasks in one prompt.

The verifier is uhhm/iris, not code in this repo: it reads portal’s own content types and ships in every portal release, so a generated repo’s CI keeps enforcing its needs after handover.

LLM backends

A backend is any command that takes a prompt on stdin and prints text.

BackendUse
engines/ollama.shLocal Ollama. MODEL=qwen3:8b answers a task in about two seconds on ergo’s GPU and is the default visitor. Generation wants a larger model (MODEL=qwen3.6:27b), which runs mostly on CPU there.
engines/claude.shClaude through Claude Code’s own login (a Max or Pro subscription), no API key. MODEL=sonnet or opus, EFFORT=low by default. It drops any ANTHROPIC_API_KEY from the environment, because a key outranks the login and a stale one fails every call. Log in once with claude auth login.
engines/anthropic.shThe Anthropic API, given a working ANTHROPIC_API_KEY from the Console. That is separate, prepaid billing; a subscription does not cover it.
engines/mailbox.shWrites each prompt to a directory and waits for a reply file: drive the generator from an agent session, or by hand.

The visitor and the generator should be different models. A model grading its own wording agrees with itself.

The onboarding site

portal-portal/ is the front end of this: a portal content repo whose front door is the intake interview. A submission lands in onboardings (open → drafted ⇄ revising → live | declined) and is worked from an owners’ desk. Today the desk step is a person running generate and sending the link; the intake record already carries every field the generator reads, so the step that remains is a listener on portal.answers.submitted that runs it and pushes the draft to a new repo. It passes its own needs.yaml.

The whole card, not the heading

A visitor never sees a bare heading and description. Portal renders a card: a button, an encouragement, one or more sections each with an icon, a title and a line of guidance, and the fields with their labels and placeholders. Since 2026-09-23 the builder fills all of it and the scorers read all of it.

  • Slots. Beside the heading and description, each form alternative gets a button, an encouragement, an icon, a form title and guidance line, and a placeholder per field. A door gets an icon and a “for whom” line. Icons are picked from a fixed list of Iconify lucide names; anything else is dropped.
  • Context from the plan. Every form alternative also carries two sections built from what the plan already knows, before the form: what happens next (the stages, in order, as one sentence) and who reads it (the handling group and the contact). Their fallbacks are true before a model has written a word; the model rewords them.
  • Scoring. answer --repo <content dir> shows the visitor each option as its whole card (headings-only tasks stay headings only). rate --repo does the same and asks a fourth question, “where”: do you know where you are, who this is for, and what the button does. All four are controlled against the gutted twin.

Scoring a built site

portal score --site out/nordvik --out out/scored \
    --iris /var/local/cargo-target/release/iris \
    --llm "MODEL=qwen3:8b engines/ollama.sh"

One command for the whole grading, which every run under runs/ before 2026-09-23 carried as its own score.sh: iris writes the visitor tasks and the simulation model, the blind visitor answers them with descriptions and with headings only, iris scores the answers (the misroutes are printed), and rate reads the cards. The two routing numbers land in summary.json, the rating in rated/rate.json. The pieces still run alone: answer --tasks for one visitor model against tasks iris wrote, and:

iris check --path out/nordvik/questions --needs-sim model.json
portal rate --model model.json --out out/rated --repo out/nordvik \
    --llm "MODEL=qwen3:8b engines/ollama.sh"

Four yes-or-no questions per persona (which option is mine, where am I and what does the button do, what will I be asked, what happens after I send it), each also put to a gutted twin of the page with headings, descriptions, the form and the followup removed. A question the twin still gets a yes on is not measuring the page, and rate says which. The followup page is read from the repo (--repo); without it “what happens next” is reported as not measured rather than as zero. LIX is over the prose only (question, descriptions, followup): headings and labels are not sentences and would pull every page down to “very easy”. The rate.json in runs/contractor-build-2026-09-21/ predates this and its asked and next columns are the old, pinned ones.

rate --repo also scores the desks, the inside of the business, which reads button labels and nothing else: for every desk and bucket with two or more buttons, a worker is told what has just happened to a case (in the state’s own words, never the label’s) and picks the button among the ones a row in that state shows, in a rotated order. Scored per desk as right of total, with what each miss was mistaken for. Measured on 2026-09-24 on the live vel: 24 of 26, the misses one button each on the neighbour-help desk and the notice board. Two limits: the outcome is put from the state id, and a Norwegian id with its accents stripped (ikke_lost) reads as English even to a strong model; and a first version that offered every button of the bucket, not the row’s, scored the same desks 17 of 26 and sent a rewrite chasing an instrument fault. A desk that scores full marks has labels a stranger to the business could press right. prefer has a positive control on record (runs/tomtervel-opener-2026-09-23, “The friendliness judge”): the live vel front door against a bureaucratic twin, friendlier 9 of 10 to the live page and 0 to the twin. A pair that comes back “no preference” throughout is alike in tone, not an instrument that cannot see; and the rejected door rewrite that was judged clearer on 5 of 10 needs routed worse across three seeds, so clarity as judged and routing as measured are different things, and routing decides.

Measured so far

Generated: a building contractor (runs/contractor-2026-09-21/), from intakes/contractor.yaml, a Sonnet-class model generating.

  • Phase 1 derived six needs from the owner’s prose, including one nobody asked for: a project manager logging a phone enquiry, taken from “enquiries arrive by phone, email and text and get lost”.
  • Phase 2 was green on the first round: 14 files, four state machines (enquiries: open → quoted → won | lost, subcontractors with an awaiting_papers loop, applicants, site_issues), two desks.
  • Phase 3, its own personas: 10 of 10 in both modes. That number is flattering, because in that run the pages were written with the personas in view. Against ten held-out personas, with qwen3:8b as the visitor:
Visitor could readRouted right
Headings and descriptions10 of 10
Headings only8 of 10

A scaffolding firm asking “how do I get to do jobs for you?” chose “Ask about working with us” (the hiring door). An attic conversion chose “Get someone to look at a building”. The generator has since been changed to hide personas from the page-writing phase.

Built: the same contractor from a plan (runs/contractor-build-2026-09-21/), nemotron writing words only, and re-scored one task per call the day after (rescore-2026-09-22/ in it): 8 of 8 with descriptions, 7 of 8 on headings alone, the miss being an apprentice bricklayer reading “Join our carpentry team” as not for them. rate on the same site, with the followup shown and every question controlled, passes its controls and puts the pages at LIX 18 to 35.

Built behind doors, 2026-09-22. Five plans with a first question before the forms, each built and then scored by qwen3:8b one task per call, against held-out personas. Behind a door a persona is asked twice, on the front door and on the door’s page.

CaseWords byWith descriptionsHeadings onlyNotes
Machine hire (hire)nemotron10 of 1010 of 10doors that ask a question the visitor already knows the answer to
Physio clinic (physio)nemotron8 of 109 of 10“Book your first appointment” fails a daughter booking for her mother
Football club, Norwegian (fotballklubb)nemotron14 of 1611 of 16three parents walk past “Er det på tide å melde inn?”
Football club, NorwegianClaude Sonnet13 of 1612 of 16better sentences, same driver/sponsor blur
Whole contractor (contractor-full)nemotron18 of 2215 of 22“Apply to work with us” and “Find a job with us” cross over for all four
Tomter Vel, Norwegian (tomtervel), first planClaude Sonnet25 of 3022 of 30“utstyr” for a swing, hearing input behind the wrong door
Tomter Vel, plan revised 2026-09-23, whole cardClaude Sonnet31 of 3830 of 38committees as groups, public lists, hearing input on the front door

Each run’s README names the misroutes and what in the plan or the heading caused them. A stronger word model did not route better on the one plan built with both: routing is decided by which distinctions sit on one page, and the visitor grades that the same whoever wrote the words.

Planned: the contractor intake read by a local model (runs/contractor-plan-2026-09-22/), nemotron, one round, 41 s. The plan passed every check build makes and a reader finds three structural misses in it the checks cannot: the crew’s report made public, the “papers missing” stage dropped, and no need for the project manager logging a call. The person reading the plan is not optional.

Existing site: uhhm.no (runs/uhhm-2026-09-21/), eight tasks.

Visitor modelWith descriptionsHeadings only
small LLM, blind7 of 85 of 8
qwen3:8b (Ollama)7 of 87 of 8
granite4.1:8b (Ollama)7 of 86 of 8

All three send a paying client through “Work happens through dialogue, not tickets”, the collaborator door. Different models, same defect: that is what a wording problem looks like.

The noise floor, 2026-09-23 (runs/spread.sh, answer --seed). The same tasks answered by qwen3:8b with four seeds (0 is the run every score above used):

RunWith descriptionsHeadings only
uhhm.no round 6, 15 tasks11, 10, 10, 1012, 12, 7, 7
Tomter Vel with the opener, 38 tasks30, 29, 29, 3126, 24, 26, 23

With descriptions the visitor is steady: one task in fifteen, two in thirty-eight. On headings alone it is not: five of fifteen on uhhm.no between seeds, three of thirty-eight on the vel. A headings-only difference smaller than that is the visitor’s mood. Compare descriptions-mode scores; read headings-only as a sanity check, and run several seeds before believing it. Small numbers, one run each. Leads, not measurements.

What else is here

  • skill/SKILL.md - the same procedure for an agent or a person working by hand.
  • sectors/ - needs specs to start from. construction/ doubles as the held-out persona set above.
  • generator/reference.md - everything the generating model is told about portal content. When the lint rejects something the reference did not warn about, the fix belongs there.

Toward a trained model

Nothing here trains weights yet. What would make it possible is accumulating: every run leaves the intake, the needs, the pages at each revision and the scores. With a fast local visitor as the scorer, generating many variants per need and keeping the winners is cheap, and those winners are the fine-tuning set for a small generator. The verifier has to be trusted before any of that: a generator trained against a judge that is often wrong learns the judge.