Training a site
A business describes itself in plain words: what it does, who comes to
it, what it keeps track of and through which stages, and who handles
what. The trainer, portal plan and portal build, turns that into a
working content repo: a question series for every role, the state
machines behind them, and a desk per handling role. No software is
written for the business; the only thing produced is content, and
every line of it is held to the lint (portal lint), which is the
same code as the site, at the same version.
The trainer was its own repo, uhhm/portal-trainer, until portal 0.6;
its runs (runs/, the lab notebook of every site trained so far)
stay there.
Build: structure from the plan, words from the model
portal build --plan plans/contractor.yaml --out out/nordvik \
--lint "portal lint" --llm "MODEL=sonnet engines/claude.sh"
A plan.yaml is a needs spec that also says what each form collects
and, per need, max_fields: how many required fields that person
will tolerate, taken from what the business said about them, never
from the form. build refuses a form over its budget, and hands the
budget to the verifier, so the check still means something after
handover. Everything structural is derived from the plan with no
model involved: the state machines (read off stages: walk the line,
be turned down at any point, succeed only from the end, pass over
skippable stages), the front door, a followup page per need, a gated
directory per group with its desk and the pages only that group may
open, and every button the machine allows. The verifier accepts the
result by construction; what it still checks is the budget.
The model writes words only, one page at a time, into a JSON object
whose shape is fixed. Its strings are merged over built-in fallbacks
key by key, so a bad reply costs wording, never a working site. --llm-for <path prefix> points the model at some pages only
(aggregates.yaml for the mails alone); a reply in the wrong language
for the plan is refused, and a mail that lost the submitter’s link
gets it back.
Onboarding from the plan
invites_by: [board] on a plan, or invited_by: [...] on a group,
puts an invite form on the inviting group’s desk for each group it may
fill, and a list of the invites sent. Portal does the rest (an account
in Kanidm, the group, a one-time credential link) through invite: on
the alternative. Every group with a desk invites into itself by
default.
Rebuilding a live site
--words-from <repo> lifts the words a repo already has back into the
words map by what each card is (its bucket, the page it leads to, the
group it invites into), across the whole repo, so a rebuild keeps
hand-edited wording even when a card moves page (a door taken away, a
need brought onto the front door) and
adds only the new slots. Cards are matched by what they are (the
bucket they record into, the page they lead to, a field’s name, the
group they invite into), not by position, so a rebuild that reorders
or reshapes a page still hands each card its own words; on uhhm.no
the live front door’s three headings follow their buckets onto the
new page. --words <json> overrides slots by page and may repeat, each file layered over
the last (the words a build used, then a round of hand changes);
words.json in the output is the words as used. See
runs/tomtervel-invites-2026-09-23 and runs/uhhm-2026-09-23.
| Same contractor, same local model (nemotron 30B-A3B on ergo) | Free-form generate | build |
|---|---|---|
| A site the verifier accepts | never, in 6 rounds and 16 min | first try, 130 s |
| Held-out personas, with descriptions | - | 8 of 8 |
| Held-out personas, headings only | - | 7 of 8 |
The wording numbers are from the re-score of 2026-09-22, one task per
call (runs/contractor-build-2026-09-21/rescore-2026-09-22/). The
first score, all tasks in one prompt, said 8 of 8 on headings only;
the apprentice picking the building survey over “Join our carpentry
team” only showed once no other visitor was in the prompt.
For comparison, a Sonnet-class model writing everything free-form
scored 8 of 10 on headings only and silently dropped a stage the owner
had described. build cannot drop a stage: it never decides one. The
denominators differ: the build site gives the two crew personas a
one-option page, so they generate no task, and 8 of 8 is eight public
personas where the generate number is ten. Held out means held out
from the model: the plan’s who and wants lines and the personas
were written by the same person about the same needs, and the fallback
heading is the need’s wants, so the two are cousins.
The front door opens modestly
A front door with more buttons than text leaves a visitor paralysed.
So the built front door opens with what the business does, in one or
two features and no button at all (does: on the plan, at most two,
each for: a need so it speaks to one role and is headed by what
that role gets); then a look-around card, things a visitor may
inspect before asking for anything (shows:, each a portal resource
block passed through verbatim, plus every public view the plan
opens); and only then the offers, as more alternatives on the same
page. Nothing is layered: doors still exist for a plan that wants
them, but a second tier is what this shape is there to avoid, and on
uhhm.no a flat front door with an opener routed better than the
doored one.
The house and its players: a role the opener invites in has to see
its win, and the machine has to give it something to do after it
sits down. build warns when a featured need gives the submitter
neither a move (moves_by: {<stage>: [submitter, ...]}) nor a public
view; a player with no move is a spectator.
Doors: a first question before the forms
Five or six needs on one page is more than a visitor sorts through
(the full contractor plan had six). A plan can put public needs
behind doors: each door is one alternative on the front door
leading to a page of its own, where the needs that name it sit.
doors:
- id: patients
who: someone with pain or an injury, or a patient who already has a time here
needs:
- id: new-patient
door: patients
...
The door’s page is a nav page, public, holding the forms. The spec
gives such a need max_steps: 2 and pins via to the door’s heading,
so the verifier walks both pages. A door nobody is behind, a door on
an internal need, or a door named like a group is refused, and a
front door over five alternatives is warned about. The visitor is
scored on both pages: a persona who picks the wrong door is a
misroute on /, one who picks the wrong form is a misroute on the
door’s page.
Cases with doors: plans/physio.yaml (a physiotherapy clinic,
patients and referrers), plans/fotballklubb.yaml (a football club,
in Norwegian, parents and helpers), plans/hire.yaml (a machine hire
yard, hiring and firms), and plans/contractor-full.yaml (three
doors over six public needs). Each has an intake under intakes/ in
the business’s own words and held-out personas.
What a plan can say about a form and its readers
multiple: trueon aselect: pick any number (portal renders toggles). What a volunteer can do, which events, which committees.groups:at the top, a list of{ id, label }: the groups with a name the public knows, such as a vel’s committees. A select withoptions_from: groupsoffers exactly those labels, so the application form and the desks cannot drift apart.watched_by: [public]: an ungated read-only page of the bucket at/status/<bucket>, linked from the followup page. Everyone sees what was sent in and where it stands.emailandtelfields, and any field markedprivate: true, are projected away before it is shown.events: { open: applied, active: onboarded }on a need: the event name a stage is entered by, where the builder’s default (the stage’s own name,receivedfor the first) is not what a live site already has. Portal replays every record by event name, so a rebuild of a site that holds records keeps the names its aggregates.yaml has, or every record vanishes from its desk.plans/uhhm.yamlrebuilds uhhm.no this way.mails: { <stage>: submitter | <address> }on a need: who hears when a case enters that stage, sent by portal the moment it moves (portal 0.3.43, delivered by uhhm/gdo on the host).submittergoes to the email the form collected, so the form must collect a requiredemail; an address is the treasurer, the board. The plan then needsmail: { from, reply_to }at the top, the sender every mail carries intosite.yaml. The words model writes each mail’s subject and body (aggregates.yamlin the words map); the fallback says what happened and, where the submitter has a move from that stage, the link to make it.stages_text: { <stage>: <words> }on a need: each stage in the plan’s language, where the ASCII id will not do (ikke_lostreads as English). Used by the fallback mails, the “what happens next” line, the desk buttons’ fallback labels, and written into needs.yaml asstage_wordsso the desk score describes an outcome in the site’s own words. With them, the vel’s desks score 26 of 26.turned_down_from: [<stage>, ...]on a need: the unfinished stages a case can still be turned down from; omitted, every one. An applicant is declined before they start and departs after, which is what uhhm.no’s live machine says and what two near-synonymous desk buttons had been standing in for.reached_from: { <stage>: [<stage>, ...] }on a need: where a stage is entered from, when the line a case walks does not say it. Two shapes need it and nothing else does - a case SENT BACK to a stage it was already past (a building application returned for the drawing it was missing, after the hearing turned something up), and a GOOD ENDING PARTWAY ALONG (an authority that changes its own decision while it is looking at the appeal, rather than passing it up to the board). A stage named there is entered from exactly those stages and nowhere else, so saying where a refusal may come from also says where it may not: an applicant may withdraw from anywhere, and be refused only once the papers are in, whichturned_down_fromcould not express because it was one list for every bad ending at once. Found by replaying a building authority’s real cases through portal’s engine, which refused three legal paths and accepted two illegal ones until the plan could say this; see runs/permits-2026-09-25/README.md.attended_by:on a need: who moves cases that no desk button moves, an automation listening on the bucket’s events, written into aggregates.yaml as portal’s own annotation for the reader.
Plan: the intake, read by a model, checked by a person
portal plan --intake intakes/contractor.yaml --out out/nordvik \
--personas sectors/construction/needs.yaml \
--llm "MODEL=nemotron-3.5-lightning:30b-a3b engines/ollama.sh"
This is the one step where a model decides structure: which needs the
intake implies, the stages, who handles what, the form, the budget.
Its YAML is parsed and validated exactly as build will, the errors go
back to it, and what passes is written verbatim for a person to read
before building from it. Personas are never asked of this model; they
are attached from a file, and where their need: names an id the plan
does not have, plan says so and leaves the mapping to the reader.
generate (below) is the older path, a model writing pages free-form
against the lint. It is kept for the record of what it measured.
How it works
intake (plain words)
│
▼ phase 1 frozen after this phase: the generator
needs.yaml ─────────────────────── may not weaken the spec to pass
│
▼ phase 2 ◄── iris findings, until it exits 0
aggregates.yaml + questions/** + desks
│
▼ phase 3 ◄── misroutes from a blind visitor model
wording
- Needs. The intake becomes
needs.yaml: one need per kind of person and thing they want, including the roles inside the business. Each names the bucket the answer lands in, the group that handles it, the states that count as finished, and the effort the visitor will tolerate. - Structure. Pages, state machines and desks are generated, then
revised against the lint,
iris check, until every need is met: a path within budget, a desk held by the right group, every finished state reachable, no record stranding. The page-writing model sees the needs and never the personas. - Wording. A separate model plays each persona, blind, first with headings only and then with descriptions. A misroute is a wording defect and goes back to the generator. A rewrite that breaks structure is reverted.
Model output is only ever written to aggregates.yaml, site.yaml
and questions/**.yaml inside the output directory.
portal generate --intake intakes/contractor.yaml --out out/nordvik \
--lint "portal lint" --llm engines/ollama.sh
portal answer --tasks tasks.jsonl --llm engines/ollama.sh > answers.jsonl
answer puts one task per call to the visitor model, with the options
shuffled per task and the call seeded from the task id: no visitor
sees another visitor’s situation, a heading’s place on the page cannot
carry the answer, and a rerun on the same site is the same experiment.
The earlier runs under runs/ answered all tasks in one prompt.
The verifier is uhhm/iris, not code in this repo: it reads portal’s own content types and ships in every portal release, so a generated repo’s CI keeps enforcing its needs after handover.
LLM backends
A backend is any command that takes a prompt on stdin and prints text.
| Backend | Use |
|---|---|
engines/ollama.sh | Local Ollama. MODEL=qwen3:8b answers a task in about two seconds on ergo’s GPU and is the default visitor. Generation wants a larger model (MODEL=qwen3.6:27b), which runs mostly on CPU there. |
engines/claude.sh | Claude through Claude Code’s own login (a Max or Pro subscription), no API key. MODEL=sonnet or opus, EFFORT=low by default. It drops any ANTHROPIC_API_KEY from the environment, because a key outranks the login and a stale one fails every call. Log in once with claude auth login. |
engines/anthropic.sh | The Anthropic API, given a working ANTHROPIC_API_KEY from the Console. That is separate, prepaid billing; a subscription does not cover it. |
engines/mailbox.sh | Writes each prompt to a directory and waits for a reply file: drive the generator from an agent session, or by hand. |
The visitor and the generator should be different models. A model grading its own wording agrees with itself.
The onboarding site
portal-portal/ is the front end of this: a portal content repo whose
front door is the intake interview. A submission lands in
onboardings (open → drafted ⇄ revising → live | declined) and is
worked from an owners’ desk. Today the desk step is a person running
generate and sending the link; the intake record already carries
every field the generator reads, so the step that remains is a
listener on portal.answers.submitted that runs it and pushes the
draft to a new repo. It passes its own needs.yaml.
The whole card, not the heading
A visitor never sees a bare heading and description. Portal renders a card: a button, an encouragement, one or more sections each with an icon, a title and a line of guidance, and the fields with their labels and placeholders. Since 2026-09-23 the builder fills all of it and the scorers read all of it.
- Slots. Beside the heading and description, each form alternative gets a button, an encouragement, an icon, a form title and guidance line, and a placeholder per field. A door gets an icon and a “for whom” line. Icons are picked from a fixed list of Iconify lucide names; anything else is dropped.
- Context from the plan. Every form alternative also carries two sections built from what the plan already knows, before the form: what happens next (the stages, in order, as one sentence) and who reads it (the handling group and the contact). Their fallbacks are true before a model has written a word; the model rewords them.
- Scoring.
answer --repo <content dir>shows the visitor each option as its whole card (headings-only tasks stay headings only).rate --repodoes the same and asks a fourth question, “where”: do you know where you are, who this is for, and what the button does. All four are controlled against the gutted twin.
Scoring a built site
portal score --site out/nordvik --out out/scored \
--iris /var/local/cargo-target/release/iris \
--llm "MODEL=qwen3:8b engines/ollama.sh"
One command for the whole grading, which every run under runs/
before 2026-09-23 carried as its own score.sh: iris writes the
visitor tasks and the simulation model, the blind visitor answers
them with descriptions and with headings only, iris scores the
answers (the misroutes are printed), and rate reads the cards. The
two routing numbers land in summary.json, the rating in
rated/rate.json. The pieces still run alone: answer --tasks for
one visitor model against tasks iris wrote, and:
iris check --path out/nordvik/questions --needs-sim model.json
portal rate --model model.json --out out/rated --repo out/nordvik \
--llm "MODEL=qwen3:8b engines/ollama.sh"
Four yes-or-no questions per persona (which option is mine, where am
I and what does the button do, what will I be asked, what happens
after I send it), each also put to a gutted
twin of the page with headings, descriptions, the form and the
followup removed. A question the twin still gets a yes on is not
measuring the page, and rate says which. The followup page is read
from the repo (--repo); without it “what happens next” is reported
as not measured rather than as zero. LIX is over the prose only
(question, descriptions, followup): headings and labels are not
sentences and would pull every page down to “very easy”. The
rate.json in runs/contractor-build-2026-09-21/ predates this and
its asked and next columns are the old, pinned ones.
rate --repo also scores the desks, the inside of the business,
which reads button labels and nothing else: for every desk and bucket
with two or more buttons, a worker is told what has just happened to
a case (in the state’s own words, never the label’s) and picks the
button among the ones a row in that state shows, in a rotated order.
Scored per desk as right of total, with what each miss was mistaken
for. Measured on 2026-09-24 on the live vel: 24 of 26, the misses
one button each on the neighbour-help desk and the notice board. Two
limits: the outcome is put from the state id, and a Norwegian id with
its accents stripped (ikke_lost) reads as English even to a strong
model; and a first version that offered every button of the bucket,
not the row’s, scored the same desks 17 of 26 and sent a rewrite
chasing an instrument fault. A desk that scores full marks has labels
a stranger to the business could press right.
prefer has a positive control on record
(runs/tomtervel-opener-2026-09-23, “The friendliness judge”): the
live vel front door against a bureaucratic twin, friendlier 9 of 10
to the live page and 0 to the twin. A pair that comes back “no
preference” throughout is alike in tone, not an instrument that
cannot see; and the rejected door rewrite that was judged clearer on
5 of 10 needs routed worse across three seeds, so clarity as judged
and routing as measured are different things, and routing decides.
Measured so far
Generated: a building contractor (runs/contractor-2026-09-21/),
from intakes/contractor.yaml, a Sonnet-class model generating.
- Phase 1 derived six needs from the owner’s prose, including one nobody asked for: a project manager logging a phone enquiry, taken from “enquiries arrive by phone, email and text and get lost”.
- Phase 2 was green on the first round: 14 files, four state machines
(
enquiries: open → quoted → won | lost,subcontractorswith anawaiting_papersloop,applicants,site_issues), two desks. - Phase 3, its own personas: 10 of 10 in both modes. That number is flattering, because in that run the pages were written with the personas in view. Against ten held-out personas, with qwen3:8b as the visitor:
| Visitor could read | Routed right |
|---|---|
| Headings and descriptions | 10 of 10 |
| Headings only | 8 of 10 |
A scaffolding firm asking “how do I get to do jobs for you?” chose “Ask about working with us” (the hiring door). An attic conversion chose “Get someone to look at a building”. The generator has since been changed to hide personas from the page-writing phase.
Built: the same contractor from a plan (runs/contractor-build-2026-09-21/),
nemotron writing words only, and re-scored one task per call the day
after (rescore-2026-09-22/ in it): 8 of 8 with descriptions, 7 of 8
on headings alone, the miss being an apprentice bricklayer reading
“Join our carpentry team” as not for them. rate on the same site,
with the followup shown and every question controlled, passes its
controls and puts the pages at LIX 18 to 35.
Built behind doors, 2026-09-22. Five plans with a first question before the forms, each built and then scored by qwen3:8b one task per call, against held-out personas. Behind a door a persona is asked twice, on the front door and on the door’s page.
| Case | Words by | With descriptions | Headings only | Notes |
|---|---|---|---|---|
Machine hire (hire) | nemotron | 10 of 10 | 10 of 10 | doors that ask a question the visitor already knows the answer to |
Physio clinic (physio) | nemotron | 8 of 10 | 9 of 10 | “Book your first appointment” fails a daughter booking for her mother |
Football club, Norwegian (fotballklubb) | nemotron | 14 of 16 | 11 of 16 | three parents walk past “Er det på tide å melde inn?” |
| Football club, Norwegian | Claude Sonnet | 13 of 16 | 12 of 16 | better sentences, same driver/sponsor blur |
Whole contractor (contractor-full) | nemotron | 18 of 22 | 15 of 22 | “Apply to work with us” and “Find a job with us” cross over for all four |
Tomter Vel, Norwegian (tomtervel), first plan | Claude Sonnet | 25 of 30 | 22 of 30 | “utstyr” for a swing, hearing input behind the wrong door |
| Tomter Vel, plan revised 2026-09-23, whole card | Claude Sonnet | 31 of 38 | 30 of 38 | committees as groups, public lists, hearing input on the front door |
Each run’s README names the misroutes and what in the plan or the heading caused them. A stronger word model did not route better on the one plan built with both: routing is decided by which distinctions sit on one page, and the visitor grades that the same whoever wrote the words.
Planned: the contractor intake read by a local model
(runs/contractor-plan-2026-09-22/), nemotron, one round, 41 s. The
plan passed every check build makes and a reader finds three
structural misses in it the checks cannot: the crew’s report made
public, the “papers missing” stage dropped, and no need for the
project manager logging a call. The person reading the plan is not
optional.
Existing site: uhhm.no (runs/uhhm-2026-09-21/), eight tasks.
| Visitor model | With descriptions | Headings only |
|---|---|---|
| small LLM, blind | 7 of 8 | 5 of 8 |
| qwen3:8b (Ollama) | 7 of 8 | 7 of 8 |
| granite4.1:8b (Ollama) | 7 of 8 | 6 of 8 |
All three send a paying client through “Work happens through dialogue, not tickets”, the collaborator door. Different models, same defect: that is what a wording problem looks like.
The noise floor, 2026-09-23 (runs/spread.sh, answer --seed).
The same tasks answered by qwen3:8b with four seeds (0 is the run
every score above used):
| Run | With descriptions | Headings only |
|---|---|---|
| uhhm.no round 6, 15 tasks | 11, 10, 10, 10 | 12, 12, 7, 7 |
| Tomter Vel with the opener, 38 tasks | 30, 29, 29, 31 | 26, 24, 26, 23 |
With descriptions the visitor is steady: one task in fifteen, two in thirty-eight. On headings alone it is not: five of fifteen on uhhm.no between seeds, three of thirty-eight on the vel. A headings-only difference smaller than that is the visitor’s mood. Compare descriptions-mode scores; read headings-only as a sanity check, and run several seeds before believing it. Small numbers, one run each. Leads, not measurements.
What else is here
skill/SKILL.md- the same procedure for an agent or a person working by hand.sectors/- needs specs to start from.construction/doubles as the held-out persona set above.generator/reference.md- everything the generating model is told about portal content. When the lint rejects something the reference did not warn about, the fix belongs there.
Toward a trained model
Nothing here trains weights yet. What would make it possible is accumulating: every run leaves the intake, the needs, the pages at each revision and the scores. With a fast local visitor as the scorer, generating many variants per need and keeping the winners is cheap, and those winners are the fine-tuning set for a small generator. The verifier has to be trusted before any of that: a generator trained against a judge that is often wrong learns the judge.