Every pass sweeps the TUI as a user sees it, next to the prior-art TUIs, run side by side in herdr panes. The objective is one consistent interaction model: the same key does the same thing on every surface, every state has a visible way out, and nothing gent draws is worse than what a prior-art agent draws for the same moment. The code-reading tui area finds structure; this area finds what only a rendered screen shows.
The references are the ui rows of PRIOR_ART.md: Claude Code first for the transcript (a clear transcript, a north star), vercel-labs/fx for input and layout, then pi and opencode, each with what to read and how to run it. Never the curl … | bash installer. Sources live under okra repo path <slug>; read the TUI code and e2e captures for any screen the running binary cannot reach without a model.
- gent:
bun run --cwd <main checkout>/apps/tui dev --debugwithGENT_DATA_DIRandHOMEunder the scratch directory.--debuguses the scripted model, so a conversation, streaming, tool calls and errors render with no paid call. A message that holdsdebug toolsplays a six-step tool turn with reasoning: bash, three reads, a grep, an edit, and a bash that exits 2, run for real undergent-debug-tools/in the session cwd. Its steps report cache writes and reads, so the status row shows the cache timer (ANTHROPIC_PROMPT_CACHE_TTL=5mshortens it). A message that holdsdebug askasks one background question withask_user_async(the tray and/answer), then runs a 20-second bash step, so an answer can join the running turn. A message that holdsdebug threadsstarts two threads withthread.start(one playsdebug tools), and one that holdsdebug handoffasks for a handoff, so the Sessions pane shows a started thread and a thread of two sessions. A message that holdsdebug delegatestarts one child withdelegate.start(it playsdebug tools), so the tray shows a running child and its calls, and the child's completion row lands in the transcript when it ends. A message that holdsdebug thinkopens each step's reasoning with a bold heading, runs one bash step, then thinks four seconds before it answers, so the live line names the heading (✻ Investigating rendering code) while it waits. A message that holdsdebug usage limitfails its turn on a usage limit that resets in five hours, or in the time it names (debug usage limit 2m; over 30 s, or it is retried), so the error row shows the reset time and, withwake.autoResumein the scratch home's.gent/config.json, the resume tray row and its fire. - A reference TUI runs with
HOME(andXDG_*) under the scratch directory and no credentials, so it can never reach a paid model. Screens that need a model come from its source and test captures instead. A login prompt is a screen to compare, not a step to complete. - One herdr tab per comparison, split into panes of the same size:
herdr pane split,herdr pane run <pane> '<cmd>',herdr pane send-keys/send-textto drive,herdr pane wait-outputto settle,herdr pane readto capture. Neverherdr agent. Close every pane and tab the sweep opened before the report. - A repeatable check (a key sequence, a resize, a capture at each step) can run as a drive script instead:
bun packages/e2e/src/drive.ts <script.json>runs any command on the e2e pty fixture (zigpty, a real controlling terminal;Bun.Terminalin Bun 1.4.2 is not one, so it delivers no SIGWINCH or SIGINT) with a live emulator and saves each capture (format inpackages/e2e/README.md). The same command and scratch environment rules hold. - Capture each screen at two sizes (a normal pane and a narrow one near 60×20) and after a resize.
- Launch and empty state: what the first screen says, where the cursor is, the time to first input.
- Input: multiline entry, paste (large and bracketed), history recall, editor keys (word jump, kill line), submit vs newline.
- Streaming: text, reasoning, tool calls, long tool output, diffs; scrollback kept or lost; follow vs manual scroll.
- Interrupt and exit: Esc, Ctrl+C ladder, Ctrl+D; what a cancelled turn leaves on screen.
- Slash commands and pickers: discovery, filtering, empty result, Esc out; model and session pickers.
- Asks and prompts: approval, question, handoff; docked pane, never modal (owner rule).
- Errors and notices: provider error, retry, overflow/compaction notice; wording and placement.
- Status: model, context gauge, cost, cwd, busy state.
- Selection and copy (OSC 52), links, mouse.
- A clear transcript: a user message, an agent reply, a child agent, a queued message and a tool call each look different at a glance; finished work is one summary line with its detail on demand.
- Consistency inside gent: the same key or word means the same thing on every surface; each pane shows its way out.
Beyond the usual candidate table: a matrix of the checklist rows × (gent, fx, pi, opencode) with one line each, the herdr pane read capture paths for every gent row, and for each candidate the reference screen it borrows from. Owner rules hold: docked panes, not modals. A candidate that only matches a reference's taste, with no user-facing gain, goes in "not worth a pass".