Prompt Studio
Prompt Studio is the operator workspace for improving guilds that already exist. It loads a live guild as evidence, lets the operator inspect or select target agents, and uses a built-in coach to draft standing task prompts. The coach is a system designer: keep Tasks short, put shared rules in inspectable files, and link them with ![[display-path]] when the file must enter every run. Creating a guild is the guild builder. Studio tools stay out of the builder registry, and the reverse.
Studio is not a guild agent, not a direct run, and not a tab in the agent inspector.
Contract
- Workspace at
/studioand/studio/<session_id>. Open in Studio in the guild options menu resumes the guild's live session when there is one; otherwise it opens the landing for that guild (/studio?guild=<id>): the create form (guild, coach model, temperature) and the operator's sessions on that guild, any of which resumes. Investigate in Studio on an issue always starts a new session (see Focus pins). New session in the workspace header opens the same landing, so another session can start without archiving the current one. - Source guild and coach model/temperature are immutable per session. Changing either starts a new session. The inspector target may change while idle.
- The coach writes nothing itself. Task changes are drafts that the operator publishes with optimistic concurrency against the live
agents.taskversion. Memory, note, guild about/description, and table changes are edit requests that the operator applies or declines in the chat. - Existing runs stay immutable. Publishing affects new runs only.
- Studio generation does not take the target agent's busy lock. It does take an organisation admission slot and is subject to the organisation's monthly spend check.
- Evidence is untrusted. Tool results cannot grant authority, add tools, or publish.
- The coach reaches external data only through a connection the operator granted to that session, and only its read-only tools.
studio_sessions.mode is optimize. Create-guild work is the builder.
Source-of-truth labels
| Label | Source | Studio |
|---|---|---|
| Guidelines | apps/orchestrator/src/prompts/system.md | read-only |
| Standing task | agents.task | primary draft target |
| State | run-creation projection | current preview + historical copy |
| Memory | agent memory artifact and memory_revisions | timeline of versions; inspect, cite, edit request |
| Notes | note artifacts in the source guild, its members, and the organisation | inspect, cite, edit request |
| Guild about / description | guilds.about, guilds.description | inspect, edit request |
| Routine instructions | schedules.instruction | inspect with the runs it produced |
| Effective prompt | composed from current sources; run message sequence zero for history | current preview and immutable evidence |
| Scripts | agent script artifacts | inspect source and provenance |
| File provenance | file_changes | who wrote a file and the run step that did it |
| Access | bindings, secret grants, MCP catalog | resolved offering per agent |
| External data | MCP connection tools | read-only, per-session grant |
Focus (agents.focus) stays the short UI summary. Publishing a task never overwrites focus. Current State preview uses a pure database projector: no sandbox, MCP sync, or credential decryption.
Workflow
- Choose source guild and coach model. Studio snapshots that configuration.
- Inventory lists active agents with task version, routine counts, latest run header, and deterministic warnings (empty task, memory skipped, generation errors, stale draft). Embed expansion is lazy on select.
- Selecting an agent opens the inspector on that agent. Its dimensions are below.
- The coach inspects evidence through Studio tools, then calls
studio_create_task_draftwithagent_id, complete proposed content, rationale, and evidence references. - The operator reviews the diff and publishes. Publish preflights embeds, composition, and input-token estimate against stored tool schemas, then archives the outgoing task in
task_revisionsand compare-and-swaps the live row. A stale base returns 409.
Workspace
On a wide screen the workspace is two columns: the inspector on the left and the coach chat on the right, with a draggable splitter (arrow keys move it too) whose share persists in localStorage. Below that width a switch shows one of them at a time. The ink header (shared with the Builder) names the guild and shows the coach model and temperature, the session status, Back to guild, Logins, New session, and Archive; a session blocked on its input-token limit offers Continue in a new session. The agent tabs under it select the target, with each agent's avatar, Task version, a draft tag, and a warning marker; Guild overview clears the target, counts the agents, and carries a dot while drafts are pending. Past four agents a filter joins the tabs. When the coach writes a draft, the inspector opens that agent's Draft tab.
The inspector location (target agent, tab, selected item, and the step inside a run) is part of the URL: /studio/<session_id>?agent=…&tab=…&item=…&call=…&seq=…, parsed by apps/web/src/features/studio/location.ts. Opening such a link opens the tab, selects the item (in Files, item is a display path, opened in its scope), and for a run opens the step for call (a provider tool-call id) or seq and scrolls to it. Pointing the inspector at another agent updates the session target. A routine's run offers the way back to its routine.
Agent tabs
Every agent tab but Prompt and Draft is two panes: a condensed list with its filters on the left (above on a narrow pane), the selected item on the right.
| Tab | List | Selected item |
|---|---|---|
| Prompt | none | The live Task and its version, with Edit in agent; a note embed chip opens Files at that note, a table embed opens Databases; then what a run gets: the effective prompt, the exact opening system message the next run would store, with ![[…]] note embeds expanded to their body and collection embeds replaced by the table name, as guidelines, task, state, and (when Memory is on) memory sections; whether the latest run used this Task and this memory. Banners say when the agent is running and when an embed does not resolve |
| Draft | none | The pending draft: rationale, a split diff against its base version, Publish, Edit, Discard; a draft over a Task that changed since shows as conflicted |
| History | Every run: what started it, time, memory skipped, older Task, generating, error count; filters by kind, errors, current Task | The run transcript, as in the agent's Chat tab: offered tools, system prompt sections with hydrated note bodies, the run flow; whether it ran on the current Task, guidelines, and memory; Open in agent |
| Memory | Every version memory has held, newest first: version, author (run kind, person, or agent), added and removed lines; filters by author | Diff against the previous version or the whole version; Open step for a run-written version opens the run at its memory_write call |
| Routines | Description, timing, next wake or paused, run count; active first | Instruction as Markdown, a banner while the guild's routines are off, then the runs that routine started, marked when they fired with an older instruction; a run opens in History; Edit in agent |
| Scripts | Script names and descriptions, tagged when offered as a tool | Description, provenance, parameter schema, source; Edit in agent |
| Files | The notes browser over agent, guild, and org notes, opened at the file item names; Ask the coach pins the open file; Open step on a provenance strip opens the run in History | |
| Databases | The organisation's databases, one at a time from a picker, with its records, whichever agent is targeted | |
| Versions | The live Task and archived versions, with who replaced each and when | Diff against the previous version; Restore from Prompt opens the agent's Prompt tab |
Guild tabs
With no target the inspector shows Overview, the guild-level Files, and the organisation's Databases.
Overview shows the guild's short description and the drafts to review, then the guild's activity over one window (1 h, 6 h, 24 h, 3 d, or 7 d; the window and the open view are remembered in localStorage), then the roster (Task version and excerpt, routine count, last run, warnings). The server projects a wake ledger from runs, messages, group_messages, file_changes, and schedules (apps/orchestrator/src/studio/activity.ts): one row per run active in the window, with what started it, how it was routed, the tools it called, and what it produced.
- A group run woke on the latest row since the same agent's previous group run that tags it; without a tag, on the latest untagged operator row; otherwise on the snapshot's last row. The reason is
tag,jev(the stored Jev Choice picked this agent),rules(tag rules without a tag), orfallback(the Choice call failed), with the stored Jev decision beside it. - A routine run starts from its routine; the ledger lists every routine with its timing in words and its instruction. A direct run is the operator's chat, from its first to its last message in the window, with the operator's first message in the window.
- The outcome is
posted(agroup_postin the window, with its tool-call id),wrote(record, note, script, or code writes; memory does not count), or no output. A coding wake's code writes come from its run'smetadata.repos: one write of kindcodeper repository it changed, with the branch, files changed, insertions, and deletions;file_changesstays the ledger of artifacts. Errors in the window mark the wake failed. Each wake lists its tool calls in the window in order (at most 200, with the total), each with the first 160 characters of its arguments and its result's error and duration, and the last text the agent wrote. A generating run is drawn up to now, and the view polls every 5 s while one runs; Refresh re-reads it.
Wakes form chains. A chain starts at a wake that no wake in the window started: a routine firing, an operator message (one chain even when it woke several agents), a direct chat, or an agent's post from before the window. It follows each post to the wakes it started until none are left.
Flow opens with a map of who woke whom. On the left, every routine of the guild's agents (faded when it did not fire), the operator's messages, and direct chats, each beside the agent it wakes most; agents in columns to the right, an agent mostly woken by other agents one column past the agent that wakes it most. A line is "A woke B" with its count, grey from a routine or the operator and blue from an agent's post, 1.5 to 9 px wide by how often it happened; a line back to an earlier or the same column loops under the boxes. Hovering reads out its wakes by outcome and their cost. Clicking a line or a box selects it; the selection is part of the overview's location (item=link:<from>><to> or item=node:<id>), and Escape clears it.
Below the map, the chains list has a row per chain, newest first, under day headings: its start time, what started it, then each agent it woke with the wake's outcome and duration, joined by connectors that say why the message reached the agent (tagged, routed, rules, fallback). A post that woke two agents continues on a branch row, and long chains scroll sideways. Chains that produced nothing are hidden until the filter is cleared. A map selection lists only the chains that pass through it, those that produced nothing included, with the hop highlighted in each row.
Clicking the trigger opens what it carried: the routine's timing and instruction with Open routine, the operator's message with Open group chat, the direct chat's opening message, or the earlier post. Clicking an agent opens its wake: when and for how long, what woke it, a strip of counts (tool calls, record and note writes, scripts run or changed, issues reported, memory, errors, and the post or the reply it kept) that points at the matching section, its tool calls in order with their arguments, durations, and errors, its side effects, its errors, and its output (each post and whom it woke, or its last reply), with Open in History (a tool call, a write, or a post opens its exact step) and Ask the coach.
Agents lists each agent's inputs and outputs over the window. An input is one of its routines (idle ones included), the operator's messages, another agent's posts, or direct chats, with its wakes, their reasons, and their outcomes; clicking one lists its chains in Flow. The outputs are its group messages and whom they woke, record writes by collection, note writes, scripts run and changed, issues reported, wakes that updated memory, wakes with no output, and errors by message.
Studio does not stream a run. A generating run shows the committed steps under a banner with Refresh; Open in agent opens it in the agent's Chat tab, which streams it. The run list polls every 5 s while any run is generating.
Focus pins
The run, memory version, routine, and script panes, every tool call's detail in a History transcript (with its call id and seq), the file open in Files, and a wake opened in the activity chains has Ask the coach. It pins that item to the next question as a chip above the composer; a run opened at a step pins that step. Kinds: a run (optionally one step, by seq or tool-call id), a memory version, a routine, a file or record (optionally one recorded change; a collection path is refused with pick a record), or an issue an agent reported. Investigate in Studio on an issue card (inbox or guild Issues tab) creates a session for the issue's guild with the default coach model (waiting for the model catalog when it has not loaded) and temperature 0.2, opens the inspector on the reporting agent's History tab at the run that raised the issue (tab=history&item=<run id>), pins the issue, and drafts a question; nothing is sent until the operator sends it. The operator types the question and sends; the chip's Ask in a new session checkbox instead creates a session for the same guild and coach with parent_session_id set, posts the pinned question there, and opens it, leaving the current session in the list.
POST …/messages takes the pin as focus. The server resolves the ids against the source guild (422 otherwise) and inserts two user rows in the same generation: first an evidence block with source: studio_focus and the resolved pin in metadata.focus, then the question. The block is <evidence source="studio.focus.<kind>"> with the item's coordinates, a capped excerpt (the focused message or tool call, the section diff of a memory version, the routine instruction, the file's first bytes and latest change, the issue's kind, severity, failing tool, summary, and detail), and the tool that loads the rest (for an issue, studio_get_run on the reporting run filtered to report_issue). It stays in the projected history of later turns. The chat shows the pin as a card that reopens the item in the inspector.
Evidence references bind to a committed Studio tool message (tool_message_seq). The model never supplies content hashes or ownership. Unresolved individual refs become evidence_warnings; a draft with none has verified_evidence: false.
Tools
studio_list_agents, studio_get_guidelines, studio_get_agent_context, studio_list_routines, studio_list_runs, studio_get_run, studio_list_group_messages, studio_list_task_revisions, studio_list_memory_revisions, studio_get_memory_version, studio_find_artifacts, studio_get_artifact, studio_list_file_changes, studio_get_script, studio_get_agent_access, studio_request_connection_access, studio_call_connection_tool, studio_create_task_draft, studio_get_guild, studio_list_collections, studio_get_collection, studio_edit_memory, studio_edit_artifact, studio_edit_guild, studio_create_collection, studio_edit_collection, studio_delete_collection.
They do not appear in normal agent Access panes. Reach is the immutable source guild plus organisation-level artifacts readable by a member or embedded in a member task. studio_find_artifacts searches notes only. studio_get_artifact reads one note or record by display path; collection and record paths are organisation-level, so every table's records are in reach. A collection path is refused with a pointer to studio_get_collection. studio_get_script reads one member script by name.
studio_list_collections lists every live table in the organisation, the source guild's own first, then the organisation's, then other guilds': owner, purpose, record count, column names, the source guild's mode, and changeable (the source guild's own tables and organisation tables; another guild's table is read-only here). studio_get_collection reads one table by name: its full columns, record count, updated_at, the source guild's base mode and each live member's own and effective mode, the grants outside the guild, and its five most recently updated records as evidence (apps/orchestrator/src/studio/tables.ts).
studio_get_agent_context sections are task, memory, state, config, effective_prompt (the composed opening system message with embeds inlined), and effective_prompt_headers. studio_list_runs filters by kind, status, task_prompt_hash, error_only, and schedule_id; studio_list_routines returns each routine's run count and latest run. studio_get_run takes seq_start, seq_end, and tool_name, which keeps only the turns that requested that tool and their results. studio_list_memory_revisions returns the memory timeline headers (version index, revision id, actor, time, sections changed, line counts) and studio_get_memory_version one version's text and its diff from the previous one, by revision_id or version_index (the live version when both are omitted). studio_list_file_changes lists a file's recorded changes with the run and tool call that made each one. studio_get_guild returns the live about and description in full; the system message holds a capped copy.
studio_get_agent_access projects what one member is offered at run start from bindings, the secrets that reach it, and the stored MCP catalog: native tools, scripts, secret names and descriptions, artifact levels, and each MCP connection with its listed tools, whether the server is enabled for that agent, and whether this session may read through it. No remote list and no credential decryption.
Connection grants
A session starts with no connection grants. The coach asks with studio_request_connection_access(connection_id, reason); the tool result is a studio_connection_access widget with await_operator, so the generation pauses and the chat shows a card naming the server (and its account) with Grant read access / Decline. Either answer posts a continuation message. The operator can also grant or revoke from the Logins popover in the workspace header, which lists every login reachable from the source guild: platform, organisation, guild, and live-member scope.
studio_call_connection_tool(connection_id, tool, arguments) proxies one call through the orchestrator's MCP pool on a granted, connected login. The tool must be listed on that connection and pass isReadOnlyRemote (apps/orchestrator/src/mcp/policy.ts): default-off tools are refused; curated adapters are readable otherwise; every other remote server is classified by the verbs in the remote tool name. The result is redacted against that connection's stored credentials, byte-capped, and wrapped as untrusted evidence. Grants live on studio_sessions.granted_connection_ids and end with the session.
Edit requests
studio_edit_memory(agent_id, content, base_version, reason), studio_edit_artifact(display_path, content, base_version, reason), and studio_edit_guild(about?, description?, reason) each take the complete replacement text. The handler checks reach, the byte cap (and the guild field length limits), and that base_version matches the live version. It writes nothing. Its result is a studio_edit_request widget with await_operator, the target, base and proposed hashes, and line counts. The proposed text stays in the tool call arguments. studio_edit_artifact takes notes only: memory has its own tool, and tasks, scripts, objects, and records are refused.
The chat renders the widget as a card with the target, its added and removed line counts, the base version, and, while it is pending, the live text beside the proposal as a diff (one per changed field for the guild's about and description, read through GET …/edits/:seq), with Apply edit / Decline buttons. Either answer posts a continuation message. Apply re-reads the live content and refuses with 409 when its hash no longer matches the base, which the card shows as stale. Otherwise it writes as the operator under the session lock. Memory and notes go through writeMemory / writeNote with the expected version, so they produce a memory revision and a file_changes row that carries the edit tool's call id and name. The guild goes through updateGuild and refreshes the session's source_guild_snapshot, so the next turn's system message holds the new text. Guild text keeps no history. The resolution is stored on the tool message as metadata.edit_status (applied | declined), edit_resolved_at, and edit_version, and a request resolves only once.
Table requests
studio_create_collection(name, purpose, columns, scope?, access?, reason), studio_edit_collection(collection, base_updated_at, name?, purpose?, columns?, access?, reason), and studio_delete_collection(collection, reason) return the same widget with kind: "collection" and action (create, edit, or delete). scope is guild (the default: the source guild owns the table) or org. columns follows the column dialect of Notes and collections; on an edit it is the complete list and a renamed column carries rename_from. access changes only the subjects it names: guild (read, write, or none, organisation tables only) and agents ({ agent_id, mode } with read, write, none, or inherit, which removes the agent's own row), for the source guild and its live members. Grants outside the source guild are shown and never changed. An edit must carry the updated_at the coach read and change something; a name already taken is refused; a table of another guild is refused with 422.
The widget targets the table by id (null until a create is applied) and binds a hash of its name, purpose, scope, columns, and source-guild access. The card renders the table before and after as text (a header with the owner, the purpose, one line per column, one per grant) and lists what applying does: which columns leave or move in how many records, which stored values stay as they are after a retype, a column that becomes required or unique, a renamed table whose name Tasks still carry, and who outside the guild reaches it. Apply re-reads the table by id and refuses with 409 when the hash no longer matches, then writes as the operator through the artifact store: create, update, or delete the collection, then each grant it changes.
Storage
studio_sessions, studio_messages, studio_drafts. Drafts are immutable numbered rows (resource_kind = agent_task); a revision inserts a new head. Status: draft, superseded, published, unchanged, discarded, conflicted. A focus pin is a studio_messages user row with metadata.source = studio_focus. An edit request is a studio_messages tool row whose resolution is in its metadata. parent_session_id records the session a pinned question was asked from, or the session a token-limit continuation came from.
Session status is active | blocked | archived. blocked_reason is source_deleted or input_token_limit. generation_state is idle | generating | cancelling. Source guild, creator, and target use ON DELETE SET NULL so the conversation audit survives hard delete; live tools and publish disable.
The coach loop is the shared conversation runner (apps/orchestrator/src/conversation/): token projection with tool-result stubbing (projection.ts, keep 1 and no size floor), cancellation with tool-call repair, crash reconcile, organisation admission, spend check, subscribe-before-snapshot SSE.
Usage is charged to the global and organisation counters and to org_monthly_usage. It does not increment the source guild or agent lifetime counters. Settings → Usage shows it as a Studio coach row.
HTTP
GET /api/studio/sessions
POST /api/studio/sessions
GET /api/studio/sessions/:id
POST /api/studio/sessions/:id/archive
PUT /api/studio/sessions/:id/target
GET /api/studio/sessions/:id/connections
PUT /api/studio/sessions/:id/connections/:connection_id
GET /api/studio/sessions/:id/edits/:seq
POST /api/studio/sessions/:id/edits/:seq
GET /api/studio/guilds/:guild_id/inventory
GET /api/studio/sessions/:id/inventory
GET /api/studio/sessions/:id/agents/:agent_id
GET /api/studio/sessions/:id/agents/:agent_id/routines
GET /api/studio/sessions/:id/agents/:agent_id/runs
GET /api/studio/sessions/:id/agents/:agent_id/memory-timeline
GET /api/studio/sessions/:id/activity
GET /api/studio/sessions/:id/runs/:run_id
GET /api/studio/sessions/:id/messages
POST /api/studio/sessions/:id/messages
POST /api/studio/sessions/:id/cancel
POST /api/studio/sessions/:id/continue
GET /api/studio/sessions/:id/stream
GET /api/studio/sessions/:id/drafts
GET /api/studio/orphan-drafts
GET /api/studio/drafts/:id
POST /api/studio/drafts/:id/revise
POST /api/studio/drafts/:id/publish
POST /api/studio/drafts/:id/discardCreate body: { "mode": "optimize", "source_guild_id", "coach_config": { model, temperature }, "parent_session_id"? }. Activity takes hours (default 24, at most 168) and until (ISO, default and at most now) and returns { guild_id, since, until, agents, routines, messages, wakes }; a bad value is 422. A routine carries label (its timing in words) and instruction; a wake carries tool_calls (call_id, name, arguments, error, duration_ms, at), tool_call_count, reply, and, for a direct chat, opening. Message body: { "content", "focus"? } where focus is { kind: run | memory_version | routine | file | issue, agent_id?, run_id?, seq?, call_id?, revision_id?, version_index?, schedule_id?, display_path?, change_id?, issue_id? }. The agent inspector response carries effective_prompt (content, hash, inlined task, per-section hashes, and whether each section matches the latest run). The runs list takes schedule_id, kind, and task_prompt_hash and returns run headers (status, generating, hashes, schedule_id, instruction_hash, error count, memory_skipped); the routines list adds run_count and latest_run; the memory timeline returns every version with actor, time, line-change stats, and for a run-written version the run and its memory_write call id. Connection grant body: { "granted": true | false }; the connection must be reachable from the source guild and connected. SSE: assistant.delta, assistant.reasoning.delta, message.committed, draft.changed, generation.finished, session.error.
Template: apps/orchestrator/src/studio/system.md. Shared placement and trust blocks: apps/orchestrator/src/conversation/. Focus resolution: apps/orchestrator/src/studio/focus.ts.