Skip to content

refactor(character): centralize prompts with byte-identical Kurisu output - #161

Merged
Lucas1479 merged 6 commits into
mainfrom
codex/character-prompt-module
Oct 8, 2026
Merged

Lucas1479 merged 6 commits into
mainfrom
codex/character-prompt-module

Conversation

@Lucas1479

@Lucas1479 Lucas1479 commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

What and why

Character names and persona text are duplicated across Chat, AUIP, browser and VN request assembly. Move those values behind one prompt renderer and the packaged characters/kurisu.toml resource while keeping the existing Kurisu model inputs byte-identical.

Base: main; predecessor #160 has been squash-merged. Next: #162. Merge commit 540b983 synchronizes that ancestry without changing any source file from 43ef7f8. Rebase/retarget #162 after this PR is merged, especially after a squash merge.

Linked Issue: not required for this behavior-preserving internal refactor.

Change class

  • Routine fix, documentation, test, maintenance, or presentation-only UI
  • Product-semantic or public-contract change discussed in the linked Issue
  • Isolated, default-off experiment

Owning layer: character prompt data and request rendering. Host protocols, verification, task authority and system/user message boundaries remain owned by their existing components.

User-visible effect: none for this slice. Startup selection, conversation ownership, RAG isolation, Host phrase extraction and persona-editor propagation belong to the next PR.

Compatibility: no persisted-state migration or new startup setting here. Remove prompt copies and the unused local CLI implementation after checking for production callers. Add package-data and source-release inclusion for the new characters/ directory. Dependencies are unchanged.

Evidence

  • A subprocess harness assembles 155 actual production prompt surfaces with synthetic configuration and external I/O blocked.
  • The pre-migration fixture was independently reproduced on untouched origin/main (e56faa1); 155/155 byte-identical.
  • Independent audit of this refactor head (3fd4bcf) also found 155/155 byte-identical, with no exception and no fixture relaxation.
  • Snapshot/module tests are tests/test_character_prompt_snapshots.py and tests/test_character_prompt_module.py.
  • The final combined stack at 1dc92d8 passed the full two-shard Python runner (421 files, 5077 passed, 5 skipped, 5 existing xfailed). That later slice deliberately changes warmup; it is not part of this PR. This PR's CI validates the intermediate head separately.
uv run --locked --no-sync python -m pytest tests/test_character_prompt_module.py tests/test_character_prompt_snapshots.py -q
  • Relevant Python tests and production-boundary snapshots pass
  • CPU/model-less baseline remains supported
  • Packaging/source-release policy and prompt documentation updated
  • Electron screenshots/build and dependency audit: not applicable to this slice

Additional per-branch validation after the unused-import cleanup: Ruff passed for all tracked Python paths with repository exclusions, and python -m pytest tests/test_provider_prompt_routing.py -q passed 2/2 on each PR branch. New GitHub checks are running.

Final check

  • One coherent problem, no unrelated cleanup
  • Reuses one renderer without speculative execution or authority changes
  • No secrets, local state, model weights, voice material or restricted assets included
  • Third-party notices and provenance preserved

Lucas1479 added a commit that referenced this pull request Oct 8, 2026
…pts (#160)

## What and why

Focus validation and task-reference lookup are non-speaking control
queries, but the legacy single-prompt API could replace their
instructions with a character persona on Gemini/Bedrock and failed to
dispatch local/hybrid requests. Route both callers through the existing
messages API and remove the unused legacy query interface.

Merge order: #160 → #161 → #162. This first slice targets `main` at
`e56faa1`. After merging each predecessor, rebase/retarget its dependent
PR onto main as needed, especially after squash merges.

Linked Issue: not required for this control-query bug fix.

## Change class

- [x] Routine fix, documentation, test, maintenance, or
presentation-only UI
- [ ] Product-semantic or public-contract change discussed in the linked
Issue
- [ ] Isolated, default-off experiment

Owning layer: control-query construction and the shared LLM messages
interface. This does not change chat persona generation or execution
authority.

User-visible effects:

- Task lookup receives neutral instructions at temperature 0 instead of
persona instructions at 0.7. Candidate bounds and selection validation
remain unchanged.
- Focus validation keeps its narrow SET/CLEAR/NONE contract and
unavailable result on failure.
- Task-lookup infrastructure failures reach the existing
`reason="error"` path and generic unresolved-task response instead of
being mistaken for an ambiguous selection.
- Both callers request a 10-second budget and 900 output tokens.

Compatibility: internal legacy API and unused prompt constants removed;
in-tree tests and diagnostic probes use the supported messages API. No
persisted-state migration.

Known limitation: Gemini's shared messages adapter does not currently
enforce the supplied timeout. This PR intentionally does not alter all
Gemini calls, including main chat, or promise a ten-second wall-clock
bound.

## Evidence

The relevant control suites (`test_focus_policy.py`,
focus/handoff/acceptance suites, `test_task_lookup.py`, and
`test_llm_message_query_backends.py`) passed in the final stacked
validation at `1dc92d8`.

```powershell
uv run --locked --no-sync python -X utf8 tools/run_tests.py --shard-count 2 --shard-index 0
uv run --locked --no-sync python -X utf8 tools/run_tests.py --shard-count 2 --shard-index 1
```

That combined stack run covered 421 files: **5077 passed, 5 skipped, 5
existing xfailed**, both commands exited 0. This is combined-stack
evidence, not a separate full-suite claim for this intermediate head;
this PR's CI will validate its own head. No dependency install or live
model calls.

- [x] Relevant Python tests pass in the validated stack
- [x] CPU/model-less baseline remains supported
- [x] In-tree probes and tests are updated
- Electron build, screenshots and dependency audit: not applicable to
this slice

## Final check

- [x] One coherent problem, no unrelated cleanup
- [x] No new speculative API, fallback or compatibility path
- [x] No secrets, local state, model weights, voice material or
restricted assets included
- [x] Third-party notices and provenance preserved
Base automatically changed from codex/character-control-queries to main October 8, 2026 03:36
@Lucas1479
Lucas1479 marked this pull request as ready for review October 8, 2026 04:06
@Lucas1479
Lucas1479 merged commit 2d1f692 into main Oct 8, 2026
17 checks passed
@Lucas1479
Lucas1479 deleted the codex/character-prompt-module branch October 8, 2026 04:39
Lucas1479 added a commit that referenced this pull request Oct 8, 2026
main at 2d1f692 has exactly the same tree as the already-merged PR #161 head 540b983. Preserve the validated character-runtime tree while recording the squash ancestry; no source changes are required.
Lucas1479 added a commit that referenced this pull request Oct 8, 2026
…identity (#162)

## What and why

A second character must not inherit Kurisu's conversation history,
reference corpus or accepted-task identity. Pin the prompt character at
backend startup, bind each conversation to its owner, and preserve the
accepted conversation identity across Work replay and recovery. Keep
retained Projects, Work and artifacts shared through their existing
references and permissions.

Base: `main`. Prerequisites #160 and #161 are merged. This is the final
integration slice; the control-query fix and byte-preserving extraction
were reviewed separately.

Linked Issue: pending. The new startup setting and conversation/Work
contracts require a linked product/public-contract discussion before
merge.

## Change class

- [ ] Routine fix, documentation, test, maintenance, or
presentation-only UI
- [x] Product-semantic or public-contract change (linked Issue pending)
- [ ] Isolated, default-off experiment

Owning layers: startup character loading; Host-owned conversation and
accepted Work facts; prompt/phrase presentation. No new execution
permission or shared-artifact authority.

### Resulting behavior and compatibility

- `AMADEUS_CHARACTER_ID` defaults to `kurisu` and uses existing setting
persistence. The backend fixes the character at startup; changing it or
its prompt file requires restart. There is no hot switch, character
list, new/import/delete-pack GUI, startup-role selector, or
per-character persona editor. The existing GUI editor remains Kurisu's
Japanese override only. New roles still require a manual TOML file and
startup setting. Voice and visual resources are not automatically
selected.
- New conversations persist their character. Otherwise valid legacy
sessions without this field belong to Kurisu. Foreign-role
open/continue/rename/delete reject at the backend before changing files,
history, selection, project bindings or projections. List titles remain
visible, with a role label only for foreign conversations; rename/delete
controls are disabled with restart guidance. There is no cross-role
transcript reader. Startup with no owned chat creates one; deleting the
last owned chat leaves an empty chat, even when foreign chats remain.
Selection, rename and deletion errors use separate notices. Successful
rename updates list data without clearing history or streaming. This
character-ownership rule is not a separate local-user ACL or global
session-management interface.
- Saving an unreadable/invalid existing session fails without
overwriting it. Explicit invalid ownership is not treated as legacy
absence.
- The existing RAG index remains Kurisu-only: other roles return no
reference without querying it. Per-role knowledge-base management is not
implemented. Retained Projects, Work and artifacts remain shared;
unretained drafts do not enter another conversation's natural-language
lookup candidates. Following up on shared work does not import the
original conversation.
- Conversation-accepted Work snapshots identity in immutable admission
evidence. Replay, Retry, Resume and automatic repair recover it from the
Operation's accepted source. Old admissions use historical Kurisu **on
read**, without rewriting accepted plans. New accepted amendments keep
their own identity; caller/Provider metadata cannot replace it.
- Independent Work-page tasks have **no conversation-role field** on
initial execution, amendments, Retry or Resume. They never gain the
role-reference paragraph that would reinterpret “yourself” as the chat
character. Execution guidance may name the fixed startup character,
without inventing a conversation source.
- Accepted Work addressing a Cooperative native context can validate its
persisted recipient after restart without opening the old chat loop.
Revision, state, workspace and native-identity checks remain; this adds
no automatic retry or duplicate dispatch permission.
- AUIP actor values used for matching, verification and winner
attribution stay `kurisu`. Only verified, model-facing copies render the
presentation identity.
- Of 107 Japanese phrase keys, 93
execution/status/verification/receipt/permission/focus templates are
fixed Host wording with reviewed Kurisu variants and neutral defaults.
Role files cannot override these keys, even with identical variables or
the builtin id. Only 14 VN story-commentary keys remain configurable,
and are not Host execution evidence. Earlier experimental packs must
remove Host fact entries from their voice_lines table. All original
Kurisu phrase outputs remain byte-identical. The existing Japanese
Kurisu persona editor now also affects AUIP/browser branches, including
action, authorization and planning judgments; this is more than a tone
change.

### Deliberate model-visible changes

Kurisu retains **154/155** pre-migration prompt surfaces byte-for-byte.
The declared exception is llama warmup using the real character prefix;
the neutral AUIP skill-trigger wording is checked separately. Control
Work recovery now retains the role-reference paragraph already present
on initial execution, where the old code lost it. Independent Work
receives no such paragraph. Original snapshot fixtures are unchanged.

## Evidence

Validation covers the boundary fixes committed as `8cb88cb` and
`1221e93`. Baseline-sync merges `1dcf4d1` and `4488a47` preserve exactly
the same file tree. For `4488a47`, main at `2d1f692` was verified
tree-identical to the already-integrated #161 head `540b983`; recording
its squash ancestry resolves all 13 conflicts without changing source,
tests, fixtures or documentation. The existing validation therefore
applies to the same file contents:

```powershell
uv run --locked --no-sync python -X utf8 tools/run_tests.py --shard-count 2 --shard-index 0
uv run --locked --no-sync python -X utf8 tools/run_tests.py --shard-count 2 --shard-index 1
uv run --locked --no-sync python -m pytest tests/test_provider_request_role_identity.py tests/test_work_continuation_role_identity.py tests/test_auip_work_recovery.py -q
# From electron/
npm test
npm run build
```

- Full Python: **421 files, 5248 passed, 5 skipped, 5 existing
xfailed**; both shards exited 0. Focused validation: session boundary
**145 passed**, phrase/loader/startup/snapshot **313 passed**,
hostile-persona boundary **45 passed**. The identity/recovery suites
also remain in the full run. Default Kurisu and experimental emotion
routing disabled, short external temp paths, existing dependencies only.
- Electron: **247 passed**, build passed; existing bundle-size warning
remains.
- Official model-less Electron launcher plus eight isolated
role-selection/management checks: **19/19 passed**, including disabled
foreign controls and direct backend mutation refusal. Same-role rename
is checked through its real API and reloaded rendering. Actual UI and
backend list confirm no active/new/foreign chat after deleting the last
owned chat. Electron/backend exited cleanly; no chat turn or model
request was sent.
- Ruff and diff checks passed. Prompt fixtures unchanged.

OpenClaw Resume test limitation: its native prompt test supplies a
synthetic restored Session and non-writer `plan` mode. Existing
default-agent workspace-lease and native Session restoration wiring is
outside this PR; the test does **not** establish complete OpenClaw
Resume availability. Hostile-persona tests use simulated model output.
They verify tool/candidate/ref contracts, fixed actor/payload,
stance/version and receipt ownership, while still allowing valid
decisions. They do not establish live-model semantic correctness or
identical legal choices across personas; application-specific action
legality stays with app receipt checks where the protocol assigns it.
Real-model persona fidelity and voice experience are also not claimed.

<details>
<summary>Actual isolated UI: before and after deleting the last owned
conversation</summary>

These are two states of the current code, not a pixel comparison with an
older revision. Only the foreign conversation shows a role label. After
deletion the chat is empty and the real backend reports no active/new
session.

![Before
deletion](https://raw.githubusercontent.com/Code-Amadeus/Amadeus/1dcf4d166afae9db87f23960b1bf229965af1418/docs/evidence/character-identity-2026-10-08/before-delete.png)

![After
deletion](https://raw.githubusercontent.com/Code-Amadeus/Amadeus/1dcf4d166afae9db87f23960b1bf229965af1418/docs/evidence/character-identity-2026-10-08/after-delete.png)

Synthetic conversations; no character artwork, personal data,
credentials or machine paths. See
`docs/evidence/character-identity-2026-10-08/README.md`.

</details>

Capture-run observation: one additional screenshot run timed out waiting
for the chat input to become enabled after the backend connected (no
renderer console/page errors). Repeating with unchanged code, timeout
and assertions passed 15/15, as did the prior validation run. The
initial timeout remains an unresolved startup-timing observation and is
not counted as a pass.

Current native-dialog limitation: a same-role double-click rename probe
did not show a title change before timeout and reported no renderer
errors. It is not counted as a pass; the final 19-check journey uses the
real rename API plus reloaded rendering, with separate component
callback coverage.

- [x] Relevant Python tests pass; CPU/model-less baseline remains
supported
- [x] Electron tests and build pass
- [x] Isolated UI screenshots included
- [x] Settings/ownership/recovery documentation updated in
`docs/character-runtime.md`
- Dependency audit: not applicable; no dependency changes
- [ ] Linked public-contract Issue completed before merge

Deferred: role-management GUI, per-role knowledge bases, remaining
hardcoded UI character names, display-name deduplication,
emotion-example polish, and voice/visual resource integration. No
cross-role history viewer or new sharing ledger.

Additional per-branch validation after the unused-import cleanup: Ruff
passed for all tracked Python paths with repository exclusions, and
`python -m pytest tests/test_provider_prompt_routing.py -q` passed
**2/2** on each PR branch. All 15 checks on pre-sync head `1dcf4d1`
passed. Fresh checks for ancestry-only merge `4488a47` have been
triggered; their result is pending.

## Final check

- [x] One coherent character-runtime integration without unrelated
feature work
- [x] No speculative identity schema, model pass, retry policy or
compatibility fallback
- [x] No secrets, real conversations, local runtime state, model weights
or restricted assets included
- [x] Third-party notices and provenance preserved
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant