Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
25 commits
Select commit Hold shift + click to select a range
c737116
fix(mcp): classify local decision advice conservatively
claude Sep 29, 2026
84c4e9a
docs: note the Command Code hook routing upgrade path
claude Sep 29, 2026
29a2567
refactor(store): drop the empty workspace branch in edges_for
claude Sep 29, 2026
07b3c70
chore(evidence): bind review follow-ups to immutable v116 evidence
claude Sep 29, 2026
f17d1d6
fix(mcp): flag forced checkouts and require a shared subject for supe…
claude Sep 29, 2026
d302365
chore(evidence): bind review fixes to immutable v117 evidence
claude Sep 29, 2026
75e7814
fix(mcp): honor git global options and negation equivalence in local …
claude Sep 29, 2026
b179129
chore(evidence): bind round-two review fixes to immutable v118 evidence
claude Sep 29, 2026
e72f142
fix(mcp): compare negation targets and flag path-specific discards
claude Sep 29, 2026
c8d2dc2
chore(evidence): bind round-three review fixes to immutable v119 evid…
claude Sep 29, 2026
315d501
fix(mcp): defer conflicting values and flag worktree and force-create…
claude Sep 29, 2026
bb85823
chore(evidence): bind round-four review fixes to immutable v120 evidence
claude Sep 29, 2026
a5c9ead
fix(mcp): oppose only ruled-out values and widen credential and disca…
claude Sep 29, 2026
4cb4909
chore(evidence): bind round-five review fixes to immutable v121 evidence
claude Sep 29, 2026
b8e5455
fix(mcp): uncap git global options and fold clustered flags and cannot
claude Sep 29, 2026
13ac391
chore(evidence): bind round-six review fixes to immutable v122 evidence
claude Sep 29, 2026
2adee81
fix(mcp): read contrasts, negated failures and terse negations correctly
claude Sep 29, 2026
7d1fa9a
chore(evidence): bind round-seven review fixes to immutable v123 evid…
claude Sep 29, 2026
86c0fe3
Merge the 1.7.9 release preparation (#242) into these follow-ups
claude Sep 29, 2026
769b728
fix(mcp): read "no longer" as a negation in completion checks
claude Sep 29, 2026
d42c413
fix(llm): make Anthropic client work with Claude 5.x models
claude Sep 29, 2026
1c473e2
test(e2e): use the current Anthropic default model in the LLM status …
claude Sep 29, 2026
c5c27b8
fix(mcp): keep program-running and file-writing options out of read_only
claude Sep 29, 2026
1b1a0b3
chore(evidence): bind the restacked tree to immutable v124 evidence
claude Sep 29, 2026
fbb088e
fix(decide,llm): close review gaps in model detection and completion …
claude Sep 29, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .claude-plugin/skill-assets.sha256
Original file line number Diff line number Diff line change
Expand Up @@ -3,4 +3,4 @@ c5d0c26f28c9ee14092f9deaf24c98dd8bef49d971fef2b7a537ffb1ab9f2887 .claude-plugin
aeee7a94671ceb306fe2d24c5acc9f2d96ad8a8e7410536566799eea6265f080 skills/engraphis-memory/SKILL.md
055655db84af07561d002f0c69744313d8413c39f3e873f941f0fa0b1e76dc66 skills/engraphis-memory/references/CONVENTIONS.md
9d090a03f5b3f36a34d91f66b72c3844591f6915755ac3a6c6ba5f1b16977de5 skills/engraphis-memory/references/SCOPING.md
210e42da31fc68f41c83b3ab9eb1f91a415dbc1ef148f75b6b170bddc35d5752 skills/engraphis-memory/references/TOOLS.md
07de31349fc135895377cfcf852c490f732c19e83a627c311d724c5d2146dbb5 skills/engraphis-memory/references/TOOLS.md
5 changes: 4 additions & 1 deletion .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -196,7 +196,7 @@ ENGRAPHIS_RETENTION_SUPERVISOR=none
ENGRAPHIS_LLM_PROVIDER=openai
# Model name (provider-specific):
# openai: gpt-4o-mini, gpt-4o, gpt-4.1-mini, o4-mini ...
# anthropic: claude-3-5-haiku-20241022, claude-3-5-sonnet-20241022 ...
# anthropic: claude-sonnet-5-5, claude-opus-5-5, claude-haiku-4-5 ...
# google: gemini-1.5-flash, gemini-2.0-flash ...
# openrouter: anthropic/claude-3.5-sonnet, openai/gpt-4o-mini ...
# custom: any model name your OpenAI-compatible endpoint accepts
Expand All @@ -210,6 +210,9 @@ ENGRAPHIS_LLM_MODEL=gpt-4o-mini
# ENGRAPHIS_LLM_BASE_URL=https://openrouter.ai/api/v1
# Optional: extra headers (JSON string) for custom providers.
# ENGRAPHIS_LLM_EXTRA_HEADERS={"HTTP-Referer":"https://myapp.com","X-Title":"engraphis"}
# Optional: reasoning effort (low|medium|high|xhigh|max) for Claude models that think by
# default (Opus 5+, Sonnet 5+, Fable). Default medium; ignored by other providers and models.
# ENGRAPHIS_LLM_EFFORT=medium

# ── Hosted Pro / Team customer client ───────────────────────────────────────
# Cloud Sync, Analytics, Automation, Auto Dreaming, Auto Consolidation, and Team
Expand Down
10 changes: 5 additions & 5 deletions BENCHMARKS.md
Original file line number Diff line number Diff line change
Expand Up @@ -94,14 +94,14 @@ interpretation and do not count as additional benchmark-quality gains.
### Public numeric evidence registry

Every exact public aggregate retained below comes from the checked-in, public-safe
[`offline-fixtures-v117.json`](docs/benchmark-evidence/offline-fixtures-v117.json) artifact. Its
[`offline-fixtures-v125.json`](docs/benchmark-evidence/offline-fixtures-v125.json) artifact. Its
SHA-256 is
`c64eadffbc7f87938e0e822b7aab33dcfece6fffa5175932ac208465d36df192`, also recorded in the
`1f74971d6213a188b31cf58f6ff6132a487da36d22455292b028dadc202a8feb`, also recorded in the
adjacent `.sha256` file. The artifact contains no raw questions, answers, prompts, customer data,
or per-record content fingerprints.

The fixture-suite digest is
`77b140068cd036330b421a0d7668847d196f1bbeaa36c965d68943457002e6eb`. The artifact defines
`42cb269b867a1e3210d0da77d6a040abc974cd44317b91ad10a2d483d07c1586`. The artifact defines
the digest algorithm and records the SHA-256 of every suite and dataset file. Each evidence ID
also binds its exact command through `sha256(UTF-8 exact command)`:

Expand All @@ -123,10 +123,10 @@ Historical LoCoMo, graph, handoff, consolidation, and security figures remain pr
source artifacts but are omitted from the current chart until each has a matching immutable,
public-safe artifact. The chart labels coding outcomes, external datasets, and operational
capacity as pending evaluation tracks rather than implying scores. Regenerate it with
`python scripts/render_benchmark_report.py --report docs/benchmark-evidence/offline-fixtures-v117.json --output docs/images/context-efficiency.svg` after selecting the report to publish.
`python scripts/render_benchmark_report.py --report docs/benchmark-evidence/offline-fixtures-v125.json --output docs/images/context-efficiency.svg` after selecting the report to publish.

The companion examples are also generated from that artifact with
`python -m scripts.render_benchmark_examples --report docs/benchmark-evidence/offline-fixtures-v117.json --output docs/images/evidence-backed-agent-examples.svg`.
`python -m scripts.render_benchmark_examples --report docs/benchmark-evidence/offline-fixtures-v125.json --output docs/images/evidence-backed-agent-examples.svg`.
The historical-to-executable mapping is in
[`docs/BENCHMARK_CHANGE_COVERAGE.md`](docs/BENCHMARK_CHANGE_COVERAGE.md).

Expand Down
39 changes: 38 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,42 @@ All notable changes to Engraphis are documented here. Format loosely follows

## [Unreleased]

- Classify local `engraphis_decide` command advice conservatively. Chained, piped,
substituted or redirected commands, and options that run programs or write files
(`rg --pre`, `pytest --basetemp`), are never labeled read-only. Recursive deletes in any
flag order, raw device writes, history-rewriting or work-discarding Git operations, SQL
and infrastructure teardown, downloaded-script execution, exfiltrating pipes and uploads,
and well-known credential files are labeled destructive or leaking. Screening is bounded
so adversarial input stays cheap.
- Match whole words, ignore zero counts and treat negated success as failure in local
completion checks; contrasts ("not only passed") and negated failures ("did not fail") are
not failures. Local support and contradiction checks compare content words rather than
shared stopwords, and support keeps one-character terms such as "C". Supersession requires
a shared subject plus a replacement cue the existing fact lacks, or a negation (including
"cannot") whose ruled-out clause the other fact asserts ("does not use port 80" opposes
"uses port 80", not "uses port 443"); between terse facts one shared word is the subject,
so "No SQLite" opposes "Use SQLite". Reinforcement requires matching cues and one fact
containing the other's content words, so conflicting values defer. Any number of Git global
options and clustered short flags no longer bypass destructive-command checks, and
path-specific or pathspec-file checkouts, worktree restores and force-created branch resets
count as discarding work. Credential paths match either path separator.
- Fixed Anthropic connections for current Claude models. Opus 4.7 and later, Sonnet 5 and
later, and Fable no longer receive `temperature`, which they reject. Replies are read from
text blocks, so a leading thinking block no longer fails with "Unexpected Anthropic response
format". Models that think by default get `ENGRAPHIS_LLM_EFFORT` (default `medium`) and at
least 4096 output tokens so reasoning cannot crowd out the reply.
- Replaced the retired `claude-3-5-sonnet-20241022` Anthropic default in the dashboard picker,
the API defaults, `.env.example`, and the provider guide with `claude-sonnet-5-5`.
- Documented how to back the experimental Jev decision adapter with Claude: pin an exact model
id, keep fallback disabled, and avoid sampling parameters and forced tool choice.
- Recognized `claude-mythos-preview` and other unversioned Fable/Mythos ids, so they no longer
receive `temperature` (an HTTP 400) and get the thinking-model output headroom.
- A remote `verify_completion` probability between the certainty bound and the 0.85 completion
bar is now `uncertain` with a null `is_complete`, instead of a decisive failure.
- Local completion checks recognize `N passing` (Mocha/Jest style), `not passing`, `0 passing`
and `errors: none`.
- Refreshed the public offline fixtures and source bindings in immutable v125 evidence.

## [1.7.9] - 2026-09-29

- Added the saved-session managed Jev transport and Classic/Smart MCP decision route
Expand Down Expand Up @@ -53,7 +89,8 @@ All notable changes to Engraphis are documented here. Format loosely follows
session or repo mapping, report the resolved destination, and reject session mismatches.
- Command Code's SessionStart hook now uses the nearest Git root's repo name, honors saved
workspace mappings unless explicitly overridden, and labels recalled context with the
server's resolved workspace.
server's resolved workspace. Save a mapping to keep recalling memories stored under the
earlier folder-named workspace default.
- Added a previewed selective move workflow for organizing mixed workspaces while retaining
source history and enforcing move eligibility and workspace access.
- Hardened the experimental Cloud decision client with validated destinations,
Expand Down
7 changes: 5 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,9 +84,9 @@ neither is an end-to-end question-answer score. Coding outcomes, external datase
operational capacity remain separate pending evaluation tracks until their artifacts are selected.

These values are evidence IDs `offline-chunking` and `offline-performance` in
[`offline-fixtures-v117.json`](https://github.com/Coding-Dev-Tools/engraphis/blob/main/docs/benchmark-evidence/offline-fixtures-v117.json),
[`offline-fixtures-v125.json`](https://github.com/Coding-Dev-Tools/engraphis/blob/main/docs/benchmark-evidence/offline-fixtures-v125.json),
SHA-256
`c64eadffbc7f87938e0e822b7aab33dcfece6fffa5175932ac208465d36df192`.
`1f74971d6213a188b31cf58f6ff6132a487da36d22455292b028dadc202a8feb`.
[`BENCHMARKS.md`](https://github.com/Coding-Dev-Tools/engraphis/blob/main/BENCHMARKS.md#public-numeric-evidence-registry)
records the matching suite digest, exact commands, and per-command config digests. The offline
fixture registry intentionally excludes external, model-dependent, consolidation, productivity,
Expand Down Expand Up @@ -442,6 +442,8 @@ open on timeout and is installed via `python scripts/install_cc_hook.py`.
The hook sends the nearest Git root's name as `repo` and lets the server apply a saved workspace
mapping. Set `ENGRAPHIS_HOOK_WORKSPACE` only for an explicit override; a previous `default`
override must be cleared to use the mapping. Its context header shows the resolved workspace.
Earlier versions defaulted to a workspace named after the project folder; save that mapping to
keep recalling those memories at session start.

### prime-agent fleet

Expand Down Expand Up @@ -848,6 +850,7 @@ file. It never searches the working directory for `.env`, and explicit process v
| `ENGRAPHIS_LLM_MODEL` | `gpt-4o-mini` | Model name (provider-specific) |
| `ENGRAPHIS_LLM_API_KEY` | Not set | API key for chat/synthesis, `llm` / `llm_structured` extraction, and structured consolidation |
| `ENGRAPHIS_LLM_BASE_URL` | Not set | Base URL for openrouter / custom OpenAI-compatible endpoints |
| `ENGRAPHIS_LLM_EFFORT` | `medium` | Reasoning effort (`low \| medium \| high \| xhigh \| max`) for Claude models that think by default (Opus 5+, Sonnet 5+, Fable); ignored by other providers and models |
| `ENGRAPHIS_DECISION_BACKEND` | `none` | `none` or `local` keeps advisory decisions local; `managed` uses the saved Cloud session and included allowance; `auto` selects managed when configured and never switches to BYOK; explicit `byok` uses a personal TypeSafe key. Legacy `typesafe`, `jev`, and `system1` mean BYOK. Remote calls also require per-call consent. |
| `ENGRAPHIS_DECISION_MODEL` | `jev-1.13.0` | Pinned model accepted by the Jev transport; other model identifiers are rejected. |
| `TYPESAFE_API_KEY` | Not set | Personal credential for explicit BYOK decisions; `JEV_API_KEY` is a fallback alias. Managed decisions use the saved Cloud session instead. |
Expand Down
20 changes: 19 additions & 1 deletion docs/LLM_PROVIDERS.md
Original file line number Diff line number Diff line change
Expand Up @@ -54,6 +54,7 @@ Every LLM setup uses these variables:
| `ENGRAPHIS_LLM_API_KEY` | Credential for the provider. It is never returned by the dashboard. |
| `ENGRAPHIS_LLM_BASE_URL` | Needed only to override a default or configure a compatible endpoint. |
| `ENGRAPHIS_LLM_EXTRA_HEADERS` | Optional JSON object of headers required by a compatible endpoint. |
| `ENGRAPHIS_LLM_EFFORT` | Reasoning effort for Claude models that think by default: `low`, `medium` (default), `high`, `xhigh`, or `max`. Ignored elsewhere. |

The sample names below are Engraphis runtime defaults, not provider recommendations. Replace them
when your account or deployment uses a different model.
Expand Down Expand Up @@ -91,13 +92,30 @@ it as `custom`; the native mode applies Anthropic's required request shape and h

```dotenv
ENGRAPHIS_LLM_PROVIDER=anthropic
ENGRAPHIS_LLM_MODEL=claude-3-5-sonnet-20241022
ENGRAPHIS_LLM_MODEL=claude-sonnet-5-5
ENGRAPHIS_LLM_API_KEY=<anthropic-api-key>
```

Leave `ENGRAPHIS_LLM_BASE_URL` unset for the public API. For model and credential details, see
the [Anthropic API documentation](https://docs.anthropic.com/).

Choosing a model:

- `claude-sonnet-5-5` is the default. Extraction, consolidation summaries, and grounded synthesis
are bounded tasks, and it costs half as much per token as Opus.
- `claude-opus-5-5` suits the hardest consolidation and conflict-review work. Compare it against
the Sonnet default on your own data before paying for it, because the two are close on everyday
tasks.
- Retired ids such as `claude-3-5-sonnet-20241022` and `claude-3-5-haiku-20241022` are rejected by
the API; the connection test reports them as an HTTP 404.

Opus 4.7 and later, Sonnet 5 and later, Fable and Mythos (including `claude-mythos-preview`)
reject `temperature` and similar sampling parameters, so Engraphis omits them for those models.
Opus 5 and later, Sonnet 5 and later, Fable and Mythos also think before answering by default. Engraphis sends `ENGRAPHIS_LLM_EFFORT` (`low`,
`medium`, `high`, `xhigh`, or `max`; default `medium`) for those models and keeps at least 4096
output tokens available so hidden reasoning cannot crowd out the reply. Other providers and older
Claude models ignore the setting.

## Google Gemini

Google Gemini uses the native `google` mode and the Gemini `generateContent` API. The native mode
Expand Down
4 changes: 4 additions & 0 deletions docs/MCP_TOOLS.md
Original file line number Diff line number Diff line change
Expand Up @@ -190,6 +190,10 @@ and `query`; `verify_completion` needs `state` and `goal`, with optional `recent
before backend lookup, with unknown/null conclusions and no remote allowance consumed.
Command decisions always return `allow_auto=false` and `escalate_to_user=true`, including
successful remote answers. Provider probability and category are advice, not shell authorization.
Local command labels are coarse: only one simple inspection command, without chaining, pipes,
substitution or file redirection, is `read_only`. Recognized destructive, history-rewriting,
exfiltrating or credential-file commands are `destructive_or_leak`; anything else, including a
command too long to screen completely, is `state_change`.

The classic recall, grounded, and answer tools (`engraphis_recall`,
`engraphis_recall_grounded`, and the `engraphis_answer` alias) accept `planning="off"|"auto"`,
Expand Down
5 changes: 5 additions & 0 deletions docs/WORKSPACE_ORGANIZATION.md
Original file line number Diff line number Diff line change
Expand Up @@ -111,6 +111,11 @@ A nonblank `ENGRAPHIS_HOOK_WORKSPACE` is an explicit override. Clear a previous
override to use project mappings. The recalled-context header names the workspace returned by
the server. The hook remains silent on errors or empty recall results.

Earlier hook versions used a workspace named after the project folder when no override was set.
Without a saved mapping, the hook now starts in `default` instead, so those memories stop
appearing at session start. To keep using them, save the mapping once, for example
`engraphis_set_workspace_routing(workspace="website", repo="website")`, or move the memories.

## Organize existing memories

Routing changes future writes. Existing memories stay where they are until explicitly moved.
Expand Down
Loading
Loading