Conversation
…ntext A scorer can declare a third parameter and receive a ScorerContext. The context holds the tool calls in start order. Two-parameter scorers work as before. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
| required = [ | ||
| parameter for parameter in positional if parameter.default is parameter.empty | ||
| ] | ||
| if len(required) > 3: | ||
| raise ValueError("scorer fn must accept at most (row, output, context)") | ||
| return len(positional) >= 3 |
There was a problem hiding this comment.
🟡 Unsupported scorers fail every evaluation row
A scorer with three positional parameters and a required keyword-only parameter passes _accepts_context validation. The runner omits that parameter, so every generated row receives a scorer_raised result instead of an upfront validation error.
Learn more
Scorers run once per generated evaluation row. The new signature check accepts any function with at most three required positional parameters, but does not examine required keyword-only parameters. The runner calls the scorer using only two or three positional arguments in _run_scorer_for_result. Thus the invalid scorer survives construction and produces an error for each row after the evaluation run has begun.
Example: def score(row, output, context, *, rubric): return 1.0 passes construction, but calling it with (row, output, context) raises a missing rubric error on every row.
Recommended fix: Reject required keyword-only parameters in _accepts_context, or bind the selected two- or three-argument call with inspect.Signature.bind at construction to reject callbacks that cannot be invoked by the runner.
Was this helpful? React with 👍 or 👎 to provide feedback.
Summary
A deterministic scorer can now grade the tool trajectory. Before, a scorer received only
(row, output).A scorer declares a third positional parameter and receives a
ScorerContext.tool_calls: the tool calls in the order they started. The list is empty when no tool ran.tool_calls_omitted: the number of calls made after the recording limit.A scorer with two parameters works as before. The SDK reads the signature once, in
Scorer.__post_init__. A function that requires more than three arguments raisesValueErrorat construction.Future inputs are new fields on
ScorerContext. The scorer signature does not change again.The spec change is in launchdarkly/ai-sdks-monorepo, TESTING.md §8.8.4, branch
claude/tool-trajectory-scorers-284fbe.Notes
argumentsandresulton each call are bounded text, the same values a judge reads.ScorerContextis exported fromlaunchdarkly_ai_server.evaluations.Test plan
uv run pytest packages/clientuv run mypy packages/*/srcuv run ruff check .🤖 Generated with Claude Code
Note
Overview
Deterministic evaluation scorers can now grade the handler’s tool trajectory, not just
(row, output).Opt in by adding a third positional parameter typed as
ScorerContext, which carriestool_calls(orderedToolInvocations from generation, empty if none) andtool_calls_omitted(count beyond the recording cap).Scorerdetects the extra parameter once at construction (accepts_context); two-argument scorers are unchanged. The runner builds context from the sametool_calls/tool_calls_omittedfields already attached to generated rows and passes it only when needed.ScorerContextis exported fromlaunchdarkly_ai_server.evaluations.Reviewed by Cursor Bugbot for commit 4456f28. Bugbot is set up for automated code reviews on this repo. Configure here.