Repository navigation
Conversation
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
josjeon
approved these changes
Oct 7, 2026
josjeon
self-requested a review
October 7, 2026 16:28
Port the structured logging and execution-context identity log from feature/SAO-16858-add-evaluator-for-LLM-thin (_shared/evaluator.py) to the separate llm/evaluator.py and luna/evaluator.py files on this branch. Adds dispatcher, skip, pre-invoke, result, and error logs to _evaluate, and an identity log to _execution_context_from_extensions. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This reverts commit ff00a4f.
wrisa
force-pushed
the
feature/SAO-16858-add-evaluator-for-LLM
branch
from
October 7, 2026 20:47
a378e06 to
27d8ea0
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
galileo.llmevaluator that invokes Galileo's LLM-as-judge scorer API directly from Agent Control. The evaluator follows the same architecture as the existinggalileo.lunaevaluator — it uses an internal JWT forauthentication, supports all the same comparison operators and threshold types, accepts
evaluate_with_contextandevaluate_with_extensionsfor dual-write and authenticated execution context, and is registered as a plugin entry point.Intentional divergence from luna evaluator introduced in this PR
Feature flag gates the entire evaluator (
is_available()), not just runtime recordsLuna's
GALILEO_FEATURE_FLAG_SCORER_INVOKE_RUNTIMEflag only suppresses thestructured recordandexecution_contextfields — the evaluator still loads and invokes the scorer using legacy inputs.The LLM evaluator has no legacy fallback: Orbit requires
execution_contextto fetch LLM credentials, so there is no inputs-only path that can succeed.GALILEO_FEATURE_FLAG_LLM_INVOKE_RUNTIME=enabledgatesis_available()itself — when the flag is absent the evaluator is not loaded at all, rather than silently degrading.stepandexecution_contextare required, not optionalLuna proceeds with inputs-only when
step=Noneorexecution_context=None. The LLM evaluator fails fast in_evaluate()with a clear error result for either condition, andclient.invoke()raisesValueErrorif either isabsent. The engine always supplies both viaevaluate_with_extensions, so this only affectsout-of-enginedirect usage.llm/config.py—ScorerInvokeConfigis leaner than Luna'sLuna's
ScorerInvokeConfigcarries two legacy fields for backward compatibility with existing Orbit callers:The LLM evaluator was written after these were identified as legacy, so
galileo.llm'sScorerInvokeConfigomits them entirely.galileo.lunaretains them to avoid breaking existing deployments.llm/client.py—invoke()takesconfig: ScorerInvokeConfig | None, notScorerInvokeConfig | JSONObject | NoneLuna's
invoke()still accepts a raw dict for config (with amodel_validatecompat branch) to support callers that were passing dicts beforeScorerInvokeConfigexisted. The LLM evaluator has no such callers, so it accepts only the typed model.Scope
User-facing / API changes:
galileo.llmavailable whenagent-control-evaluator-galileois installed. Configured withscorer_id, threshold, operator, scorer_version_id and optional scorer_label, scorer_config, payload_field, timeout_ms.agent_control.evaluators import LlmEvaluator, LlmEvaluatorConfignow available with galileo extras.galileo.llmregistered inagent_control.evaluators.Internal changes:
llm/client.py—GalileoLLMClientwith connection pooling, JWT internal auth, CA file support, and the same connection tuning env vars as Luna (reusing GALILEO_LUNA_* env var names).llm/config.py—LlmEvaluatorConfigwith identical schema toLunaEvaluatorConfig.llm/evaluator.py—LlmEvaluatorwithevaluate, evaluate_with_context, evaluate_with_extensions, and _execution_context_from_extensions.evaluators/__init__.pyin the SDK updated to optionally re-export LLM types.llm/config.py—GALILEO_FEATURE_FLAG_LLM_INVOKE_RUNTIMEenv var andllm_invoke_runtime_enabled()function gateis_available().llm/evaluator.py—LlmEvaluatorwithevaluate_with_extensionsas the primary entry point (engine-facing);evaluateandevaluate_with_contextfail fast sincestepandexecution_contextare both required.is_available()checks the feature flag in addition to import availability.Out of scope:
Risk and Rollout
Risk level: low
galileo.llm. No existing Luna controls are affected.galileo.lunaentry point in the same package; if the package isn't installed the import silently falls back.Rollback plan:
galileo.llm. No schema migrations or persistent state changes involved.Testing
test_llm_evaluator.py— integration-style tests for execution context, extensions, org ID validation, and feature flag (GALILEO_FEATURE_FLAG_LLM_INVOKE_RUNTIME) gating is_available().test_llm_coverage_gaps.py—94tests covering all utility helpers, operator branches, payload routing, connection config validation, JWT auth, CA file fallback, client pool rotation, HTTP/request error propagation, config threshold validation, step/execution_context required guards, and factory-based record construction. Bringsllm/client.pyto97%,llm/evaluator.pyto97%,llm/config.pyto98%patch coverage.make checkace-demo.Checklist
evaluators/contrib/README.md, sdks/python/ARCHITECTURE.md,SDK __init__.pydocstring updated to referencegalileo.llm.SDK __init__.pyandGalileoLLMClientenv var docstring updated to document theGALILEO_FEATURE_FLAG_LLM_INVOKE_RUNTIMErequirement.~80%identical code; the remaining divergence is intentional (feature flag scope, required vs optional step/execution_context, leaner ScorerInvokeConfig). A future refactor to extract a shared base (_shared/) is tracked separately and should account for these differences.