Skip to content

feat(evaluators): Add evaluator for LLM - #280

Draft
wrisa wants to merge 10 commits into
mainfrom
feature/SAO-16858-add-evaluator-for-LLM
Draft

wrisa wants to merge 10 commits into
mainfrom
feature/SAO-16858-add-evaluator-for-LLM

Conversation

@wrisa

@wrisa wrisa commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Added a new galileo.llm evaluator that invokes Galileo's LLM-as-judge scorer API directly from Agent Control. The evaluator follows the same architecture as the existing galileo.luna evaluator — it uses an internal JWT for
    authentication, supports all the same comparison operators and threshold types, accepts evaluate_with_context and evaluate_with_extensions for dual-write and authenticated execution context, and is registered as a plugin entry point.

Intentional divergence from luna evaluator introduced in this PR

  • Feature flag gates the entire evaluator (is_available()), not just runtime records

    Luna's GALILEO_FEATURE_FLAG_SCORER_INVOKE_RUNTIME flag only suppresses the structured record and execution_context fields — the evaluator still loads and invokes the scorer using legacy inputs.

    The LLM evaluator has no legacy fallback: Orbit requires execution_context to fetch LLM credentials, so there is no inputs-only path that can succeed. GALILEO_FEATURE_FLAG_LLM_INVOKE_RUNTIME=enabled gates is_available() itself — when the flag is absent the evaluator is not loaded at all, rather than silently degrading.

  • step and execution_context are required, not optional

Luna proceeds with inputs-only when step=None or execution_context=None. The LLM evaluator fails fast in _evaluate() with a clear error result for either condition, and client.invoke() raises ValueError if either is absent. The engine always supplies both via evaluate_with_extensions, so this only affects out-of-engine direct usage.

  • llm/config.py — ScorerInvokeConfig is leaner than Luna's

    Luna's ScorerInvokeConfig carries two legacy fields for backward compatibility with existing Orbit callers:

threshold: float | None = None        # legacy Orbit field
score_threshold: float | None = None  # legacy Orbit field

The LLM evaluator was written after these were identified as legacy, so galileo.llm's ScorerInvokeConfig omits them entirely. galileo.luna retains them to avoid breaking existing deployments.

  • llm/client.py — invoke() takes config: ScorerInvokeConfig | None, not ScorerInvokeConfig | JSONObject | None

Luna's invoke() still accepts a raw dict for config (with a model_validate compat branch) to support callers that were passing dicts before ScorerInvokeConfig existed. The LLM evaluator has no such callers, so it accepts only the typed model.

Scope

User-facing / API changes:

  • New evaluator type galileo.llm available when agent-control-evaluator-galileo is installed. Configured with scorer_id, threshold, operator, scorer_version_id and optional scorer_label, scorer_config, payload_field, timeout_ms.
  • from agent_control.evaluators import LlmEvaluator, LlmEvaluatorConfig now available with galileo extras.
  • New entry point galileo.llm registered in agent_control.evaluators.

Internal changes:

  • llm/client.py — GalileoLLMClient with connection pooling, JWT internal auth, CA file support, and the same connection tuning env vars as Luna (reusing GALILEO_LUNA_* env var names).
  • llm/config.py — LlmEvaluatorConfig with identical schema to LunaEvaluatorConfig.
  • llm/evaluator.py — LlmEvaluator with evaluate, evaluate_with_context, evaluate_with_extensions, and _execution_context_from_extensions.
  • evaluators/__init__.py in the SDK updated to optionally re-export LLM types.
  • llm/config.py — GALILEO_FEATURE_FLAG_LLM_INVOKE_RUNTIME env var and llm_invoke_runtime_enabled() function gate is_available().
  • llm/evaluator.py — LlmEvaluator with evaluate_with_extensions as the primary entry point (engine-facing); evaluate and evaluate_with_context fail fast since step and execution_context are both required. is_available() checks the feature flag in addition to import availability.

Out of scope:

  • Shared base class / deduplication of Luna and LLM code (intentionally kept as independent implementations to avoid coupling).

Risk and Rollout

Risk level: low

  • The LLM evaluator is opt-in — it only activates when a control is explicitly configured with type: galileo.llm. No existing Luna controls are affected.
  • The new entry point is registered alongside the existing galileo.luna entry point in the same package; if the package isn't installed the import silently falls back.

Rollback plan:

  • Remove or disable any controls using type: galileo.llm. No schema migrations or persistent state changes involved.

Testing

  • Added or updated automated tests
  • test_llm_evaluator.py — integration-style tests for execution context, extensions, org ID validation, and feature flag (GALILEO_FEATURE_FLAG_LLM_INVOKE_RUNTIME) gating is_available().
  • test_llm_coverage_gaps.py — 94 tests covering all utility helpers, operator branches, payload routing, connection config validation, JWT auth, CA file fallback, client pool rotation, HTTP/request error propagation, config threshold validation, step/execution_context required guards, and factory-based record construction. Brings llm/client.py to 97%, llm/evaluator.py to 97%, llm/config.py to 98% patch coverage.
  • Ran make check
  • lint and typecheck pass.
  • Manually verified behavior
  • scorer invocation tested against a live Galileo environment via ace-demo.

Checklist

  • Linked issue/spec
  • SAO-16858
  • evaluators/contrib/README.md, sdks/python/ARCHITECTURE.md, SDK __init__.py docstring updated to reference galileo.llm. SDK __init__.py and GalileoLLMClient env var docstring updated to document the GALILEO_FEATURE_FLAG_LLM_INVOKE_RUNTIME requirement.
  • Included any required follow-up tasks
  • LLM and Luna share ~80% identical code; the remaining divergence is intentional (feature flag scope, required vs optional step/execution_context, leaner ScorerInvokeConfig). A future refactor to extract a shared base (_shared/) is tracked separately and should account for these differences.

@codecov

codecov Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

@josjeon
josjeon self-requested a review October 7, 2026 16:28
wrisa and others added 6 commits October 7, 2026 13:41
Port the structured logging and execution-context identity log from
feature/SAO-16858-add-evaluator-for-LLM-thin (_shared/evaluator.py)
to the separate llm/evaluator.py and luna/evaluator.py files on this
branch. Adds dispatcher, skip, pre-invoke, result, and error logs to
_evaluate, and an identity log to _execution_context_from_extensions.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@wrisa
wrisa force-pushed the feature/SAO-16858-add-evaluator-for-LLM branch from a378e06 to 27d8ea0 Compare October 7, 2026 20:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants