Skip to content

RubricBasedEvaluator drops per-invocation rubric verdicts when invocations carry different rubrics #7301

Description

@erauner12

Describe the bug

RubricBasedEvaluator keeps the effective rubric list as instance state: format_auto_rater_prompt calls create_effective_rubrics_list(actual_invocation.rubrics), which sets self._effective_rubrics_list, and convert_auto_rater_response_to_score later reads get_effective_rubrics_list() to map parsed verdicts back to rubrics.

LlmAsJudge.evaluate_invocations formats the prompt for every invocation first (the loop at the top of the method), then gathers all sample tasks. So by the time any response is converted, _effective_rubrics_list is whatever the last invocation's rubrics were. A verdict for a rubric that is on an earlier invocation but not on the last one is discarded with the warning Rubric ... not found in the rubrics provided to the metric.

Observed on 2.9.0; the same code is on main (2.10.0): rubric_based_evaluator.py lines ~374–417 and ~445, llm_as_judge.py lines ~215 and ~238.

To Reproduce

  1. An eval case with two invocations, where invocation 0 carries rubric a and invocation 1 carries rubric b (per-invocation rubrics, type matching the metric), and no criterion-level rubrics.
  2. Run RubricBasedFinalResponseQualityV1Evaluator.evaluate_invocations with any judge model that answers both prompts.
  3. Invocation 0's result has no rubric score for a (the warning above is logged); invocation 1's has b.

With the same rubric on both invocations the bug is invisible, which is why criterion-level rubrics never hit it.

Expected behavior

Each sample's response is converted against the rubrics of the invocation it was sampled for.

Suggested fix

Make the effective list per invocation rather than per evaluator: return it from create_effective_rubrics_list and carry it alongside the LlmRequest into _evaluate_single_sample, or resolve it inside convert_auto_rater_response_to_score from the invocation that produced the response. The per-rubric majority in MajorityVotePerInvocationResultsAggregator is unaffected.

Desktop

  • Python 3.13, google-adk 2.9.0 (also present on main)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions