Describe the bug
RubricBasedEvaluator keeps the effective rubric list as instance state: format_auto_rater_prompt calls create_effective_rubrics_list(actual_invocation.rubrics), which sets self._effective_rubrics_list, and convert_auto_rater_response_to_score later reads get_effective_rubrics_list() to map parsed verdicts back to rubrics.
LlmAsJudge.evaluate_invocations formats the prompt for every invocation first (the loop at the top of the method), then gathers all sample tasks. So by the time any response is converted, _effective_rubrics_list is whatever the last invocation's rubrics were. A verdict for a rubric that is on an earlier invocation but not on the last one is discarded with the warning Rubric ... not found in the rubrics provided to the metric.
Observed on 2.9.0; the same code is on main (2.10.0): rubric_based_evaluator.py lines ~374–417 and ~445, llm_as_judge.py lines ~215 and ~238.
To Reproduce
- An eval case with two invocations, where invocation 0 carries rubric
a and invocation 1 carries rubric b (per-invocation rubrics, type matching the metric), and no criterion-level rubrics.
- Run
RubricBasedFinalResponseQualityV1Evaluator.evaluate_invocations with any judge model that answers both prompts.
- Invocation 0's result has no rubric score for
a (the warning above is logged); invocation 1's has b.
With the same rubric on both invocations the bug is invisible, which is why criterion-level rubrics never hit it.
Expected behavior
Each sample's response is converted against the rubrics of the invocation it was sampled for.
Suggested fix
Make the effective list per invocation rather than per evaluator: return it from create_effective_rubrics_list and carry it alongside the LlmRequest into _evaluate_single_sample, or resolve it inside convert_auto_rater_response_to_score from the invocation that produced the response. The per-rubric majority in MajorityVotePerInvocationResultsAggregator is unaffected.
Desktop
- Python 3.13, google-adk 2.9.0 (also present on main)
Describe the bug
RubricBasedEvaluatorkeeps the effective rubric list as instance state:format_auto_rater_promptcallscreate_effective_rubrics_list(actual_invocation.rubrics), which setsself._effective_rubrics_list, andconvert_auto_rater_response_to_scorelater readsget_effective_rubrics_list()to map parsed verdicts back to rubrics.LlmAsJudge.evaluate_invocationsformats the prompt for every invocation first (the loop at the top of the method), then gathers all sample tasks. So by the time any response is converted,_effective_rubrics_listis whatever the last invocation's rubrics were. A verdict for a rubric that is on an earlier invocation but not on the last one is discarded with the warningRubric ... not found in the rubrics provided to the metric.Observed on 2.9.0; the same code is on
main(2.10.0):rubric_based_evaluator.pylines ~374–417 and ~445,llm_as_judge.pylines ~215 and ~238.To Reproduce
aand invocation 1 carries rubricb(per-invocationrubrics, type matching the metric), and no criterion-level rubrics.RubricBasedFinalResponseQualityV1Evaluator.evaluate_invocationswith any judge model that answers both prompts.a(the warning above is logged); invocation 1's hasb.With the same rubric on both invocations the bug is invisible, which is why criterion-level rubrics never hit it.
Expected behavior
Each sample's response is converted against the rubrics of the invocation it was sampled for.
Suggested fix
Make the effective list per invocation rather than per evaluator: return it from
create_effective_rubrics_listand carry it alongside theLlmRequestinto_evaluate_single_sample, or resolve it insideconvert_auto_rater_response_to_scorefrom the invocation that produced the response. The per-rubric majority inMajorityVotePerInvocationResultsAggregatoris unaffected.Desktop