Skip to content

knowledge: improve review precision from BCApps PR 11063 feedback - #162

Open
Wenjie Fan (gggdttt) wants to merge 1 commit into
mainfrom
self-improvement/bcapps-11063-quality
Open

knowledge: improve review precision from BCApps PR 11063 feedback#162
Wenjie Fan (gggdttt) wants to merge 1 commit into
mainfrom
self-improvement/bcapps-11063-quality

Conversation

@gggdttt

@gggdttt Wenjie Fan (gggdttt) commented Sep 8, 2026

Copy link
Copy Markdown
Collaborator

Summary

Improves BCQuality knowledge based on maintainer thumbs-down feedback from explicit BCApps PR review runs.

Source feedback

Validation

  • Frontmatter and article structure validation completed.

Generated by the BC-ALAgentsInternal self-improvement workflow.

Offline evaluation: regression

Candidate correctness failed: unexpected or missed findings remain. Existing ignored gold comments retain their neutral scoring semantics.

Common engine code: ecf8e31759d6ddd6d78e3a0b7836b40134368009; dataset SHA256: ACDC50C36CC88DD31A5AB628A7C3D06A44E5F7628DA9DA79A0916E6AF0F46222.
Model: gpt-5.6-luna; judge: gpt-5.3-codex. Exact D: synthetic__perf-watermark-folder-sync-01.
Dataset base: 231538937b19317eddc3ac654459ca29b3ef4170; candidate: 46b6ad94497fcb272417b170b2cbb7b3f96d7209.

Arm Evaluation commit Engine pin Knowledge pin
baseline 27da12f4d5dfec9862f23219b9d1e80c465cc18a 4407f20b44fe92e3693d0444d25cc3430f1d5262 8584217c7506eea7eef27a9db353d78c0cd6a9a1
candidate 09695a2fecf379e823762ccbc425a09b9259a823 35b7bcdecf431b9b1a18ca83830a38c2fc353344 e74e2a6601ec87a1ffec6e3d7eb074aa67437ef2

Baseline

Run: https://github.com/microsoft/BC-Bench/actions/runs/34221111997; conclusion: success; wall clock: 12.6 minutes.
Raw findings and runtime pin proofs are retained in the self-improvement-feedback artifact and the linked evaluation run.

Entry Expected Generated Missed Unexpected F1 AI credits Agent seconds
synthetic__perf-watermark-folder-sync-01 0 2 0 2 0 unavailable 572.746033486

Observed AI credits: unavailable; coverage 0/1 entries. Missing entries are not extrapolated into a total.

Aggregate metric Value
total 1
expected_comment_count 0
generated_comment_count 2
matched_comment_count 0
missed_comment_count 0
incorrect_comment_count 2
precision 0
recall 1
f1 0

Candidate

Run: https://github.com/microsoft/BC-Bench/actions/runs/34222250479; conclusion: success; wall clock: 7 minutes.
Raw findings and runtime pin proofs are retained in the self-improvement-feedback artifact and the linked evaluation run.

Entry Expected Generated Missed Unexpected F1 AI credits Agent seconds
synthetic__perf-watermark-folder-sync-01 0 1 0 1 0 unavailable 234.327608426

Observed AI credits: unavailable; coverage 0/1 entries. Missing entries are not extrapolated into a total.

Aggregate metric Value
total 1
expected_comment_count 0
generated_comment_count 1
matched_comment_count 0
missed_comment_count 0
incorrect_comment_count 1
precision 0
recall 1
f1 0

Missing telemetry is unavailable, not zero. Evaluation-only metrics exclude candidate generation and are not the full-cycle cost.
Human review must verify source-patch fidelity, gold correctness, and target attribution. No automatic merge or branch-protection claim is made.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant