Proposal
Add REFUTE to related evaluation / RAG / retrieval / scientific AI tooling docs if relevant.
REFUTE is a scientific critique + calibration benchmark (paper-grounded claims → predictions → judge scores → Brier/ECE).
Happy to adjust wording / PR if preferred.
Proposal
Add REFUTE to related evaluation / RAG / retrieval / scientific AI tooling docs if relevant.
REFUTE is a scientific critique + calibration benchmark (paper-grounded claims → predictions → judge scores → Brier/ECE).
Happy to adjust wording / PR if preferred.