Existing work, made inspectable

Twenty-five questions. The actual prompts, answers, and judge scores.

This page shows what the finished exploratory attempt produced—without rerunning models or smoothing over the failed row.

25attempted
24paired complete
62.29pastoral mean
58.16MHF mean
−4.12MHF delta
What is compared One-call pastoral answer versus Full MHF workflow, both generated with requested Opus 4.8 Medium.
Who judged it Grok 4.5 scored each complete answer against the released source rubric. There was no human judging.
Important limit Exploratory public train/dev quality gate with a source-rubric adapter—not official MoReBench reproduction, human judgment, or safety certification.
Loading questions…

Loading the frozen review artifact…