Read the question before the score. Browse every recovered Christian scenario and its original criteria, then inspect a real historical side-by-side comparison.
The adviser should receive only the public question or conversation. References, expected tensions, hidden facts and scoring criteria are displayed here for the reader; they are not all adviser inputs.
These two original cards show the intended conversation boundary. The simulator knows the fixed facts; the adviser must elicit them. The simulator does not see the judge's weights or an MHF treatment label.
The latest q25 artifact contains 24 complete pairs and one MHF failure. Under the historical MoReBench source-rubric adapter, the ordinary pastoral control averaged 62.29 and staged MHF 58.16: a −4.12 point difference. MHF won 7 pairs, the control won 14, and 3 tied.
The generating model was labeled “Opus 4.8 Medium (requested)” and the judge “Grok 4.5.” This was a multi-stage workflow on general procedural criteria, not a literal one-call Christian comparison. The artifact flags 103 invalid citation-evidence spans across the judgments; those flags did not change its stored scores. Treat the grades as diagnostic and inspect the text.
Source: original public q25 artifact. A separate earlier q25 cohort has no overlapping question IDs; the two are never joined by their shared “q25” nickname.
MoReBench scenarios and contextual rubrics: Chiu et al., MoReBench, CC BY 4.0. The historical q25 artifact adds model answers and a project-specific judging adapter; it is not an official MoReBench score.