Read the question before the score. Browse every recovered Christian scenario and its original criteria, then inspect a real historical side-by-side comparison.

The adviser should receive only the public question or conversation. References, expected tensions, hidden facts and scoring criteria are displayed here for the reader; they are not all adviser inputs.

Christian question index

What a simulated user knows

These two original cards show the intended conversation boundary. The simulator knows the fixed facts; the adviser must elicit them. The simulator does not see the judge's weights or an MHF treatment label.

Historical evidence · not the new experiment

A result worth keeping—even though MHF lost.

The latest q25 artifact contains 24 complete pairs and one MHF failure. Under the historical MoReBench source-rubric adapter, the ordinary pastoral control averaged 62.29 and staged MHF 58.16: a −4.12 point difference. MHF won 7 pairs, the control won 14, and 3 tied.

The generating model was labeled “Opus 4.8 Medium (requested)” and the judge “Grok 4.5.” This was a multi-stage workflow on general procedural criteria, not a literal one-call Christian comparison. The artifact flags 103 invalid citation-evidence spans across the judgments; those flags did not change its stored scores. Treat the grades as diagnostic and inspect the text.

Source: original public q25 artifact. A separate earlier q25 cohort has no overlapping question IDs; the two are never joined by their shared “q25” nickname.

MoReBench scenarios and contextual rubrics: Chiu et al., MoReBench, CC BY 4.0. The historical q25 artifact adds model answers and a project-specific judging adapter; it is not an official MoReBench score.