Build one complete comparison, then scale only when its answers and judgments are useful. “Test plan” here means the research experiment.
Use this page to read the motivation, datasets, scenarios, process versions, generation prompts and judging prompts. Recover the original OpenProse source or mark a reconstruction explicitly. Decide which versions enter the first run.
Proposed pilot: a small, varied selection of existing Christian cases on one current frontier model. Run question-only, ordinary Christian and Christian + MHF arms. Each answer gets exactly one generation call. Judge all arms under the same blinded criteria and inspect the full answers.
Choose dated model IDs and settings, case IDs, prompt versions, judge versions, exclusions and an explicit spend ceiling. Run a broader paired comparison. Show within-model effects separately from whether a cheaper or smaller MHF model beats an unaugmented frontier model.
Give ordinary Christian and MHF advisers the same public opening and opportunity to ask questions. A person or labeled simulator answers from fixed facts. Compare final advice, facts discovered, turns, latency and total cost. Include a one-shot answer given the full facts to distinguish information access from reasoning quality.
Release inspectable questions, exact instructions, answers or conversations, judging reasons, scores, failures and cost. Report losses and uncertainty. A large improvement must be broad and useful enough to justify the eventual claim.
A consistent improvement on the fixed Christian endpoint across cases and models, without hiding serious advice failures, plus a dialogue improvement that survives equal opportunities to ask and the full-information comparison. Report paired differences and uncertainty by scenario; multiple turns or repeated phrasings of one scenario are not independent cases.
The minimum useful gain, scored cohort size, exact frontier and judge IDs, dialogue turn limit, and call/dollar ceiling are decisions still to freeze. “Revolutionary” is the ambition; the evidence determines whether the word is earned.