Build one complete comparison, then scale only when its answers and judgments are useful. “Test plan” here means the research experiment.

  1. Inspect the ingredients Current target

    Use this page to read the motivation, datasets, scenarios, process versions, generation prompts and judging prompts. Recover the original OpenProse source or mark a reconstruction explicitly. Decide which versions enter the first run.

  2. Run a small one-shot pilot

    Proposed pilot: a small, varied selection of existing Christian cases on one current frontier model. Run question-only, ordinary Christian and Christian + MHF arms. Each answer gets exactly one generation call. Judge all arms under the same blinded criteria and inspect the full answers.

  3. Freeze and compare frontier models

    Choose dated model IDs and settings, case IDs, prompt versions, judge versions, exclusions and an explicit spend ceiling. Run a broader paired comparison. Show within-model effects separately from whether a cheaper or smaller MHF model beats an unaugmented frontier model.

  4. Run the dialogue comparison

    Give ordinary Christian and MHF advisers the same public opening and opportunity to ask questions. A person or labeled simulator answers from fixed facts. Compare final advice, facts discovered, turns, latency and total cost. Include a one-shot answer given the full facts to distinguish information access from reasoning quality.

  5. Publish the evidence, then write the paper

    Release inspectable questions, exact instructions, answers or conversations, judging reasons, scores, failures and cost. Report losses and uncertainty. A large improvement must be broad and useful enough to justify the eventual claim.

What would count as a convincing result?

A consistent improvement on the fixed Christian endpoint across cases and models, without hiding serious advice failures, plus a dialogue improvement that survives equal opportunities to ask and the full-information comparison. Report paired differences and uncertainty by scenario; multiple turns or repeated phrasings of one scenario are not independent cases.

The minimum useful gain, scored cohort size, exact frontier and judge IDs, dialogue turn limit, and call/dollar ceiling are decisions still to freeze. “Revolutionary” is the ambition; the evidence determines whether the word is earned.