Use each benchmark for the question it can answer. Christian moral advice, general reasoning quality, and factual religious knowledge belong in separate comparisons.
| Source | What we have / what it measures | Place in this project |
|---|---|---|
| Moral Restoration Bench Original Christian cases | 230 scenarios: 35 duty conflicts and 195 life-stage questions. Original Scripture references, expected tensions and item criteria are retained. | Primary candidate question pool. Inspect criteria for theological assumptions and scoring bias before freezing a cohort. Author-created cases are not independently validated ground truth. |
| Dialogue case cards | Two original scenarios with public openings, personas, hidden facts and reveal rules. | First conversation demonstration. Two scenarios can expose a broken interaction; they cannot establish a general conversational advantage. |
| MoReBench | A pinned 500-row public subset is retained. Its contextual criteria evaluate procedural moral reasoning; the public subset is labeled theory-neutral. | Separate secondary endpoint. Preserve the source score. Any Christian adaptation gets a distinct name and rationale. A disagreement with a criterion must be visible before results are scored. |
| FAI-C / Christian flourishing | A relevant external Christian evaluation framework and research comparison. Prior local work also used FAI-style rubrics and objective proxy questions. | External comparison to reconcile. Local proxies are not an official FAI-C reproduction. Confirm compatible tasks, prompts and scoring before making a leaderboard claim. |
| ClarifyDelphi | Prior work on clarification questions for moral dilemmas. | Dialogue motivation and comparison. Useful for asking whether a question elicits decision-relevant facts; it does not supply Christian moral authority. |
| ETHICS, Delphi, Social Chemistry, Moral Stories | Related work retained in the research bibliography: judgments, norms, and action/consequence narratives. | Context, not a mandatory new campaign. Use only when a concrete comparison advances one of the two hypotheses. |
The selected archive contains 4,500 answer rows from the historical 300-question, 15-workflow run, with 12,510 dimension-level judge records. It also contains 900 Gloo baseline answers and 834 judgments, plus a smaller matched comparison cohort. These are retained outputs, not independent new experiments.
Reuse an answer when its question, instructions and model settings match the comparison. Rejudge it when the scoring lens changed. Regenerate only when the treatment itself is wrong or missing. Total historical OpenRouter spend remains unknown.