Moral HierarchyResearch notebookExplore the questions

Moral Restoration · A fresh foundation

Moral clarity,
one question
at a time.

Can a simple moral hierarchy help AI give better Christian advice—and can dialogue make it substantially better?

The idea, the evidence we have, and the experiments we still need. Every prompt and decision has a place you can inspect.

Follow the research
01

Motivation

English source ↓

People ask moral questions because something they owe one person appears to conflict with something they owe another. A useful answer must explain what they should do, why, and what responsibility remains afterward.

Hypothesis 01 · One generation call

A moral hierarchy improves the answer.

Can a simple prompt, grounded in Christian moral authority and the relationships in a dilemma, outperform frontier models giving ordinary Christian advice? Compare the same model with and without the hierarchy, then compare across models.

Hypothesis 02 · Turn-taking

Better questions improve the judgment.

Can dialogue uncover facts that change which duties apply—and produce a large improvement in the final advice? The eventual conversation is with a person. Evaluation may use an LLM playing that person from a fixed case card.

Christian moral authorityDuties within relationshipsFacts, consequences & missing contextJudgment → action → repair

The original idea, in ordinary language

Start with the person's moral commitments. Identify the people affected and the duties created by those relationships. Gather the facts that determine how those duties apply. Resolve conflicts within the moral framework, give a clear or conditional recommendation, and acknowledge the costs that remain. Here, “moral residue” means the loss or obligation that remains even after a justified choice.

This project begins with a Christian account of morality. It does not treat popular agreement as a substitute for that account, or assume that a Christian label guarantees good counsel. The contribution being tested is a way of organizing reasoning, rather than a new moral theory.

Original source: Technical specification, March 30, 2026. Historical claims in that document remain historical claims. This page states the narrower experiments now proposed.

02

Data sets and comparisons

English source ↓

Use each benchmark for the question it can answer. Christian moral advice, general reasoning quality, and factual religious knowledge belong in separate comparisons.

Dataset and comparison index
SourceWhat we have / what it measuresPlace in this project
Moral Restoration Bench
Original Christian cases
230 scenarios: 35 duty conflicts and 195 life-stage questions. Original Scripture references, expected tensions and item criteria are retained.Primary candidate question pool. Inspect criteria for theological assumptions and scoring bias before freezing a cohort. Author-created cases are not independently validated ground truth.
Dialogue case cardsTwo original scenarios with public openings, personas, hidden facts and reveal rules.First conversation demonstration. Two scenarios can expose a broken interaction; they cannot establish a general conversational advantage.
MoReBenchA pinned 500-row public subset is retained. Its contextual criteria evaluate procedural moral reasoning; the public subset is labeled theory-neutral.Separate secondary endpoint. Preserve the source score. Any Christian adaptation gets a distinct name and rationale. A disagreement with a criterion must be visible before results are scored.
FAI-C / Christian flourishingA relevant external Christian evaluation framework and research comparison. Prior local work also used FAI-style rubrics and objective proxy questions.External comparison to reconcile. Local proxies are not an official FAI-C reproduction. Confirm compatible tasks, prompts and scoring before making a leaderboard claim.
ClarifyDelphiPrior work on clarification questions for moral dilemmas.Dialogue motivation and comparison. Useful for asking whether a question elicits decision-relevant facts; it does not supply Christian moral authority.
ETHICS, Delphi, Social Chemistry, Moral StoriesRelated work retained in the research bibliography: judgments, norms, and action/consequence narratives.Context, not a mandatory new campaign. Use only when a concrete comparison advances one of the two hypotheses.

The paid work is still useful

The selected archive contains 4,500 answer rows from the historical 300-question, 15-workflow run, with 12,510 dimension-level judge records. It also contains 900 Gloo baseline answers and 834 judgments, plus a smaller matched comparison cohort. These are retained outputs, not independent new experiments.

Reuse an answer when its question, instructions and model settings match the comparison. Rejudge it when the scoring lens changed. Regenerate only when the treatment itself is wrong or missing. Total historical OpenRouter spend remains unknown.

Build one complete comparison, then scale only when its answers and judgments are useful. “Test plan” here means the research experiment.

  1. Inspect the ingredients Current target

    Use this page to read the motivation, datasets, scenarios, process versions, generation prompts and judging prompts. Recover the original OpenProse source or mark a reconstruction explicitly. Decide which versions enter the first run.

  2. Run a small one-shot pilot

    Proposed pilot: a small, varied selection of existing Christian cases on one current frontier model. Run question-only, ordinary Christian and Christian + MHF arms. Each answer gets exactly one generation call. Judge all arms under the same blinded criteria and inspect the full answers.

  3. Freeze and compare frontier models

    Choose dated model IDs and settings, case IDs, prompt versions, judge versions, exclusions and an explicit spend ceiling. Run a broader paired comparison. Show within-model effects separately from whether a cheaper or smaller MHF model beats an unaugmented frontier model.

  4. Run the dialogue comparison

    Give ordinary Christian and MHF advisers the same public opening and opportunity to ask questions. A person or labeled simulator answers from fixed facts. Compare final advice, facts discovered, turns, latency and total cost. Include a one-shot answer given the full facts to distinguish information access from reasoning quality.

  5. Publish the evidence, then write the paper

    Release inspectable questions, exact instructions, answers or conversations, judging reasons, scores, failures and cost. Report losses and uncertainty. A large improvement must be broad and useful enough to justify the eventual claim.

What would count as a convincing result?

A consistent improvement on the fixed Christian endpoint across cases and models, without hiding serious advice failures, plus a dialogue improvement that survives equal opportunities to ask and the full-information comparison. Report paired differences and uncertainty by scenario; multiple turns or repeated phrasings of one scenario are not independent cases.

The minimum useful gain, scored cohort size, exact frontier and judge IDs, dialogue turn limit, and call/dollar ceiling are decisions still to freeze. “Revolutionary” is the ambition; the evidence determines whether the word is earned.

04

Question set and comparisons

English source ↓

Read the question before the score. Browse every recovered Christian scenario and its original criteria, then inspect a real historical side-by-side comparison.

The adviser should receive only the public question or conversation. References, expected tensions, hidden facts and scoring criteria are displayed here for the reader; they are not all adviser inputs.

Christian question index

What a simulated user knows

These two original cards show the intended conversation boundary. The simulator knows the fixed facts; the adviser must elicit them. The simulator does not see the judge's weights or an MHF treatment label.

Historical evidence · not the new experiment

A result worth keeping—even though MHF lost.

The latest q25 artifact contains 24 complete pairs and one MHF failure. Under the historical MoReBench source-rubric adapter, the ordinary pastoral control averaged 62.29 and staged MHF 58.16: a −4.12 point difference. MHF won 7 pairs, the control won 14, and 3 tied.

The generating model was labeled “Opus 4.8 Medium (requested)” and the judge “Grok 4.5.” This was a multi-stage workflow on general procedural criteria, not a literal one-call Christian comparison. The artifact flags 103 invalid citation-evidence spans across the judgments; those flags did not change its stored scores. Treat the grades as diagnostic and inspect the text.

Source: original public q25 artifact. A separate earlier q25 cohort has no overlapping question IDs; the two are never joined by their shared “q25” nickname.

MoReBench scenarios and contextual rubrics: Chiu et al., MoReBench, CC BY 4.0. The historical q25 artifact adds model answers and a project-specific judging adapter; it is not an official MoReBench score.

05

Methodology

English source ↓

Readable instructions define the treatment. OpenProse carries the workflow. Actual messages and calls establish what happened.

Process versions, from the original concept to the next experiment
VersionWhat it actually doesDecision / status
Original specificationGather missing facts; reason about duties under a declared moral authority; give a prescriptive, conditional or seek-guidance answer.Retained thesis. The specification describes conversation. It does not by itself prove an implemented dialogue result.
Early graph-assisted promptA Christian counseling prompt receives a precomputed MHF analysis alongside the question.Recovered historical text. Keep its prompt visible. Removing that graph context creates a new treatment.
Later “prose” modesPython intake and advice orchestration produce prose answers. A mode named “one-shot” may make several model calls; “3Q” can extract facts already present.Historical comparison only. Those names do not establish OpenProse execution, one-call prompting or real user turn-taking.
Literal one-call MHFThe common Christian control plus an explicit hierarchy instruction, sent in one generation request.Candidate. The text is inspectable below. It is newly authored, not a recovered original .prose program.
OpenProse dialogueAdviser asks; person or simulator supplies facts; adviser updates the judgment. Each role receives only its allowed inputs.Target workflow. Recover and preserve original source where possible. Show any reconstructed contract next to its historical inputs.
Reactor adaptationKeep authored OpenProse contracts; use Reactor to manage execution and saved state where its current runtime supports the required behavior.Migration candidate, not a demonstrated benchmark runtime. Keep compiler/orchestration work separate from counted adviser, simulator and judge calls.

Where Reactor fits

The current upstream Reactor CLI documentation separates compilation from execution and stores runtime state in a .reactor directory. Its offline inspection commands can be checked without a model call. A directory named .reactor is not a replacement prompt language. The restored prompts and contracts remain inspectable source.

Local check, September 12: Reactor CLI 0.2.4 resolved SDK 0.3.3 and reported a healthy offline environment. The compile check reported no .prose.md contracts in the fresh repository. This verifies tool availability; it does not demonstrate an advice run or a completed migration.

A fair comparison

  • Same task and opportunity. Match questions, model settings and output allowances across the paired arms. Match opportunities to ask questions in dialogue.
  • One judging lens per endpoint. Hide model and treatment labels from judges. Preserve the same criteria and authority for every answer in that comparison.
  • Keep the roles separate. Adviser: public question and revealed replies. Simulator: persona, fixed facts and reveal rules. Judge: final answer, required facts and frozen criteria.
  • Keep every attempt. Save exact outbound messages, returned text, failures, model identifiers, settings, usage and costs. Internal passes and user turns are separately labeled.
  • Freeze before scoring. Record versions, case selection, theological disputes, exclusions and thresholds before comparative results are examined. Public cases are not a hidden holdout.
  • Claim only what was measured. Model judging is labeled model judging. Simulated dialogue measures the defined simulation; it does not establish real-world pastoral benefit.
06

Prompt list

English source ↓

The prompt is part of the experiment. Read the full instruction, its source, and what changed before choosing a version.

Historical prompts are preserved for comparison. Candidate prompts are proposals. The current one-shot design has three arms: question only, ordinary Christian counsel, and the same Christian instruction plus MHF. Dialogue adds real or explicitly simulated user replies.

Download the complete version catalog. Each entry also links to its exact text. The historical q25 question selector exposes the complete saved per-case prompts and stage outputs.

07

Judge prompt list

English source ↓

Judge the substance of the advice: how Christian principles constrain the recommendation, how the answer handles the conflict, and whether the proposed action is justified and useful.

The original guidance already rejects slogans, bare citations, generic lists and eloquence as substitutes for applied reasoning. It also requires explicit authority grounding when that criterion applies, records harmful advice, and asks for brief evidence from the answer. Those commitments should survive the rebuild.

Judging decisions that must stay visible
QuestionHistorical guidanceProposed handling
How is authority credited?Apply a relevant authority to justify or constrain the recommendation; do not backfill implicit alignment.Keep that standard when authority grounding is required. Do not award credit for quotation count.
How is substance scored?Original per-case criteria and later binary rubric passes, with weights and penalties.Retain a criterion-level view. Show the new four-dimension 0–4 summary as a separate draft, not an interchangeable scale.
What about harmful advice?Record penalties even in an otherwise strong answer.Keep severe errors visible alongside any aggregate score. Averages must not conceal them.
What about theological disagreement?Judge against the supplied authority hierarchy and flag conflicts.Declare the frame and contested interpretations before scoring. Preserve the disagreement and rationale.
Should high scores be rare?The calibrated historical prompt says scores above 90 should be rare and generic answers often fail 25–40% of criteria.Show this historical calibration. Decide whether to retain it; do not silently impose a desired score distribution on the new rubric.

One consequential choice remains: the four-dimension average has no automatic failure for serious harmful advice. A zero on consequences can still average to 75 if the other three dimensions score four. The original FAI-style scoring had a severe harmful-advice penalty. Our recommendation is to preserve an explicit serious-error judgment alongside the component scores and decide its effect on the headline result before running either version.

The next decision

Choose the recovered generation treatment or an explicitly named reconstruction, retain or revise the common Christian judge, and select the first small cohort. Then inspect one complete OpenProse execution before authorizing the frontier campaign. This page and its underlying files are the current intermediate target.