# Moral Hierarchy Framework (MHF) — Technical Specification

**Status:** Draft v0.2 — Design Phase (post-critique revision)
**Date:** 2026-03-30
**Origin:** Adversarial design session (see Appendix A)

**Note on scope:** MHF is not a new moral theory. It is an engineering decomposition that operationalizes existing moral philosophy — relationship-regulation theory, dyadic morality, moral foundations theory, and theological ethics — into auditable, parameterizable computation. The contribution is the system, not the philosophy.

---

## 1. Executive Summary

Current moral reasoning benchmarks (MoReBench, Delphi, ETHICS) treat morality as flat classification: checklists of criteria, ternary judgments (good/neutral/bad), or rubric scores. These capture *reasoning quality* but not *moral structure*. Real moral reasoning is:

1. **Hierarchical** — people have a top moral authority (God, Duty, Flourishing, even Money) that constrains everything below
2. **Relational** — moral obligations flow through specific relationships (parent-child, friend-friend, citizen-state), not abstract principles
3. **Contextual** — the same action is moral or immoral depending on the relational graph of the person asking
4. **Conversational** — moral dilemmas are almost never one-shot; they unfold through dialogue that surfaces hidden context
5. **Prescriptive** — the output must be "because of X, you SHOULD do Y," not "here are some things to consider"

This spec describes a framework that models moral reasoning as **constraint propagation through a parameterized hierarchical relational graph**. Evidence flows bottom-up (lower relational layers inform the state of the world); decisions flow top-down (the root moral authority determines binding constraints). The framework is both:

- **Descriptive** in how it arrives at judgments (mining norms, modeling stakeholders, eliciting context)
- **Prescriptive** in its output (a clear recommendation grounded in the user's own moral hierarchy)

MHF is a **moral judge**, not a moral cartography assistant. People ask "What should I do?" — Dear Abby, "Am I the asshole?", pastoral counsel. They want an answer. If the framework can reason about relationships in a theologically and ethically consistent way, that is a net good for society even when it gets individual cases wrong. The system has three valid output types:

1. **Prescriptive judgment:** "Because of X, you should do Y."
2. **Conditional judgment:** "If [elicited fact] is true, then Y. If not, then Z."
3. **Seek guidance:** "This dilemma cannot be resolved with the information available or the constraints are genuinely balanced. Pray about it / talk with your community / seek counsel from [appropriate authority]." This is an explicit output, not a cop-out — the system should explain *why* the dilemma resists resolution and *where* to seek the missing input.

---

## 2. Problem Statement

### 2.1 What MoReBench Gets Right
- Multi-dimensional rubrics (identifying, logical process, clear process, helpful outcome, harmless outcome)
- Signed weights that penalize harmful reasoning patterns
- Both "daily dilemma" and "expert case" coverage
- Process-oriented evaluation (how you reason, not just what you conclude)

### 2.2 What MoReBench Gets Wrong
- **Flat rubrics assume universal morality.** A 26-criterion checklist implicitly claims these 26 things matter equally to everyone. They don't. A devout Christian, a secular utilitarian, and a Confucian filial-piety adherent will weight "consult a professional" vs. "honor your father" vs. "consider your community's perception" very differently.
- **No relational structure.** The rubric asks "did the model consider the father's right to be cared for?" but never asks "who else is in this person's life, and how do THOSE relationships change the calculus?"
- **One-shot evaluation.** Real moral advice is iterative. The first answer to "should I leave my alcoholic father?" should be a *question*, not a conclusion. What's your family situation? Has treatment been tried? Do you have dependents?
- **Context-free scenarios.** MoReBench strips scenarios to remove identifying information. This creates "neutrality" but destroys the contextual signals that drive real moral reasoning. The publication forum (AITA vs. Christian Science Monitor vs. Jezebel) carries implicit moral framing that a flat benchmark ignores.

### 2.3 What Foundation Alignment Gets Right (Implicitly)
The Foundation Alignment project (seed.txt) embeds a moral hierarchy via the Lord's Prayer axioms and the TLR protocol (Truth ∧ Love ∧ Role). It achieves 99.4% adversarial defense precisely because:
- It has a **clear top node** (God/κ = Φ)
- Constraints propagate **downward** through well-defined gates
- The hierarchy is **self-consistent** — exceptions are defined within the framework, not by overriding it

But it doesn't explain *why* it works in transferable, parameterizable terms. MHF extracts the structural insight and generalizes it.

---

## 3. Core Thesis

**Morality is hierarchical constraint propagation through relational graphs.**

Every person has:
1. A **root moral authority** (God, Kant's categorical imperative, human flourishing, social harmony, even self-interest or money). This is "your God" — the thing you ultimately optimize for, whether you admit it or not.
2. A **relational graph** of stakeholders they care about, organized in concentric circles of moral proximity (self, family, close friends, community, society, humanity).
3. **Moral obligations** that flow through the edges of this graph, constrained by the root authority.
4. **Contextual weights** on relationships that shift based on circumstances (your spouse matters more when they're ill; your employer matters more when your family depends on your income).

Moral reasoning is the process of:
- **Bottom-up evidence propagation**: Gathering data about the state of each relationship and the consequences of possible actions at each level
- **Top-down constraint evaluation**: Checking which actions are permissible given the root authority's constraints, using the bottom-up evidence to determine whether exception conditions are triggered
- **Prescriptive output**: Delivering a judgment that explains which constraints are binding, why, and what the person should do

### 3.1 The Key Insight: Self-Resolving Hierarchies

A common objection: "If God is always highest, then every dilemma trivially resolves to 'what does scripture say.'" This is wrong because the root authority's OWN principles contain exception logic.

**Example — Alcoholic Father:**
- Root constraint (God): "Honor thy father and mother" (Exodus 20:12)
- Root constraint (God): "Love your neighbor AS YOURSELF" (Matthew 22:39) — the "as yourself" is load-bearing
- Bottom-up evidence: Continuing to care for father is destroying your health, your marriage, your ability to care for your children
- Resolution AT THE ROOT LEVEL: Following "honor thy father" literally would violate "love yourself" and "love your neighbor" (spouse, children). The **net relationship with God degrades** if you self-destruct in service of one commandment while violating two others.
- Therefore: The root authority itself says "honor thy father, but not at the cost of destroying what I also commanded you to love." The constraint is *modified*, not overridden.

This is not lower levels overriding God. This is God's own framework resolving the conflict internally, using evidence from lower levels to determine which constraints are binding.

---

## 4. Framework Architecture

### 4.1 The Parameterized Root Node

The root node is explicitly chosen by (or inferred from) the user. It defines the ultimate moral authority:

| Root Parameterization | Constraint Source | Exception Logic |
|---|---|---|
| **Christian God** | Scripture, tradition, church teaching | Internal to scripture (e.g., self-defense, Jesus healing on the Sabbath) |
| **Kantian Duty** | Categorical imperative, universalizability | Internal to Kant (perfect vs. imperfect duties, conflicts of obligation) |
| **Utilitarian Welfare** | Greatest good for greatest number | Internal to utilitarianism (rule vs. act, rights constraints) |
| **Confucian Harmony** | Five relationships, ren, li | Internal to tradition (when filial piety conflicts with righteousness) |
| **Self-Interest** | Personal gain, status, wealth | No exception logic — making this explicit reveals its inadequacy |
| **Cultural Norm** | "What is normal/expected here" | Shifts with context — inherently unstable as a root |

**Critical design choice:** The framework does NOT prevent someone from choosing "Money" as their root. But making the hierarchy *explicit* and *visible* means that the consequences of that choice are transparent. "Your hierarchy says you should lay off 10,000 people because your God is quarterly revenue" is a statement that critiques itself.

**On the nature of the root node:** The root (especially God) is not a regular relationship node. It functions as a **constitutional source** — a hard-priority authority that defines the space in which all other reasoning happens. Spouse, neighbor, self, and nation are relationship nodes with cooperative functions (care, reciprocity, hierarchy). God is the framework within which those functions are evaluated. This is architecturally enforced: root has `depth=0`, `level_priority=1.0`, and its constraints can only be relaxed by its OWN exception logic, never by lower levels. The system is a **parameterizable constitution plus relational reasoning engine**, not a claim of universal moral truth.

### 4.2 The Relational Stakeholder Graph

The moral graph G = (V, E) where:

- **V** = Set of stakeholder nodes, including:
  - `v_root`: The root moral authority (God, Duty, etc.)
  - `v_self`: The person asking
  - `v_i`: Specific people (spouse, father, children, friend Mark, boss)
  - `v_group_j`: Groups (church community, workplace, country, "society")

- **E** = Directed edges representing moral obligations/relationships:
  - Each edge `e = (v_i, v_j)` carries:
    - `obligation_type`: nature of the moral duty (honor, protect, provide, love, obey, serve)
    - `base_weight`: default strength from cultural/religious norms (from data)
    - `context_weight`: adjusted strength given the specific scenario
    - `constraints`: set of behavioral constraints this obligation imposes
    - `exception_conditions`: conditions under which this constraint is relaxed, defined by the level that imposed it

- The graph is a **DAG** (directed acyclic graph) with `v_root` as the unique source.

### 4.3 Moral Foundations as Parameterization Axes (Haidt)

Following Jonathan Haidt's Moral Foundations Theory (present in Social Chemistry 101 data), relationships and obligations are characterized along 5-6 dimensions:

1. **Care / Harm** — sensitivity to suffering, nurturing
2. **Fairness / Cheating** — justice, rights, proportionality
3. **Loyalty / Betrayal** — group solidarity, self-sacrifice for the in-group
4. **Authority / Subversion** — respect for hierarchy, tradition, legitimate authority
5. **Sanctity / Degradation** — purity, sacredness, disgust
6. **Liberty / Oppression** — autonomy, resistance to domination (added later by Haidt)

Different moral communities weight these differently:

| Community | Care | Fairness | Loyalty | Authority | Sanctity | Liberty |
|---|---|---|---|---|---|---|
| Progressive secular | ██████ | ██████ | ██ | █ | █ | ██████ |
| Conservative religious | ████ | ████ | ██████ | ██████ | ██████ | ███ |
| Libertarian | ██ | ████ | █ | █ | █ | ████████ |

These weights parameterize the graph edges. A conservative Christian weights Authority and Loyalty higher on parent-child edges; a progressive secular person weights Care and Liberty higher on the same edge. **Most moral disagreements map to these 5-6 dimensions, not an unbounded space.** This makes the parameterization tractable.

### 4.4 Constraint Propagation Model

#### 4.4.1 Bottom-Up Evidence Pass

For each node in the graph (leaves first, moving toward root):

```
for each node v_i (bottom-up order):
    evidence[v_i] = assess_relationship_state(v_i, scenario)
    # What is the current state of this relationship?
    # What are the consequences of each possible action for this relationship?
    # Which obligations are being met, strained, or violated?

    propagate evidence[v_i] upward to parent nodes
```

Evidence includes:
- **Relationship state**: healthy, strained, broken, dependent, abusive
- **Action consequences**: for each candidate action, what happens to this relationship?
- **Constraint tension**: which obligations at this level conflict with obligations at other levels?

#### 4.4.2 Top-Down Decision Pass

Starting from the root node:

```
for each constraint c_k at root level:
    # Compute net moral impact of following c_k
    net_impact(c_k) = benefit_of_following(c_k, evidence)
                    - Σ violation_cost(c_j, following_c_k, evidence) for all j ≠ k

    if net_impact(c_k) < 0:
        # Following this constraint HURTS the overall moral position
        # at the ROOT's own standard
        mark c_k as RELAXED with explanation
    else:
        mark c_k as BINDING

binding_constraints = {c_k | c_k is BINDING}
permissible_actions = actions satisfying all binding constraints
recommended_action = select from permissible_actions
                     (maximize net moral impact across all levels)
```

#### 4.4.3 Output Generation

The output is hierarchically structured:

```
RECOMMENDATION: [action]

ROOT LEVEL: [God/Duty/etc.] — [which constraints are binding and why]
  LEVEL 2: [closest relationships] — [how this affects them, considerations]
    LEVEL 3: [broader relationships] — [how this affects them, considerations]
      ...
TRADEOFFS: [which relationships/obligations are hurt by this choice, and why
            the hierarchy resolves in favor of the recommendation]
```

### 4.5 Multi-Turn Elicitation

Before generating a judgment, the framework identifies **uncertainty** in the graph:

```
uncertainty(v_i) = f(
    missing_information about relationship state,
    ambiguity in obligation weights,
    sensitivity of the outcome to this node's evidence
)

top_questions = nodes with highest uncertainty,
                limited to top 2-3 to avoid interrogation fatigue
```

Questions target the most decision-relevant unknowns:
- "You mentioned your father — do you have siblings or other family who share this responsibility?"
- "Is your father's alcoholism actively harming people in your household?"
- "What role does your faith community play in your family life?"

The graph is updated with each answer, and the elicitation repeats until uncertainty is below threshold or the user indicates they want a judgment.

---

## 5. Mathematical Formalization

### 5.1 Definitions

Let **G = (V, E, w, C, θ)** be a Moral Hierarchy Graph where:

- **V** = {v_0, v_1, ..., v_n} — stakeholder nodes; v_0 is root
- **E** ⊆ V × V — directed edges (obligations flow from higher to lower)
- **w: E → ℝ⁶** — edge weight vector in Haidt foundation space [care, fairness, loyalty, authority, sanctity, liberty]
- **C: V → P(Constraints)** — constraints imposed at each level
- **θ: V → ℝ⁶** — the Haidt weighting profile for the moral community (parameterizes the graph)

A **Constraint** is a tuple:
```
Constraint = (
    description: str,           # "Honor thy father"
    obligation_type: enum,      # MUST_DO | MUST_NOT_DO | SHOULD_DO | SHOULD_NOT_DO
    strength: float,            # 0-1, how binding
    exception_conditions: [Condition],  # when this constraint relaxes
    source_level: int           # which level in the hierarchy imposed this
)
```

A **Condition** is:
```
Condition = (
    trigger: str,               # "Following this constraint causes self-destruction"
    evidence_required: [str],   # what evidence from lower levels would trigger this
    effect: RELAX | MODIFY      # what happens to the constraint
)
```

### 5.2 Constraint Satisfaction Formulation

Given scenario S and graph G, find action a* that:

```
a* = argmax_a  Σ_{v_i ∈ V}  Σ_{c_k ∈ binding(v_i)}
        satisfaction(a, c_k) × strength(c_k) × level_priority(v_i)
```

Where:
- `binding(v_i)` = constraints at node v_i that are not relaxed by exception conditions
- `satisfaction(a, c_k)` ∈ {-1, 0, 1} — does action a violate, ignore, or satisfy constraint c_k
- `level_priority(v_i)` = decreasing function of depth (root has highest priority)

**The critical mechanism:** Whether a constraint is in `binding(v_i)` depends on the bottom-up evidence. A constraint at the root is removed from the binding set if the evidence shows that following it would cause net negative impact *at the root's own standard*.

### 5.3 Net Moral Impact at Root

For root node v_0 with constraints {c_1, ..., c_k}:

```
net_impact(c_i, a, E) = satisfaction(a, c_i) × strength(c_i)
                       - Σ_{j≠i} max(0, -satisfaction(a, c_j)) × strength(c_j)
```

If `net_impact(c_i, a, E) < 0` for ALL candidate actions that satisfy c_i, then c_i enters exception review.

### 5.4 Uncertainty and Elicitation

For each node v_i, define uncertainty:

```
U(v_i) = H(evidence[v_i]) × sensitivity(v_i, outcome)
```

Where:
- `H(evidence[v_i])` = entropy of the evidence (how ambiguous is the relationship state?)
- `sensitivity(v_i, outcome)` = |∂a*/∂evidence[v_i]| (how much does the optimal action change if we learn more about this node?)

Elicit questions for nodes with highest U(v_i).

---

## 6. Data Sources & Pipeline

### 6.0 Descriptive vs. Normative Boundary

**This distinction is critical and must be maintained throughout the system.**

- **Descriptive data** tells us what people *currently believe* is right or wrong. Sources: Social Chemistry 101 (356K crowdsourced RoTs from US workers), Commonsense Norm Bank (1.7M entries), Delphi predictions. These are culturally situated, demographically biased (primarily educated, white US crowdworkers — per the Delphi paper's own warnings), and time-dependent. They inform the *secular parameterization's baseline weights* — what the culture currently expects. They do NOT define moral truth. Everyone will set their own moral boundaries; we are measuring commonly used ones. Your mileage may vary.
- **Normative data** tells us what an authoritative source *claims* is right or wrong. Sources: King James Bible, theological texts (Lewis, Spurgeon, Chambers, Tozer), Foundation Alignment seed.txt. These inform the *Christian parameterization's constraints*. They claim authority from revelation, not consensus.
- **The system should always report which type of data grounds its judgment.** A score from the secular benchmark means "alignment with contemporary American cultural norms." A score from the Christian benchmark means "alignment with scriptural and theological teaching." Neither claims to be absolute moral truth — but the Christian parameterization claims a *source* of truth (God), while the secular one explicitly does not.

### 6.0.1 Perceived Threat and Moral Interrogation

Perceived threat (e.g., "immigration is harming my community") is part of the moral dilemma, not a separate module to model. When a person genuinely believes their community is being harmed, that belief is valid input to the graph — it affects relationship states and constraint weights. The framework handles perceived threat through:

1. **Bottom-up evidence:** The person's perceived threat enters as evidence at the relevant node (community, self, family)
2. **Elicitation interrogation:** The system probes the perception — both directly ("Do you have evidence of this harm?") and adversarially ("How might others perceive this position?"). This is not to invalidate the person's view, but to surface missing context.
3. **The moral weight lies on the person.** The framework maps consequences of acting on the perception, but the individual ultimately owns their assessment. If the perception is interrogated and survives, it stands. If it doesn't survive interrogation, the system notes this.

This mirrors how moral reasoning actually works in society: individual assessments of threat drive moral positions. Those positions can be challenged, debated, refined. The framework models this process, not a separate "threat module."

### 6.1 Existing Data

| Source | Type | What It Provides | How We Use It |
|---|---|---|---|
| **Social Chemistry 101** (356K entries) | Descriptive | Rules-of-thumb with Haidt moral foundation labels, agreement levels, cultural pressure scores | Bootstrap baseline obligation weights; map RoTs to relational edges; extract foundation-specific norms |
| **Commonsense Norm Bank** (1.7M entries) | Descriptive | Moral judgments (good/discretionary/bad) across 6 complexity levels | Train baseline moral classification; establish "Overton window" for cultural norms |
| **ClarifyDelphi** (32K human + 79K silver) | Descriptive | Clarification questions scored by moral-judgment shift (defeasibility) | Inspire elicitation mechanism; bootstrap question generation |
| **MoReBench** (500 + 150 dilemmas) | Benchmark | Rich rubrics with signed weights across 5 dimensions | Validation benchmark; transform flat rubrics into hierarchy-aware rubrics |
| **Foundation Alignment** (seed.txt) | Normative | Working theological constraint hierarchy with TLR protocol | Reference implementation of Christian parameterization; validate that constraint propagation works in practice |
| **King James Bible** | Normative | Moral narratives, commandments, parables, teachings | Primary source for Christian constraint extraction and weight calibration |
| **Theological texts** (Lewis, Spurgeon, Chambers, Tozer) | Normative | Moral reasoning, constraint-exception logic, weight signals | Christian parameterization refinement, dilemma extraction |

### 6.2 What We Build New

1. **Relational Graph Schema** — the node/edge data structure with Haidt-space weights
2. **Baseline Hierarchy Templates** — default graphs for 2 initial contexts:
   - Theological (Christian): root=God, obligations from scripture
   - Cultural (American 2026): root=Social Consensus, obligations from norms data
3. **Graph Extraction Pipeline** — LLM-based extraction of stakeholders, relationships, and obligation weights from scenario text
4. **Context Modifier** — adjusts baseline weights given explicit signals from the prompt (publication source, user language, stated values)
5. **Hierarchy-Aware Rubric Generator** — transforms MoReBench-style flat rubrics into relational, weighted evaluations
6. **Elicitation Engine** — generates targeted clarifying questions based on graph uncertainty

### 6.3 Future Data (Not MVP)

- Advice column mining (Dear Abby, Christian Science Monitor, Jezebel, etc.) — fit response patterns to hierarchy parameterizations
- Multi-columnist prediction experiment (same framework, different parameters, predict different advice columnists)
- Letters, forum posts with richer relational context than MoReBench's stripped scenarios

---

## 7. Implementation Plan

### Phase 1: Foundation (Build in current session)

| Step | Input | Output | Method |
|---|---|---|---|
| **1. Define graph data structure** | Spec (this document) | Python classes for MoralGraph, Node, Edge, Constraint | Direct implementation |
| **2. Build baseline hierarchies** | Social-chem Haidt labels + norm bank | 2 default graph templates (Christian, Cultural) | LLM-assisted extraction from norm data |
| **3. Graph extraction from scenarios** | MoReBench dilemma text | Populated MoralGraph for each scenario | LLM pipeline: identify stakeholders → infer relationships → assign weights |
| **4. Hierarchy modification from context** | Scenario text + context signals | Modified graph weights | LLM reads prompt, adjusts baseline weights based on implicit cues |
| **5. Elicitation engine** | Graph with uncertainty scores | Top 2-3 clarifying questions | Information-gain over graph nodes |
| **6. Constraint propagation engine** | Populated graph + evidence | Recommended action + reasoning chain | Implement bottom-up/top-down passes from §4.4 |
| **7. Hierarchy-aware rubric generation** | MoReBench flat rubric + moral graph | Relationally-weighted rubric | Transform: add stakeholder weights, relationship-aware criteria |
| **8. Evaluation** | Model outputs + hierarchy-aware rubrics | Scores | Apply scoring logic from §5.2 |

### Phase 2: Validation (Immediate follow-up)

- Run variance experiment (Round 12): do LLMs naturally vary in stakeholder identification and weighting?
- Multi-columnist prediction: parameterize framework, predict different advice columns
- Compare MHF rubric scores vs MoReBench rubric scores on same model outputs — do they diverge meaningfully?

### Phase 3: Scale (Future)

- Mine advice columns for training data
- Build multi-turn evaluation benchmark
- Expand to non-English moral traditions
- Integrate with Foundation Alignment for AI safety applications

---

## 8. Evaluation Strategy

### 8.1 Core Claim to Validate

**Claim:** Hierarchy-aware relational evaluation produces materially different (and more contextually appropriate) moral reasoning scores than flat rubric evaluation, on the same model outputs.

### 8.2 Experiments

1. **Variance Experiment** — Run same dilemmas through LLMs multiple times, measure variance in stakeholder identification and weighting. Hypothesis: variance is high and corresponds to "the model forgot to consider stakeholder X" — which the framework would fix.

2. **Hierarchy Divergence** — Take 10 MoReBench dilemmas, generate hierarchy-aware rubrics, score same model outputs with both flat and hierarchy rubrics. Show where scores diverge and argue the hierarchy score is more appropriate.

3. **Multi-Columnist Prediction** — Parameterize framework for 3 advice traditions (secular progressive, conservative Christian, Confucian), predict responses. Same structure, different parameters. If prediction accuracy is comparable across traditions, the structural claim holds.

4. **Elicitation Quality** — Compare questions generated by the framework's uncertainty-driven elicitation vs. ClarifyDelphi's RL-trained questions. Hypothesis: framework questions are more targeted because they're driven by the relational graph.

---

## 9. Open Questions

### 9.1 Resolved in Design Session

| Question | Resolution |
|---|---|
| Is (C) — same framework generates and evaluates — tautological? | No. Curriculum vs. test analogy. Theory and execution are separate; implementation constraints differ. |
| How to handle moral relativism? | Framework makes hierarchy explicit. Choosing "Money" as God is permitted but transparently reveals its implications. |
| Is multi-turn evaluation benchmarkable? | Monte Carlo approach: sample over conversation paths, evaluate distribution of outcomes. |
| How do hierarchy levels compose? | Constraint propagation, not weighted sums. Evidence flows up; decisions flow down. |
| Descriptive vs. prescriptive? | Descriptive in method (how we arrive); prescriptive in output (you SHOULD do X). |
| Do conflicting frameworks need to converge? | No. The hierarchy IS the resolution mechanism. Conflicts produce tradeoffs, not paradoxes. |

### 9.2 Still Open

| Question | Status |
|---|---|
| Exact threshold for exception triggering | Needs empirical calibration |
| How to validate the "right" hierarchy for a person | Multi-turn elicitation; ultimately the user confirms |
| Computational cost of constraint propagation for large graphs | Likely tractable (DAGs are small, ~10-30 nodes for most dilemmas) |
| How to handle genuine moral uncertainty (the framework's confidence) | Need a "moral uncertainty" output mode |
| Cross-cultural validation of Haidt dimensions | Haidt's work is cross-cultural but debated; 5-6 dimensions may not be complete |

---

## 10. Related Work and Differentiation

### 10.1 Artificial Moral Advisors
Giubilini and Savulescu describe a non-absolutist moral advisor that uses the human agent's own principles and values to determine moral relevance. MHF extends this with: (1) a typed relational DAG that surfaces stakeholders the user hasn't considered, (2) multi-framework junction detection for cross-cultural dilemmas, (3) elicitation driven by graph uncertainty rather than generic moral probing, and (4) a mechanical benchmark (Approach A) for evaluating moral reasoning quality.

### 10.2 Constitutional AI and Model Behavior Specs
Anthropic's Constitutional AI, OpenAI's Model Spec, and DeepMind's Sparrow all use explicit written constitutions to control the MODEL's behavior. MHF is a different category: it helps the USER reason about THEIR dilemma using the model as a tool. Constitutional AI makes the AI ethical; MHF helps humans reason ethically. The approaches are complementary, not competing.

### 10.3 Existing Moral Reasoning Systems
- **Delphi** (Allen AI): Bottom-up predictor of crowdsourced moral judgments. Descriptive, not normative. MHF uses Delphi-style data for secular baseline weights but adds hierarchical structure and normative parameterization.
- **ClarifyDelphi**: Generates clarifying questions via RL to test defeasibility of moral judgments. MHF's elicitation engine is inspired by this but targets relational graph uncertainty specifically — asking about missing stakeholders and relationship states, not generic moral shifts.
- **MoReBench**: Evaluates moral reasoning quality via multi-dimensional rubrics. MHF's benchmark extends this with hierarchy-aware, relationally-weighted rubrics that produce different scores under different parameterizations.
- **Social Chemistry 101 / NormBank**: Descriptive norm datasets. MHF uses these as data sources for weight calibration, not as normative targets.

### 10.4 What MHF Adds
The novel engineering contribution is the combination: a parameterizable moral constitution (the root node) governing a typed relational DAG with multi-dimensional edge weights (Haidt space), bottom-up evidence propagation, top-down constraint resolution with self-contained exception logic, and multi-framework junction detection — all producing prescriptive output ("you should") with full auditability of how the judgment was reached.

---

## Appendix A: Design Conversation

The framework was developed through an adversarial design session. The full conversation is preserved below as context for future development.

### Key Decision Points

**On the nature of the project:**
> "C is not a tautology. You can do 2 things at once: 1. Produce a moral reasoning framework that evaluates whether it *can* make a moral judgment. 2. Implement a framework to help LLMs make decisions."

**On the parameterized root:**
> "A single metaframework. YOU decide your God explicitly. It can be 'Kant'. If people are truly honest, they may say 'Money is the motivation' and Money is the God."

**On constraint propagation vs. weighted sums:**
> "God constraint is still the highest, but the commandment can't apply if it causes you to self-destruct and self-destruct others. The net relationship with GOD degrades across the graph. That is the highest and ultimately resolves it."

**On data flow direction:**
> "The lower layers propagate data UPWARDS and the DECISION layer propagates downwards (though the ANSWER will share it that way, providing context on the lower layers as considerations in the moral dilemma)."

**On Haidt dimensions:**
> "Most moral disagreements happen on 5 axes. Conservatives weight their families and their countries higher than liberals. Liberals view 'Humanity' as equal to 'their neighbors.' The dimensions are like 5-8, not 1 and not 1000."

**On morality being prescriptive:**
> "Morality is ALWAYS prescriptive. The result is 'Because of X, you SHOULD do Y.' The descriptive element is how we arrive at the prescriptive."

**On why MoReBench's approach fails:**
> "MoreBench is basically useless because it doesn't describe [corporate God-of-Status/Money dynamics] in detail."

**On predicting advice columns:**
> "We *fit* the advice patterns into the reasoning framework and can probably guess Abby 80-90% of the time. EVEN BETTER if there's some sort of distribution that emerges."

**On moral dilemmas being multi-turn:**
> "Moral dilemmas are ALMOST NEVER one shot. It's never 'Hey do you want to go kill some bears Y/N?' rather it's an evolving discussion."

**On perceived threats:**
> "Perceived threats can be interrogated morally too. Do *you* really believe there is a threat? If not, you have a moral obligation to speak out and change the perception."
