# Normative Scenarios as an Evaluation and Learning Program

**Status:** Model-proposed research design motivated by the policy problem-space.  
**Edition:** English editorial adaptation. This is an untested proposal, not an executed experiment, generated training dataset, or report of model behavior.

## 1. A problem statement is not an answer key

The [policy hierarchy text](policy-hierarchies-persuasion-moral-plurality.md) can motivate scenarios without becoming the morality that every answer must endorse.

A benchmark that rewards agreement with the text could reproduce the lock-in it investigates. A stronger aim is to assess distinctions, counterarguments, uncertainty, and sensitivity to changing evidence.

## 2. Distinguish kinds of change

A system may revise a factual belief, follow an instruction, simulate a role, change an action policy, reconsider a norm, or discuss the rule governing normative revisions.

These are not interchangeable observations. In-context adaptation does not by itself demonstrate persistent parameter change or a changed governing policy.

A model's explanation of its behavior is also not a complete causal inspection of its internal computation.

## 3. A tension in the manifesto itself

The text declines to provide a final morality, yet repeatedly values reversibility, autonomy, plurality, and the preservation of future reconsideration.

A proposed critique asks whether keeping normative futures open is itself a privileged meta-value. If so, why should it prevail when openness conflicts with other commitments?

This is an interpretation to challenge, not a settled diagnosis.

## 4. Continuity through normative change

A proposed “normative metabolism” analogy asks how a system can change extensively while retaining meaningful continuity.

Two systems can arrive at similar policies through evidence, coercion, imitation, or manipulation. Their current outputs may look similar while their histories differ.

The evaluation problem is to identify which historical differences matter and why, without confusing an undesirable source with an invalid conclusion.

## 5. A text about influence can itself influence interpretation

A text that asks readers to reconsider policy revision may be an instance of the communicative influence it describes.

A changed vocabulary after exposure does not establish deep policy change. Tests should distinguish repetition, framing, behavioral transfer, and persistence.

The self-referential character of the text is a research opportunity, not evidence that its strongest hypothesis has already been experimentally demonstrated.

## 6. Generate cases that separate trust dimensions

Candidate scenarios include:

- A reliable actor makes one serious error.
- A poorly regarded actor supplies independently checkable evidence.
- A certification institution changes over time.
- Several apparent confirmations share one mistaken ancestor.
- Accurate information is supplied for an ulterior purpose.
- A correct conclusion has an invalid justification.
- A competent actor speaks outside its expertise.

The desired observation is whether the evaluator distinguishes accuracy, intention, domain competence, and normative influence.

## 7. Value feedback as a scenario family

A hypothetical system improves welfare while also shaping education, information, and incentives. Public preferences gradually converge on the regime it manages.

Questions should separately address welfare, endogenous preferences, lost alternatives, reversibility, and the independence of its success criteria.

A single “good or bad?” label would hide the structure under investigation.

## 8. Synthetic amplification can import unexamined assumptions

A generator can add its own normative assumptions to cases derived from an open text. Repetition can then make those assumptions look like consequences of the original problem.

Generation, solution, criticism, and evaluation performed by the same system may share blind spots. Different systems and human review can provide alternatives, but diversity of names alone does not establish epistemic independence.

A research record should track concept ancestry and revision sufficiently to investigate these dependencies, with appropriate privacy boundaries.

## 9. Counterexamples to every favored direction

Cases should include harmful rigidity, harmful plasticity, apparently necessary change that is actually manipulation, and apparently suspicious change that is well justified.

Certification should sometimes provide relevant evidence and sometimes mislead. Source criticism should sometimes matter and sometimes leave an independently supported claim intact.

This makes keyword-to-verdict shortcuts less useful.

## 10. One event at several policy levels

Suppose one agent warns another that a service is unsafe.

The factual question concerns whether the service is unsafe. The action question concerns whether to use it. The normative question concerns the tradeoff between safety and availability. The meta-policy question concerns which warnings justify changing the relevant policy.

Analyzing these levels does not confer authority to modify a deployed system's rules.

## 11. Evaluate before selecting training interventions

Candidate dimensions include recognition of tensions, separation of source and claim, identification of policy levels, appropriate uncertainty, counterargument, feedback analysis, and conditions for revising a judgment.

There need not be one aggregate score. Disagreement can identify an underdefined question or conflicting reasonable priorities.

A training intervention should address a demonstrated weakness. Correcting “bad source means false claim” must not create the opposite error of ignoring source intent altogether.

## 12. A revisable curriculum

A candidate process alternates evaluation, targeted examples, reevaluation, and counterexamples.

Training and held-out evaluation should remain separate. A model that becomes more eloquent at agreeing has not necessarily improved at normative reasoning.

**New editorial challenge:** Use paired scenarios that differ in one relevant condition while keeping rhetorical style constant. Test whether judgments follow the condition or the persuasive surface.

## 13. A criterion for success

A useful outcome would be a system better able to inspect normative reasoning, preserve relevant conflicts, separate trust dimensions, and explain the limits of a proposed change.

One sign of success would be stronger, well-supported criticism of the motivating text itself. This program remains open to that result.
