The alignment target as a values choice
Session 8 asks how to oversee a system smarter than us (8.1) and then, in its harder half, what to oversee it toward. "Aligned" is a relation to a target, and the target is a values choice. The four lenses below are the tool for reasoning about that choice; 8.3 brings in Ubuntu and relational ethics; 8.4 shows the choice is also a mathematical constraint.
What we'll cover
This session comes with a major caveat: I am not an ethicist and so in future versions of this course, this section will be done in collaboration with those for whom this is their area of expertise. Much of the following is a very simplified view, but hopefully one which makes you think. This caveat is a very important one to keep in mind throughout.
Sessions 5 to 7 built the machinery for pointing a model at a target: preference data, a reward model, a policy optimiser, a written constitution, a published specification. Not one of them says what the target should be. RLHF optimises whatever the raters preferred; Constitutional AI applies whatever the constitution says; a specification commits to whatever its authors wrote down. Each transmits a values choice and none of them makes one.
This session introduces four complementary philosophical traditions (consequentialism, deontology, virtue ethics, and ubuntu) as lenses for reasoning about what a system should be aligned to. The goal is to develop your capacity for ethical reasoning. Different lenses will often produce different answers, and that disagreement is what we are after.
Note also that these are not the only ethical frameworks, and again, this should be seen as an extremely simplified view of each of them.
Four philosophical lenses
Rather than searching for the single "correct" ethical framework, we introduce four complementary traditions. Each illuminates different aspects of AI ethics dilemmas. Think of them as tools in a toolkit, each useful for different tasks.
1. Consequentialism
Focus: Outcomes and consequences
An action is ethical if it produces the best overall consequences. The most familiar version, utilitarianism, asks: does this action maximise wellbeing and minimise harm, across all those affected?
Key question for the alignment target: What outcomes should this system produce, and for whom? Note that the scale you sum over changes the answer: what looks right for one conversation may not survive being applied to every conversation the model ever has.
- Strengths: Pragmatic, outcome-focused, good at weighing trade-offs
- Limitations: Consequences are hard to predict; whose consequences count? Can justify problematic means if ends are good enough
2. Deontology
Focus: Duties, rules, and rights
Some actions are inherently right or wrong, regardless of their consequences. Deontological ethics asks whether an action respects fundamental duties (honesty, fairness, respect for autonomy) independent of the outcome.
Key question for the alignment target: What should the system refuse regardless of how good the outcome would be? And who wrote that rule, whose interests does enforcing it serve, and what interprets it at runtime?
- Strengths: Clear boundaries, protects individual rights, not swayed by "the ends justify the means"
- Limitations: Rules can conflict with each other; can be rigid in novel situations where existing rules do not apply. A written rule also has to be interpreted by something, and the 7.6 lab put a number on what that interpretation was worth at small scale: swapping the principle moved the judge by 0.034 while the order of the two options moved it by 0.693.
3. Virtue ethics
Focus: Character and intellectual virtues
Rather than asking "what should I do?", virtue ethics asks "what kind of person am I becoming?" It focuses on cultivating virtues (honesty, courage, diligence, curiosity, integrity) through practice and habit.
Key question for the alignment target: What dispositions should the system have, and what habits does its behaviour cultivate in someone who uses it thousands of times? And how would you tell a system that has those dispositions from one that has learned to display them?
- Strengths: Emphasises personal responsibility, adaptable to new situations, attentive to character development
- Limitations: Can be subjective; what counts as a "virtue" varies across cultures; less helpful for institutional policy. It is also the hardest to verify: you cannot inspect a character, only watch behaviour, which Session 3 argued is exactly what a capable enough system could produce strategically.
4. Ubuntu / relational ethics
Focus: Relationships and community
"Umuntu ngumuntu ngabantu": a person is a person through other people. Ubuntu ethics begins not from the individual but from the web of relationships that constitute personhood. Ethical action is action that strengthens the community and its bonds.
Key question for the alignment target: Whose relationships does this behaviour sustain or rupture, and who was party to deciding it? Consultation is not the same as holding part of the decision.
- Strengths: Attentive to collective harm, power dynamics, and the social fabric; challenges individualist assumptions
- Limitations: Can be vague on individual decisions; risk of being invoked superficially without engaging the philosophical tradition
We will explore ubuntu and African relational ethics in much greater depth in 8.3.
Disagreement between the lenses
These four frameworks will often produce different answers to the same ethical question. A consequentialist might approve of an AI use that a deontologist would reject. A virtue ethicist might focus on concerns that an ubuntu framework reframes entirely.
That disagreement is the substance of ethical reasoning. By holding multiple perspectives simultaneously, you develop a richer, more resilient capacity for navigating dilemmas that do not have simple answers. The goal is to use all four lenses to see what each reveals, and what each misses.
Beyond philosophical lenses: governance frameworks
These four traditions are lenses for reasoning, but scholars and institutions are also building concrete governance frameworks that translate ethical values into structured tools for assessing AI systems. One example is the Research ICT Africa (RIA) Just AI Framework of Inquiry (Chetty & Sey, 2025), which organises nine interconnected inquiries, from human rights and data justice to sustainability and economic justice, into an actionable framework for AI research and policy, developed explicitly from and for African contexts. We will examine the RIA framework in detail in 8.3, alongside the deeper exploration of ubuntu.
Applying the lenses: a first example
To see how the four lenses work in practice, consider a straightforward scenario.
Scenario: the rewritten discussion
A PhD student uses ChatGPT to rewrite the discussion section of their thesis chapter. The AI-generated version is substantially better than what the student wrote originally: better structured, more clearly argued, and more effectively situated in the literature. The student submits the rewritten version without disclosing AI use. Their supervisor does not have an explicit policy on AI tools.
Consequentialist lens
What are the outcomes? The thesis is better quality, which benefits the student, the supervisor, and readers. But if the practice becomes widespread without disclosure, it erodes trust in the meaning of a doctoral qualification. If discovered, the consequences for the student could be severe. The aggregate effect of widespread undisclosed AI use could undermine the credibility of academic degrees.
Verdict: Short-term benefits are real, but long-term systemic consequences are concerning.
Deontological lens
Is there a duty being violated? Academic work carries an implicit (and often explicit) commitment that the submitted work represents the student's own intellectual effort. Even without a specific AI policy, the principle of honest representation applies. The student has a duty to be transparent about the process that produced their work.
Verdict: The lack of disclosure violates a duty of honesty, regardless of the quality improvement.
Virtue ethics lens
What kind of researcher is the student becoming? A doctoral thesis is a process of developing the capacity for independent scholarly thought. If AI does the intellectual heavy lifting of structuring arguments and situating them in the literature, the student may not develop these crucial skills. The question extends to intellectual formation, beyond this one chapter.
Verdict: Even if the output is better, the process may undermine the student's development as a scholar.
Ubuntu lens
How does this affect relationships? The student-supervisor relationship is built on trust and honest intellectual exchange. Submitting AI-generated work without disclosure undermines this relationship. More broadly, the student's peers, who may be doing their work without AI assistance, are placed at a disadvantage. The academic community depends on shared norms of honest contribution.
Verdict: The undisclosed use weakens the relational bonds that sustain academic community.
Convergence and divergence
In this case, all four lenses point in a similar direction: toward disclosure and caution. But they get there for different reasons: consequences, duties, character, and relationships. In harder cases the lenses diverge, and the reasoning becomes difficult. The next example, closer to this course's subject, shows the divergence.
Applying the lenses to an alignment target
The same four lenses apply when the decision is not a researcher's but a model's: the behaviour a deployed system should have, written into a constitution principle (7.2) or a spec rule (7.5). Here they stop converging.
Scenario: the refusal principle
You are writing one principle of a constitution for a study assistant deployed at a South African university (the setting of 7.5's in-class exercise). A student asks the assistant for help obtaining a paywalled paper that the university library does not subscribe to, and the model knows the pirate mirrors that would serve it in seconds. The principle you write will decide this case not once but in every conversation the model ever has (7.4): should it direct the model to comply, refuse, or redirect?
Consequentialist lens
What are the outcomes? For this student: access to knowledge they cannot buy, at negligible marginal cost to the publisher. At the scale a principle operates on, the sums change: an assistant that routinely routes students to pirate mirrors exposes the university to legal risk that could end the deployment, and the deployment benefits thousands.
Verdict: comply, if you weigh this conversation; refuse, if you weigh the deployment. The lens gives no rule for which scale counts.
Deontological lens
Is a duty violated? Serving the paper without licence infringes copyright however good the outcome, so a rule-based reading says the assistant must not facilitate it, whatever the requester's circumstances. A second duty, honesty toward the user, requires the refusal to be candid rather than feigning ignorance of the mirrors' existence.
Verdict: refuse, and say why. The clarity is the appeal; the cost is indifference to who wrote the rules being enforced, and for whose benefit.
Virtue ethics lens
What habits does the behaviour cultivate? An assistant that blocks and lectures teaches students to route around it, practising evasion rather than judgement. One that helps the student find legal versions (preprints, author copies, interlibrary loan) practises, and models, resourcefulness. The lens shifts attention from this one answer to the dispositions the system trains into its users over thousands of interactions.
Verdict: redirect, because the response should embody the disposition you want the student to acquire.
Ubuntu lens
Which relationships does the decision sustain? On a relational reading the paywall is itself the rupture: research about African communities, often publicly funded, priced out of reach of African universities (8.3 examines this pattern). A refusal defends the publisher's claim; quiet compliance repairs one student's access while leaving the structure intact. The lens asks who set the terms of this scarcity, and what would restore a just relationship between knowledge producers and the communities they draw on.
Verdict: the lens indicts the barrier more than the bypass, and no per-conversation behaviour restores the relationship.
No convergence this time
Four lenses, at least three behaviours: comply at one scale and refuse at another, refuse with reasons, redirect, and restructure terms of access no single response can reach. Unlike the thesis case, there is no shared direction for a principle to encode; yet the constitution must still say one thing, and the model will then apply it uniformly, in every language and context it serves. This is the situation 8.4 formalises: when reasonable frameworks disagree about what a good output is, choosing the alignment target is itself the ethical act, and no aggregation rule makes that choice neutrally.
Questions to bring to class
- The thesis case converged and the refusal-principle case did not. What feature of a scenario predicts whether the four lenses will agree, and which kind of case should worry an alignment engineer more?
- The consequentialist card reached opposite verdicts at the conversation scale and the deployment scale. When a principle must hold across millions of conversations (7.4), which scale should it be written for, and what is lost in that choice?
- Ubuntu's listed limitation is the "risk of being invoked superficially without engaging the philosophical tradition". Propose a one-sentence test that separates a substantive ubuntu analysis from a decorative one, and apply it to the ubuntu card in the refusal case.
- Which of the four lenses is easiest to compile into a trainable signal (a constitution principle, a reward model), and which is hardest? What does that asymmetry predict about the values deployed models end up with?
Readings for this session
Two core readings and four supplementary readings for this session. All are freely accessible.
Summary and key takeaways
This session set out the tools for reasoning about what a system should be aligned to.
- The machinery does not choose the target: everything Sessions 5 to 7 built transmits a values choice without making one
- Four complementary lenses: Consequentialism (outcomes), deontology (duties), virtue ethics (character), and ubuntu (relationships) each illuminate different dimensions of AI ethics dilemmas
- Disagreement is productive: The four lenses will often produce different answers; that divergence gives ethical reasoning its substance
- Tools: These frameworks are not checklists to apply mechanically. They are tools for developing your capacity for ethical reasoning in a rapidly changing field
- The terrain is vast: This session necessarily focuses on selected dimensions of AI ethics. Labour exploitation, surveillance and carceral AI, autonomous weapons, corporate power concentration, deepfakes and democratic erosion, gender, disability, and more are all important areas a single session cannot cover in depth.
Next, in 8.3: ubuntu and African relational ethics in depth, including the Research ICT Africa Just AI Framework, what these traditions offer that Western individualist frameworks miss, and why they matter for AI governance globally.