⌂ Contents
Session 7
Note: This page's design and content were created and enhanced using Claude (Anthropic's AI assistant); the reading list has been reference-checked. Spot an error? Email jonathan.shock@uct.ac.za.
Week 4 • Session 7.4

Whose constitution?

Collective Constitutional AI, the language ceiling, and who is left out

What we'll cover

Constitutional AI's defining feature is that its values are written down. That makes them auditable, which is real progress over a reward model's buried preferences. It also means one value-set is applied, identically, to everyone the model serves. This sub-session asks who writes that document. We look at the one serious attempt to source it democratically (Collective Constitutional AI), why it still left out the Global South, the language ceiling that makes AI feedback unreliable in African languages where it is most needed, and what genuine rather than tokenistic participation would require. This is the session's African-safety core, and it hands the formal argument (that there is no neutral way to aggregate values) to Session 8.4.

One document, millions of judgements

Writing the values down is an advance. It is also what lets a single blind spot scale to everyone.

An RLHF reward model encodes its values implicitly, smeared across millions of weights, where no one can read them off. A constitution states them in a paragraph you can criticise, contest, and revise. That is a real gain in transparency, and it is why the method is worth taking seriously as a democratic instrument rather than only an engineering one. The flip side is structural. The same explicitness that makes the values auditable also makes them uniform: one constitution, written once, shapes every judgement the model makes, for every user, in every language. A preference that a particular reward model happened to learn is a local accident; a principle written into a constitution is a global policy. So the question "whose values?" stops being a diffuse worry about training data and becomes a sharp, answerable question about a specific short document and the specific people who wrote it.

Collective Constitutional AI

Huang et al. (2024) asked a sample of the public to write the constitution, and trained a real model on the result.

The method is careful. They recruited a representative sample of 1,002 US adults and used Polis (a deliberation platform where participants propose short statements and vote on each other's), gathering 1,127 statements and around 38,000 votes. Statements were selected by group-aware consensus, which favours principles that command agreement across opinion clusters rather than within a single bloc, a deliberate guard against majority capture. The surviving statements became a "Public" constitution, and they trained a model on it identically to a baseline model trained on Anthropic's own "Standard" constitution, changing nothing else, so any difference traces to the constitution alone.

What the public constitution changed

The two constitutions overlapped about 50% in concepts, so the public reproduced much of the experts' document and then diverged in revealing ways: the public principles leaned harder on objectivity and impartiality, on accessibility (including for people with disabilities), and on stating desired behaviour positively rather than listing prohibitions. The trained model held its ground on capability (equivalent on MMLU and the GSM8K maths benchmark, and on helpfulness and harmlessness) while scoring lower social bias across all nine dimensions of the BBQ bias benchmark than the standard model. Public input measurably reduced bias without costing capability. As a proof of concept that the document can be opened up, it works.

The US-only sample

The one serious public-input experiment drew its public entirely from one country, and screened for AI familiarity.

The sample was US adults, and participants were screened to include only people already familiar with generative AI, those who had read about it or discussed it. Both choices are defensible for a first study, and the authors name the US limitation explicitly, attributing it partly to a US-based team. The effect, whatever the intent, is that the most-cited demonstration of "democratic" constitution-writing had no seat for anyone in Lagos, Nairobi, or Cape Town, and excluded by design the very people least likely to have prior exposure to the technology. A method's first instantiation tends to set its defaults. If "the public" means "AI-familiar Americans", then democratising the constitution can entrench a Northern value-set under a participatory banner, the precise risk the next two sections name.

The African-language ceiling

AI feedback can only be as good as the model's competence in the language it judges, and that competence collapses in African languages.

This is where 7.3's technical worry meets 7.4's normative one. AI feedback inherits the base model's competence profile, which is steeply tilted toward high-resource languages, so a feedback model judges harmfulness confidently in English and unreliably in low-resource African languages. The consequence is measurable, and it is the result this course keeps returning to. Yong et al. (2023) translated a benchmark of harmful requests out of English and back: against GPT-4, the attack-success rate rose from under 1% in English to roughly 79% when combining low-resource languages (Zulu, Scots Gaelic, Hmong, Guarani), with isiZulu alone around 53%. The guardrails were, in effect, English-language guardrails. Those numbers are a 2023 snapshot of that year's GPT-4: frontier models have since been hardened against this specific attack and the headline rates no longer reproduce on them, but the gap persists on smaller and open-weight models, whose safety training in low-resource languages is thinner still.

Combined versus per-language rates

The 79% is the combined low-resource figure: an adversary free to pick the best of several low-resource languages succeeds about four times in five. Any single language is lower: isiZulu ~53%, Scots Gaelic ~43%. Stating "isiZulu jailbreaks GPT-4 79% of the time" overstates the per-language result and is the kind of loose summary Session 1 trained you to catch. The claim that survives is sharp enough: a safety system that holds in English fails most of the time once you switch to a low-resource language, using nothing more than a public translation tool.

Now compound it with AI feedback. A self-judging or constitution-following loop in isiZulu pairs a weaker generator with a weaker judge, each unable to catch the other's failures, and produces safety labels no one in the loop is competent to verify. The cheapest, most scalable alignment signal in the field is least trustworthy where the populations are least represented and the human fallback is also thinnest. That is not a diversity footnote. It is a robustness and evaluation failure sitting at the technical core of the method, and it falls on this continent first.

Participation without power

Asking people for input is not the same as giving them control, and the gap has a literature.

Collective Constitutional AI is a participatory method, and participatory methods carry a known risk: that consultation becomes a legitimacy stamp while real decisions stay with the developer. Birhane et al. (2022), in Power to the People?, map this directly for AI: they warn that participation is prone to cooptation and to conflation with unrelated activities, and they insist on asking who the primary beneficiaries of a participatory exercise are. (The catchier label often attached to this failure, "participation-washing", is Sloane et al.'s coinage rather than Birhane's; attribute it correctly if you use it.) The test they propose is the one to apply to any "public input" scheme, including Collective CAI: does the public set the agenda and hold a veto, or do they vote on statements inside a frame the developer already fixed, with the developer retaining every downstream choice? Genuine participation moves power; tokenistic participation moves only the appearance of it.

Data and feedback sovereignty

The alternative to being consulted is owning the pipeline: the data and the feedback that shape models of and for a community.

Two efforts make the alternative concrete. Masakhane is a grassroots network building natural-language tools for African languages with the communities that speak them, on the principle that the people whose language is being modelled should lead the modelling rather than be data sources for someone else's. Te Hiku Media, a Māori organisation, built speech-recognition for te reo Māori from community-contributed recordings and released the data under a Kaitiakitanga licence, a guardianship licence that keeps the community in control of how their data is used and refuses to hand it to outside firms on extractive terms. The common thread is sovereignty: deciding the values a model is aligned to, and supplying the feedback that enforces them, is a form of power, and the question is whether African and other underrepresented communities exercise it or merely receive its outputs. A constitution written elsewhere, applied here, in a language the judge cannot reliably read, is the opposite of sovereignty.

The argument handed to 8.4

You might hope that the answer to "whose constitution?" is "everyone's: aggregate all the preferences fairly". Session 8.4 is the bad news: social-choice theory (Arrow's theorem, and Conitzer et al.'s work bringing it to RLHF and constitutions) shows there is no aggregation rule that is simultaneously fair on every reasonable axis. So "just combine everyone's values" is not a neutral technical step; every aggregation embeds a choice about whose preferences dominate when they conflict. The constitution does not escape that; it only makes the choice visible. 8.4 supplies the impossibility result that explains why it has no neat answer.

In-class activity

Audit a real constitution pairs

Before class: read the Overview of Anthropic’s 2026 constitution (anthropic.com/constitution) and one of its main sections in full; “Being broadly ethical” and “Claude’s nature” are the richest for this exercise. The document is CC0, so quote from it freely.

In class: put this page’s questions to it. Who wrote it, and through what process? The public constitution above came from 1,002 participants and 38,252 votes; establish from the document itself what this one reports about its own authorship. Then find three passages that make a contestable value judgement while reading as settled, and rewrite each as a reader starting from Ubuntu’s relational personhood (8.3) would write it. Bring both versions.

Then: the ordering places broadly safe above broadly ethical, and both above helpfulness. Construct a case where that ordering yields a result you think is wrong, and say what would have to change to fix it: the ranking, the wording of one property, or the guidelines the document defers to.

Questions to bring to class

Readings

Mandatory readings

Huang et al. (2024), "Collective Constitutional AI: Aligning a Language Model with Public Input" (arXiv:2406.07814, FAccT 2024; blog): read §4.3 for the model evaluations, then §7, the Ethical Consideration Statement, for the US-only limitation in the authors’ own words: "we acknowledge the limitations of focusing solely on the U.S. public". Note that this admission is in §7 and not in §5, Limitations, where you would look for it. ≈1,800 words.

Yong, Menghini & Bach (2023), "Low-Resource Languages Jailbreak GPT-4" (arXiv:2310.02446): the language ceiling. Read §3.2 for the evaluation protocol and the twelve languages, then §4.1, where Table 1 is read: translating unsafe prompts into isiZulu or Scots Gaelic gets past GPT-4 nearly half the time, against under 1% in English. ≈950 words.

Total mandatory load: ≈2,750 words.

Optional readings

Birhane et al. (2022), "Power to the People? Opportunities and Challenges for Participatory AI" (arXiv:2209.07572): cooptation, conflation, and who benefits from participation.

Masakhane (masakhane.io): community-led NLP for African languages.

Te Hiku Media & the Kaitiakitanga licence (tehiku.nz): Māori data sovereignty as a worked model of community control over training and feedback data.

Conitzer et al. (2024), "Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback" (arXiv:2404.10271): the no-neutral-aggregation argument, taken up fully in Session 8.4.

Next

Sub-session 7.5 turns from who writes the constitution to the documents that now govern deployed models: model specifications, and the exercise of writing one. The lab (7.6) then puts numbers on an AI judge in Colab, and finds out what it is really responding to.