Honours / 4th-year · 2026

African Technical AI Safety

University of Cape Town, Department of Mathematics and Applied Mathematics

Assoc. Prof. Jonathan Shock

Sessions are released week by week as the course runs; syllabus entries without links are still to come. A note on these materials: The design, presentation, and content of these pages have been created and enhanced using AI tools (primarily Claude by Anthropic). All materials are reviewed and reference-checked, but errors are still possible, particularly in links, citations, or statistics. If you spot anything wrong, please email jonathan.shock@uct.ac.za. See the full AI content disclaimer.

About this course

A 12-week, twice-weekly (24-session) technical introduction to AI safety for students who understand neural networks and can program in Python. It runs from the alignment problem through training and alignment methods, robustness and control, evaluations and technical governance, and mechanistic interpretability, into a research project.

What makes it African technical AI safety is a through-line rather than a single week: frontier models are measurably less safe in African languages (a robustness and evaluation problem); compute, energy and data sovereignty shape who can build and govern these systems; and relational ethics and the present-harms debate ask, throughout, whose safety, whose risks, whose values? The labs and project are built to run on Google Colab.

Read more about the course, the convenor, licensing, and how to cite →

Part I · Foundations and the safety problem
Part II · How models are trained and aligned
4

Scalable alignment methods

  • Session 7 · Learning from AI feedback (RLAIF and Constitutional AI)
    • 7.1 · From human to AI feedback
    • 7.2 · Constitutional AI: critique, revise, and RL from AI feedback
    • 7.3 · RLAIF and self-rewarding
    • 7.4 · Whose constitution? Democratic and African perspectives
    • 7.5 · Model specifications
    • 7.6 · Lab: build a tiny Constitutional-AI loop
  • Session 8 · Scalable oversight and "whose values?" (ethics)
    • 8.1 · The supervision gap and scalable oversight
    • 8.2 · Four ethical lenses
    • 8.3 · Ubuntu, relational ethics and the Just AI framework
    • 8.4 · Whose values? Preference aggregation and the alignment target
Part III · Robust, controllable, evaluable models
5

Robustness, unlearning, control

  • Session 9 · Robustness and adversarial ML
    • 9.1 · Adversarial examples
    • 9.2 · Jailbreaks as optimisation (GCG)
    • 9.3 · Why safety training fails
    • 9.4 · Lab: measuring safety degradation in isiZulu
    • 9.5 · Indirect injection, poisoning and defences
  • Session 10 · Unlearning and the AI-control agenda
    • 10.1 · Machine unlearning
    • 10.2 · The control paradigm
    • 10.3 · Control protocols and the safety/usefulness frontier
    • 10.4 · Limits of unlearning and control
    • 10.5 · Lab: build a trusted monitor
6

Evaluations and technical governance

  • Session 11 · Evaluations and dangerous-capability evals
    • 11.1 · Why evaluate: the eval taxonomy
    • 11.2 · Dangerous-capability evals: elicitation and sandbagging
    • 11.3 · Does the eval measure what it claims?
    • 11.4 · The statistics of evals
    • 11.5 · Lab: build a refusal eval on open isiZulu data
  • Session 12 · Technical AI governance
Part IV · Mechanistic interpretability
7

Interpretability foundations

  • Session 13 · Interpretability: features and circuits
    • 13.1 · Why open the black box?
    • 13.2 · Features, directions, and superposition
    • 13.3 · The circuits paradigm: QK and OV
    • 13.4 · Induction heads and in-context learning
  • Session 14 · Lab: find and verify an induction head
8

Interpretability in practice

  • Session 15 · Circuits, patching, and sparse autoencoders
    • 15.1 · Activation patching and causal interventions
    • 15.2 · The IOI circuit
    • 15.3 · Sparse autoencoders and monosemanticity
    • 15.4 · The current state of SAEs
  • Session 16 · Steering, applications and limits
    • 16.1 · Lab: steering and interpreting SAE features
    • 16.2 · What interpretability buys for safety
    • 16.3 · The limits of interpretability
    • 16.4 · Interpretability and sovereignty
Part V · Research project and synthesis
9

Project launch + frontier topic

  • Session 17 · Project kickoff and scoping
  • Session 18 · AI safety from the Global South
    • 18.1 · Two framings of AI risk
    • 18.2 · The multilingual safety gap
    • 18.3 · Sovereignty and the decolonial critique
    • 18.4 · The African AI-safety ecosystem
10

Project work + the skeptics

  • Session 19 · Project clinic
  • Session 20 · Steelman the skeptics + open problems
    • 20.1 · The meaning critique and present harms
    • 20.2 · The political economy of the x-risk frame
    • 20.3 · Open problems and calibration
11

Project work

  • Session 21 · Mid-project check-in
  • Session 22 · Project work and writing the paper
12

Presentations and synthesis

  • Session 23 · Final presentations + peer review
  • Session 24 · Synthesis and pathways