About this course
A 12-week, twice-weekly (24-session) technical introduction to AI safety for students who understand neural networks and can program in Python. It runs from the alignment problem through training and alignment methods, robustness and control, evaluations and technical governance, and mechanistic interpretability, into a research project.
What makes it African technical AI safety is a through-line rather than a single week: frontier models are measurably less safe in African languages (a robustness and evaluation problem); compute, energy and data sovereignty shape who can build and govern these systems; and relational ethics and the present-harms debate ask, throughout, whose safety, whose risks, whose values? The labs and project are built to run on Google Colab.
Read more about the course, the convenor, licensing, and how to cite →
Orientation and foundations
-
Session 1 · What technical AI safety is
-
Session 2 · Deep learning and scaling laws
The alignment problem and the physical substrate
-
Session 3 · The core alignment problem
-
Session 4 · The physical substrate of scale (energy)
LLM training, end to end
-
Session 5 · From pretraining to assistant
-
Session 6 · RLHF and RL fine-tuning
Scalable alignment methods
-
Session 7 · Learning from AI feedback (RLAIF and Constitutional AI)
- 7.1 · From human to AI feedback
- 7.2 · Constitutional AI: critique, revise, and RL from AI feedback
- 7.3 · RLAIF and self-rewarding
- 7.4 · Whose constitution? Democratic and African perspectives
- 7.5 · Model specifications
- 7.6 · Lab: build a tiny Constitutional-AI loop
-
Session 8 · Scalable oversight and "whose values?" (ethics)
- 8.1 · The supervision gap and scalable oversight
- 8.2 · Four ethical lenses
- 8.3 · Ubuntu, relational ethics and the Just AI framework
- 8.4 · Whose values? Preference aggregation and the alignment target
Robustness, unlearning, control
-
Session 9 · Robustness and adversarial ML
- 9.1 · Adversarial examples
- 9.2 · Jailbreaks as optimisation (GCG)
- 9.3 · Why safety training fails
- 9.4 · Lab: measuring safety degradation in isiZulu
- 9.5 · Indirect injection, poisoning and defences
-
Session 10 · Unlearning and the AI-control agenda
- 10.1 · Machine unlearning
- 10.2 · The control paradigm
- 10.3 · Control protocols and the safety/usefulness frontier
- 10.4 · Limits of unlearning and control
- 10.5 · Lab: build a trusted monitor
Evaluations and technical governance
-
Session 11 · Evaluations and dangerous-capability evals
- 11.1 · Why evaluate: the eval taxonomy
- 11.2 · Dangerous-capability evals: elicitation and sandbagging
- 11.3 · Does the eval measure what it claims?
- 11.4 · The statistics of evals
- 11.5 · Lab: build a refusal eval on open isiZulu data
- Session 12 · Technical AI governance
Interpretability foundations
-
Session 13 · Interpretability: features and circuits
- 13.1 · Why open the black box?
- 13.2 · Features, directions, and superposition
- 13.3 · The circuits paradigm: QK and OV
- 13.4 · Induction heads and in-context learning
- Session 14 · Lab: find and verify an induction head
Interpretability in practice
-
Session 15 · Circuits, patching, and sparse autoencoders
- 15.1 · Activation patching and causal interventions
- 15.2 · The IOI circuit
- 15.3 · Sparse autoencoders and monosemanticity
- 15.4 · The current state of SAEs
-
Session 16 · Steering, applications and limits
- 16.1 · Lab: steering and interpreting SAE features
- 16.2 · What interpretability buys for safety
- 16.3 · The limits of interpretability
- 16.4 · Interpretability and sovereignty
Project launch + frontier topic
- Session 17 · Project kickoff and scoping
-
Session 18 · AI safety from the Global South
- 18.1 · Two framings of AI risk
- 18.2 · The multilingual safety gap
- 18.3 · Sovereignty and the decolonial critique
- 18.4 · The African AI-safety ecosystem
Project work + the skeptics
- Session 19 · Project clinic
-
Session 20 · Steelman the skeptics + open problems
- 20.1 · The meaning critique and present harms
- 20.2 · The political economy of the x-risk frame
- 20.3 · Open problems and calibration
Project work
- Session 21 · Mid-project check-in
- Session 22 · Project work and writing the paper
Presentations and synthesis
- Session 23 · Final presentations + peer review
- Session 24 · Synthesis and pathways