The shrinking list
I published a piece that contradicted itself: it argued for assessing what machines cannot do, and reported the objection that dismantles it. Seven parts to settle that debt.
Episode 1 — The shrinking list
On 25 August, I published here a text that contradicted itself. I reported an objection someone had raised, I acknowledged it was valid, and I concluded by saying I did not know how to answer it.
Seven days later, I still had no answer. So I decided to look for one over seven texts.
Here is the debt, as it stands.
What I had written, and why it was insufficient
I had proposed a shift that seems reasonable. Since artificial intelligence writes, summarises, translates, programmes and explains, school should stop evaluating mainly recall, and instead evaluate what resists automation: critical reasoning, argumentation, novel problems, ethical reflection.
That is the reasoning found in just about every report published on the subject over the past two years. It has the advantage of being clear, and that of being reassuring: it assumes there is a properly human territory, and that it suffices to retreat to it.
The objection fits in one sentence: what if that territory did not exist?
More precisely: what if the list of things a machine cannot do were not a foundation, but a mobile frontier that recedes with every version?
The short history of a list that empties
I am not doing futurology. I am describing what has already happened.
For a long time, automatic translation served as an example of what a machine could never do. People explained that a text carries connotations, a register, a culture, and that those things cannot be computed. That was true of the systems of the time. It is no longer true of current systems, whose output is, in many professional use cases, judged sufficient by those who previously paid a translator.
Then people said Go would remain out of reach, because it required a form of intuition. It fell in 2016, and it fell in a way particularly embarrassing for the argument: the machine produced moves that the best players initially judged poor, then recognised as superior. It was not merely playing well. It was playing things human intuition did not reach.
Then came programming. Then structured writing. Then pedagogical explanation. Then, more recently, part of mathematical reasoning.
At each stage, the same sequence: a flag is planted on a competency, it is declared human by nature, and the flag falls within a timeframe counted in years, sometimes months.
I am not claiming everything will fall. I note that the list has never stopped shrinking, and that those who drew it were wrong every time — not out of stupidity, but because they were reasoning about the systems they had before their eyes.
The real problem is not speed. It is the choice of criterion.
One might think the difficulty is a matter of pace: technology moves fast, school moves slowly, school should be sped up.
I believe this is a diagnostic error, and that it leads to anxious and ineffective reforms.
The problem lies elsewhere. By proposing to teach 'what the machine cannot do', we define the content of education by the complement of a technical object. We index a system that takes twelve years to form a pupil on a variable that revises itself every twelve months, and that no ministry controls.
Run the thought experiment. A child entering primary school this year will leave secondary school in about twelve years. If the curriculum they follow is built as the negative of what machines could do when they entered, it will be obsolete before their fourth year of secondary. Not because the curriculum was bad, but because it was defined by subtraction.
A criterion that depends on a mobile frontier produces a policy that runs behind.
And it produces something else, more insidious: it places teachers in a position of permanent technological vigilance, where every laboratory announcement becomes a threat to what they do in class. That is exhausting, demoralising, and there is no reason for it to stop.
What I am looking for, and what I am not
I am not trying to reassure. It is possible that the conclusion of this series will be uncomfortable.
I am not looking for a prediction about what systems will be able to do in five years. I do not know, nobody knows, and an education policy that rests on a technological forecast is fragile by construction.
What I am looking for is more modest and, I believe, more useful: a criterion that does not depend on the frontier.
In other words, a way of deciding what a school must continue to teach that remains valid whether the machine progresses or stalls, whether it clears the next hurdle or hangs on it for ten years.
Does such a criterion exist? I believe so, and I believe there are several. They are not technical. That is precisely what makes them stable.
The itinerary of the seven episodes
This series is a single line of reasoning, cut into seven parts.
1. The shrinking list — this text: showing that the usual criterion is unstable, and why it is a method problem before it is a pace problem. 2. We still teach mental arithmetic — a first stable criterion: there are competencies we need not to produce a result, but to judge a result produced by another. 3. What cannot be delegated — a second criterion: accountability. Some acts matter because a human being answers for them, and that does not depend on who performs them. 4. School does not measure, it sorts — the function nobody names in this debate, and the invisible cost of its degradation for those who have only their transcript. 5. The cost of the shortcut — the most serious objection against episodes 2 and 3, taken seriously: you acquire the foundation only by executing it, and nobody will execute it anymore. 6. Where the line falls, and who draws it — the instruments of a reform are not commutative. The order in which they are activated decides the outcome. 7. What I would do with a curriculum and three years — the synthesis, and who bears the burden of proof.
One clarification, to avoid a misunderstanding that always returns in this debate.
I am putting no one on trial. Not teachers, who do what they can with crowded classes and curricula they did not write. Not pupils, who use the available tool as every generation before them did. Not ministries, which arbitrate under budget constraint with incomplete information.
I simply believe we chose the wrong criterion, collectively and in good faith, and that it is still time to change.
An argument I refuse to use
There is a convenient way of sweeping away the concern, and I see it circulating widely. It consists in recalling that the panic is always the same.
Plato, in the Phaedrus, has Socrates say that writing will destroy memory: those who learn to write will stop exercising their recollection and content themselves with external signs. People said of the printing press that it would drown minds in bad books. Of the calculator, that it would produce generations unable to count. Each time, the announced disaster did not happen.
The argument is seductive, and I will do without it. For two reasons.
The first is logical. The fact that previous alarms proved excessive says nothing about this one. A string of past errors does not establish a rule; it merely establishes that we were wrong before. If we reasoned that way in medicine, no epidemic would ever be taken seriously.
The second is more interesting. These alarms were not entirely wrong. Socrates was right: literate memory has indeed receded. Entire societies used to transmit vast corpora by heart, and that capacity was lost — it survives only among specialists. What Socrates had not seen was that the loss would be compensated by a greater gain, and that society would accept the trade.
That is the real lesson of the precedents. Not that nothing is ever lost, but that something is always lost, and the question is whether the trade was consensual or imposed.
Writing was a good trade. Nobody decided it collectively, however: it imposed itself over centuries, without arbitration.
We, we have an opportunity Socrates did not. The movement is happening on the scale of a decade, before our eyes, in institutions that know how to deliberate. We can choose what we accept to lose. That is a rare historical privilege, and squandering it would be a shame.
Episode 2 — We still teach mental arithmetic
The calculator entered classrooms half a century ago. It performs the four operations infinitely better than a human being, without fatigue and without error.
And yet, fifty years later, we still teach mental arithmetic. In every education system in the world, without notable exception.
That is not a holdover. It is not pedagogical conservatism. It is the trace of a reasoning we did collectively without articulating it, and which I want to articulate here — because it contains, I believe, the first stable criterion this series seeks.
Two competencies confused under the same word
When we say a pupil 'knows how to calculate', we are actually designating two different things.
The first is instrumental: producing a correct result. How much is 47 times 23. On that ground, the machine won in 1642, the game is over, and nobody asks for a rematch.
The second is formative: having built, through manipulating numbers, a representation of what a number is. Knowing that 47 times 23 is a little over a thousand. Feeling that a result is aberrant before even checking. Estimating an order of magnitude in one second, without a tool.
This second competency does not serve to calculate. It serves to judge a calculation.
And it is for that, precisely for that, that we continue to teach mental arithmetic.
The first stable criterion
There are competencies we need not to produce a result, but to judge a result produced by another.
That is the criterion. It is stable because it does not depend on the frontier. No matter how good automatic translation becomes, you need to be able to tell whether a translation is adequate. No matter how good a diagnostic system becomes, you need to be able to tell whether its recommendation fits the patient. No matter how good an analytical tool becomes, you need to be able to tell whether its conclusions follow from its data.
The competency that lets you do this is not the one that produces the result. It is the one that evaluates it. And that competency is built by practice — by writing, by calculating, by translating, by constructing, before you ever learn to judge what a machine produces.
The implication for education is direct: the subjects that build this competency are not being protected by the fact that machines are bad at them. They are being protected by the fact that we need people who are good at judging them.
What this criterion selects
It selects competencies that are supervisory in nature: the ability to understand a result well enough to accept or reject it, and to explain why.
It deselects competencies that are purely productive: the ability to produce a result that a machine can produce faster and more reliably.
And it does so permanently, regardless of the state of the technology. Because the need to judge a machine's output does not depend on whether the machine is good or bad at producing it. It depends on whether a human being is called upon to answer for it.
Why this criterion does not depend on the frontier
A criterion that depends on what the machine cannot do is, by definition, revised every time the machine improves. A criterion that depends on what a human being must answer for is revised only when our idea of responsibility changes — which happens on the scale of decades, not of quarters.
That is the difference between a policy that runs and a policy that holds.
What this criterion does not settle
It tells you what to keep teaching. It does not tell you how to teach it, or how to evaluate it in a world where the tools exist to avoid doing it yourself.
Those are the subjects of episodes 3 through 6. The criterion gives direction. The rest is about what it costs to follow it.
The point of observation
I write from a specific place: a continent where the school-age population is the youngest in the world. That is generally presented as a burden. In this particular matter, it is an advantage.
Education systems reforming elsewhere today do so against a heavy existing structure: constituted bodies, established tracks, century-old exams, textbook publishers, frozen hiring practices. Every change is negotiated against an interest.
Here, a significant part of what will exist in fifteen years has not yet been built. The classrooms are not erected, the tracks not opened, the hiring procedures not frozen.
That guarantees nothing. A less-written page can be filled with anything, and often with a late copy of what others are abandoning.
But it leaves a real possibility: not to repeat the sequence in the wrong order. Not to start by equipping. Not to wait for the signal to degrade before discovering what it was worth. Not to discover in twenty years that we trained a generation of users without supervisors.
Open question: when you hire, what is the signal you actually give the most weight to? And how long has it been since you re-examined it?
Episode 3 — What cannot be delegated
I have proposed one criterion so far: we teach what we need in order to judge the machine's output. Now I want to add a second, because it addresses a dimension the first leaves out.
There are acts whose value depends not on who performs them, but on the fact that a human being answers for them.
The concept
In law, in medicine, in engineering, in finance, there are moments where a signature means: I have checked, I take responsibility, and if it is wrong, it is on me. That is not a ritual. It is a load-bearing structure.
The word for this is accountability. Not in the management sense — being held to account by a superior — but in the deeper sense: being the person who can be called to answer.
A medical diagnosis, a judicial decision, a financial audit, a building plan, a child's school report: in each case, the value of the act depends on the fact that a named human being stands behind it.
This is the second criterion: we teach what must be understood in order to be accountable for what we sign.
Why this criterion is different from the first
The first criterion says: teach what is needed to judge the machine. The second says: teach what is needed to answer for a decision.
The first is about competence. The second is about responsibility. They overlap in many cases — a doctor who cannot judge a diagnostic system cannot answer for a treatment either — but they are not the same thing.
A student can be perfectly capable of judging an AI-generated text and still refuse to sign it, because the act of signing creates an obligation that the act of judging does not.
And conversely, there are cases where accountability precedes competence: a junior doctor must sign a prescription they do not fully understand, because the law requires it. The signature is not the proof of mastery; it is the acceptance of consequence.
What this means for what we teach
If the second criterion is valid, then the content of education is not only what is hard for machines to do. It is what a person must understand in order to stand behind a professional act.
That changes the list.
Writing a first draft is delegable. Reviewing it is not, because review creates accountability.
Producing a translation is delegable. Validating it is not, because validation creates accountability.
Running a calculation is delegable. Interpreting the result is not, because interpretation creates accountability.
The common thread is not difficulty. It is the presence of a moment where a human being says: this is mine.
The aviation precedent, in the right direction
I used aviation in episode 1 to illustrate the shrinking list. I now want to use it to illustrate what was preserved.
Automation has not removed pilots or their training. It has done two things.
It has shifted the content of the licence. Training focuses less on routine manual flying and much more on managing automatic modes, detecting degradation, and above all on handover — that is, the moment when the system hands back control in conditions that are rarely favourable.
It has maintained accountability. No matter what happens, the captain remains responsible for the aircraft. No manufacturer, no regulator has seriously proposed transferring that responsibility to the system. Not out of nostalgia: because a responsibility that rests on nobody answers to nobody.
The result is instructive. The profession did not disappear, it did not remain identical, and it was not redefined by the question 'what does a human do better than the system'. It was redefined by the question 'what must one answer for, and what must one know in order to answer for it'.
That is a reversal, and it is transferable.
The criterion, formulated
What a school must teach is what one must understand in order to be accountable for what one signs.
The order of operations changes completely.
We no longer start from the machine and deduce what remains for humans. We start from the question: in this profession, which acts commit someone? Then we ask: what must one understand to assume them? And that, and only that, becomes the incompressible core of education.
The rest — drafting, formatting, research, writing the first version — can be delegated without damage, and has always been, to assistants, interns, software.
What cannot be delegated is the understanding sufficient to say: I have checked, I sign, and if it is wrong, it is on me.
The objection that comes immediately
It is good, and I want to address it now.
'Responsibility is a legal construct. It changes by law. A single statute created product liability; another will transfer it to a software publisher, or create a no-fault compensation regime. Your criterion is therefore as mobile as the first one; it simply moves at the pace of parliaments rather than laboratories.'
That is factually correct, and I believe it strengthens the criterion rather than weakening it.
The first criterion was about what machines can do. It moved with technology. The second is about what humans must answer for. It moves with law — which moves, but at a pace that allows adaptation.
A parliament that transfers responsibility to a software publisher is making a deliberate, debatable, reversible decision. A laboratory that publishes a new model is not. The first is politics. The second is physics. A criterion anchored in politics is more stable than one anchored in physics, because politics can be influenced.
The real question this raises
If accountability is the right frame, then the question for education is not 'what will machines not be able to do?' but 'where will society still demand that a human being answers?'
And that question has a more concrete form: which professions, in twelve years, will still require a named person to stand behind their output?
A teacher, probably. A doctor, certainly. A judge, certainly. An engineer signing a structural calculation, certainly. A journalist signing an investigation, probably. A programmer signing production code, perhaps.
In each case, the criterion is not 'can a machine do this?' but 'will a human being still be called to answer for it?'
Open question: in your profession, which act creates accountability that no machine can assume?
Episode 4 — School does not measure, it sorts
I have proposed two criteria so far. The first says we teach what is needed to judge the machine. The second says we teach what is needed to answer for what we sign.
Both leave aside a function of school that is rarely named in this debate, and that is precisely the one most at risk from the changes we are discussing.
The function nobody names
School measures. That is the part everyone talks about.
But school also sorts. That is the part nobody talks about, and it is the one that matters most for social mobility.
A transcript is not only an assessment of knowledge. It is a portable, impersonal signal. It tells an employer, a university, a scholarship committee: this person passed a standardised test at a given date, under controlled conditions, and scored X. That signal has a value independent of who the person knows, where they come from, or how they present themselves.
That is its entire virtue. And it is the virtue that generative AI undermines.
How the signal degrades
When a tool can produce homework, essays and answers that are indistinguishable from a student's own work, the transcript ceases to be a reliable signal of the student's capacity. It becomes a signal of the student's capacity minus the tool's contribution, and nobody can tell how much is which.
The traditional response to signal degradation is to intensify the signal: more exams, harder exams, proctored exams. That is one option. It has a cost.
The other option is to move the signal elsewhere: portfolios, supervised projects, oral examinations, probationary periods. That is also an option. It also has a cost.
And the third option, which is the default, is to do nothing: the signal degrades slowly, employers stop trusting it, and they fall back on the second option on their own — without any public decision being made, and without any safeguard.
The invisible cost, and who bears it
Here is the point I consider the most important of this entire series.
The transcript is a portable and impersonal signal. That is its cardinal virtue, and you only realise it when you lose it.
A young man from Ouesso or Kolda who gets a mention on a national exam holds information that nobody can contest, that travels to the capital, and that requires of him neither connections, nor recommendations, nor a known name.
He has nothing else. That is often, very concretely, the only thing he owns.
The son of a capital-city businessman does not need that signal. He has others: his school's name, his father's address book, an internship obtained through a family friend, a ease that shows in the first ten seconds of an interview. If the transcript ceases to be worth anything, he loses almost nothing.
The degradation of the school signal does not cost everyone the same. It is in fact exactly proportional to what you do not have elsewhere.
It is a silent transfer of value that appears in no statistic, that is subject to no explicit arbitration, and whose losers will never know they lost something. They will merely notice that places continue to go to the same people.
And it is particularly serious on a continent where school remains, for a large part of youth, the only mechanism of mobility that does not pass through a network.
Why this will not be settled in a classroom
One conclusion imposes itself, and it is unpleasant for those who write circulars.
If the main problem is a signal problem, then it is not solved in school. It is solved where the signal is read: among employers, in entrance examinations, in hiring procedures.
A pedagogical reform that leaves intact the way employers sort will merely displace the disorder.
That is why episode 6 will not talk about curricula. It will talk about the order in which to act, and the actors who are never in the room when education is discussed.
An honesty required: the signal was already damaged
I would be dishonest to let it be thought that this problem dates from generative systems.
The school signal was already under attack before, and by older, better-known mechanisms. Exam fraud exists everywhere an exam matters. Subject leaks are a documented plague in several systems on the continent. Compliance certificates are for sale. And where class sizes explode without marking resources keeping pace, evaluation degrades on its own, without any technology being to blame.
That must be said for two reasons.
First because it is true, and an analysis that attributes to a novelty what existed before is aiming at the wrong target and proposing the wrong remedies.
Second because it changes the diagnosis. What AI does is not create a flaw: it makes free and universal a flaw that was previously costly and reserved.
Classical fraud requires a network, money, a risk. It was therefore, paradoxically, an additional signal — it benefited mainly those who already had the means. Automated homework generation, by contrast, costs nothing and requires no connections.
One could see this as a democratisation of cheating, and smile. That would be missing the point: when everyone can produce the signal, nobody emits it. What disappears is not the cheaters' advantage, it is the information contained in the diploma — and therefore its value for those who had only it.
The remedy, then, cannot be only an anti-cheating remedy. Better surveillance does not restore a signal whose imitation cost has become zero. We need a device that produces information, not just one that prevents fraud. That is exactly the difference between a better-guarded exam and a better-designed exam.
Open question: when you hire, what is the signal you actually give the most weight to? And how long has it been since you re-examined it?
Episode 5 — The cost of the shortcut
I have proposed two criteria. The first says we teach what is needed to judge the machine. The second says we teach what is needed to answer for what we sign.
Both assume the same thing, and they do not say it: they assume that people capable of judging and answering still exist.
That is what this episode's objection attacks, and I think it is right on the essential.
The objection, at its strongest
It holds in three steps.
One. The capacity for supervision is not acquired through explanation. It is acquired through the repeated practice of the gesture one will later supervise. You do not become capable of sensing that a text is hollow by reading a course on hollow texts: you become so by having written two hundred texts yourself, many of them hollow.
Two. This practice is costly, slow, and unpleasant. It was never maintained by virtue, but by necessity: people wrote because there was no way not to write.
Three. This necessity is disappearing. Yet no pupil, no student, no professional will voluntarily choose the slow version of a task when the fast version produces an acceptable and immediately rewarded result.
Conclusion: in a generation, there will be no more supervisors. There will be people who believe they are supervising, and who validate what seems well-written to them.
This objection is better than everything I have written against it in the four previous episodes. So I will start by strengthening it.
What learning research says about difficulty
In cognitive psychology there is a body of work, notably associated with Robert Bjork, on what are called desirable difficulties.
The idea is counter-intuitive and well established. Certain learning conditions that slow immediate performance improve long-term retention. Testing yourself rather than rereading. Spacing revisions rather than massing them. Having to produce an answer rather than recognising it in a list.
The corollary is even more interesting: the learner is a poor judge of their own learning. They evaluate their mastery by the fluency of the experience. Yet the most fluent conditions — rereading a well-written course, following a clear explanation — produce a sense of understanding markedly superior to actual learning.
Apply that to a tool whose main property is precisely to make the experience fluent, and you get the worst imaginable device for self-assessment. The pupil who has their homework produced does not feel they are cheating. They feel they have understood, because the text seems right when they reread it.
That is an illusion of competence, and it is particularly hard to fight, because nothing in the subjective experience signals it.
The aviation precedent, in the other direction
I used aviation in episode 1 to show what automation preserved. I must now use it to show what it damaged.
The industry has documented, over several decades, a phenomenon it calls automation dependency: as systems take over routine flight, crew manual skills erode for lack of opportunities to exercise them. Post-accident investigations have pointed to handover difficulties in situations where the autopilot disengaged.
The industry's response was not to renounce automation. It was not to trust pilots' goodwill either.
Restoring cost through structure
The industry restored cost through structure. Simulators, imposed failure scenarios, periodic manual-flying requirements. The friction is maintained because it is necessary, scheduled because it would otherwise vanish, and made costly because cheap friction disappears.
The result is instructive. The profession did not disappear, it was not redefined by the question 'what does a human do better', and it was not reduced to supervising a system that does everything. It was redefined by the question 'what must one be able to do in order to remain answerable'.
That is the reversal I am trying to transfer to education.
The cost of the shortcut, in the classroom
If desirable difficulties are real, and if automation dependency is real, then the shortcut of having AI produce schoolwork has a cost that is invisible at the moment and devastating over time.
The pupil saves time. The teacher loses the signal. The employer receives an uninterpretable transcript. And the social function of school — sorting by merit rather than by access — degrades without anyone deciding it should.
The cost is not borne by the pupil who cheats. It is borne by the pupil who does not cheat and whose signal is now worth the same as the cheater's. It is borne by the employer who can no longer distinguish. It is borne by the society that loses its mechanism of impersonal evaluation.
Open question: what is the hardest thing in your training that you now realise built you? Would you still impose it on someone starting out?
Episode 6 — Where the line falls, and who draws it
Six episodes of diagnosis. Now I must propose, and I want to avoid the usual form.
Education reform documents all contain roughly the same list: reform curricula, train teachers, equip schools, integrate technology, assess differently. The list is not wrong. Its defect is to present each item as independent, as though each would have its own effect that adds to the others.
Yet these instruments do not have the same power, and above all — this is the point of this episode — they are not commutative. The order in which they are activated decides the outcome.
I am going to argue that the usual order is almost exactly the reverse of the right one.
Layer 1 — The criterion, and who defines it (maximum effect)
That is the decision that contains all the others: what is taught, what is evaluated, and on what basis.
It is today made on a single criterion — what the machine cannot do — by people who do not control the technology. Episodes 1 through 5 have shown why this mechanically produces the wrong result: it selects competencies by subtraction from a moving target.
Acting on this layer means two things, and they are inexpensive.
Defining the criterion by accountability rather than by complementarity: what must a person understand in order to answer for a professional act? This is the criterion of episodes 2 and 3.
And applying the test from episode 4: is the signal this evaluation produces still interpretable by an employer? A beautiful exam that produces an uninterpretable transcript has failed at its social function.
An institution that does only this, and nothing else from this list, already achieves the essential.
Layer 2 — The signal, rebuilt (strong effect, rarely done)
Episode 4 established this: a signal whose imitation cost has become zero is a signal that no longer sorts.
Rebuilding the signal means restoring cost to evaluation: controlled conditions, no aids, oral components, time pressure, anti-fraud measures that work.
Three possible forms, cumulative: supervised examinations under exam conditions; continuous assessment through verified projects with checkpoints; and evaluation by a third party — an employer, a professional body, an external examiner.
The important word is verified. A project that is never checked is not an assessment. It is a performance with no feedback.
Layer 3 — Maintaining the skill (deferred effect, invisible)
That is what episode 5 made mandatory: if the formative friction does not maintain itself, it must be scheduled.
Concretely: periodic exercises without tools, on the gestures that ground professional judgement. And, to verify that supervision still works, deliberate injection of known errors into workflow — a document containing a false datum in a case review, a deliberately flawed calculation note in a verification.
That is not a trap set for teams and should never be presented as one. It is a measurement instrument, and it is the only one I know that produces information before the accident rather than after.
Layer 4 — The tool and its purchase (real effect, unique effect)
I put it last, and this is not a value judgement: without a tool, nothing that precedes it has an object.
But alone, it produces nothing.
Buying without having decided the criterion gives the same work to the same people, faster, with the same people more tired. Buying without a rebuilt signal gives a transcript nobody can interpret. Buying without skill maintenance gives a chain of validation that empties in silence.
The tool is a multiplier. Multiplied by zero, it makes zero.
The order, and why it is always reversed
Criterion → signal → skill → tool.
The real order is exactly reversed: the tool is bought first, a training plan is announced, the criterion is discovered while walking, and the signal is never discussed.
That is not incompetence. It is a rational consequence of management constraints. The tool is visible, decided in a quarter, appears in the annual report, and reassures a board worried about missing something. The criterion is invisible, takes months, requires listening to people who are not usually listened to, and produces its effects after the mandate of whoever decided it.
In other words: the most effective instrument is the one that yields the least political return inside an organisation.
Naming it does not solve it. But as long as it is not named, deployment failures will continue to be attributed to change management, when they come from a sequencing problem.
Babel, and why the image is more precise than it seems
Magnifica humanitas opens on an alternative: building a new tower of Babel, or building a city where people live together.
I first found the image too convenient. I now believe it is exactly in its place, provided we remember what, in the narrative, makes the tower fail.
It is not its height. It is not the technology — the bricks and bitumen work perfectly, the text specifies. It is that the builders cease to understand each other.
Applied to education, the image says something precise: the tower fails not because the tools are bad, but because the people operating them no longer share a common framework of judgement. An education system that equips teachers with AI without rebuilding the criterion and the signal is building a tower whose bricks work and whose builders do not understand each other.
Open question: of the four layers described here, which one could your country address before the end of the school year — and what really prevents it?
Episode 7 — What I would do with a curriculum and three years
Six episodes of diagnosis. Now I must synthesise, and I want to do it without grandeur.
What I would do first
I would change the criterion. Replace 'what can machines not do?' with 'what must a person understand in order to be accountable for a professional act?' That is a one-line decision that changes everything downstream, because it selects different competencies, justifies different subjects, and legitimises different evaluations.
I would rebuild the signal. Restore cost to evaluation: supervised exams, time pressure, oral components, anti-fraud that works. Not because assessment should be harder, but because a signal that costs nothing to imitate is a signal that does not sort.
I would maintain the skill. Schedule periodic friction: exercises without tools, deliberate error injection, manual reconstruction of what the machine produces in three seconds. Not as punishment, but as the only way to preserve the capacity to judge.
And I would do these three things before buying a single tool.
What I would not do
I would not launch a national AI-in-education programme. Not because it is a bad idea, but because it is the wrong first step. A tool without a criterion is a multiplier applied to zero.
I would not train teachers to use AI. Not because they should not learn, but because training in the tool without rebuilding the signal and the criterion is training people to produce results nobody can interpret.
I would not create an AI literacy curriculum. Not because literacy does not matter, but because teaching people to prompt a system while the system's output has no evaluative framework is teaching a skill with no context of judgement.
I would not wait for the technology to stabilise. It will not stabilise. A policy that waits for stability is a policy that never acts.
The sequence
Three things, in this order, over three years.
Year one: change the criterion and rebuild the signal. This is cheap in resources and expensive in political will. It requires one decision at ministerial level and a dozen pilots in willing schools. The effect is immediate: every evaluation criterion changes, and the signal begins to recover interpretability.
Year two: maintain the skill and design the friction. This requires a curriculum of periodic exercises, a network of evaluators trained in deliberate-error detection, and a framework for third-party assessment. The effect is deferred but real: within two years, the capacity to judge begins to rebuild.
Year three: deploy the tool. Buy the technology, but deploy it inside the framework built in years one and two. The tool then multiplies a non-zero value. The effect is visible: faster production, broader access, lower cost — but within a structure that allows evaluation and accountability.
Why this is not the usual recommendation
The usual recommendation starts with the tool. This one starts with the criterion.
The usual recommendation treats technology as the solution. This one treats it as the multiplier of a solution that must be built first.
The usual recommendation is popular because it is visible. This one is unpopular because it is structural. Changing a criterion is invisible. Rebuilding a signal is painful. Maintaining a skill is expensive and seems backward-looking. Buying a tool is photogenic.
But a photogenic tool applied to the wrong criterion is a well-lit descent in the wrong direction.
Three things this series did not address
An honest series says what it left aside. There are three gaps, and I know them.
Unequal access to the tool. I reasoned as if all pupils had access to these systems. That is false, and the gap is considerable — between countries, between cities and rural areas, between schools in the same city. There is therefore, for a period whose duration I cannot estimate, a double penalty: those without the tool are disadvantaged in daily work, and they will be disadvantaged a second time if evaluation shifts to formats that assume practice with it. I have no solution to offer that does not reduce to a wish.
Language. That is the gap that bothers me most, because it is the one I know best. These systems are vastly more performant in English and French than in Lingala, Kikongo, Wolof, Bambara or Amharic. A pupil who thinks and speaks first in a poorly represented language does not have access to the same tool as others — they have access to a tool that works badly for them, and that forces them to pass through a second language to be understood. The whole question of supervision then arises differently: how do you judge a system's output in a language where it is poor, when you have no benchmark for knowing how poor it is? This subject deserves a series of its own, and it will have one.
Dependency. Nothing I have proposed questions the fact that these systems belong to a very small number of actors, outside the continent, and that their availability, price and behaviour can change without notice. Building education policy on an infrastructure one does not control is a real risk. It is not the school's problem, but it concerns the school.
These three gaps do not invalidate the three criteria. They indicate where the next series must begin.
Open question, the last: if you had to transmit a single competency to someone entering school this year and leaving in twelve years — which one, and why that one?