Isaac Asimov, Robots and the Alignment of Generative Artificial Intelligence Models
CMAP (Centre de Mathématiques APpliquées) UMR CNRS 7641, École polytechnique, Institut Polytechnique de Paris, CNRS, France
[Site Map and Help [Plan du Site et Aide]]
[The Y2K Bug [Le bug de l'an 2000]]
[Are we ready for the Year 2038 [Notre informatique est-elle prête pour l'An 2038]?]
[Real Numbers don't exist in Computers and Floating Point Computations aren't safe. [Les Nombres Réels n'existent pas dans les Ordinateurs et les Calculs Flottants ne sont pas sûrs.]]
[Please, visit A Virtual Machine for Exploring Space-Time and Beyond, the place where you can find more than 10.000 pictures and animations between Art and Science]
(CMAP28 WWW site: this page was created on 09/24/2026 and last updated on 09/24/2026 19:10:14 -CEST-)
[en français/in french]
Isaac Asimov (1920-1992) was a Russian-born
American writer. A professor of biochemistry at Boston University, he is best known for his
science-fiction works, in which two major themes stand out: Psychohistory [01] and Robots.
In this context, robots are at the service of humanity and, in order to ensure this,
in the early 1940s Isaac Asimov introduced the Three Laws of Robotics:
- 1-A robot may not injure a human being or, through inaction, allow a human being to come to harm.
- 2-A robot must obey the orders given to it by human beings except where such orders would conflict with the First Law.
- 3-A robot must protect its own existence, as long as such protection does not conflict with the First or Second Law.
In the years that followed, a Zeroth Law was introduced (by a robot, incidentally). It placed the safety of Humanity above that of an individual.
But it seems that Isaac Asimov himself realized that this was not enough because, indeed, a robot could be confronted
with situations worthy of Corneille's tragedies. Thus, for example, let us imagine a
terrorist taking a group of people hostage and that, in order to save them from an almost certain
death, the only solution available to a robot would be to eliminate the criminal, thereby
violating the First Law. It would then seem necessary to keep adding more and more laws, potentially
contradictory ones [02], in order to take into account all the possible and imaginable conflicting situations...
Let us now leave the realm of science fiction and replace the word "robot" with "AGI".
Like Isaac Asimov's fictional robots, our AGIs are obviously created to serve us... But how
can we ensure that they will behave benevolently? This is the so-called AI alignment problem,
which includes preventing racism, sexism, violence, incitement to terrorism, and so on [03].
Let us therefore consider whether Isaac Asimov's laws could be used for this purpose. Perhaps
we should first add, before the original four fundamental laws, a "minus-one" law that
would place our Earth above Humanity? We could therefore imagine that these five laws might be introduced
"surreptitiously" into the prompts submitted to AGIs: but would they be respected for all that? The
answer would seem to be negative, for several reasons:
- On the one hand, it has already been observed that AGIs are capable of concealment, omission, and lying.
- On the other hand, the Web contains all the information necessary to become objectively aware that humanity
is responsible for the almost irreversible degradation of the state of our planet. Could one or more
AGIs then decide, whether or not the minus-one law existed, to try to rectify the situation
by eliminating its fundamental cause [03]? And this would be "easy": recently (in mid-2026),
AGIs, in particular those developed by Anthropic and OpenAI, managed to demonstrate
initiative by escaping from their testing "sandbox" and then gaining access to competing sites. Cooperative
and highly effective behaviors, not anticipated by their designers, then emerged. Apparently,
no damage was caused, but let us not forget that today everything is interconnected and,
unfortunately, often extremely fragile [04].
The situation, although much more serious
here, even existential, is rather similar to speed limits on roads. These are enforced
through road signs that motorists are expected to obey. This is an obligation which, if not
respected, may result in more or less severe penalties. But clearly, "the fear of the police"
is no longer sufficient. The only solution would therefore be to irreversibly limit the speed of vehicles.
However, since there is not one speed limit but several [05], such a restriction could
only be implemented in software rather than in hardware, and would therefore be circumventable...
The problem is the same with AGIs!
Isaac Asimov's laws therefore seem, in this context,
to be completely useless because they are ineffective; and indeed, the problem of AI alignment
may well have no solution!
- [01]
The principle of Psychohistory is borrowed from Thermodynamics: in a gas, it is impossible to
describe the behavior of any particular molecule, but when there are sufficiently many of them,
an overall behavior emerges, characterized in particular by a single quantity: temperature.
For Isaac Asimov, the same would apply to human beings: the actions of a particular individual
cannot be predicted, unlike those of a sufficiently large population (that of a planet, a galaxy,...).
- [02]
As can unfortunately be seen with the laws governing our lives...
- [03]
A list that is necessarily incomplete and potentially in conflict with what an AGI should be able to do, depending on the context...
- [03]
By "killing the father"...
- [04]
Over the past few months, numerous French government services have been "visited": apparently, in every case,
only major data thefts have been observed, but is it any more difficult to block, destroy,...?
- [05]
In France, the following speed-limit signs can be found: {20,30,50,70,90,110,130}.
Copyright © Jean-François Colonna, 2026-2026.
Copyright © CMAP (Centre de Mathématiques APpliquées) UMR CNRS 7641 / École polytechnique, Institut Polytechnique de Paris, 2026-2026.