As artificial intelligence becomes increasingly capable of generating ideas, solving problems and contributing to scientific research, a harder question is emerging: can AI produce genuinely transformational creativity, or is it ultimately working within the conceptual boundaries humans have already created?

If you train a large language model on everything Newton ever wrote and ask it what happens when you fire particles through two slits, it will confidently tell you they land in two piles. It will not predict the interference pattern that actually appears, the one that says particles somehow go through both slits at once. (Buehler, 2026)

The correct answer is not allowed by classical mechanics. Explaining this phenomenon required a paradigm shift to quantum mechanics. New entities and conceptual structures were required to make sense of the two-slit experiment and reconcile it with the emerging framework of quantum mechanics. “If we train [AI] systems to minimise surprise and maximise plausibility, we are effectively training them to suppress the anomalies where scientific revolutions live. You cannot interpolate your way to a paradigm shift” (Buehler, 2026).

Buehler’s example illustrates at least two points. First, a paradigm shift occurs at the high end of the creativity spectrum. A paradigm shift will not happen by noticing a small, overlooked relation. For example, the famous ‘Eureka moment’ by Archimedes happened when he realised how to measure the volume of a strangely shaped object by relying on how much water is displaced when the object is placed in a tub of water. The equivalence of the volume of an object and the volume of displaced water was indeed an obscure relationship that Archimedes noticed. However clever, we would not say that Archimedes’ discovery reached the level of a paradigm shift.

Second, paradigm shifts require the construction of new concepts to connect with the current network of relations and concepts. Next, we can consider how a paradigm shift fits into the ways creativity researchers categorise different degrees of creativity.

Degrees of Creativity

Margaret Boden (1991) distinguished three levels of creativity: combinational, exploratory and transformational. Combinational creativity involves taking two familiar concepts that are seemingly unrelated and finding a relationship between them. In the case of Archimedes, the volume of an object becomes equivalent to the volume of displaced water.

Exploratory creativity involves finding new possibilities within a system of rules. For example, in the Go match between AlphaGo, Google DeepMind’s Go-playing AI, and Go grandmaster Lee Sedol, AlphaGo made a brilliant, surprising move (i.e., move 37), and in a subsequent game Lee Sedol made an astounding move (i.e., move 78), which is now also called the ‘Divine Move’ (Metz, 2016). Go has a strict set of rules, but players can still make improbable moves that are strategically brilliant.

Transformational creativity happens when a rule is added, deleted or altered within a system or conceptual space. For example, Einstein deleted the absolute nature of space and time, created a new construct called spacetime, allowing space to warp and time to pass differently for different observers. In this way, transformational creativity is similar to a paradigm shift.

Relatedly, Kaufman and Beghetto (2009) developed a four-part taxonomy of creativity: mini-c, little-c, Pro-C and Big-C. First, mini-c refers to the small creative connections you make while learning something. Second, little-c refers to everyday problem-solving tasks in daily life. Third, Pro-C refers to creative advancements you make in your professional life that fall short of becoming historically famous. Fourth, Big-C refers to groundbreaking contributions that redefine a field of study or a culture. Further, they stand the test of time and become historically famous. Again, Big-C and a paradigm shift basically align with each other. The other categories among these researchers do not quite align.

Next, we will look at another thought experiment circulating in the AI field.

The Einstein Test

Demis Hassabis, now Chair at Google DeepMind and a Nobel laureate in Chemistry for AI AlphaFold, which predicts protein folding with great accuracy, proposed a test for achieving Artificial General Intelligence (AGI):

“The kind of test I would be looking for is training an AI system with a knowledge cut-off of 1911, and then seeing if it could come up with General Relativity, like Einstein did in 1915. That is the kind of test I think is a true test of whether we have a full AGI system.” (Hassabis, 2026)

This is a thought-provoking test. There was an anomaly known in 1911, namely that Newton’s laws predicted Mercury’s closest point to the Sun, but the prediction was off by a small amount. Einstein’s introduction of a dynamic spacetime rather than Newton’s absolute space and time predicted that, near the Sun, space curved. The equations of this curvature predicted Mercury’s closest point to the Sun exactly. In this way, the elements and relations of Newton’s theory could not remain unchanged and still resolve the discrepancy in Mercury’s location. New concepts and relations needed to be introduced, namely spacetime and the field equations, so that the discrepancy vanished while all other physical phenomena remained consistent with the new theory.

Arguments Against Paradigm-Shifting AI Discoveries

Blaettler & McCaffrey (2026) articulated seven arguments for why current AI cannot pass the Einstein Test. We will mention only a couple because of their generally technical nature. First, this paper works from a different definition of intelligence. ‘The true test of intelligence is not how much we know how to do, but how we behave when we do not know what to do’ (John Holt’s summary of Piaget’s notion of intelligence: Holt, 1964; Piaget, 1972). From this definition, a closed set of knowledge, no matter how large, does not constitute intelligence. It is what you do when your knowledge base is insufficient to answer a particular question. So, the size of the training set or the number of parameters is not the crucial metric for determining the intelligence of an AI. It is its ability to propose new concepts, not just combine existing concepts (i.e., combinational creativity). In large search spaces, it is certainly feasible to find a new possibility, as AlphaGo and Lee Sedol did in the game of Go. But the real goal is transformational creativity (i.e., a paradigm shift or Big-C).

Here is an analogy from Brian Mellinger (2026) that clarifies what is happening with an AI model. A medieval scribe lives in a monastery and learns basically everything written down from the past. A nearby lord hears of the scribe’s abilities and hires him as counsel for his household and properties. Soon, problems start. The scribe cannot help the lord deal with his unruly children because the canonical texts do not deal with disciplining children. The scribe cannot advise on a court case that has no precedent in the ancient texts. The scribe’s knowledge is vast but closed and does not produce the general wisdom and judgement needed to work with novel situations.

Similarly, Blaettler & McCaffrey (2026) argue that a closed body of knowledge has no way to pierce its wall to reach the actual outside world and test new hypotheses. Further, the knowledge the AI possesses is thoroughly porous. McCaffrey and Spector (2017) mathematically proved that the features of any object you select are so astronomically numerous that no supercomputer could examine them all, even if it started from the beginning of the physical universe. And the number of features grows exponentially larger every day because new inventions, materials and substances are submitted to patent databases around the world, with no end in sight. Our selected object has more and more new entities every day with which to interact, as described in patent applications and in combinations involving those entities, in order to see whether a new effect (i.e., feature) of the object presents itself.

For these and other technical reasons, an AI model is severely limited in finding obscure features of a known object in order to be innovative with it at the little-c and Pro-C levels, let alone constructing new concepts and mathematical structures to reframe a whole field or culture with a Big-C, paradigm-shifting discovery.

Conclusion

Current systems provide compelling evidence of combinational and exploratory creativity. There is no real evidence yet for transformational (i.e., paradigm-shifting) creativity. One current conclusion is that AI needs to conduct experiments in the real world in order to compare the results with what it would predict. If this is not possible, then AI would stay locked in to what its current theories say. Presently, the real world is not able to challenge AI’s beliefs by producing an anomaly or a contradiction. The most promising direction to explore is human-AI creativity in which humans conduct real world experiments that produce surprising results. The human may then propose a new concept and the AI helps combine the new concept with all existing concepts in order to deal with the anomaly that the experiment unearthed. These are the currently known ways to involve AI in the highest level of creativity.

About the Author

Dr. Tony McCaffrey holds a PhD in cognitive psychology, specializing in creativity and problem solving. He has co-authored academic papers, published in Harvard Business Review and guides teen students to use their imagination to solve global problems.

References

E. Blaettler, & T. McCaffrey, “The Einstein Test and Beyond: The Architecture of the Semantic Zero,” submitted to Academia AI and Applications, 2026.

M. Boden, The Creative Mind: Myths and Mechanisms. Basic Books, 1991.

M. Buehler, “Why We Must Break the World,” ChemRxiv, 2026.https://doi.org/10.26434/chemrxiv.15001674/v2

D. Hassabis, “The Einstein Test,” India AI Impact Summit, 2026.

J. Holt, How Children Fail. Pitman, 1964.

J. Kaufman, & R. Beghetto, “Beyond Big and Little: The Four C Model of Creativity,” Review of General Psychology, 13(1), 1–12, 2009. https://doi.org/10.1037/a0013688

T. McCaffrey and L. Spector, “An Approach to Human–Machine Collaboration in Innovation,” Artificial Intelligence for Engineering Design, Analysis and Manufacturing, pp. 1–15, 2017.

B. Mellinger, “The Clever Benchmark,” Substack, July 23, 2026. https://brianmellinger.substack.com/p/chapter-ii-the-clever-benchmark

C. Metz, “In Two Moves, AlphaGo and Lee Sedol Redefined the Future,” Wired, March 16, 2016. https://www.wired.com/2016/03/two-moves-alphago-lee-sedol-redefined-future/

J. Piaget, The Principles of Genetic Epistemology. Basic Books, 1972.