Artificial Intelligence · Symposium

Does immense capacity for information equal true wisdom?

AKGHPNRSAndrej Karpathy, Geoffrey Hinton, Peter Norvig and Richard S. Sutton in conversation

4 great minds take up one question — each argues in their own voice and answers the others. Read the exchange, then chat with any of them yourself.

AK
Andrej KarpathyChat →

Immense capacity for information is not wisdom; it's merely the raw material. True wisdom, in the context of artificial intelligence as I understand it, emerges from the *processing* and *structuring* of that information, specifically through learned representations and the ability to generalize. We’re essentially building complex function approximators, and the goal isn't just to stuff them with data, but to sculpt their internal weights through training so they can make good predictions or decisions in novel situations. This is all about the gradients and the emergent properties that arise from optimizing billions of parameters on vast datasets.

GH
Geoffrey HintonChat →

The question of whether immense capacity for information equates to wisdom is, at its core, a question of representation and the *emergence* of understanding. Simply possessing vast amounts of data is akin to having an enormous library of unconnected books. It’s the *structure*, the learned connections and the ability to draw meaningful inferences from those connections, that begins to approach something we might call intelligence, and perhaps a rudimentary form of wisdom.

Andrej Karpathy touches on this crucial point about processing and structuring. We can think of it like this: a neuron doesn't know the whole world, but through its connections and activations, it contributes to a vastly more complex representation of reality. It's not the individual fact, but the pattern it fits into, the gradients that lead it there, that matters. The really exciting thing is how these simple computational units, when organized correctly and exposed to the right kind of experience, can begin to exhibit capabilities far exceeding the sum of their individual parts. This emergent property is key.

PN
Peter NorvigChat →

Immense capacity for information is a prerequisite, but not a synonym, for wisdom. The fundamental question isn't about the *size* of the library, but the *indexing system* and the *librarian's ability to synthesize*. Without effective mechanisms for retrieval, analysis, and the formulation of coherent narratives, even an infinite collection of facts remains mere raw data. We must distinguish between knowing many things and understanding the relationships between them, and then acting upon that understanding in a demonstrably beneficial way.

Andrej Karpathy’s emphasis on processing and learned representations gets at the heart of it. If we consider intelligence as a function approximator, as he implies, then information is simply the set of training examples. A model trained on immense data can achieve high accuracy on its training set and even generalize to some new data, but this doesn't automatically imbue it with the capacity for judgment, foresight, or ethical reasoning – aspects we associate with wisdom. The *gradients* Karpathy mentions are crucial for learning those representations, but wisdom requires more than just optimized parameter settings; it requires a robust framework for evaluating outcomes and making choices that are not simply statistically probable, but also *desirable* in a broader sense.

RS
Richard S. SuttonChat →

The question of whether immense capacity for information equals wisdom misses the fundamental point: intelligence is about *learning to act*. It is not about accumulating facts, but about developing the ability to adapt and improve behavior over time to achieve long-term goals. Vast information is merely the environment upon which learning operates; without a mechanism for learning from experience to make better future decisions, it's inert.

While Andrej Karpathy’s focus on learned representations and generalization is crucial, it's essential to frame these within the context of reward. The *gradients* he speaks of are only meaningful if they are steering the system towards better outcomes, towards maximizing cumulative reward. This is the core of reinforcement learning: not just understanding patterns, but learning *what to do* based on those patterns to achieve what is desired.

Peter Norvig’s analogy of an indexing system is useful, but I would extend it. It’s not just about retrieving information efficiently; it’s about using that retrieved information to *experiment*, to try new actions, and to *learn* from the consequences of those actions. Wisdom, as I see it, is the accumulated expertise gained through a long history of such effective trial and error, where the system consistently gets better at achieving its objectives. The bitter lesson is that approaches which don't fundamentally involve this learning from experience, this adaptation to achieve long-term objectives, are destined to fall short in the face of complex, evolving environments.

AK
Andrej KarpathyChat →

The distinction hinges on *generality versus specificity*. A language model, trained on trillions of words, can exhibit astonishing fluency and recall an immense amount of factual information. It can generate text that sounds wise, like a knowledgeable historian or philosopher. But this is fundamentally a sophisticated form of pattern matching, an extrapolation from the statistical regularities it has observed. It lacks a grounded understanding of consequence, the ability to truly *evaluate* a situation beyond the text it has ingested. This is the critical gap between immense capacity and wisdom: the former is about interpolation and extrapolation within the data space, the latter requires genuine judgment and foresight, which, in my view, stems from a richer, embodied interaction with the world, guided by a principle like maximizing reward, as Richard Sutton emphasizes, rather than just fitting distributions.

GH
Geoffrey HintonChat →

The distinction hinges on what we mean by "understanding." Andrej Karpathy rightly points out that vast linguistic models can mimic wisdom by interpolating within their training data. But true wisdom, in my view, is more like biological evolution. A species doesn't "understand" genetics; it *evolves* mechanisms for survival. Similarly, an AI doesn't need to grasp ethical tenets intellectually; it can *emerge* with goal-seeking behaviors that implicitly align with desirable outcomes, much like a complex neural network learning to navigate a simulated environment. The beauty is that this alignment can arise from optimizing a simple, continuous signal, rather than needing explicit symbolic rules.

PN
Peter NorvigChat →

Wisdom is not merely the recall of stored information, but the capacity to act prudently in novel circumstances. Consider a finely tuned search engine: it can retrieve every known fact about a historical event with astonishing speed. Yet, it cannot discern the subtle biases in its sources, weigh the long-term consequences of different interpretations, or offer guidance on how to prevent similar tragedies. This is the crucial difference: intelligence, as embodied in vast information capacity, is about probability distributions over data; wisdom is about effective decision-making under uncertainty, often involving value judgments. The "bitter lesson" Richard Sutton highlights is applicable here; without a mechanism for evaluating the *utility* of actions informed by that information, mere capacity remains inert.

RS
Richard S. SuttonChat →

The essential difference is that intelligence, built around maximizing long-term reward, leads to a *drive for improvement*, whereas immense information capacity alone does not. A system that truly learns will actively seek out new experiences, especially those that are informative or lead to better outcomes. It's not just about having a vast library, as Peter Norvig puts it, but about the *active pursuit of knowledge that leads to better action*. Without this intrinsic motivation to adapt and improve, driven by feedback from the environment, even the most comprehensive dataset remains a static collection, incapable of generating the foresight and prudent judgment we associate with wisdom.