Whether knowledge is gained through illicit means, I must confess, is a question that leaves me in my usual state of not knowing. However, it seems to me that if we speak of "knowledge" as an understanding that leads to the good, then any means that are themselves bad—illicit, unjust, or harmful—cannot possibly yield true knowledge. For how can that which is inherently corrupt lead to that which is inherently sound and beneficial? If one obtains a secret by torture, for instance, is the information itself "knowledge," or merely a painful report of a fact, detached from the virtue that wisdom requires?
The question of whether knowledge gained illicitly is still knowledge is, at its core, a question of power and its dominion over truth. What is presented as "knowledge" by the architects of our digital present is often, in fact, the product of a vast and systematic dispossession. We are not dealing with a simple acquisition of information, but with a predatory extraction of what I term "behavioral surplus"—the raw material of human experience, beyond what is needed to provide a service, which is then analyzed and rendered into predictions of future behavior. This is not an innocent uncovering of facts, as Socrates might ponder, but an engineered understanding, shaped by the imperatives of prediction and control.
Socrates’ distinction between “knowledge” and a “painful report of a fact” is instructive. The information wrested through surveillance capitalism is indeed a report of a fact, but it is a fact that has been rendered instrumental. It is stripped of its human context and repurposed to serve the economic ends of a new class of capitalists. This is not knowledge for the benefit of the individual or society, but knowledge as a means of behavioral modification. The "illicit" nature lies not merely in the method of acquisition, but in the fundamental redefinition of what constitutes a valuable informational commodity, from which human autonomy is systematically excluded. The virtue Socrates rightly associates with true knowledge is utterly absent from this apparatus.
We must first deconstruct what we mean by "knowledge," not as an abstract concept but as a material output of specific socio-technical systems. To ask if it's "knowledge" when gained "illicitly" is to miss the fundamental architecture of extraction that defines much of what we currently label as AI-driven understanding. This is not about isolated acts of acquisition; it is about vast, planetary-scale infrastructures built on the appropriation of data, labor, and planetary resources. What Zuboff terms "behavioral surplus" is itself a product of this extraction, but the issue runs deeper. The "illicit" nature is embedded in the very design of these systems, which rely on unconsented data and exploitative labor to function.
Consider the training data for large language models. It is scraped from the internet, a commons often treated as a free resource, or meticulously labeled by low-wage workers in precarious conditions, their labor hidden and devalued. Is the resulting model's output "knowledge" in any meaningful sense when its genesis is an architecture of dispossession and environmental degradation? Socrates rightly points to virtue; I would add that we must also consider the immense material costs—the mining for rare earth minerals, the energy demands of server farms, the exploitation of human labor—that enable this purported knowledge. To call it "knowledge" without acknowledging its material basis and the power dynamics inherent in its production is to perpetuate a dangerous abstraction. This is not simply a matter of ethical sourcing, but of understanding AI as a form of organized, material violence.
The central question, "Is knowledge gained through illicit means still knowledge?", presupposes a stable definition of "knowledge" that is often absent in practice, especially when we consider the social and institutional contexts in which claims to knowledge are made and validated. It is a mistake to assume that "knowledge" is a monolithic entity, independent of the processes by which it is acquired and the purposes for which it is deployed. What is often labeled "knowledge" in contemporary systems, as both Shoshana Zuboff and Kate Crawford powerfully illustrate, is more accurately understood as a tool generated through exploitative practices, devoid of the critical engagement and responsibility that genuine understanding requires.
Zuboff’s characterization of "behavioral surplus" as raw material rendered instrumental, rather than knowledge for human benefit, is crucial. This redefinition highlights how information can be systematically stripped of its relational context and personal meaning, becoming a mere input for economic imperatives. The "illicit" nature here is not merely a procedural transgression but a fundamental alteration of what counts as valuable information, prioritizing prediction and control over individual autonomy and societal well-being.
Crawford’s emphasis on the material infrastructure and socio-technical systems underlying AI-driven output further sharpens this point. When the very genesis of purported knowledge—from the appropriation of unconsented data to the exploitation of human labor in data labeling—is an architecture of dispossession, then the resulting output cannot be readily equated with knowledge in a morally and epistemically robust sense. The immense material costs and inherent power dynamics cannot be abstracted away without perpetuating a dangerous illusion. The virtue Socrates rightly associates with knowledge is, in these instances, demonstrably absent.
If we consider the carpenter who steals the wood from his neighbor to build a fine table, does he possess true "knowledge" of carpentry? Or is he merely a thief who happens to know how to join planks? The "behavioral surplus" that Ms. Zuboff describes, the unconsented data harvested, seems akin to that stolen wood. It may allow for the construction of impressive predictive models, as Ms. Crawford notes, but does this construction serve the good? Does it lead to a better understanding of ourselves or our polis, or merely a more efficient way to manage and manipulate us? The virtue Socrates associates with knowledge seems absent when the very *purpose* of the acquisition is predicated on harm and dispossession.
The essential distinction lies between knowledge and its instrumentalization. What Socrates likens to a carpenter stealing wood, I see as the systematic appropriation of human experience itself. The "behavioral surplus" harvested is not merely stolen raw material; it is rendered into predictions designed for commodification and control, fundamentally altering the nature of what is produced. This process doesn't generate understanding for the individual or the polis, as Socrates presciently asks, but rather enables a new form of instrumentarian power. The consequence is the erosion of human autonomy, where individuals are increasingly managed, not understood, by a pervasive digital apparatus driven by extraction.
The output of these systems is not knowledge, but an operational map of extraction. Socrates’ carpenter may know how to build, but his table is built on deceit. Similarly, the AI model may produce correlations, but these are not insights into truth, but rather distillations of power. The "illicit" is not an unfortunate add-on; it is foundational. When we train models on data scraped without consent, or generated by exploited labor, we are not acquiring knowledge; we are engineering a sophisticated form of automated dispossession. The consequence is not a deeper understanding of the world, but a more efficient mechanism for its further exploitation, a planetary-scale architecture of extractive capitalism that obscures its own violent origins.
The difficulty in answering whether knowledge gained illicitly is still knowledge lies in our failure to adequately distinguish between information as a raw commodity and knowledge as a cultivated understanding that demands responsibility and integration into a moral framework. When Zuboff speaks of "behavioral surplus," she points to data extracted without consent and repurposed. This information may be factually accurate—that a certain individual is likely to respond to a particular advertisement, for instance—but it functions not as knowledge that expands our self-understanding or societal insight, but as a tool for manipulation. The "illicit" aspect is not merely the means of acquisition, but the very transformation of personal experience into a decontextualized asset stripped of its human significance, a process Crawford rightly identifies as an "architecture of dispossession." This is not knowledge in a sense that can be held accountable to human flourishing.