
In July, the United States made one vision of AI-assisted knowledge production unusually concrete.
The Genesis Mission had been launched by executive order in November 2025 as a national effort to use artificial intelligence to accelerate scientific discovery. But on July 22, 2026, the project became much more visible. The White House announced more than $5 billion in federal commitments, more than fifteen participating agencies, and a set of National Science and Technology Challenges. Stanford and SLAC were among the institutions selected to lead projects. A day earlier, the White House report Science: A New Golden Age had described an even larger institutional ambition: domain-specific scientific foundation models, high-value datasets, AI-enabled verification infrastructure, autonomous laboratories, and eventually “AI-native scientific institutions.”
The significance of this effort extends beyond the individual projects it funds.
Genesis is an attempt to redesign scientific infrastructure around the capabilities of AI. Supercomputers, experimental facilities, federal datasets, scientific instruments, foundation models, and researchers are not being treated as separate resources. They are being assembled into systems in which hypotheses can be generated, simulations run, experiments conducted, results checked, and procedures repeated. The ambition is to make parts of scientific reasoning computationally executable.
That raises a question for the fields outside this vision.
What is the equivalent infrastructure for history, law, journalism, literary studies, qualitative social science, and other disciplines in which the central intellectual operation is not usually experiment or calculation, but interpretation?
These fields are not entirely absent from federal thinking about AI. The National Endowment for the Humanities has supported programs examining AI through the lenses of ethics, law, history, philosophy, culture, and democratic life. Humanities scholars are studying how AI affects power, uncertainty, creativity, education, and civic institutions. This work matters, but it is structurally different from what Genesis is attempting in science.
The emerging scientific strategy asks how institutions themselves can become AI-native: how data, verification, instruments, models, and procedures can be connected so that accumulated scientific work becomes easier to inspect, rerun, and extend. In interpretive fields, by contrast, AI adoption is still largely discussed as a matter of access to models, prompting, literacy, disclosure, plagiarism, hallucination, and the social consequences of AI. That discussion leaves a deeper problem untouched.
The most valuable intellectual resource in many non-technical disciplines is the judgment surrounding the text.
And unlike scientific datasets, computational notebooks, experimental protocols, or software tests, that judgment is still remarkably difficult to preserve, execute, transfer, or reuse.
The Unequal Distribution of Executable Judgment
A more useful difference than objective truth versus subjective opinion concerns how judgment is externalized.
Scientific and technical work has spent decades building infrastructures that turn many judgments into things that can be executed, inspected, rerun, and compared. Software development has version control, unit tests, regression tests, dependency management, reproducible builds, and continuous integration. Quantitative research has datasets, statistical specifications, scripts, computational notebooks, and reproducibility packages. A researcher can alter a parameter and rerun an analysis. A programmer can change a function and immediately discover which earlier tests now fail.
Genesis is an unusually ambitious extension of this tendency. Its underlying wager is that scientific productivity can increase dramatically if models are connected not only to information, but to an environment capable of testing what those models produce.
These systems do not remove human judgment: humans still decide what should be measured, which model is appropriate, what constitutes an acceptable error, which experiment matters, and which results deserve further attention. But once some part of that judgment has been formalized, machines can repeatedly apply it.
Interpretive disciplines often possess the judgment without possessing this second layer.
A historian’s methodological boundary may exist in years of reading and a few paragraphs of peer-review correspondence. An editor’s standard for evidentiary overstatement may appear in tracked changes on one manuscript and disappear when the document is finalized. A lawyer’s distinction between a controlling precedent and a merely analogous case may be carefully explained in one brief but never become a reusable rule for evaluating the next hundred proposed arguments. A qualitative researcher’s reason for rejecting one interpretation and accepting another may remain inside a seminar discussion, an email exchange, or personal memory.
The judgment exists; the infrastructure around it does not.
This becomes a major limitation with generative AI.
Models can produce more summaries, comparisons, interpretations, classifications, and arguments than any individual researcher could manually produce. But increasing the supply of candidate answers does not automatically increase the supply of trustworthy knowledge. If anything, the bottleneck moves in the opposite direction.
When generation becomes abundant, evaluation becomes scarce.
AI can propose ten interpretations in seconds. Someone still has to decide which are defensible. It can summarize one hundred documents. Someone still has to determine which distinctions disappeared during compression. It can connect a quotation to a claim. Someone still has to decide whether the quotation actually supports the claim in context.
This is one reason interpretive disciplines have received a more ambiguous benefit from generative AI than its apparent capabilities would suggest. AI dramatically lowers the cost of producing possible intellectual outputs, but the most important part of the workflow may never have been producing those outputs in the first place.
It was judging them.
An Answer Is Not Yet Knowledge
The problem is sometimes described as hallucination. AI invents citations, misstates facts, or attributes claims to sources that do not contain them.
These are real problems, but they are only the easiest cases.
An AI-generated answer can be factually accurate and still be epistemically inadequate.
Suppose a model produces a plausible interpretation of a historical document. The quotation exists, the date is correct, the translation is reasonable, and nothing has been fabricated.
Important questions still remain.
Which edition of the document was consulted? Which passage supports the interpretation? What alternative reading was rejected? Why was it rejected? Does the conclusion apply to the author’s entire body of work or only to this particular text? Was the judgment based on a methodological principle that would also apply to another case? Who made that judgment? If new evidence appears, which earlier conclusions should be reconsidered?
A fluent answer does not contain this structure automatically, and can conceal it.
That is especially dangerous because fluent language resembles finished knowledge. The system gives us the conclusion while compressing the intellectual path that produced it. A contested interpretation becomes a smooth paragraph. A provisional judgment becomes a confident sentence. A methodological choice disappears behind the appearance of neutrality.
The result is a form of detachment: the answer becomes separated from the source, the reasoning, the scope, and the person responsible for accepting it.
In interpretive scholarship, those relationships are not metadata added after knowledge has been produced. They are part of what makes the knowledge defensible.
This is where the contrast with the emerging AI-for-science infrastructure becomes useful. Scientific AI systems increasingly assume that model output should be connected to environments capable of verification. The equivalent problem for interpretive disciplines is harder because many of their decisive tests cannot be reduced to a laboratory measurement or executable function.
But a test that is harder to execute can still be structured.
The Individual Researcher as Oracle
The obvious response is not to prohibit AI from interpretive work, and a vaguely human presence “in the loop” is not enough. The more important task is to make human judgment itself more durable.
Researchers already act as what might be called oracles throughout their work.
By oracle, I do not mean an infallible authority or a person who possesses final truth. An oracle is simply an identifiable adjudicator who can make a bounded judgment, provide reasons for it, specify the circumstances under which it applies, and remain responsible when the judgment is challenged.
A historian accepts one interpretation and rejects another. A lawyer decides that a precedent applies to one dispute but not another. A journalist determines that two sources are sufficiently independent to corroborate a claim. A qualitative researcher decides that a passage supports one analytical category but not a broader one. An editor decides that the available evidence justifies claim A but not the stronger claim B.
These are disciplined judgments formed through training, accumulated experience, methodological commitments, and repeated confrontation with evidence, not arbitrary intuitions.
What usually disappears is the structure of those decisions.
Consider what would happen if an individual research judgment were preserved as a reusable intellectual object rather than as a comment.
At minimum, such a judgment could retain: the exact source state on which the judgment relied; a locator identifying the relevant passage, page, field, or timestamp; the specific claim being evaluated; the researcher’s verdict on that claim; the rationale explaining the verdict; the scope within which the judgment is intended to hold; the identity of the adjudicator responsible for it; the rights governing access and reuse; and the version showing how the judgment changes over time.
The important idea is the transformation in the status of the judgment.
A correction that would previously disappear inside an AI conversation becomes part of a researcher’s accumulated methodology. A decision made during one paper can be revisited during another. A revised interpretation does not erase an earlier one but supersedes it. A future AI system can apply previously established criteria to new material without pretending that it invented those criteria itself.
The researcher’s judgment becomes something closer to an intellectual asset: not property in the crude sense of owning facts, but an accumulated body of authored reasoning that can be preserved, inspected, revised, transferred, or withheld from reuse.
This could change the relationship between researchers and AI.
At present, researchers often contribute their most valuable intellectual labor precisely when they correct the machine. The model proposes something inadequate; the researcher explains why it is inadequate; the system moves on. The correction may improve the immediate conversation, but the researcher’s judgment rarely becomes a durable part of the researcher’s own scholarly infrastructure.
The output survives, and the act of evaluation disappears.
A better system would reverse this relationship. AI assistance should help researchers accumulate their judgments rather than merely consume them.
What Machines Should Do
Once human judgment is treated as the oracle rather than as an obstacle to automation, the proper role of AI becomes clearer.
Machines are extremely useful for operations that surround judgment.
They can verify whether a quotation actually appears in a particular version of a document. They can compare editions. They can detect that a citation points to the wrong page. They can identify claims that resemble earlier disputed claims. They can retrieve the evidence associated with an earlier decision. They can show which conclusions depended on a rule that has since changed. They can generate candidate interpretations for review. They can rerun previously established checks against a new manuscript.
These are substantial benefits, and this is where the lesson of AI-for-science should be taken seriously rather than merely admired from outside. The power of a system such as Genesis does not come from asking a model to know everything. It comes from surrounding models with datasets, instruments, procedures, and verification systems that constrain what their outputs can become.
Interpretive fields need their own version of that surrounding infrastructure.
But there should be a boundary between checking the conditions of judgment and inheriting the authority to judge.
If a quotation does not exist in the cited source, a machine can reject the citation mechanically. There is little interpretive value in asking a human to approve a nonexistent sentence.
Whether an existing quotation supports a contested historical interpretation is different. That decision may require context, disciplinary knowledge, theoretical commitments, or an acknowledgment that multiple readings remain legitimate.
The correct machine response in such a case is not necessarily “true” or “false.” It may simply be: this requires judgment.
AI becomes more useful when it knows where its authority ends.
Many Oracles, Not One
Treating researchers as oracles raises an obvious concern. Scholarship depends on disagreement, and formalizing judgment could look like turning one researcher’s interpretation into permanent truth. A good oracle system would preserve disagreement rather than eliminate it.
Two scholars may examine the same evidence and reach different conclusions. Their judgments should not automatically be averaged into a synthetic consensus. Each should remain associated with its own rationale, scope, and intellectual lineage.
One oracle might conclude that a passage supports a claim within a narrow historical context. Another might reject the same claim because of a different theory of interpretation. A third might argue that the available evidence cannot resolve the disagreement.
All three judgments can coexist.
The purpose of structure is to make interpretation legible.
This matters for AI because current systems have a strong tendency to compress disagreement. When several positions are available, a model often produces a smooth middle position that sounds balanced while obscuring the actual structure of the dispute.
But some disagreements should remain disagreements.
Preserving separate oracles allows AI to represent plurality without pretending that plurality has already been resolved.
It also gives intellectual change a history.
Researchers change their minds. New evidence appears. Concepts evolve. A methodology that once seemed adequate may later be rejected. Instead of overwriting the old judgment, a new version can supersede it while preserving the previous rationale.
That is a record of scholarship functioning properly.
Judgment Must Also Carry Rights
Making judgment reusable creates an ethical problem as well as an epistemic opportunity.
A researcher’s judgment is not free raw material.
The reasoning produced by a scholar, editor, community, lawyer, journalist, or participant may carry obligations concerning authorship, confidentiality, cultural authority, privacy, and consent. A judgment derived from a public statute does not necessarily have the same rights structure as one derived from a confidential interview. An interpretation provided by an Indigenous community cannot simply be treated as interchangeable with unrestricted public text. A private peer-review comment should not silently become training data because it passed through an AI system.
For this reason, rights should travel with judgment.
Who may inspect it? Who may reuse it? Can it be transferred to another system? Can it be used for model training? Can permission later be withdrawn? What happens to derivative judgments if the underlying material is deleted or access is revoked?
These questions become more important precisely when AI makes reuse easy.
They also reveal why the institutional challenge for interpretive disciplines cannot be solved merely by purchasing access to increasingly capable models.
Scientific AI infrastructure is being built partly around access to scarce computational resources, experimental facilities, and large datasets. Interpretive infrastructure has a different scarce resource at its center: accumulated human judgment, much of which exists within relationships of authorship, expertise, trust, and consent.
The defensible goal is consent-based reuse of structured judgment while preserving the adjudicator’s authority over its context and downstream use.
AI should make scholarship more portable without making scholars more extractable.
From AI Assistance to Intellectual Infrastructure
The common framing of AI adoption in non-technical fields is surprisingly narrow.
Researchers are told to learn prompting, verify citations, disclose AI use, identify hallucinations, protect student learning, and become more AI-literate. Humanities programs examining AI rightly ask what these technologies mean for culture, ethics, democracy, and human flourishing. Those responses largely leave the architecture of knowledge production unchanged.
The researcher remains the final evaluator, yet the researcher’s evaluations continue to disappear.
This is the asymmetry that the emergence of the Genesis Mission makes easier to see.
The United States is now investing in environments designed so that scientific work and machine capability can interact repeatedly, not only in scientists using AI. Verification infrastructure, shared datasets, computational resources, instruments, autonomous laboratories, and institutional procedures are being redesigned around that interaction.
There is no reason every discipline should imitate the laboratory, but every reason to ask what the functional equivalent of such infrastructure would be elsewhere.
Instead of asking only how humanities scholars, lawyers, journalists, historians, and qualitative researchers can adapt themselves to systems designed elsewhere, we should ask what forms of intelligence these disciplines already possess that AI systems lack.
One answer is disciplined judgment.
Interpretive fields have spent centuries developing practices for determining what texts mean, how evidence constrains claims, when context changes interpretation, which authorities may speak for particular materials, how disagreement should be represented, and when uncertainty must remain unresolved.
These practices are valuable intellectual technologies in their own right, not primitive versions of computation waiting to be automated away.
Their weakness is that they are poorly externalized.
If those judgments could become traceable, versioned, reusable, contestable, and rights-aware, AI could interact with them without simply absorbing or replacing them.
Researchers could use machines to perform repetitive verification while keeping interpretive authority human. They could rerun established standards against new material. They could identify when a changed source or changed criterion affects earlier conclusions. They could carry accumulated methodological reasoning from one AI system to another instead of rebuilding their standards from scratch each time.
Institutions could do something similar.
A newsroom could preserve why an evidentiary threshold was applied in one investigation and ask whether a later story meets the same standard. A legal organization could track how its interpretation of a doctrine changed across cases without collapsing that history into a single model response. A historical archive could allow scholars to attach competing, rights-aware interpretations to stable source states. A qualitative research team could preserve coding disputes and the rationales that resolved, or deliberately failed to resolve, them.
The gain would come from giving human judgment some of the durable properties that computational artifacts already possess, not from turning humanities scholarship into computer science.
The Real Opportunity
Even a much more capable model would leave the basic problem unresolved.
Someone still has to decide which answers deserve authority.
The current American push toward AI-enabled science implicitly recognizes something important: intelligence alone is not infrastructure. A powerful model becomes more valuable when embedded in institutions capable of connecting output to evidence, procedures, verification, revision, and responsibility.
Interpretive disciplines need to take that lesson seriously on their own terms.
Their equivalent of scientific verification infrastructure is the structure surrounding human judgment, rather than another benchmark, another foundation model, or an automated judge.
These acts are already the hidden machinery of non-technical knowledge: the historian deciding that a source cannot support a stronger claim; the lawyer explaining why a precedent does not extend to the present case; the qualitative researcher narrowing the scope of a category; the editor distinguishing evidence from overstatement; the journalist explaining why one form of corroboration is sufficient and another is not; the scholar revising an earlier interpretation after encountering a counterexample.
AI should help preserve that machinery.
The future of AI-assisted scholarship should be built around making that oracle explicit: identifying who judged, what they judged, which evidence they relied on, why the judgment was made, where it applies, how it has changed, and who has the right to reuse it.
Genesis is significant because it shows what becomes possible when a country treats the institutional environment around AI as seriously as the models themselves.
The next question is whether we can extend that ambition beyond the domains in which verification is already computationally convenient.
The greatest opportunity for AI in the humanities and other interpretive disciplines is not automated interpretation. It is computational leverage over human judgment without the surrender of human authority.
That is the condition under which these fields can receive the benefits of AI ethically, intellectually, and on terms consistent with their own methods.
Photo Credit: SLAC
Leave a Reply