I’m writing1 this from Montreal, where I’m attending a meeting I’m co-organizing2 on AI and research integrity. We’ve been working on this meeting for months, and we are bringing together a truly excellent group of academics, fraud sleuths, startup founders, AI tech company representatives, science philanthropy funders, US science policy experts, and an acclaimed science journalist! The meeting will consist of working sessions where we are going to identify blockers to progress and devise solutions and launch larger-scale (hopefully funded) collaborations. We hope to solve some nascent problems in this space: teams duplicating efforts, struggling to share datasets, insufficient evaluation standards, and a lack of tool interoperability. My objective in bringing these incredible people together is nothing short of aligning efforts to build the future of science I’ve been writing about these last few months.3
Put simply, what I’m trying to midwife into existence is this: an open evidence graph, which is a shared updating public record that connects scientific claims to the evidence that supports or challenges them. This graph is meant to preserve how a result was obtained, the checks it passed and the limits of what it demonstrates. This graph would make scientific evidence transparent - it can aid humans or AIs in assessing what we think is true and why and ultimately guide how we should act in light of that evidence4.
Developing this kind of structured representation at the scale I’m proposing would be prohibitively labor-intensive without recent advances in AI. And still, the most challenging technical problem in developing this evidence graph by far is verification — how do you know that the right theory and evidence are being extracted from papers, that the knowledge is accurately represented and aggregated, and ultimately that the evaluation correctly weights the evidence? At every level we need datasets and evaluation - my recent attempt to demonstrate an eval for error finding in papers is just a tiny first step in that direction.
In what follows I want to tell you a little bit more about why I’m doing this — what futures I’m trying to forestall, the looming question of capabilities enhancement in the age of unaligned AI, and why I think this graph needs building despite the potential risks.
Avoiding feudal science
I don’t want to live in a world in which access to the most useful scientific understanding depends on powerful patrons who control the data, models, compute, and interfaces needed to produce and use it. That is the default world we are headed to. That is a world of feudal science - to gain the best scientific understanding of the world, we will have to pay or (as a scientist) work for a large model provider. A scientist will be like a serf, pledging fealty to a model overlord to get a patch of science to work on.
Some of this isn’t new. Pharmaceutical companies often fail to disclose negative trial results and keep much of their drug-development knowledge proprietary. What’s changed is the scale. AI companies (both foundation model and specialized AI-scientist companies) are already building up their own internal understanding of scientific disciplines. There is a real danger of this spreading across science, particularly as public science funding continues to be withdrawn and the best academics are increasingly lured away to industry. I’m not blaming individual academics who’ve left for greener pastures, but it’s undeniable that academia is already experiencing a brain drain.
Not only is there a danger of the sharp increase in hidden and paywalled knowledge, there is the risk of bias — as funding for research shifts toward private actors, independent research that questions or undermines those actors’ narratives will become harder to sustain. This is Henry Farrell’s concern about social science research on the impact of AI increasingly being conducted by AI companies themselves. Creating and maintaining a true public shared knowledge commons is thus essential to fight the risks of feudal science. It won’t replace independent funding or access to compute and proprietary experiments, but it would give us a shared basis for assessing claims without relying entirely on a model provider.
Preserving cognitive capacity
Another risk I’m increasingly worried about is the loss of human cognitive capacities in the face of both the undermining of training and the offloading of judgment onto AI. As an expert, I find AI incredibly useful in extending my span, doing by myself what previously would have taken an entire team of knowledge workers to accomplish. As a result of this, however, I’m passing less of my knowledge on to more junior employees. And I’m hardly the only one to both notice and contribute to this problem. I have had countless conversations with engineers about the problem of how future senior engineers are going to get made. I haven’t heard a convincing answer to this. This risk is also accelerated by the brain drain from academia. The more our best minds migrate away from the universities, the fewer researchers are available to mentor the next generation. We are eating our seed corn.
Not only are we facing a training crisis, we are facing a crisis of dependence amongst particularly younger users of these tools. Faced with the pressures of advancement, younger students and workers are increasingly using AI as a crutch without having developed the skills from doing things by hand. This, when combined with the lack of opportunities for mentorship, risks eroding human cognitive capability. This would be bad in a world of unbounded growth in AI intelligence - risking an intellectual monoculture, for example. But in a world where there are clear blind spots in what AI can do, this risk is even greater.
I’ve written about thinkers who question whether current AI has the situated understanding needed for reliable judgment. On this view, AI’s lack of embodiment may leave limitations that the existing paradigm can’t resolve. I have some faith that eventually these difficulties will be surmounted. But how long will that take? Imagine a world where we outsource our judgment to an AI that lacks the situated understanding needed to truly replace that judgment (i.e. our current world). If that limitation persists and our cognitive capacities erode, then we will not experience a successful handover to AI. Growing output and apparent competence would conceal a decline in our ability to recognize and correct mistakes5. A related risk arises in institutional decision-making. In another wonderful essay by Henry Farrell about “robot solutionism”, automated complaints and appeals prompt institutions to automate their responses, creating an escalating bureaucratic arms race. Even individually effective tools can produce a collectively harmful system. Cognitive outsourcing and model inscrutability could make those dynamics harder to recognize and correct. I’m very concerned about this.
Doing research currently requires the researcher to learn something, if not from the result, then from the process of doing it. Doing things the old-fashioned way produced all the expertise in the world. We need to preserve this process for the next generation to ensure we aren’t at the peak of human cognitive skills. I believe that creating a public knowledge graph could be a crucial tool in avoiding this future. Junior researchers could trace claims back to the experiments behind them, check the graph’s assessments, and defend proposed corrections to experienced researchers. Maintaining the graph could become part of how we train scientists, provided institutions make time for that work and reward it. The graph itself won’t solve the problem of situated understanding, but it would give people a shared record against which to exercise and develop their judgment.
AI capability risk
Finally, the risk everyone is currently writing and panicking about is the risk of misaligned super-powerful AI. The above two risks arguably would be dwarfed by this one, with all of us turned into grey goo or zoo animals. If AI truly advances science far beyond the ability of any humans to understand or reason about scientific findings, then not only will science cease to be legible to us, we’ll be putting ourselves at risk from a misaligned AI. In this world, the creation of an open evidence graph could be actually harmful if it significantly aids AI in the process of scientific discovery and understanding. Therefore this might be a reason we shouldn’t be building what I’m working on.
Imagine the open evidence graph in the hands of a superintelligent AI. Obviously it (or it plus a malicious human actor) would be even more able to wreak havoc when considering biosafety alone. But it could also potentially harm us through better social science (don’t laugh!). Supercharged social science also could be used to manipulate the structure of society or persuade us in harmful ways. A polished, apparently comprehensive evidence system might make people more willing to defer even when its conclusions are wrong and meant to mislead.
Now I’ll admit that although this is something I’ve thought about over the last year, I have mostly pushed these concerns aside. This is partly because I’ve thought the AI companies and the nonprofit safety ecosystem are full of incredibly thoughtful and intelligent people who are more than able to resolve AI risk. Recent events like the Hugging Face incident have called that into question for me. I’ve also sidelined these concerns in my mind because I’m still unsure what to think about current AI systems6 — are they permanently hamstrung by lack of being-in-the-world? Are they simply a normal technology? Or are we birthing superintelligence, as many employees and leaders of these AI companies would have us believe? The answer to this question matters a lot for which risk I’ve outlined is most severe and I genuinely don’t know the answer.
We must understand, we will understand
If we don’t build an open graph, I fear feudal science and the erosion of our cognitive commons will only accelerate. But if we do build it, we might risk accelerating dangerous AI capabilities. Somehow we have to balance these risks in deciding what to do. As you might guess, I still land on the side of building an open evidence graph. Private actors have strong incentives to build their own internal representations of scientific knowledge. That doesn’t make the additional danger from public release negligible. I’m making a judgment under uncertainty: I think the benefits of giving scientists and the public an independent record they can inspect and challenge justify building it, while evaluating what is safe to release. The other dangers I’ve outlined would persist under many possible AI futures.
There are a number of factors that could change my mind. One is evidence that releasing particular parts of the graph would materially increase dangerous capabilities.7 Another is whether a partly AI-generated open evidence graph increases unwarranted trust in AI and erodes human intellectual independence or the willingness of the scientific community to correct errors. Incorporating human expertise into the process of building and validating these systems is crucial. It’s not all-or-nothing. For example, I’d postpone releasing a part of the graph if testing showed that it substantially helped an AI carry out a dangerous biological task. My goal in designing these systems transparently is to preserve our capacity for validation. A big focus of the conference is precisely on how we can scale our validation.
Finally, in convening an amazing cross-disciplinary group in Montreal, I’m hoping that we can address these problems in a unified manner. I’m hopeful that we can create a better version of science in the age of AI, one in which the scope of human understanding of the universe, not just machine understanding, is significantly improved. We must understand, we will understand.
The title of this post is an allusion to David Hilbert’s famous quote “we must know, we will know,” and, in light of recent advances in AI theorem proving, a possible alternative.
With the Center for Open Science and the Institute for Replication. Sponsored by Coefficient Giving! This is burying a lede a bit, but I do want to also disclose that CG is also funding my research. Thank you CG!
Some of these folks are aware of my objectives, while others may not be. Muahaha rubs hands together.
Guidance could, for example, include areas like what experiment a human or AI should run next, what public policy should be adopted that best achieves societal objectives, or what health interventions a person should undertake.
Many academics are raising the alarm about this. This essay by Harry Rickerby about autonomous lab failures highlights what happens when an AI-driven bad assumption gets embedded in the research process. This wonderful piece by Daniel Litt issues a clarion call for the preservation of mathematical expertise. And Kevin Munger points out that research produces expertise in addition to a paper.
And I think and read about AI a lot. I marvel at people’s ability to be so certain.
As I’ve written about before, it’s unlikely that this knowledge graph will have immediate coverage across all of science; it’s more likely to start as a patchwork. That gives us opportunities to assess later releases, though it won’t make earlier ones reversible.

