Biohub's open biology data push grows to $1.8 Billion.
Federal agencies, Google DeepMind, Meta and Isomorphic Labs joined a nonprofit backed by Priscilla Chan and Mark Zuckerberg. Paying companies see the data first.
By The Editors · The Game of Giving
On October 7, Biohub, a 501(c)(3) nonprofit research institute, announced that the U.S. Department of Energy, the National Institutes of Health and three companies were joining its effort to produce open biological data for training artificial intelligence models. Together, the organizations are investing $1.8 billion in funding, data, computation and new measurement technology, which Biohub calls the largest coordinated commitment to generating AI-ready biological data to date.
Reuters describes Biohub as a philanthropic venture of Meta chief executive Mark Zuckerberg and his wife, Dr. Priscilla Chan. The effort, called the Virtual Biology Initiative, was announced in April and anchored by a founding $500 million commitment from Biohub itself.
For anyone who funds research, the deal is worth reading closely. It shows how a philanthropy can use its own money to bring in government and corporate partners, and what those partners may get in return.
Who is putting in what
- Biohub. The founding $500 million. Of that, $400 million supports new technologies that expand what biologists can measure, and a further $100 million funds research outside Biohub.
- The Department of Energy. More than $500 million over five years in lab measurement, modeling and computation, through its Genesis Mission.
- The National Institutes of Health. Coordination of datasets, repositories and knowledge bases built with more than $500 million in earlier federal investment. Biohub will work with NIH to standardize them for AI model training.
- Google DeepMind, Isomorphic Labs and Meta. $300 million, invested collectively.
By our arithmetic, those four lines add up to the $1.8 billion total, although the NIH figure describes earlier federal spending on existing data rather than new money. The announcement also names the Allen Institute, Broad Institute, Gladstone Institutes, the Human Cell Atlas, the Human Protein Atlas and the Wellcome Sanger Institute as groups working toward the same goal. It says NVIDIA will support the initiative with computing infrastructure and that Renaissance Philanthropy is helping to expand funding for data generation.
What the money is meant to build
The aim is data detailed enough to train models that predict how cells behave, so that some experiments can be run on a computer. "An accurate predictive model of biology could dramatically accelerate scientific discovery by enabling scientists to perform experiments digitally," Alex Rives, Biohub's head of science, said in the announcement.
Rives told Reuters that current cell datasets run to hundreds of millions of cells, while an accurate predictive model will require billions and eventually trillions. He said the work would normally take decades, that the partners aim to compress it into five years, and that a first dataset should be ready in about a year.
Chan described the data as a shared resource. "We have always held this as a community asset, not just for one group, so that it can build upon itself over time," she told Reuters.
Open, but not all at once
The datasets will eventually be released publicly, but companies that fund them get a head start, Rives told Reuters. "With commercial funders we have embargo periods where there's a period of time where the groups can work on the data, and then it becomes available as a public scientific resource," Rives said.
Rives told Reuters that the government-funded work running in parallel will carry no such restrictions, and that Biohub plans to approach pharmaceutical companies and philanthropies next. Reuters also noted that the OpenAI Foundation has started a grant program of more than $125 million to fund biological and medical datasets for AI research.
Questions for funders weighing an open-data partnership
Few foundations can put $500 million on the table, but the questions this deal raises apply at any size. In our view, a funder joining a shared data effort should get clear answers to four.
- Who sees the data first, and for how long? Biohub gives commercial funders an embargo period. Ask for the length in writing, and whether it differs from one funder to the next.
- Which parts are restricted? Biohub says the government-funded work carries no embargo. Know which datasets your grant pays for and which rules apply to them.
- What does open mean here? The announcement describes the result as an open resource for the research community. Ask which license, which access process and which release date will make that true.
- Who maintains it afterward? Biohub says it is building shared standards, common identifiers and a single point of access. Infrastructure like that has to be paid for after the first grants are spent.
Our pieces on keeping restricted gifts and on how foundations have changed course cover why the terms attached to money matter as much as the amount.
General information for donors and nonprofit leaders, not legal or tax advice. How we report is set out in our editorial guide.