Nearly 70% of the samples in the Human Cell Atlas, one of biology's flagship reference maps, carry no record of the donor's ancestry at all. Among the samples that do, people of European ancestry appear roughly six times more often than global population figures would predict.
That is the result of an audit led by researchers at the Icahn School of Medicine at Mount Sinai, published July 20 in Cell Genomics. The team examined more than 13,500 samples across three major consortia and compared each against global population data, US cancer incidence and disease-specific references.
The timing is the point. These atlases are increasingly used as training data for AI foundation models of human biology.
The Same Skew in All Three Resources
The Human Cell Atlas, the Human Tumor Atlas Network, and the PsychAD Consortium all leaned in the same direction.
Tumor samples in the Human Tumor Atlas Network were about 69% European. The PsychAD brain dataset was nearly two-thirds European. Several cancer types also showed differences by sex that went beyond what disease incidence would predict.
The detail that closes the most obvious escape hatch: the European overrepresentation in the Human Cell Atlas held up even under the most conservative assumptions about the missing records, meaning it cannot be explained away as an artifact of incomplete metadata.
"These atlases are becoming the reference maps for biology and medicine, and they are increasingly used to train the AI models that will shape future research and care," said senior corresponding author Kuan-lin Huang, associate professor of genetics and genomic sciences and of artificial intelligence and human health. "We wanted to ask a simple but important question: do these maps represent people from different populations fairly, or are some groups being left out?"
The Mount Sinai announcement notes that in this study, "Latino" refers to ethnicity as reported in available data rather than a single genetic ancestry.
Why This Matters More Now Than Five Years Ago
Genomics has been here before, expensively.
Genome-wide association studies were run overwhelmingly in people of European ancestry. An analysis in Nature Communications found that 67% of polygenic scoring studies over one decade included exclusively European-ancestry participants, and that European-derived scores predicted significantly worse in samples of African ancestry. The National Human Genome Research Institute has since funded a dedicated consortium to fix the resulting portability problem, describing early scores built on mostly European data as simply not effective in diverse populations.
Single-cell consortia risk reproducing that at a moment when the stakes compound. Their outputs are described as universal reference maps and are being used to train models that then get applied to interpret new data from anyone. A skewed reference does not just miss rare variation; it can encode the majority pattern as the working definition of normal.
Huang's own summary of what unsettled him is narrower and more fixable than the ancestry gap itself. "The most surprising finding was how much ancestry information was simply missing from some of these costly studies," he said. "These gaps can be passed into AI models trained on the datasets, often without users of the AI models realizing it."
What the Authors Say Their Own Analysis Cannot Show
The researchers are explicit about the limits.
The analysis relied on demographic information as reported in public datasets rather than direct measurement of genetic ancestry, and broad ancestry categories cannot capture the complexity of identity. Because the work covered three major public consortia, the findings may not describe every single-cell dataset in the field.
Most importantly, this study did not test whether AI models trained on these data actually perform worse for underrepresented groups. It documents the input problem and infers risk downstream. Establishing the harm is the next experiment, and the team says it plans to run that evaluation alongside expanded audits and tracking of whether representation improves over time.
The Part That Is Solvable This Year
The paper is not only a complaint. The team published a field-specific checklist covering how to plan recruitment, record demographic information, balance samples, and report whether AI models perform equally well across ancestry and sex groups.
Notably, the analysis was carried out largely by student researchers, with co-first authors at the University of Oxford, the University of North Carolina at Charlotte and Saint Louis University working with Huang at Mount Sinai. The audit did not require new sequencing, only a systematic look at what the field had already published about itself.
The immediate practical implication is unglamorous. Recruiting more diverse donors is slow, expensive, and genuinely hard. Recording ancestry metadata is neither, and nearly 70% of Human Cell Atlas samples are missing it.
Key Questions Answered
What are single-cell atlases?
Large reference datasets profiling human biology one cell at a time, built by international consortia and intended to serve as universal maps for research, medicine and AI training.
What did the audit find?
Across more than 13,500 samples in three consortia, European ancestry was consistently overrepresented, and nearly 70% of Human Cell Atlas samples had no ancestry recorded at all.
How large is the imbalance?
Among Human Cell Atlas samples with ancestry data, Europeans appeared roughly six times more often than global population figures would predict, a gap that held even under conservative assumptions about the missing records.
Does this mean AI medical tools are inaccurate for some groups?
The study did not test that. It documents bias in the training data and raises the risk, but establishing downstream harm requires separate research the authors plan to conduct.
Why is this compared to genetics research?
Genome-wide association studies had the same European skew, and the polygenic risk scores derived from them transfer poorly to other populations, a documented failure the field is still correcting.
What is the most immediate fix?
Recording ancestry metadata, which nearly 70% of Human Cell Atlas samples lack. That is a reporting problem, separate from the harder challenge of recruiting more diverse donors.