I looked for every health data system that shares features with my Medical Snapshot proposal, something I’ve been working on since 2007. Such programmes medically measure people over time, store what they find, and don’t usually tell the individuals what their own data says. This article summarises what I have learned by looking at 25 existing systems, covering 90 years from the Tuskegee Syphilis Study ↗ (1932) to Our Future Health ↗ (2022). I have tried to indicate my biases and assumptions but the purpose of these notes is to inform my own thoughts about the Snapshot concept. I also came across built-in biases in search engines.
Measurements
I scored each system against 18 features of the Snapshot, which is not intended to be a fair assessment of how good a system is. UK Biobank, for example, scores 7.5 out of 18 and yet is one of the foremost voluntary biobanks in the world, so clearly my review is from one very narrow viewpoint!
This article is about the process of generating the full scoring table, which is as exhaustive as it is boring. I hope it will help in some future experimental study design.
Headline results

Plot 1: Research mechanics (x) vs Snapshot participant-benefit features (y, max 6). Research biobanks cluster at moderate mechanics and about 2 of the 6 participant-benefit features, regardless of country or scale. Above that cluster the plot is empty apart from the Snapshot.

Plot 2: Research mechanics (x) vs Snapshot governance features (y, max 11). The bracket marks eight governance features that no operational system has combined with high mechanical capability.
The R source for both plots, with data embedded goes with the scoring page.
The Western forensic databases (NDNAD ↗, CODIS ↗) score 3/18, the same as the China MPS DNA database ↗. They share the same pattern: biological banking, no individual feedback, clinical care independence, with inversions on population type, state control, and voluntariness. It might seem strange to both Western and Eastern sensibilities to give them the same score, but such sensibilities are not the point of this review.
Searches and biases
The indexed English-language bioethics literature is large and easy to search. My comparison started there with Framingham ↗, UK Biobank ↗, deCODE ↗, and FinnGen ↗. We always expect bias, but this corpus is very evidently focussed on finding the answers they expect to the questions they have asked, and quite lacking in curiosity. I tried some translations of English queries into various languages, but various built-in biases meant that there were very few results because other scientific traditions view this problem very differently. As an example: translating (via both machine and obliging native Chinese speaker) the phrase “biobank observer study no feedback ethics” into Chinese produces a relative paucity of results because there is not much Chinese work which discusses English-framed questions. That is really interesting but completely unhelpful to my enquiry.
Therefore I made subsequent searches in Chinese, French, Spanish, and Russian using native-language framing, which seemed to address the bias inherent in translated English queries. Thus, searching for 健康医疗数据相关研究的伦理审查 (“ethics review of health data research”) reaches the PRC legal and health-policy literature on its own terms, and returns a wealth of discussion. There is another layer of bias here, because it was me who chose the ad hoc methods for “native-language framing” and I have no specialist knowledge of translation or applied linguistics - but the improved results at least suggest I seem to be on the right track.
In a similar vein, the Francophone cohort governance literature is almost invisible in the English-language sources despite being the most developed framework I found anywhere for healthy volunteer ethics. This is the INSERM VolREthics charter ↗ and a two-volume IGAS report on cohort studies ↗. These texts have substantially formed my thinking on the general topic, being the collective wisdom of dozens of specialists over many years on this very topic. The Spanish-language public health literature gave me Gaceta Sanitaria ↗ on big-data health research ethics.
The Russian-language sources I reached were mostly regulatory and compliance-focused, with little of the argument and dissent the other four turned up, which is itself a result.
Source credibility
Not all sources are equal, and some widely-cited ones are unreliable due to very obvious conflicts of interest, especially one widely-cited and seemingly credible Australian source.
Sources are grouped into three tiers. Tier A is government primary documents: legislation, official statistics, institutional reports. Tier B is independent organisations with documented methodology, such as Citizen Lab ↗ and Human Rights Watch (giving limitations on interview-based claims). Tier C is for material adjacent to advocacy, where institutional funding creates structural conflicts of interest.
Evidently Tier C is to be regarded with suspicion.
The special case is the Australian Strategic Policy Institute (ASPI) who produced a 2020 Genomic Surveillance Report which is widely cited in the English literature regarding Chinese DNA collection. ASPI is funded by the Australian Department of Defence, the US State Department, NATO, and weapons manufacturers including BAE Systems, Lockheed Martin, and Raytheon. It is easy to see why ASPI’s other China-focused outputs consistently align with those funders' interests, and the genomic report was part-funded by a US government strategic promotional organisation. Even the Australian government found ASPI had published “op-ed overreach” and “partisan commentary.” Nevertheless some claims made in ASPI reports are supported independently by organisations such as Human Rights Watch, Citizen Lab, or Chinese government sources so I have accepted them. There’s a lot more work to be done to get a fair overview of this. I have tried to make a reasonable first pass.
Discussion of results…
The pattern is consistent across both plots. Research biobanks cluster around mechanics 3–5.5 and two participant-benefit features. Framingham ↗ (1948, USA), UK Biobank ↗ (2006, UK), and the China Kadoorie Biobank ↗ (2004, China) cluster together despite being separated by six decades and with radically different governance contexts. That is an unexpected result so I feel that this narrow viewpoint has been useful despite its obvious problems.
The DoD Serum Repository ↗ scores highest among operational systems (8/18). It predates the Snapshot proposal by 22 years and I conclude it is equivalent in many respects. According to my Snapshot criteria, it fails on voluntariness (compulsory military), state exclusion, benefit-sharing, commercial funding, and public release, but on the whole it contains valuable lessons for Snapshot investigation.
The novelty of the snapshot system is where there are sparse or empty columns. These are column C (retroactive release), D (directionality), C9 (clinically triggered release) and D11 (participant can initiate release).
Generation Scotland ↗ returns basic clinical measurements to participants and, with permission, to their GP, which is why it gets 5.5/18.
More assumptions and biases
Dimensions A and B are fairly neutral towards the point of view of any assessor, including me, because they are about structural facts.
On the other hand my assumptions are really clear in Dimensions D, E, and F. You can see that at least at present I tend to the Western idea of rights-based traditions:
D13 (voluntary consent) is a binary choice to join or not. In contexts where community obligation rather than individual choice governs participation norms (Confucian healthcare settings, some African community consent frameworks), much more nuance is needed. The [C] notation suggests where these considerations may apply, but it’s only a hint that we are out of our depth and specific knowledge is needed.
E14 (benefit-sharing) assumes that the Snapshot’s idea of payment is the ideal, for which there is no evidence but it does support the point of view of the entire review. This view agrees with the French healthy-volunteer tradition set out in the Global Ethics Charter for the Protection of Healthy Volunteers ↗. On the other hand it is regarded as extractive and exploitative in the African tradition, where payment above expense reimbursement may compromise the liberty to volunteer. There is a lot to consider in the general African rejection of traditional WHO style funding, something covered in the appendices in my paper on One Health. Perhaps the Snapshot should reject this idea entirely.
An example of how my assumptions give some pretty weird outcomes is Majengo ↗, which funds around US$1,088 of healthcare per woman per year but I scored it as 0 because in-kind community benefit is a different thing from wages. All of Us ↗ pays US$25 for an in-person visit and so I gave it 1. A reasonable person might decide that Majengo shares more with its participants than All of Us does.
I found a helpful scoping review of benefit-sharing in biobanking ↗ that shows disagreement over
what benefits should count and who should receive them. Fifteen of the 25 systems are marked [?] because I reached my personal limit of what I could check. If
someone wants to help me do this more rigorously then I will gladly accept all assistance!
F17 (governance accountable to participants) embeds English trust law. The concept exists in some form in most legal systems, but the specific model the Snapshot assumes is not universal. To the extent an African system operates under modern law, Eastern African countries may tend to have English trust law somewhat frozen in time from the colonial era.
Other designs emerging from other traditions might replace payment with community benefit-sharing, or replace trustee governance with state oversight backed by strong individual rights provisions.
The outrageous NZ experiment
This review mostly does not consider the ethical context in which systems were deployed, so for example we don’t expect a well-developed idea of informed consent 90 years ago. However one cervical cancer study begun in 1966 at Auckland’s National Women’s Hospital stands out because it is a famous observer-only system in the worst possible way. Poor practice was enabled by a male-centric medical system where unethical and sometimes abusive practices were tolerated. I include it here because it has been the focus of intensive research for some decades based on the observer-only nature of the study. It is a distressingly clear example of how a badly designed observer-only system is disastrous, in this case with at least 14 deaths and at least an additional 35 people contracting cancer, with quite possibly more in the cohort of 948 women. The exposé was titled “An unfortunate experiment” which is how it is widely known. This tragedy was also a feminist moment in NZ, with numerous far-reaching changes being made in the way medical systems are run and how women are treated.
Modern overviews are available from the preserved Ministry of Health record ↗ and Cartwright Inquiry article ↗. Its repeated cytology, histology and long-term follow-up provided a fair degree of clinical rigour, and we do now have some grim knowledge of what happens if precursors to cervical cancer are not treated.
Overall
This review is my best effort to put my Snapshot concept in context. I find it heartening that tens of millions of people have passed through such systems already, because this suggests that the problem I identified is relevant universally. Whether the Snapshot could improve on other systems is unclear but there are now some interesting new ways the idea can develop.