This page carries the workings behind the Health Observer Systems comparative review : feature definitions, search and source notes, and the full scoring table. The R source reproduces the plots.
What I Measured and Why
This analysis compares existing systems against a hypothetical system, in a way that is exploratory rather than solid science.
The 18 features in the scoring matrix are derived from the Medical Snapshot’s design. That means the matrix measures how close each system comes to the Snapshot’s architecture rather than some kind of quality metric. This matrix asks a narrow question: which of the Snapshot’s features has each existing system implemented?
How the measuring works
The 18 features fall into six dimensions, ordered logically like this:
What is gathered?"] --> B["B: Information Embargo
Who sees it?"] B --> C["C: Retroactive Release
When does it come back?"] C --> D["D: Directionality
Who holds power?"] D --> E["E: Participation Ethics
What's the deal?"] E --> F["F: Governance
Who enforces the rules?"]
A: Data Collection (4 features). Does the system test broadly across many conditions, test healthy people, test the same people repeatedly over time, and retain biological samples? These are the operating mechanics.
B: Information Embargo (3 features). Does the system withhold results from participants, operate independently from their clinical care, and keep the specific tests conducted unknown to participants?
C: Retroactive Release (3 features). Does a formal pathway exist for releasing stored data when something happens later? Is it triggered by the treating doctor’s diagnosis rather than by researchers? Does the released data primarily benefit the individual?
D: Directionality (3 features). Can the participant or their doctor initiate release? Is state access structurally excluded? Did the participant choose to join?
E: Participation Ethics (2 features). What flows back to the participant? E14 records the benefit-sharing variant: reimbursement of expenses, wage-like payment for time, a data dividend, community benefit, or no compensation at all. The Snapshot pays for time, so E14 scores 1 for wage-like payment, ½ where expenses alone are reimbursed, and 0 for the rest. Does commercial early-access funding support the operation?
F: Governance (3 features). Is eventual public data release mandatory? Do trustees have a legal duty to participants rather than to funders or governments? Is data stored across multiple legal jurisdictions?
A score of 1 means the system shares that feature with the Snapshot. A score of 0 means it doesn’t. A score of −1* means the system does the opposite of what the Snapshot intends: compulsory participation where the Snapshot requires voluntary, or state control where the Snapshot requires exclusion. Inversions count as 0 in the total but are marked separately because they carry information.
What I searched, and the biases I found
The indexed English-language bioethics literature is large and easy to search. The comparison started there: Framingham, UK Biobank, deCODE, FinnGen, the standard landscape.
Search only that literature and you find the questions it has chosen to ask. We expect bias, but in these cases the biases were so significant I needed an additional strategy. Therefore I made subsequent searches in Chinese, French, Spanish, and Russian, using native-language framing rather than translated English queries.
As an example: translating (via both machine and obliging native speaking scientist) the phrase “biobank observer study no feedback ethics” into Chinese produces work responding to English-framed questions. That is interesting but completely unhelpful here.
However searching for 健康医疗数据相关研究的伦理审查 (“ethics review of health data research”) reaches the PRC legal and health-policy literature on its own terms, and returns a wealth of discussion.
In a similar vein, the Francophone cohort governance literature holds something almost invisible in the English-language sources, which is the INSERM VolREthics charter ↗ and a two-volume IGAS report on cohort studies ↗ , the most developed framework I found anywhere for healthy volunteer ethics. The Spanish-language public health literature gave me Gaceta Sanitaria ↗ on big-data health research ethics.
The Russian-language sources I reached were mostly regulatory and compliance-focused, with little of the argument and dissent the other four turned up, which is itself a finding.
Source Credibility
Not all sources are equal, and some widely-cited ones are unreliable.
Sources are grouped into three tiers. Tier A is government primary documents: legislation, official statistics, institutional reports. Tier B is independent organisations with documented methodology, such as Citizen Lab ↗ and Human Rights Watch (with named limitations on interview-based claims). Tier C is advocacy-adjacent material where institutional funding creates structural conflicts of interest.
As a special case, The Australian Strategic Policy Institute (ASPI) is excluded at all tiers, which matters because ASPI’s 2020 genomic surveillance report is widely cited in the English literature regarding Chinese DNA collection. ASPI is funded by the Australian Department of Defence, the US State Department, NATO, and weapons manufacturers including BAE Systems, Lockheed Martin, and Raytheon. ASPI’s China-focused outputs consistently align with those funders’ interests, and the genomic report was part-funded by US government strategic promotional organisation. The Australian government found ASPI had published “op-ed overreach” and “partisan commentary.” I have cited claims made in ASPI reports where they appear to be supported independently by organisations such as Human Rights Watch, Citizen Lab, or Chinese government sources. There’s a lot more work to be done to get a fair overview of this. I have tried to make a reasonable first pass.
The Scoring
Systems are ordered by total score descending, then by year of establishment.
| System | A1 | A2 | A3 | A4 | B5 | B6 | B7 | C8 | C9 | C10 | D11 | D12 | D13 | E14 | E15 | F16 | F17 | F18 | Total |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Medical Snapshot (reference) | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1[wage] | 1 | 1 | 1 | 1 | 18 |
| DoD Serum Repository (USA, 1985) | 0 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 0 | 1 | 0 | 0 | 0 | 0[?] | 0 | 0 | 0 | 0 | 8 |
| EPIC (Europe, 1992) | 1 | 1 | 1 | 1 | 1 | ½ | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0[?] | 0 | 1 | 0 | 0 | 7.5 |
| UK Biobank (2006) | 1 | 1 | 1 | 1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | ½[reimb] | 0 | 1 | 0 | 0 | 7.5 |
| China “Physicals for All” (2013–) | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 0 | 0 | −1*[C] | −1*[C] | −1* | −1*[C] | 0[?] | 0 | 0 | −1* | 0 | 7* |
| China Kadoorie Biobank (2004) | 1 | 1 | 1 | 1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0[?] | 0 | 1 | 0 | 0 | 7 |
| Taizhou (China, 2009) | 1 | 1 | 1 | 1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0[?] | 0 | 1 | 0 | 0 | 7 |
| All of Us (USA, 2018) | 1 | 1 | 1 | 1 | −1* | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 1[wage] | 0 | 1 | 0 | 0 | 7 |
| ALSPAC (UK, 1991) | 1 | 1 | 1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | ½[reimb] | 0 | 1 | 0 | 0 | 6.5 |
| Framingham (USA, 1948) | 1 | 1 | 1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0[?] | 0 | 1 | 0 | 0 | 6 |
| Whitehall I (UK, 1967) | 1 | 1 | 0 | 1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0[?] | 0 | 1 | 0 | 0 | 6 |
| Nurses’ Health Study (USA, 1976) | 1 | 1 | 1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0[?] | 0 | 1 | 0 | 0 | 6 |
| BioBank Japan (2003) | 1 | 0 | 1 | 1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0[?] | 0 | 1 | 0 | 0 | 6 |
| CNHBM (China, 2017) | 1 | 1 | 1 | 1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0[?] | 0 | 0 | 0 | 0 | 6 |
| Guthrie Card (global, 1963) | 0 | 1 | 0 | 1 | ½ | 1 | 0 | 1 | 0 | 1 | 0 | 0 | 0 | 0[none] | 0 | 0 | 0 | 0 | 5.5 |
| Generation Scotland (2006) | 1 | 1 | 0 | 1 | ½ | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0[?] | 0 | 1 | 0 | 0 | 5.5 |
| deCODE (Iceland, 1998) | 1 | 1 | 0 | 1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | −1* | 0[none] | 1 | 0 | 0 | 0 | 5 |
| Estonian Biobank (2001) | 1 | 1 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0[?] | 0 | 1 | 0 | 0 | 5 |
| FinnGen (Finland, 2017) | 1 | 1 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0[?] | 1 | 0 | 0 | 0 | 5 |
| Our Future Health (UK, 2022) | 1 | 1 | 0 | 1 | −1* | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0[?] | 0 | 0 | 0 | 0 | 4 |
| China CNGB / BGI GeneBank (2016) | 0 | 1 | 0 | 1 | 1 | 1 | 0 | 0 | 0 | 0 | 0 | −1* | 0 | 0[?] | 0 | 0 | 0 | 0 | 4 |
| Majengo Cohort (Kenya, 1985) | 0 | 0 | 1 | 1 | ½ | 0 | 0 | 0 | 0 | 0 | 0 | 0 | ½ | 0[comm] | 0 | 0 | 0 | 0 | 3 |
| NDNAD / CODIS (UK/USA, 1995/1998) | 0 | −1* | 0 | 1 | 1 | 1 | 0 | 0 | 0 | 0 | 0 | −1* | −1* | 0[none] | 0 | 0 | 0 | 0 | 3 |
| China MPS DNA database (2003) | 0 | −1* | 0 | 1 | 1 | 1 | 0 | 0 | 0 | 0 | 0 | −1* | −1* | 0[none] | 0 | 0 | −1* | 0 | 3 |
| Tuskegee (USA, 1932) | 0 | −1* | 1 | 0 | −1* | −1* | 1 | 0 | 0 | −1* | −1* | −1* | −1* | 0[comm] | 0 | 0 | −1* | 0 | 2 |
| NZ Unfortunate Experiment (1966) | 0 | −1* | 1 | 0 | −1* | −1* | 0 | 0 | 0 | −1* | −1* | −1* | −1* | 0[none] | 0 | 0 | −1* | 0 | 1 |
Notation: 1 = feature present. 0 = absent. ½ = partial. −1* = inverted (system does the opposite); counts as 0 in total. −1*[C] = inverted but evidence rests on advocacy-adjacent inference only. ASPI excluded at all tiers.
E14 benefit-sharing variants: [wage] wage-like payment for the participant’s time, which is what the Snapshot does. [reimb] expenses reimbursed and nothing further. [div] a data dividend. [comm] community or in-kind benefit. [none] nothing. [?] not yet verified, and scored 0 by default, so fifteen of these are placeholders rather than findings. Individual in-kind inducements are recorded under [comm] for want of a closer variant, which is where Tuskegee’s free meals and burial insurance sit.