A diagnostic-device study can be well-designed and still not tell you what you think it tells you. That happens when the people, settings, or data in the study don’t match the situation the device is actually being used for. Before trusting a claim about how well a connected diagnostic tool “works,” ask who was studied, where, and how those results were checked against a reliable standard. This guide gives you a plain framework for asking those questions.
Why representation gaps matter
A device tested mostly on one age group, one skin tone, one clinical setting, or one disease stage may perform differently for someone outside that group. This is not a flaw unique to any single product. It is a known pattern in health technology research generally, and it is one reason regulators now ask device makers to be transparent about who was included in the studies behind a device’s claims.
Representation gaps do not mean a device is unsafe or useless. They mean the evidence has a boundary. Knowing where that boundary sits is what lets you judge whether a study result is likely to apply to your situation.
Terms to know
- Study population: the specific group of people included in a study — their ages, health status, and other traits.
- Reference standard: the trusted method used to check whether the device’s result was actually correct (for example, a lab test or a specialist’s diagnosis).
- Sensitivity: how often the device correctly flags people who truly have the condition being tested for.
- Specificity: how often the device correctly clears people who truly do not have the condition.
- Risk of bias: a formal judgment about whether flaws in how a study was designed or run could have skewed its results.
- External validity: whether a study’s findings are likely to hold up in a different group of people or a different setting than the one studied.
Study population versus real-world use: what to compare
When you read about a connected diagnostic device, line up what the study actually did against what you are actually asking the device to do. A few comparisons matter most:
- Who was tested: Were participants a mix of ages, skin tones, and health backgrounds, or a narrow group?
- Who took the images or readings: Were they collected by trained clinicians in a controlled setting, or by ordinary users at home, the way the device is meant to be used day to day?
- What the device was compared against: Was every result checked against a reliable reference standard, or only some of them?
- How many people had the condition versus did not: A study built mostly around people already known to have a condition can make a device look more accurate than it will seem in a general population, where the condition is less common.
- How many studies back the claim: One small study is weaker evidence than several independent studies that agree.
A field guide: questions to ask of any diagnostic-device study
- Who is in the study, and who is left out? Look for a description of participants’ ages, sex, skin tone, or other relevant traits. If this is missing, treat that as a gap, not a neutral omission.
- Was the result checked against a reliable reference standard? A device’s own output being compared to itself, or to an unverified impression, is weaker evidence than a lab-confirmed diagnosis or specialist review.
- Were images or readings collected the way the device is actually used? A device tested only with expert-captured, high-quality images may not perform the same way when a first-time user captures the input at home.
- How many people, and how many had the condition? Small samples, or samples weighted heavily toward people who already have the condition, limit how far a result can be generalized.
- What do sensitivity and specificity actually mean here? A high sensitivity number sounds reassuring, but check what it would mean for someone who gets a wrong result — missed cases, false alarms, or both.
- Was risk of bias formally assessed? Some systematic reviews use a structured tool to rate each study’s risk of bias. If most or all included studies carry a high risk of bias, treat the pooled result as an early signal, not a settled fact.
- Does the device maker disclose known limitations? Under current FDA, Health Canada, and UK MHRA guiding principles for machine learning-enabled medical devices, disclosing known gaps in who was studied, and other performance limitations, is considered good transparency practice.
What the evidence shows about this pattern
A 2017 systematic review and exploratory meta-analysis in BMJ Open screened 3,296 references and identified 11 studies, most evaluating melanoma-screening smartphone apps, that together reported 17 two-by-two accuracy tables. Formal quality assessment found a high risk of bias across all of the included studies. The pooled studies covered 1,048 people in total, most of whom already had the condition being tested for, with only 290 healthy volunteers included as a comparison group. Even so, the pooled estimates looked favorable on the surface, with summary sensitivity around 0.82 and summary specificity around 0.89.
This is a useful real-world example of the gap between a headline accuracy number and the population and design behind it. A result built on a small, heavily condition-weighted sample, with high risk of bias throughout, is a different kind of evidence than a large, independently repeated study across a broad and representative population — even when the two produce similar-looking percentages.
What regulators say transparency should cover
Guiding principles jointly developed by the FDA, Health Canada, and the UK’s MHRA call for machine learning-enabled medical devices to clearly communicate information that could affect patient risks and outcomes to the people who use or rely on them. Among the categories of information these principles identify as good practice to disclose are known biases or failure modes, confidence intervals for reported results, and known gaps in the data used to develop or test the device — including whether certain groups of people were not well represented and may therefore be at higher risk of a biased result.
These principles apply directly to the software and algorithms inside many connected diagnostic tools. If that information is not disclosed anywhere you can find it, that absence is itself worth noting when you weigh a claim.
Limits of this guide
This article explains how to evaluate representation and bias in a study, following the sourcing standards described in our Editorial Policy. It does not evaluate any specific product, and it does not tell you whether a particular device is accurate enough for your situation. It also cannot replace a conversation with a qualified healthcare professional about your own health or a specific diagnostic result. If you are facing an urgent health concern, contact your local emergency services or a healthcare provider directly rather than relying on a device or an article to guide that decision.
Where to go next
If you are new to reading device and health-imaging evidence, Start Here lays out the basic ideas this site uses across its guides. For a closer look at how we select and evaluate sources like the ones used in this article, see How We Research.
Medical information disclaimer
This article is for general educational purposes only and is not medical advice. It does not diagnose, treat, or recommend any product, device, or course of action. Connected Diagnostics Evidence is an independent editorial publication and is not affiliated with, and does not continue the products, research, or services of, the former CellScope company. Always talk with a qualified healthcare professional about a specific health concern or diagnostic result.
By Connected Diagnostics Evidence Editorial Team. Updated September 9, 2026.
[…] Bias and Representation in Connected Diagnostic Evidence: Questions to Ask of a Study […]