• Skip to main content

Everyday Imaging Evidence

See what an image can show—and what it cannot.

  • Home
  • Start Here
  • How We Research
  • Editorial Policy
  • About
  • Contact

How to Read a Diagnostic-Accuracy Study of a Connected Medical Device

posted on September 7, 2026

A diagnostic-accuracy study measures how often a health app or connected device gets it right — and how often it doesn’t. Before trusting a diagnostic claim from a smartphone sensor, camera-based screening tool, or other connected medical device, check four things: the reference standard used, who was tested, the reported accuracy, and the disclosed limitations.

What Counts as a Reference Standard in a Diagnostic-Accuracy Study?

A reference standard is the trusted, established method against which a new device’s results are measured — the closest thing researchers have to “the truth.” For a skin-lesion app, that’s typically a biopsy read by a pathologist. For an ear exam or hearing device, it might be an in-person exam with a specialist. Without a named reference standard, or with one that’s just another unproven app, a study’s accuracy numbers don’t mean much.

Studies also report two other figures worth learning: sensitivity (how often the device correctly flags a true case) and specificity (how often it correctly clears a healthy person). A device can have a high overall “accuracy” while still missing many real cases or triggering many false alarms — sensitivity and specificity separate those two failure modes so you can judge which one matters more for a given condition.

How Reliable Is the Evidence? An Evidence Ladder

Diagnostic-accuracy evidence comes in tiers, and higher tiers are harder to produce and generally more trustworthy. Our How We Research page covers this approach in more depth; the table below applies it specifically to connected diagnostic devices.

Evidence Tier What It Shows Real-World Example Key Limitation
Case reports or demos The device produced a correct result in a small number of examples A handful of before/after screenshots Cannot show how often the device is wrong
Single study Accuracy measured against a reference standard in one defined group One clinic compares an app’s readings to biopsy results Small samples and single settings limit generalizability
Systematic review Multiple studies gathered and assessed together for quality and bias Buechi et al., BMJ Open, 2017 Can only be as strong as the individual studies it reviews
Regulatory classification Independent government review of submitted data for a specific claim FDA device software function determination Covers the specific claim reviewed, not every marketing claim made

What Did the Best Available Review Actually Find?

A 2017 systematic review in BMJ Open is a useful real-world example of tier-three evidence. The authors searched MEDLINE, Scopus, Web of Science, and Business Source Premier for all published studies evaluating a health app that used a smartphone’s built-in sensors — such as its camera — for diagnosis. They screened 3,296 references and found only 11 qualifying studies, most of which assessed melanoma (skin cancer) screening apps and reported 17 usable two-by-two accuracy tables across 1,048 total subjects (758 with the condition being tested for, 290 healthy volunteers).

The reviewers assessed study quality with the QUADAS-2 tool, a standardized checklist for spotting bias in diagnostic-accuracy research, and reporting quality with the STARD statement, a separate checklist for whether a study disclosed enough detail to be trusted and reproduced. Every one of the 11 included studies came back high risk of bias.

What this establishes: as of the end of 2016, published accuracy evidence for smartphone-sensor health apps was thin, concentrated in one condition category, and methodologically weak even where it existed.

What this does not establish: it says nothing about the accuracy of any specific product available today, since apps and detection algorithms have changed since the search cutoff; it doesn’t cover categories like digital otoscopy or ECG-style sensors, since too few qualifying studies existed to review; and it doesn’t measure performance outside a controlled study setting. Our corrections page is where we note whether new evidence changes the picture.

Is a Connected Diagnostic Device FDA-Regulated?

Separately from accuracy, a product’s FDA status depends on its stated intended use. Under Section 201(h) of the Food, Drug, and Cosmetic Act, a product is generally a medical device if it’s intended to diagnose, treat, cure, mitigate, or prevent a disease or condition. This applies whether the product is a physical instrument or software running on a phone — the FDA calls diagnostic software a “device software function,” which can include what’s known as “Software as a Medical Device.” A product marketed solely for general wellness, with no diagnostic claim attached, is handled differently and may not require the same level of review.

Devices can reach the market through different regulatory pathways — such as 510(k) clearance, De Novo classification, or Premarket Approval — depending on risk level. A product going through one of these pathways for one specific claim isn’t automatically cleared for a different diagnostic claim added later in its marketing.

What Questions Should You Ask Before Trusting a Diagnostic Claim?

  1. Find the reference standard. What was the device checked against, and is that a recognized clinical method?
  2. Check the sample size and population. A small sample, or a group that doesn’t resemble you in age, skin type, or condition severity, limits how far the results apply.
  3. Look for sensitivity and specificity, not just “accuracy.” One combined percentage can hide a device that misses many true cases or raises many false alarms.
  4. Check who reviewed the study’s quality. Was it peer-reviewed, and has an independent systematic review or a regulator assessed its risk of bias?
  5. Read the limitations section. Disclosed weaknesses are a sign of a study you can actually evaluate.
  6. Check the specific FDA status, if any, for the specific claim being made — not just for the device category in general.

A simple decision path: if a device’s marketing says it can “diagnose” or “detect” a condition, look for a peer-reviewed accuracy study on that exact product — not just its device category. If no such study exists, treat the diagnostic claim as unverified, regardless of how confidently it’s marketed.

Why Does Evidence Lag Behind for Connected and Smartphone-Based Devices?

Connected and app-based diagnostic tools are newer than most clinical-grade equipment, so their evidence base is often thinner — the 2017 review’s total of 11 studies illustrates this. That doesn’t make these tools unreliable by default; it means the burden falls on the reader to check whether independent evidence of accuracy exists for the specific product and claim, rather than assuming a smartphone sensor performs like dedicated clinical equipment. Our editorial policy explains how we handle sourcing for claims like these.

When Should You Contact a Clinician Instead of Relying on a Device?

A connected device’s result — positive or negative — is not a diagnosis. If a symptom is severe, worsening, or accompanied by warning signs the device doesn’t ask about, contact a healthcare provider or local emergency services for anything urgent, rather than relying on an app’s result alone.

Frequently Asked Questions

What is a reference standard?

It’s the established, trusted method used to check whether a new test or device got the right answer — such as a biopsy for a skin condition or an in-person clinical exam. A diagnostic accuracy study without a clearly defined reference standard is difficult to evaluate.

What do sensitivity and specificity mean?

Sensitivity is how often a test correctly identifies people who actually have the condition. Specificity is how often it correctly clears people who don’t have it. A device can appear “accurate” overall yet perform poorly on one of these two measures.

What does “high risk of bias” mean in a study?

It means the study had design or reporting weaknesses — such as an unclear reference standard, a small or unrepresentative sample, or missing details — that could make its reported accuracy look better or worse than it actually is. Reviewers use standardized tools, such as QUADAS-2, to score this systematically rather than on impression.

Does FDA clearance mean a device is accurate?

FDA clearance or classification means a specific submission was reviewed for a specific intended use — it isn’t a general guarantee of accuracy for every claim a product’s marketing might make. Checking exactly what a claim a clearance covers is part of reading the evidence carefully.

Can I trust an app’s stated accuracy percentage?

A single accuracy percentage, on its own, is hard to verify without knowing the reference standard, the study population, and whether the figure has been independently reviewed. Looking for the underlying published study — not just the marketed number — is the more reliable approach.

Learn More

New to this topic? Start Here for an overview of how we approach consumer health imaging and diagnostic evidence. For background on our independence from the former CellScope company, see our domain-history notice.

Sources

  • Buechi R, Faes L, Bachmann LM, et al. Evidence assessing the diagnostic performance of medical smartphone apps: a systematic review and exploratory meta-analysis. BMJ Open. 2017;7(12):e018280. PubMed
  • U.S. Food and Drug Administration. How to Determine if Your Product is a Medical Device. Content current as of 09/29/2022.

Medical information disclaimer: This article is for general education and does not diagnose any condition or replace advice from a qualified healthcare provider. Everyday Imaging Evidence is an independent editorial publication and is not affiliated with the former CellScope company. Last reviewed: September 2026.

By Everyday Imaging Evidence Editorial Team

Filed Under: diagnostic device evidence and safety

Reader Interactions

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *