Methods & Provenance
This index is built to be auditable: every figure traces to a source record, and the retrieval and classification are locked and versioned. Two tracks are published — research papers and patents — and they are counted separately, in different units.
Research Papers
The Field Definition
A locked, versioned PubMed query (q2-2026-06-30) retrieves records involving a
biofluid-derived sample and a recognized analyte class (ctDNA/cfDNA, CTCs, EVs/exosomes, cfRNA,
circulating proteins/metabolites), across oncology, prenatal, transplant, and other applications.
Classification
Each record is scope-gated and tagged against a locked rubric (v4.1-2026-08-18) by Claude Haiku at temperature 0, along four independent axes: application, oncology breadth, analyte, and study purpose. The application axis has six categories ordered by the care pathway — screening, diagnosis, prognosis, therapy selection, treatment-response monitoring, and minimal residual disease — and oncology breadth has three values: one named cancer, two or more named cancers studied together, or a pan-cancer / tissue-agnostic approach.
Two rules are enforced outside the model. Every record tagged diagnosis gets a second, focused call that asks only which population the abstract describes; the tag is kept only for diagnostic workup and review articles, and dropped on case-control and other study designs. And the rule that application tags attach only to oncology records is applied in code after the model answers, not left to the prompt. Records the classifier is unsure about are flagged for review rather than silently kept or dropped. Before any bulk run, a candidate prompt must pass two mandatory gates: a per-record calibration gate on a hand-labeled gold set, and a population-level check that compares tag rates on a pinned sample against thresholds — because a prompt can be right record by record and still wrong in aggregate.
Trend lines count review articles and primary research papers together, for every tag. The mix varies by tag, and it runs most review-heavy for diagnosis: under the strict definition, primary studies of diagnostic workup are rare, so much of what carries that tag is review literature discussing the use.
Coverage
Coverage is complete: every record the query returns for 2008–2026 is classified individually — 49,212 records, of which 35,481 are in scope. Counts are direct. The counting unit is one paper, counted once, in its earliest year.
Patents
How Patents Are Retrieved
Patents are not found by keyword. Retrieval runs on a frozen family of CPC classification codes (v1.1, edition 2026.01) applied to the USPTO's bulk patent data, which pulls 46,410 candidate US-granted patents. The code family is deliberately a recall filter — it decides what gets read, not what counts.
The Same Scope Test As The Research Track
Every candidate is then read by the same Claude Haiku classifier at temperature 0, against a patent-side rubric (trackB-3part-v1.4-2026-08-17) that applies the two tests the paper rubric applies: is the analyte circulating or cell-free, and is it measured as a disease biomarker? A patent passes on what it claims, not on the CPC code that surfaced it. Result: 3,157 in-scope US documents. The same calibration gate is mandatory here before any run.
The Counting Unit Is An Invention, Not A Document
One invention is typically filed many times — continuations, divisionals, and parallel filings in other patent offices. Counting documents would therefore count the same idea repeatedly. In-scope documents are collapsed onto their INPADOC patent family, so a continuation chain or a multi-office filing counts once: 3,157 documents become 2,195 inventions (30.5% collapsed). Inventions are dated by earliest priority year — when the idea was first filed, not when it granted.
Who Owns An Invention
A patent can name several owners, and many do — university with hospital, company with the lab whose work it licensed. Every organization named on any document in the family is kept, so 368 of the 2,195 US inventions have more than one owner, and 1,203 distinct organizations appear across 2,579 ownership links. A co-owned invention is therefore counted once under each of its owners: an owner ranking adds up to more than the number of inventions, because it measures share of ownership, not a division of the total. 92 inventions are held by individuals rather than an organization.
Owner names come as the patent offices recorded them, and the same organization is sometimes written two ways on two documents of one family — a misspelling, a national subsidiary, or a personal name in the other word order. Where that happens the invention reads as co-owned when it has a single owner. A conservative screen finds this on about 24 US and 27 China inventions; the pairs are listed openly and are resolved by hand against the registry, never merged automatically, because a parent and its subsidiary look identical to any automatic rule and are not the same owner.
The Separate China View
The US-anchored set only contains an invention if it was filed in the United States. Inventions filed only in China are retrieved separately from Google Patents public data and classified with the same rubric: 102,157 patent families read, 8,689 in scope. This is published as its own view rather than added to the US number, because the two are not directly comparable — the China set includes pending applications as well as grants, its abstracts are machine-translated, and it carries no citation data, so it has no impact ranking.
Applications That Have Not Granted Yet
A patent takes years to grant, so the granted-only view runs out before the present: the recent priority years are not small, they are unfinished. US pre-grant publications — applications the USPTO publishes ~18 months after filing, whether or not they ever grant — fill that gap. 15,437 still-pending applications were read by the same classifier against the same rubric, giving 1,435 in-scope documents and 1,366 pending inventions once collapsed onto patent families. Applications that already became one of the granted patents above are removed by an exact USPTO crosswalk, so nothing is counted twice, and pending is dated by the same earliest-priority rule. Pending is always shown as its own band, never merged into the granted number.
Two filters were applied before spending anything on classification. Applications filed before 2020 that never granted are abandoned, not pending — grant rate by filing cohort stops moving after about 2018 — so only 2020-and-later filings are treated as live. And a set defined by absence from a grant list inherits every other reason a record could be absent, so the absence was checked directly rather than assumed.
Patent counts look small next to a keyword search on a patent site, and are meant to. They count in-scope, de-duplicated inventions, not documents matching a search term.
Known Limitations
- Classification is reproducible to ~99.4% on the in-scope call; re-running the same records can shift a small number of borderline items.
- 2026 is a partial year. Conference abstracts are excluded (PubMed-indexed works only).
- Pre-grant data moves the recent-year truncation forward rather than removing it: an application is only published about 18 months after it is filed, so the newest filing years are still incomplete, and the current year is empty by construction. Everything on the patent page except the growth chart — platform, application, owners, impact — still counts granted inventions only.
- Owner rankings are only as good as the names the patent offices recorded. In the China data an owner is named on 49.5% of in-scope inventions, so that ranking describes the half of the set that has one; the US data names an organization on almost all of it. Neither side reconciles a company with its subsidiaries.