This page covers the data layer that turns blood, tissue, cells, genomes, proteins, and clinical records into machine-readable biology. The stack includes sequencers, flow cells, reagents, single-cell and spatial assays, molecular diagnostics, long-read systems, laboratory workflows, cloud analysis, de-identified clinical-genomic datasets, and software that ties assays to patient or experimental context. AI biology uses this layer before a model can learn useful biology: samples must be prepared, measured, quality-controlled, linked to metadata, and reused in discovery, clinical, or reimbursement decisions. The current basket read is that ILMN is the best source-backed anchor because sequencing scale, NovaSeq X utilization, consumables, and free cash flow are already visible. TEM, NTRA, TXG, and GH have clearer data-platform or clinical-genomic routes, but each still needs proof that volume becomes margin, cash flow, or paid data economics. QGEN is a molecular-workflow monitor, and PACB remains a watch name until long-read adoption can outrun financing risk.
What the stack is: instruments, assays, consumables, diagnostic tests, software, and rights-cleared datasets that convert biological samples into structured genomics, transcriptomics, proteomics, spatial, single-cell, and clinical-genomic data.
What it does: the stack prepares samples, runs sequencing or molecular measurements, controls quality, attaches metadata and consent, stores the output, and gives researchers, clinicians, pharma teams, or AI workflows a dataset that can be compared across samples and time.
Where it sits: inside research labs, clinical labs, oncology practices, diagnostic networks, pharma translational groups, cloud analysis environments, and commercial data platforms that connect test results to drug-discovery or care decisions.
How the theme uses it: AI models need measured biology with known provenance. The parent theme uses this node when sequencing, single-cell, spatial, proteomic, diagnostic, or clinical-outcome data can train models, validate targets, select cohorts, monitor response, or support pharma and clinical decisions.
Terms used later: omics means large-scale biological measurement such as genomics, transcriptomics, proteomics, or spatial biology; consumables are recurring flow cells, reagents, kits, cartridges, or test materials; ASP means average selling price per test or product; pull-through means follow-on consumables or tests after an installed instrument or customer relationship; data monetization means paid use of datasets, analytics, trial matching, or software rather than only test volume.
Report boundary: this node is a tactical report layer. It ranks the current value-chain basket, setup labels, confirmation triggers, invalidation levels, and chart provenance. Durable research, claim IDs, raw source registries, and sector thesis maintenance stay in the linked knowledge pages. Price and volume context comes from read-only discovery daily_ohlc through 2026-06-12.
Current Setup
Measurement substrateBiology data matters when measurement turns into repeat use, reimbursement, or paid data products.
The strongest names show recurring consumables, diagnostic volume, data-platform revenue, stable ASP, or cash conversion rather than only AI language around assays and datasets.
Positive proofILMN has the cleanest scale-and-cash evidence.
TEM, NTRA, TXG, and GH carry more direct data optionality, but conversion proof is still company-specific.
Conversion gateVolume must become margin, cash, or paid datasets.
Watch consumables pull-through, payer collections, Data and Applications growth, Atera adoption, and Shield economics.
Primary constraintFunding, reimbursement, and dilution can absorb the growth.
Diagnostics ASP, customer budgets, cash burn, convertibles, China, and guide credibility remain live checks.
InputSamples + metadata
MeasurementSequencing, omics, tests
ReuseClinical + pharma datasets
ProofRevenue, margin, FCF
AI biology depends on measured biology that can be reused, audited, and tied to decisions. A model benchmark matters less than the dataset behind it: sample provenance, consent and privacy controls, assay reproducibility, linked clinical outcomes, external validation, and repeat customer use decide whether biology data becomes useful evidence. This node is the measurement substrate for the broader theme. The stronger proof points are instrument utilization, consumables pull-through, paid data or software revenue, reimbursed diagnostic volume, stable ASP, margin quality, cash conversion, and partner renewals.
The useful demand sources are pharma discovery groups, academic and translational labs, oncology providers, screening programs, diagnostic labs, and biopharma data customers. ILMN has scaled sequencing and Q1 free cash flow, plus NovaSeq X and BioInsight/Billion Cell Atlas evidence. TEM has the most explicit Data and Applications revenue line. NTRA and GH have fast clinical-genomic testing growth that can create outcome-linked datasets if reimbursement and collections hold. TXG has direct single-cell and spatial biology fit, with Spatial consumables growth and Atera as the next product-cycle proof. The next positive proof is repeat utilization, disclosed data-product economics, stable diagnostic ASP, paid pharma use, and cash flow that improves without masking dilution.
The setup can fail even if AI interest stays high. Diagnostics revenue is not automatically AI revenue, and high test volume can still disappoint if payer mix, ASP, denial rates, receivables, or lab costs weaken. Single-cell, spatial, sequencing, and long-read tools can sell instruments without enough consumables pull-through. TEM still has operating cash burn and bundled AI economics; TXG has not disclosed Atera backlog, ASP, capacity, cancellations, or pull-through; GH still guides to negative 2026 free cash flow; QGEN is in a guidance reset; PACB has subscale revenue and a large convertible-debt overhang. Watch Q2 results, payer coverage, customer funding, and weekly trend breaks before treating data scale as durable equity value.
Static setup labels were built from weekly bars aggregated from discovery daily_ohlc through 2026-06-05. Local daily_ohlc now runs through 2026-06-12, so the old setup thresholds are stale and should be refreshed before use as current trading evidence. TEM has public daily bars only from 2024-06-14, so its API chart cannot show a full three-year public trading window.
Basket
This basket is inherited from the parent theme, then reranked by a 10-agent primary research tournament and one merge agent. Ranking uses source-backed node exposure first, economic conversion second, and technical timing last. Core names have the clearest evidence that biology data can convert into recurring revenue, data products, reimbursed testing, or cash flow. Option names have real exposure but weaker AI-specific conversion evidence. PACB is kept as a watch name because long-read sequencing matters scientifically, while financing and burn still dominate the equity read.
Sequencing scale, NovaSeq X utilization, recurring consumables, and free cash flow make ILMN the clearest data-generation anchor.
Market cap$21.5B
Next earningsLate Jul/Aug est.
Latest qtr revenue$1.091B
Role in stack
Research, clinical, pharma, hospital, government, academic, and diagnostic-lab customers use Illumina sequencers, flow cells, consumables, services, and analysis workflows to generate genomic data. The conversion gate is NovaSeq X utilization, recurring consumables, SomaLogic multiomics integration, and whether BioInsight or Billion Cell Atlas evidence becomes paid or decision-relevant use.
Revenue mix
Q1 2026 revenue was led by consumables at $797M, about 73% of revenue; instruments were $120M and service and other revenue was $174M. That mix makes recurring use more important than one-time system placements.
Latest qtr revenue
Q1 2026 revenue was $1.091B, up 4.8% year over year, from Illumina's Q1 2026 release and linked ILMN lane. The next check is whether Q2/Q3 placements, consumables, China, margin bridge, SomaLogic, and cash flow confirm the Q1 recovery.
Diagnostics volume and a disclosed Data and Applications segment give TEM the clearest paid-data route, with cash-burn proof still open.
Market cap$9.0B
Next earningsAug 2026 est.
Latest qtr revenue$348.1M
Role in stack
Providers generate oncology and hereditary diagnostic data, while biopharma and provider customers buy de-identified data, modeling, trial matching, analytics, and workflow products. Conversion depends on Diagnostics funding the data engine while Data and Applications revenue, renewals, margin, and cash conversion become visible.
Revenue mix
Q1 2026 revenue was $261.1M Diagnostics, about 75% of total revenue, and $87.0M Data and Applications, about 25%. The data segment is the key node evidence, but standalone AI economics remain bundled inside broader disclosures.
Latest qtr revenue
Q1 2026 revenue was $348.1M, up 36.1% year over year, from Tempus AI's Q1 2026 release and linked TEM lane. The same period still had a $125.9M net loss and negative $73.3M operating cash flow.
High-volume cfDNA testing gives NTRA scaled clinical-genomic evidence, with reimbursement and cash quality as the main filters.
Market cap$27.7B
Next earningsAug 2026 est.
Latest qtr revenue$696.6M
Role in stack
Oncology, women's health, and organ health customers use cfDNA tests that create large clinical-genomic datasets. Theme economics come through test volume, payer coverage, realized ASP, collections, lab scale, and evidence generation rather than a separately disclosed AI software product.
Revenue mix
Revenue is mostly product and test revenue in one reportable segment. Q1 2026 included more than 1.0M tests processed and 258,900 oncology tests, making Signatera/MRD volume and reimbursement quality central to the node read.
Latest qtr revenue
Q1 2026 revenue was $696.6M, including $693.9M of product revenue, from Natera's Q1 2026 release and linked NTRA lane. The next check is whether oncology volume, ASP, gross margin, receivables, and operating cash flow support the raised FY2026 guide.
Single-cell and spatial assays are direct AI-biology data tools, but Atera and consumables pull-through still need commercial proof.
Market cap$2.8B
Next earningsNot confirmed
Latest qtr revenue$150.8M
Role in stack
Academic, translational, pharma, and AI-biology customers use Chromium, Xenium, Visium, cloud analysis, and future Atera workflows to generate single-cell and spatial datasets. Conversion depends on Atera shipment quality, installed-base growth, consumables utilization, and whether spatial data demand produces recurring pull-through.
Revenue mix
Q1 2026 revenue was $129.8M consumables, $11.3M instruments, $8.8M services, and $0.9M license and royalty revenue. Consumables are the recurring line; Atera backlog, ASP, capacity, cancellations, and follow-on pull-through remain undisclosed.
Latest qtr revenue
Q1 2026 revenue was $150.8M from 10x Genomics' Q1 2026 release and linked TXG lane. Products and services revenue grew 9% excluding prior-year settlement revenue, while full-year guidance stayed $600M-$625M.
Clinical-genomic testing and Shield screening create real dataset exposure, but payer mix and free-cash-flow burn decide conversion.
Market cap$12.7B
Next earningsAug 2026 est.
Latest qtr revenue$301.7M
Role in stack
Oncology providers, screening customers, and biopharma partners use Guardant tests, companion diagnostics, Shield screening, and data services. Data conversion depends on Shield volume, payer mix, realized ASP, Screening gross margin, Reveal reimbursement, Oncology volume, and lower cash burn.
Revenue mix
Q1 2026 revenue was $205.0M Oncology, $53.0M Biopharma and Data, $41.6M Screening, and $2.1M Licensing and Other. Biopharma and Data is the clearest paid-data line, while Shield is the main screening-volume proof point.
Latest qtr revenue
Q1 2026 revenue was $301.7M, up 48% year over year, from Guardant Health's Q1 2026 release and linked GH lane. The same guide still called for negative FY2026 free cash flow of $185M-$195M.
Molecular workflow and sample-prep exposure make QGEN useful, but the current guide reset keeps it below higher-purity data names.
Market cap$6.8B
Next earningsAug 2026 est.
Latest qtr revenue$492M
Role in stack
Life-science and diagnostics customers buy sample prep, assays, QuantiFERON, QIAstat-Dx, QIAcuity digital PCR, QIAGEN Digital Insights, and Parse-linked single-cell workflows. Conversion depends on recurring consumables, product launches, Parse contribution, QuantiFERON stabilization, and guide credibility.
Revenue mix
FY2025 consumables and related revenue were 90% of sales. Q1 2026 sales spanned Sample Technologies, Diagnostic Solutions, PCR/Nucleic Acid Amplification, Genomics/NGS, and Other, with one-segment reporting limiting product-level economics.
Latest qtr revenue
Q1 2026 sales were $492M from QIAGEN's Q1 2026 release and linked QGEN lane. Sales declined 1% CER, QuantiFERON declined 5% CER, and FY2026 sales guidance was cut to 1%-2% CER growth.
HiFi long-read sequencing is scientifically important, but PACB remains a watch name until pull-through and runway improve.
Market cap$0.4B
Next earningsAug 2026 est.
Latest qtr revenue$37.2M
Role in stack
Research and applied-genomics customers buy Revio and Vega systems plus HiFi long-read consumables to generate structural-variant, rare-disease, metagenomic, and large-atlas datasets. The scientific role is direct; the equity gate is whether placements and pull-through improve before debt and cash burn dominate.
Revenue mix
Q1 2026 revenue was $21.8M consumables, $9.7M instruments, and $5.6M service and other revenue. The recurring consumables line needs stronger Revio and Vega pull-through before long-read relevance becomes durable equity evidence.
Latest qtr revenue
Q1 2026 revenue was $37.2M from PacBio's Q1 2026 release and linked PACB lane. Revenue was flat year over year, with 15 Revio and 27 Vega placements and 2026 revenue guidance of $165M-$175M.
Market caps use read-only discovery instruments.market_cap values queried on 2026-06-13. Exact next-earnings dates were not confirmed in bounded searches of official investor-event pages and financial-calendar sources; broad month estimates come from linked security lanes that state normal quarterly cadence. Latest-quarter revenue metrics use each company's Q1 2026 release or 10-Q routed through the linked security lane.
What Confirms Or Weakens
AreaWhat confirmsWhat weakens or invalidatesWatch next
01Node thesis
Measured biology becomes reusable
What confirms
Biology data shows repeat use through NovaSeq X utilization, single-cell and spatial consumables, reimbursed test volumes, data-license renewals, pharma data use, and clinical evidence tied to decisions.
What weakens or invalidates
Companies report platform language without paid usage, utilization, renewals, ASP, gross margin, cash-flow improvement, or externally validated dataset use.
Watch next
Q2 2026 results
Data-product disclosures
Atera launch evidence
Diagnostics ASP commentary
02Economics and conversion
Revenue becomes cash quality
What confirms
Data and Applications, Biopharma and Data, consumables, or diagnostic-test growth converts into margin, free cash flow, stable collections, and lower dilution risk.
What weakens or invalidates
Revenue grows while receivables, cash burn, SBC, convertibles, or operating losses worsen; diagnostics growth remains unlinked to data monetization.
Watch next
TEM Data and Applications
GH FCF burn
NTRA receivables
ILMN/TXG consumables
PACB cash runway
03Customer and funding
Budgets support usage
What confirms
Pharma, diagnostic labs, academic labs, translational customers, and oncology providers keep buying instruments, consumables, assays, data products, and workflow tools despite NIH and biotech funding pressure.
What weakens or invalidates
Academic budgets, smaller biotech funding, customer capex, or payer coverage pressure slows orders, testing volume, utilization, renewals, or partner usage.
Watch next
Health Care sector lane
Order commentary
Research customer funding
Pharma data contracts
04Reimbursement and clinical data
Tests get paid and reused
What confirms
NTRA and GH sustain oncology or screening volume with stable ASP, payer coverage, gross margin, and collections; TEM keeps Diagnostics coverage stable while Data and Applications outgrows Diagnostics.
05Operating and supply constraint
Launches scale into recurring use
What confirms
Atera shipments, NovaSeq X placements, Revio/Vega pull-through, QIAGEN product launches, and lab capacity scale without gross-margin leakage or order cancellations.
What weakens or invalidates
Instrument placements fail to generate consumables, Atera backlog or ASP remains undisclosed, lab costs rise faster than volume, or long-read orders stay below the cash-runway need.
Watch next
NovaSeq X utilization
Atera capacity and ASP
Revio pull-through
QuantiFERON and Parse
06Stale condition
Refresh trigger
What confirms
Discovery daily_ohlc remains current through 2026-06-12 and no material Q2 result, product launch, reimbursement update, financing, guidance change, or partner disclosure has arrived since the cited lanes.
What weakens or invalidates
A new trading session, earnings release, Atera update, payer decision, data-license disclosure, financing announcement, credit event, guide change, or new official earnings date arrives before refresh.
Watch next
Discovery coverage
Source lanes
Setup levels
Earnings calendar
Source Trail
Canonical Thesis
AI Biology Platform Validation for the validation ladder from platform claims to paid adoption, clinical data, reimbursed diagnostics, recurring consumables, margin quality, and cash conversion.
Bounded source checks used official company Q1 2026 releases and investor materials from Illumina, 10x Genomics, Tempus AI, Guardant Health, Natera, QIAGEN, and PacBio to verify current results, product launches, guidance, and data-platform claims before ranking.
Earnings-date checks on 2026-06-13 searched official investor-event pages for Illumina, Tempus AI, Natera, 10x Genomics, Guardant Health, QIAGEN, and PacBio, plus public financial-calendar pages. Exact next-earnings dates were not confirmed; the cards label broad month estimates from the linked security lanes or Not confirmed where no credible date was available.
Earnings Date And Revenue Sources
Next earnings labels: ILMN late July or August 2026 estimate from the ILMN lane; TEM August 2026 estimate from the TEM lane; NTRA August 2026 estimate from the NTRA lane; TXG not confirmed; GH August 2026 estimate from the GH lane; QGEN August 2026 estimate from the QGEN lane; PACB August 2026 estimate from the PACB lane.
Latest-quarter revenue sources: ILMN Q1 2026 revenue $1.091B from Illumina's Q1 2026 release and ILMN lane; TEM Q1 2026 revenue $348.1M from Tempus AI's Q1 2026 release and TEM lane; NTRA Q1 2026 revenue $696.6M from Natera's Q1 2026 release and NTRA lane; TXG Q1 2026 revenue $150.8M from 10x Genomics' Q1 2026 release and TXG lane; GH Q1 2026 revenue $301.7M from Guardant Health's Q1 2026 release and GH lane; QGEN Q1 2026 sales $492M from QIAGEN's Q1 2026 release and QGEN lane; PACB Q1 2026 revenue $37.2M from PacBio's Q1 2026 release and PACB lane.
Market-cap metrics use read-only discovery instruments.market_cap queried on 2026-06-13: ILMN $21.5B, TEM $9.0B, NTRA $27.7B, TXG $2.8B, GH $12.7B, QGEN $6.8B, and PACB $0.4B.
Discovery And Chart Provenance
The chart API route is /api/securities/{ticker}/chart?frequency=weekly&window=3y&as_of=latest. This page opts into the selected-security right rail with data-report-api="required".
Setup labels use read-only DuckDB table daily_ohlc from ../discovery/data/discovery.duckdb. Weekly bars use first open, maximum high, minimum low, final close, and summed volume by calendar week.
Indicators are 20-week and 100-week EMAs computed from full available weekly close history before visible-window clipping. Volume compares latest weekly volume with the prior 20-week average.
All seven tickers had local daily_ohlc coverage through 2026-06-12. ILMN, TXG, NTRA, GH, QGEN, and PACB had more than three years of daily bars; TEM had daily bars from 2024-06-14 through 2026-06-12.
Discovery checks used read-only DuckDB queries against instruments and daily_ohlc, plus spot discovery status and lineage-status checks for ILMN, TEM, and PACB. The local report API at 127.0.0.1:8765 was not running during validation, so representative selected-security API validation could not connect.
Static setup labels, triggers, and invalidation thresholds in the right rail metadata remain from the 2026-06-05 setup package. Because local OHLC is now newer, refresh those setup thresholds before using them as current trading evidence.
Shared-File Deferrals
No themes/ai-enabled-biology-and-scientific-discovery/nodes/node-data.json manifest is present for this theme family.
Shared files were intentionally left unchanged in this review scope. sitemap.json still carries the older 2026-06-08 / 2026-06-05 metadata for this node and should be refreshed by the serialized coordinator pass if aggregate metadata consistency is required.
Known Gaps
TEM needs more granular paid data-license, renewal, and standalone AI/data-product economics; current disclosures bundle parts of the AI story into broader segments.
TXG needs Atera orders, backlog, ASP, shipment capacity, cancellations, and consumables pull-through before the spatial product cycle can be treated as proven.
NTRA and GH need product-line economics, payer coverage, ASP, receivables, and cash-conversion evidence before diagnostics volume can be treated as durable data economics.
QGEN needs proof that the 2026 guide reset is contained and that Parse, QDI, QIAcuity, and QuantiFERON can support recurring molecular workflow growth.
PACB needs financing, credit, and recurring pull-through evidence before long-read scientific relevance can outweigh common-equity risk.