Preprints surface days after a discovery. Years before the patents, the funding rounds, the analyst coverage. Finch Signal tracks that signal across 580+ self-maintaining research topics — each scored on size, momentum, and investment potential, updated weekly.
Scientific papers appear within days of a discovery — years before patents, funding rounds, and analyst coverage. Finch Signal captures that signal systematically — updated weekly.
Every major technology wave showed up first as an early signal at the research stage — a surge in activity years before the patents, the funding, and the analyst coverage caught up.
Each topic's research intensity was already elevated in Finch's data years before the market moment. Lead times are indicative, drawn from weekly research-momentum signals since 2016.
Preprint velocity in GLP-1 receptor agonists surged 180% before any major pharma coverage. The signal was unambiguous two years before Ozempic became a household name.
Research output at the intersection of deep learning and molecular biology began compounding well before the sector attracted major venture capital attention.
Efficiency research publications in perovskite photovoltaics doubled over two years, well ahead of the commercial investment wave that followed into the sector.
Publication velocity in autonomous agents and reasoning sat at the highest tier from early 2020 — years before “agentic AI” became the defining enterprise and venture theme of 2024–25.
Electrolysis and green-hydrogen publication momentum re-accelerated through 2021–22, ahead of the US Inflation Reduction Act's hydrogen production credits and the global electrolyzer capex wave that followed.
Structured CSV and Parquet, updated weekly. A single dataset spans 580+ self-maintaining research topics — each with two independent readings, scale (how big) and momentum (how fast it's accelerating), plus an investment-potential score. Weekly readings run back to 2016, and every topic keeps its complete rise-and-fall history.
| topic | field | investment_potential | scale | momentum | momentum_direction |
|---|---|---|---|---|---|
| Reinforcement Learning in Robotics | Computer Science | 8 | 100 | 209 | rising |
| Software-Defined Networks & 5G | Computer Science | 8 | 69 | 194 | rising |
| Gaussian Processes & Bayesian Inference | Computer Science | 7 | 73 | 193 | rising |
| Mobile Crowdsensing & Crowdsourcing | Computer Science | 7 | 58 | 191 | rising |
| Cryptographic Implementations & Security | Computer Science | 8 | 40 | 190 | rising |
| ECG Monitoring & Analysis | Medicine | 8 | 52 | 185 | rising |
| Advanced Graph Neural Networks | Computer Science | 8 | 100 | 184 | rising |
| Adversarial Robustness in ML | Computer Science | 8 | 100 | 182 | rising |
| topic | scale (size) | momentum (accel) | read |
|---|---|---|---|
| Advanced Battery Materials | 100 | 95 | big · steady |
| Reinforcement Learning in Robotics | 100 | 209 | big · surging |
| Cryptographic Implementations | 40 | 190 | emerging · surging |
| Natural Language Processing | 97 | 147 | huge · accelerating |
| Mobile Crowdsensing | 58 | 191 | emerging · surging |
| Topological Materials | 88 | 117 | large · rising |
| topic | entered | field | investment_potential |
|---|---|---|---|
| Machine Learning & Data Classification | 2017 | Computer Science | 9 |
| Blockchain Technology Applications | 2018 | Computer Science | 9 |
| Software-Defined Networks & 5G | 2018 | Computer Science | 8 |
| SARS-CoV-2 Detection & Testing | 2020 | Medicine | 9 |
| AI in Law | 2021 | Computer Science | 6 |
| Poxvirus Research & Outbreaks | 2022 | Immunology | 8 |
Four views into the live index — momentum trajectories across nine topics since 2020, the current leaders, and the fastest-rising research terms. All drawn from the same weekly data we deliver.
The full index spans 580+ research topics — many of them biotech. For a technology investor we re-map those topics onto a familiar TMT taxonomy: 7 Core sub-indices a software, semiconductor, or internet fund tracks directly, 3 Adjacent deep-tech baskets you filter in on demand, and 23 biotech baskets hidden by default.
A systematic, four-stage pipeline converts raw preprint publications into structured investment signals — no manual curation, no subjective scoring.
Weekly ingestion of academic preprints from the largest open-access repositories for frontier research. Over 3M classified preprints since 2016.
Each abstract is classified into one or more of 580+ self-maintaining research topics using a proprietary taxonomy — new topics enter automatically the year their field emerges, so nothing is hand-curated. 99.3% classification accuracy.
Weekly publication velocity is normalised into a 0–100 momentum index per topic, where 100 = the topic's own trailing-year normal. A 4-week smoothing separates signal from noise.
One weekly dataset — Weekly Momentum (topic_weekly.csv) — updated weekly with a 3-week settled lag. Every topic carries two independent axes: a 0–100 scale index for absolute size and a 4-week-smoothed momentum index for pace versus its own normal, plus an investment-potential score.
The index is used by investors, strategists, and advisors who need a quantitative foundation for technology thesis work — not a narrative, a dataset.
Identify emerging categories 2–4 years before deal flow appears. Build data-backed thesis documents before the sector has a name.
Scan adjacent technical domains for threats and opportunities. Build R&D pipeline maps grounded in publication velocity, not analyst opinion.
Validate target company positioning against underlying research trends. Identify sectors approaching peak publication velocity before valuation follows.
Systematic, weekly signals across 580+ research topics. Integrate preprint momentum into quantitative models as a leading factor for technology-sector positioning.
Deliver technology landscape assessments backed by publication data, not keyword searches. Differentiate strategy reports with proprietary signal intelligence.
License a structured feed to enrich alternative data products. One clean weekly dataset, consistent schema, CSV and Parquet.
Over 3 million papers classified. 580+ self-maintaining research topics. One weekly dataset — each topic scored on scale, momentum, and investment potential.
Explore Sample Data