Open Benchmark · 413 Curated Cases · 164 Public Tickers

Can AI predict
biotech stock moves?

An open benchmark for evaluating how well LLMs reason about clinical trial results, FDA decisions, and biotech catalysts, then predict the stock price reaction.

Why biotech

The hardest domain
for stock prediction

Biotech is unusually event-driven. FDA decisions, clinical trial readouts, safety updates, or even changes in trial design can move a stock dramatically in a single session. Our curated dataset includes moves ranging from -89.27% to +268.07%.

Interpreting these catalysts requires biology expertise and clinical context. A press release that sounds “positive” can still lead to a selloff if the results don't meet the bar the market has set.

This makes biotech a uniquely challenging testbed for AI reasoning. Models must go beyond sentiment analysis and actually understand the science.

Why “positive” results can mean a selloff

The effect size is weaker than expected
Results apply only to a narrow subgroup
Safety signals appear in the data
Endpoints don't meaningfully de-risk later phases
The readout doesn't materially change approval odds

The setup

Real catalysts,
de-identified

Each case gives the model the actual catalyst (the press release announcing trial results, an FDA decision, or a safety update) along with the trial design, prior data, and related literature. The model must interpret the news and predict the stock price reaction.

To prevent models from simply recalling memorized outcomes, press releases are de-identified. Company names, drugs, tickers, dates, and executives are replaced with placeholders, so the model has to reason from the science, not its training data.

How scoring works
01

De-identified press release

Actual announcements with company, drug, ticker, trial, and date identifiers removed, plus a deterministic leakage check before publication.

02

Linked clinical trial data

Structured data from ClinicalTrials.gov: endpoints, trial design, enrollment, arms, interventions, eligibility criteria, and site locations.

03

Source evidence trail

Each event exposes its source link, evidence status, and any unresolved source, date, or trial-timing warnings.

04

Reproducible impact scoring

Every case uses the same previous-adjusted-close to event-adjusted-close return window and the same clamped -10 to +10 scoring rule.

The dataset

413 curated stock-event cases

The benchmark spans Phase 1–3 readouts, FDA approvals and rejections, and topline results across 164 public tickers.

By catalyst type

FDA Approvals
161
Phase 3 Positive
77
Phase 2 Positive
50
Phase 3 Negative
30
Topline Positive
28
FDA Rejections
17
Phase 2 Negative
13
Phase 1 Positive
10
Trial Milestones
7
Topline Negative
4
Clinical Holds Lifted
3
Programs Discontinued
3
Applications Under Review
2
Regulatory Opinions
2
Commercial Launches
2
Clinical Publications
1
Regulatory Designations
1
Regulatory Restrictions
1
Emergency Use Authorizations
1

By therapeutic area

233
Oncology
56% of dataset
180
Other areas
Cardio, neuro, rare disease, etc.

We focused heavily on oncology, where we found better generalization within a broad disease area compared to mixing unrelated indications. The dataset also spans companies of different sizes, since large-cap biotech tends to exhibit much lower volatility than small and mid-cap names.

By stock price impact

Market-cap adjusted
Very Positive
26 (6.3%)
Positive
30 (7.3%)
Slightly Positive
37 (9%)
Neutral
220 (53.3%)
Slightly Negative
36 (8.7%)
Negative
31 (7.5%)
Very Negative
33 (8%)

Auditable scoring

Every label is derived from one rule: previous-session adjusted close to event-session adjusted close, divided by five and clamped to a -10 to +10 score. Unverifiable historical market-cap estimates are not used.

De-identified press releases

Three-stage cleaning removes direct identifiers, while the benchmark input API omits exact dates, ticker-bearing IDs, price action, and labels.

413 linked clinical trials

Every case has a matched ClinicalTrials.gov entry with structured endpoints, trial design, patient populations, arms, interventions, and eligibility criteria.

Magnitude matters

Measures directional accuracy and magnitude error, not just up/down, against a single reproducible adjusted-close score.