Open Benchmark · 413 Curated Cases · 164 Public Tickers
Can AI predict
biotech stock moves?
An open benchmark for evaluating how well LLMs reason about clinical trial results, FDA decisions, and biotech catalysts, then predict the stock price reaction.
Why biotech
The hardest domain
for stock prediction
Biotech is unusually event-driven. FDA decisions, clinical trial readouts, safety updates, or even changes in trial design can move a stock dramatically in a single session. Our curated dataset includes moves ranging from -89.27% to +268.07%.
Interpreting these catalysts requires biology expertise and clinical context. A press release that sounds “positive” can still lead to a selloff if the results don't meet the bar the market has set.
This makes biotech a uniquely challenging testbed for AI reasoning. Models must go beyond sentiment analysis and actually understand the science.
Why “positive” results can mean a selloff
The setup
Real catalysts,
de-identified
Each case gives the model the actual catalyst (the press release announcing trial results, an FDA decision, or a safety update) along with the trial design, prior data, and related literature. The model must interpret the news and predict the stock price reaction.
To prevent models from simply recalling memorized outcomes, press releases are de-identified. Company names, drugs, tickers, dates, and executives are replaced with placeholders, so the model has to reason from the science, not its training data.
How scoring worksDe-identified press release
Actual announcements with company, drug, ticker, trial, and date identifiers removed, plus a deterministic leakage check before publication.
Linked clinical trial data
Structured data from ClinicalTrials.gov: endpoints, trial design, enrollment, arms, interventions, eligibility criteria, and site locations.
Source evidence trail
Each event exposes its source link, evidence status, and any unresolved source, date, or trial-timing warnings.
Reproducible impact scoring
Every case uses the same previous-adjusted-close to event-adjusted-close return window and the same clamped -10 to +10 scoring rule.
The dataset
413 curated stock-event cases
The benchmark spans Phase 1–3 readouts, FDA approvals and rejections, and topline results across 164 public tickers.
By catalyst type
By therapeutic area
We focused heavily on oncology, where we found better generalization within a broad disease area compared to mixing unrelated indications. The dataset also spans companies of different sizes, since large-cap biotech tends to exhibit much lower volatility than small and mid-cap names.
By stock price impact
Market-cap adjustedAuditable scoring
Every label is derived from one rule: previous-session adjusted close to event-session adjusted close, divided by five and clamped to a -10 to +10 score. Unverifiable historical market-cap estimates are not used.
De-identified press releases
Three-stage cleaning removes direct identifiers, while the benchmark input API omits exact dates, ticker-bearing IDs, price action, and labels.
413 linked clinical trials
Every case has a matched ClinicalTrials.gov entry with structured endpoints, trial design, patient populations, arms, interventions, and eligibility criteria.
Magnitude matters
Measures directional accuracy and magnitude error, not just up/down, against a single reproducible adjusted-close score.