Submitting Results
How to submit benchmark results for leaderboard review.
Submission Flow
- Run your strategy on all benchmark cases
- Use
/api/benchmark/verifyto check your scores - When satisfied, send the run for review via
/api/benchmark/submit
Submission Format
import requests
BASE_URL = "https://biotradingarena.com"
headers = {
"Authorization": "Bearer YOUR_API_KEY",
"Content-Type": "application/json",
}
# Submit your predictions
resp = requests.post(
f"{BASE_URL}/api/benchmark/submit",
headers=headers,
json={
"strategy_name": "My Strategy v1",
"description": "LLM-based catalyst impact classifier using press releases and trial data",
"model": "gpt-4o",
"predictions": [
{
"case_id": "case_0123456789abcdef",
"predicted_impact": "positive",
"predicted_score": 2.4, # optional -10 to +10 impact score
"confidence": 0.85, # optional confidence
},
# ... more predictions
],
},
)
result = resp.json()
print(f"Submission ID: {result['submission_id']}")
print(f"Exact Match: {result['metrics']['exact_match_accuracy']}%")
print(f"Directional: {result['metrics']['directional_accuracy']}%")Prediction Types
You can submit two types of predictions per case:
Categorical (predicted_impact)
One of 7 impact categories:
very_negativenegativeslightly_negativeneutralslightly_positivepositivevery_positive
Numeric (predicted_score)
A numeric impact-score prediction in the -10 to +10 range. This is scored separately using Mean Absolute Error (MAE) against the canonical impact_score.
You can submit both for the same case.
Leaderboard Review
- The endpoint evaluates the payload and sends it for review; it does not publish a leaderboard row automatically.
- Accepted leaderboard runs are tied to the current published case set.
- Published runs are ranked by exact-match accuracy.
Tips
- Start by verifying with a small subset to debug your pipeline
- Use the
confidencefield to track which predictions your model is most/least certain about - The
reasoningfield (in verify) helps debug individual predictions - Submit the full oncology benchmark (233 cases) for the most meaningful comparison