AgenticGenomics: pharmacogenomics benchmark dataset for "Trustworthy agentic genomics through versioned skill libraries" (five-condition, 44,550 evaluations)
Benchmark dataset for Corpas et al. (2026), "Trustworthy agentic genomics through versioned skill libraries: deterministic, auditable pharmacogenomics across nine models" (under revision at Cell Genomics; CELL-GENOMICS-D-26-00551).
Design. Nine frontier large language models x 110 CPIC Level A pharmacogenomic cases x three population contexts (European, admixed Latin American, East African) x five conditions x three replicates = 44,550 evaluations. The five conditions are: (1) free-prompted; (2) retrieval-augmented from the CPIC guideline corpus; (3) skill-reasoning (the model reasons over a versioned SKILL.md specification); (4) skill-execution (the specification's CPIC mapping is executed deterministically as code); and (5) an answer-supplied positive control. The three core conditions (1, 2, 5) contribute 26,730 evaluations; the two skill conditions (3, 4) contribute a further 17,820 (v3_armAB_fullgrid.json/.jsonl).
Version 1.2.0 (revision). Completes the five-condition design by adding the 17,820 skill-arm evaluations that the manuscript describes but that were absent from the earlier three-arm deposit (v1.1.0); refreshes the archived analysis-code snapshot; and pins the reproducibility package to repository commit 3f482d4 (https://github.com/manuelcorpas/24-AGENTIC-PGX-BENCHMARK), corrected to run on a case-sensitive filesystem and to load credentials from the environment. Includes the end-to-end executed-pipeline analysis (real-genome deterministic caller to executed skill to abstention).
Provenance. All data derive from public CPIC Level A guidelines and PharmGKB annotations. Genotypes are synthetic, text-specified canonical cases. No patient data and no new sequencing data are included. File checksums are in CHECKSUMS.sha256; no credentials are present in the archive. Scoring fields per record: A1 (phenotype identification), A2 (drug-specific recommendation), A3 (lethal-class safety action).
Keywords
agentic genomics; pharmacogenomics; large language models; CPIC; retrieval-augmented generation; benchmark; reproducibility; clinical decision support; precision medicine; skill libraries| Item Type | Dataset |
|---|---|
| Resource Type |
Resource Type Resource Description Dataset Quantitative |
| Capture method | Experiment, Simulation |
| Date | 24 July 2026 |
| Language(s) of written materials | English |
| Creator(s) |
Corpas, M; Iacoangeli, A; Bourdenx, M; Skene, N; Aldraimli, M; Fatumo, S |
| LSHTM Faculty/Department | MRC/UVRI and LSHTM Uganda Research Unit |
| Participating Institutions | London School of Hygiene & Tropical Medicine, London, United Kingdom |
| Date Deposited | 24 Jul 2026 10:26 |
| Last Modified | 24 Jul 2026 10:26 |
| Publisher | Zenodo |