Data for: "Relative Sensitivity of Blood Culture and PCR-Based TaqMan Array Card in Defining the Aetiology of Bacteraemia in Stillbirths and Deceased Children Aged <5 Years in eastern Ethiopia" – Data Codebook

Persistent identifier

10.17037/DATA.00005312

Description

This dataset comprises 979 paired blood culture and PCR results from postmortem blood samples collected from stillbirths and children aged <5 years in eastern Ethiopia between 2019 and 2024 through the Child Health and Mortality Prevention Surveillance (CHAMPS) Ethiopia platform (https://champshealth.org/site/ethiopia/). The study aims to evaluate the relative diagnostic yield of these two methods in detecting bacteraemia.

Data codebook

The dataset is comprised of three tables:

  1. Cohort wide format Dataset (Sheet1): Contains one row per participant (N=1,311) for the total enrolment, mapping baseline characteristics and specimen collection status. Of the 1,311 total enrolled participants, 979 had paired culture and PCR results available for final analysis.
  2. Sample-level long format Dataset (Sheet2): This contains two rows per sample of participant with discordant result pairs, tracking testing methods (Blood culture versus PCR). This structure helps when conducting conditional logistic regression.
  3. Pathogen-level long format Dataset (Sheet3): The full-volume evaluation database tracking every individual pathogen target across all analysed blood samples of the cases to preserve cumulative pathogen detection sums (479 isolates by culture and 574 detections by PCR), without restricting data down to discordant.

To protect participant identity, registration IDs (CHAMPS_id) have been replaced with different IDs that link measurements across the three datasets but cannot be used to trace results back to participant identifiable information.

Researchers wishing to replicate the analysis should note that Odds Ratios () evaluating diagnostic method discordance must account for the matched pair clustering design. Running analyses on either the full dataset or isolating the discordant rows using the properly constructed newid variable yields identical results. For demonstration, the Stata replication command for primary paired sample-level diagnostic yield evaluation are: clogit result met, group(newid)

Dataset-Sheet1

The cohort wide format dataset contains one row per participant (N=1,311) for the total enrolment, mapping baseline characteristics and specimen collection status. Of the 1,311 total enrolled participants, 979 had paired culture and PCR results available for final analysis.

Variable Name Variable Label Answer Label Answer Code Variable Type
bld_avail Post-Mortem Blood Specimen Availability for microbiological testing     String
    Blood specimen successfully drawn and processed CH00001  
    No blood specimen available for testing CH00002  
mits_perf Were there any circumstances that prevented the Minimally Invasive Tissue Sampling (MITS) from being performed?     String
    MITS not performed. Potential reasons may include lack of consent CH00001  
    MITS performed successfully CH00002  
mits_sex Sex of the deceased person     String
    Female Female  
    Male Male  
testdn Checking which among culture and PCR from the blood was done     String
    Both done Bothdone  
    Culture only culture only  
    PCR only Tac only  
    None None of them  
casetype1 Case type or category of age at the time of death     String
    Stillbirth stilb  
    Neonate (0–27 days) nicu  
    28 days–59 months 28day-59mo  
indtinv The interpretation of the PCR result. Categorical where the “0” represents the invalid results by the PCR. Invalid is described in the result and the figure 1.     Numeric
    Invalid result by the PCR 0  
mloc1 Location where the MITS procedure was conducted     String
    Haramaya haramaya  
    HFCSH HFCSH  
    Kersa Kersa_water  
CalcLocation1 Place of Death dichotomised as facility and community     String
    Community / Home Setting community  
    Facility / Hospital Setting facility  
outcome The test results from the two methods     String
    Both culture and PCR positive both pos  
    Both culture and PCR negative both neg  
    Culture positive only culture-only  
    PCR positive only taconly  
volcat1 Collected blood Sample Volume Category     String
    Under 5ml <5ml  
    5 to 10ml 5-10ml  
    Over 10ml >10ml  
study_id Unlinked study ID - IDs are shared across the 3 datasets Open ended   Numeric

Dataset-Sheet2

The sample-level long format Dataset contains two rows per sample of participant with discordant result pairs, tracking testing methods (Blood culture versus PCR). This structure helps when conducting conditional logistic regression.

Variable Name Variable Label Answer Label Answer Code Variable Type
mits_sex Sex of the deceased person     String
    Female Female  
    Male Male  
casetype1 Case type or category of age at the time of death. Shared matching covariate tier from wide cohort dataset (Table 1)     String
    Stillbirth stilb  
    Neonate (0–27 days) nicu  
    28 days–59 months 28day-59mo  
mloc1 Location where the MITS procedure was conducted. Shared matching covariate tier from wide cohort dataset (Table 1)     String
    Haramaya haramaya  
    HFCSH HFCSH  
    Kersa Kersa_water  
CalcLocation1 Place of Death dichotomised as facility and community. Shared matching covariate tier from wide cohort dataset (Table 1)     String
    Community / Home Setting community  
    Facility / Hospital Setting facility  
result Overall diagnostic yield for specific sample     String
    Sterile / No Growth / No Target Detected nogrowth  
    Positive Pathogen Signal / Target Isolated pos  
met Diagnostic method used     String
    Conventional Post-Mortem Blood Culture cult  
    PCR tac  
newid Matched Strata ID (Sample Pair Grouping). Stratum identifier that links the two matching pairs belonging to the same participant. Required argument for grouping in conditional logistic regression. Open ended   Numeric
volcat1 Collected blood Sample Volume Category. Shared matching covariate tier from wide cohort dataset (Table 1)     String
    Under 5ml <5ml  
    5 to 10ml 5-10ml  
    Over 10ml >10ml  
study_id Unlinked Subject ID. Public integer linkage key mapping rows back to the baseline Cohort Wide Dataset Open ended   Numeric

Dataset-Sheet3

The Pathogen-level long format dataset tracks every individual pathogen target across all analysed blood samples of the cases to preserve cumulative pathogen detection sums (479 isolates by culture and 574 detections by PCR), without restricting data down to discordant.

Variable Name Variable Label Answer Label Answer Code Variable Type
pathogen_id Pathogen target     String
    A. baumannii A. baumannii  
    Aeromonas species Aeromonas species  
    Bartonella species Bartonella species  
    Brucella species Brucella species  
    Candida species Candida species  
    E. cloacae E. cloacae  
    E.coli E.coli  
    Enterococcus species Enterococcus species  
    Few other bacterial isolates Few other bacterial isolates  
    H. influenzae H. influenzae  
    K. pneumoniae K. pneumoniae  
    Klebsiella species Klebsiella species  
    L. monocytogens L. monocytogens  
    M. catarrhalis M. catarrhalis  
    N. gonorrhea N. gonorrheae  
    N. meningitidis N. meningitidis  
    Ornitisia species Ornitisia species  
    P. auregenosa P. auregenosa  
    Pantoea species Pantoea species  
    Ricketsia species Ricketsia species  
    S. agalactiae S. agalactiae  
    S. aureus S. aureus  
    S. pneumoniae S. pneumoniae  
    S. pyogenes S. pyogenes  
    Salmonella species Salmonella species  
    Serratia species Serratia species  
    Shewanella species Shewanella species  
    Streptococcus species Streptococcus species  
    Treponema pallidium Treponema pallidium  
    Ureaplasma species Ureaplasma species  
result Overall diagnostic yield for specific sample     String
    Sterile / No Growth / No Target Detected negative  
    Positive Pathogen Signal / Target Isolated positive  
met Diagnostic method used     String
    Conventional Post-Mortem Blood Culture Culture  
    PCR TAC  
fastid Pathogen group dichotomised as the three common fastidious and other     String
    Fastidious fastidious  
    Non-fastidious non-fastidious  
grouped1 Pathogen group dichotomised as Gram negative and Gram positive     String
    Gram-Negative Bacteria Neg gram  
    Gram-Positive Bacteria gram pos  
study_id Unlinked Subject ID. Public integer linkage key mapping rows back to the baseline Cohort Wide Dataset Open ended   Numeric
newid Matched Pathogen Strata ID. Stratum grouping identifier unique to each Sample ID + Pathogen target combination. Required argument for paired discordant regression models Open ended   Numeric