CODE-15% (Telehealth Network of Minas Gerais, 15% subset)

Open Completed

Quick facts

Format12-lead · 10.24 s · 400 Hz · HDF5
Patients233,770
Records345,779
Leads12
LicenseCC BY 4.0
OriginTelehealth Network of Minas Gerais (TNMG) — Brazil

Overview

345,779 twelve-lead ECGs from 233,770 patients, and the largest dataset in this catalogue. It is an openly licensed 15% sample of the CODE cohort gathered by the Telehealth Network of Minas Gerais (TNMG), a Brazilian public telehealth service that reports ECGs recorded by non-specialist staff in primary-care units across the state.

A record is a row, not a file. 18 HDF5 parts each hold one (N, 4096, 12) array, so a signal reference names a row as well as a file — exams_part0.hdf5:tracings:417. Loading needs pip install ecgbench[hdf5]; nothing else about the API changes.

Six binary abnormality labels ship — 1dAVb, RBBB, LBBB, SB, ST, AF — derived from the TNMG’s own cardiologist reports, alongside a normal_ecg flag, age, sex, a neural-network age estimate, and mortality follow-up for 233,647 records. That last is unusual in an open ECG release and is what makes the dataset usable for survival work as well as classification.

The label trap is the important thing on this page. 308,004 records carry none of the six abnormalities, but only 134,657 records are flagged normal. The remaining 173,347 have some finding this six-class vocabulary does not name — see “About those counts” below.

Patients repeat heavily. 66,929 of the 233,770 patients contributed more than one recording, one of them 38. ECGBench’s folds are grouped on patient_id, so no patient appears in two folds.

Its limb-lead order is the standard I, II, III, aVR, aVL, aVF, which was checked rather than assumed — its sibling CODE-test, from the same cohort at the same sampling rate, uses a different one.

Label distribution

ClassRecords% of 345,779Note
RBBB9,6722.80%right bundle branch block
ST7,5842.19%sinus tachycardia
AF7,0332.03%atrial fibrillation
LBBB6,0261.74%left bundle branch block
1dAVb5,7161.65%1st-degree AV block
SB5,6051.62%sinus bradycardia
**≥1 of the six****37,775****10.92%**3,671 carry more than one
`normal_ecg`134,65738.94%explicitly flagged normal
**neither****173,347****50.13%****not flagged, and not normal**

About those counts

Every figure on this page was recomputed from the shipped exams.csv, whose md5 matches Zenodo’s published value (0107516d3f63864498fb77d15799cc95) exactly. The headline figures — 345,779 records and 233,770 patients — agree with the release description.

An empty label list is not a normal ECG, and this is the mistake the dataset invites. The six flags name six specific findings. Half the release (173,347 records, 50.1%) is neither flagged with any of them nor flagged normal_ecg: those recordings have something else — an axis deviation, an old infarct, a nonspecific ST change — that this vocabulary cannot express. A model trained on the six flags alone silently treats all 173,347 as confident negatives for every class. ECGBench exposes normal_ecg alongside abnormality_codes precisely so the two cases stay distinguishable, and the fold stratification separates NORMAL from OTHER for the same reason. Restrict your negatives to normal_ecg if you need clean ones.

The class table is multi-label and does not sum to 37,775. 3,671 records carry more than one of the six, so the six class counts total 41,636 against 37,775 records.

nn_predicted_age is a model output, not an observation. It is an age estimated from the tracing by a neural network. Exposed because it ships; training against it is training against another model.

Missing mortality follow-up means “not followed up”, not “survived”. death and timey are blank for 112,132 records and ship as an object-dtype column mixing True/False with NaN. ECGBench exposes death as a nullable boolean and adds an explicit has_followup flag, because reading NaN as False converts 112,132 unknowns into 112,132 survivors. Of the 233,647 records with follow-up, 8,341 died.

Demographics: ages run 17–100 (mean 53.2) and 40.3% of records are from men.

The class label is not a fresh expert read of each tracing. The six flags were derived from the reporting cardiologists’ free text by a combination of text mining and the automatic Minnesota coding, so they inherit that pipeline’s error rate as well as the reporters’.

Where the files come from

Zenodo serves 19 separate files — exams.csv plus 18 exams_part*.zip archives of roughly 2.7 GB each, expanding to 66 GB of HDF5. There is no single URL to auto-download from, so pass --data-path.

Three things about the layout that a naive reader gets wrong, all of which ECGBench’s splitter handles and checks:

Every part has one more row than it has records. Parts 0–16 hold 20,001 rows for 20,000 records and part 17 holds 5,780 for 5,779. The extra is an all-zero padding row with exam_id 0 that appears in no CSV.

exams.csv is not in file order. Its trace_file column says which part holds a record but not where in it, and its rows within a part do not follow that part’s own exam_id dataset. The row index has to come from that dataset; taking it from the CSV’s row number mislabels almost everything.

The samples are already in millivolts. The sibling release’s README claims a scale of “1e-4V … multiplied by 1000 in order to obtain the signals in V”, which is self-contradictory — the first clause puts a median R wave at 0.17 mV and the second at 1,750 V. Measured here: the median per-record peak is 4.27 mV and the median lead-II R amplitude 1.75 mV. signal_unit_scale is 1.0.

Zenodo publishes checksums only for the .zip archives, so an extracted copy cannot be checked against the provider directly. ECGBench verifies it structurally instead, on every run: all 18 arrays must have the expected (N, 4096, 12) shape, and each part’s exam_id set must equal the set exams.csv assigns to it. A mismatch raises rather than dropping records quietly.

The lead order was likewise derived from the arrays rather than taken on trust, over 1,200 records spanning parts 0, 5, 11 and 17: III = II − I, aVR = −(I+II)/2, aVL = I − II/2 and aVF = II − I/2 all hold to a median relative error under 2%, while every other assignment is off by more than 140%.

Validation summary (400 Hz)

VersionRecordsNote
original345,779all records, with is_valid + quality_issues
clean337,23897.53% pass rate
excluded8,5416,798 amplitude_outlier, 1,773 missing_leads, 62 flat_line

About the excluded records

amplitude_range_mv here is [-20, 20], not ECGBench’s usual [-10, 10], and the same value is used for CODE-test so the two releases of one cohort are cleaned to a single standard.

The reason is the cohort itself. These are telehealth recordings made by non-specialist staff on portable tele-electrocardiographs in primary-care units, so they carry considerably more electrode artefact and baseline wander than a hospital dataset: the median per-record peak is 4.27 mV, against 1.74 mV for the hospital-collected SPH. Measured over 20,000 records of part 3, the exclusion rate would be 38.7% at ±5 mV, 11.8% at ±10, 6.5% at ±15, 2.0% at ±20 and 0.15% at ±30. ±10 would discard roughly 40,000 legitimate recordings; ±20 keeps the exclusion near the 1–2% the rest of the catalogue sits at while still catching the railed leads — the worst record in that sample peaks at 281 mV.

What that produced over all 345,779 records: 6,798 records fail amplitude_outlier (10,161 lead-level failures — most have one or two bad leads, not twelve), 1,773 fail missing_leads with a lead recorded as exactly zero throughout, and 62 fail flat_line. Only 91 records fail more than one check. Nothing fails nan_values or truncated_signal: every array is complete and every record is exactly 4,096 samples.

Fold membership is identical between original/ and clean/clean/ is a row subset, not a re-split — so a model trained on clean/ can be scored against original/ for the same fold without re-partitioning.

Splits

SplitFoldscleanoriginal
train1–8269,733276,624
val933,73534,578
test1033,77034,577

How the folds were made

Ten folds, grouped on patient_id and stratified on the rarest of the six abnormalities each record carries, with the unflagged records split into NORMAL and OTHER rather than pooled — those two are different things, as above.

Rarest-wins is used because the release ranks nothing. Taking the first-listed flag instead would starve the small classes in favour of RBBB.

What that produced, measured on the output:

One caveat that patient grouping cannot fix. The source contains a small number of byte-identical duplicate recordings filed under different patient_ids — 47 non-degenerate records among part 0’s 20,000 (0.24%), in groups that sometimes span patients. Grouping on patient_id cannot see those, so a very small amount of duplicate-record leakage between folds is possible. It comes from the source, is documented here rather than silently accepted, and is far below the leakage that ignoring patient_id altogether would cause.

Building the splits

# All 18 parts must be extracted alongside exams.csv.
for f in exams_part*.zip; do unzip -n "$f"; done

ecgbench splits --dataset code15 --data-path /path/to/code-15/

Loading with ECGBench

# pip install ecgbench[torch,hdf5]   <- h5py is needed for this dataset
from ecgbench import ECGDataset

ds = ECGDataset(
    "code15",
    split="train",
    data_path="/path/to/code-15/",
    window=(48, 4000),       # the 10 s inside the symmetric zero padding
    labels=True,
)

len(ds)                                      # 269733
ds[0]["signal"].shape                        # (12, 4000)
ds[0]["record_id"]                           # 26

# THE TRAP: an empty list is not a normal ECG. Read both.
ds[0]["labels"]["abnormality_codes"]         # '' — none of the six
ds[0]["labels"]["normal_ecg"]                # False -> some OTHER finding
ds[0]["labels"]["stratify_class"]            # 'OTHER'  (folds only)

# Mortality follow-up, with its missingness intact. death is None for a
# record that was never followed up, not False.
ds[0]["labels"]["death"]                     # None
ds[0]["labels"]["has_followup"]              # False — "not followed up"

# Standard lead order — but its sibling CODE-test is NOT standard, so
# select by name if you use both releases.
ds.config.lead_names
# ['I','II','III','aVR','aVL','aVF','V1','V2','V3','V4','V5','V6']

both = ECGDataset("code15", split="train", data_path="/path/to/code-15/",
                  leads=["aVR", "aVL", "aVF"])
both[0]["signal"].shape                      # (3, 4096)

# Samples are already millivolts, despite the sibling README's "1e-4V".
ds.config.signal_unit_scale                  # 1.0