Relationships are declared once and mirrored automatically. A
shares records pair contains some of the same recordings, so
training on one and evaluating on the other contaminates the test set.
Verified means the overlap was checked against the actual data files;
otherwise it is taken from documentation.
The 2021 challenge reuses the whole 2020 training set and adds the Chapman-Shaoxing and Ningbo cohorts on top. The arithmetic supports that exactly: removing those two cohorts from this release leaves 88,253 − 10,247 − 34,905 = 43,101 records across the same six cohorts the 2020 challenge used, which is precisely the published size of the 2020 public training set. Every 2020 training record should therefore be assumed present here. Treat results across the two years as sharing records, and never use one year’s training set to evaluate a model trained on the other. verified: false because the count agreement was checked against the 2020 challenge description, not against a 2020 download.
INCART is one of the six source cohorts of the 2020 challenge training set, as it is of the 2021 one. The 2021 overlap is verified against the files (see that relationship, declared there); this one follows from the challenge descriptions plus the arithmetic that the 2021 release minus its Chapman and Ningbo cohorts leaves exactly the 43,101 records of the 2020 training set. Not checked against a 2020 download, hence unverified. Do not evaluate on INCART after training on either challenge year.