Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
145 changes: 145 additions & 0 deletions Demographics_starter_kit.qmd
Original file line number Diff line number Diff line change
@@ -0,0 +1,145 @@
---
authors:
- name: Briha Ansari
- name: Patrick Sadil
---

# Demographics {#sec-demographics}

The demographics form captures the descriptive and background characteristics of each A2CPS participant: age, sex, gender identity, race and ethnicity, education, employment, marital status, household income, disability status, dominant hand, height and weight, pain duration, and a few COVID-related items. Almost every A2CPS analysis draws on these fields so demographics is often the first modality a project touches. Participants also belong to one of two surgical cohorts (knee arthroplasty and thoracic surgery), and since these cohorts differ systematically in age and body composition, the demographic variables matter when we interpret cohort comparisons.

Many fields are coded categorical values (for example, `sex` is `1 = Male`, `2 = Female`), with the codes defined in the accompanying data dictionary. In addition to the raw fields, the cleaning workflow computes a validated **body mass index (BMI)** from height and weight, and attaches a set of **data-quality flags** that mark implausible or logically inconsistent values so that we can decide how to handle them.

## Starting Project

### Locate Data

Where are the relevant files?

```bash
$ {{< var release.root >}}/{{< var release.dir >}}/demographics/reformatted
```

The reformatted data are in a single file, `reformatted_demo.csv` (one row per participant), together with `demo_dict_updated.csv`. The dictionary is essential here, since it lists the coded values for every categorical field (`sex`, `genident`, `ethnic`, `edulevel`, `empstat`, `maristat`, `incmlvl`, and so on) alongside the fields added during cleaning.

### Extract Data

The file is a plain CSV, and here we use the [`tidyverse`](https://www.tidyverse.org/).

```{r setup}
library(tidyverse)
```

```{r load}
demo <- read_csv("data/pre-surgery/demographics/reformatted/reformatted_demo.csv")
```

Since the categorical fields are stored as numeric codes, using the data dictionary as a reference is highly recommended (`demo_dict_updated.csv`).

```{r dict}
demo_dict <- read_csv("data/pre-surgery/demographics/reformatted/demo_dict_updated.csv")

demo_dict |>
filter(field_name == "sex") |>
select(field_name, select_choices_or_calculations)
```

Next, a quick look at the continuous fields and the cohort label:

```{r peek}
demo |>
select(record_id, cohort, age, sex, bmi_kg_m2, paindur) |>
head()
```

`glimpse()` confirms the data types. Note that the coded categoricals come in as numbers, so we recode them to labelled factors before summarizing or plotting.

```{r glimpse}
demo |>
select(record_id, cohort, age, sex, edulevel, bmi_kg_m2, paindur) |>
glimpse()
```

### Data Quality

The reformatted file has already been cleaned. It keeps the implausible values and flags them instead of dropping them, which lets the end user decide how to handle each flagged value. The flag columns are:

- **`bmi_measurement_issue`** and **`height_weight_check`**: these mark BMI values built from missing or implausible height and weight. Most records are `"Valid Range"`, but a small number are flagged. The raw `bmi_kg_m2` column still contains extreme values (as low as 2 and as high as 64).
- **`paindur_outlier_flag`**: flags pain-duration values that are zero or negative (~200 records), extreme statistical outliers (~80), or logically impossible (pain lasting longer than the participant has been alive).
- **`age_outlier_flag`** and **`age_pain_logic_error`**: these mark missing ages and cases where the reported pain duration exceeds the participant's total months of life.
- **`covid_logic_error`**: flags internally inconsistent COVID responses (e.g., a "No" with a follow-up detail still answered).
- **`dom_hand_status`** and **`incmlvl_clean`**: cleaned, human-readable versions of dominant hand and income level.

It helps to tally a flag first and see how many records get affected:

```{r qc}
demo |>
count(bmi_measurement_issue)
```

### Cross-Modality Links

Every file includes `record_id` and `guid`, which uniquely identify a participant and let us link demographics to other A2CPS modalities. Since demographics is almost always used as covariates for another modality, the typical join attaches these background fields to a primary analysis table:

```{r cross-modality}
#| eval: false
inner_join(other_modality, demo, by = intersect(names(other_modality), names(demo)))
```

## Exploratory data analysis

Since the two surgical cohorts are recruited separately, a useful way to get oriented is to compare their basic demographics, because the differences here shape the interpretation of cohort comparisons. We look at age and BMI, and we use the quality flags to keep only the analysis-ready BMI values.

```{r prep}
demo_clean <- demo |>
mutate(
cohort = factor(cohort),
bmi = if_else(bmi_measurement_issue == "Valid Range", bmi_kg_m2, NA_real_)
)

demo_clean |>
group_by(cohort) |>
summarise(
n = n(),
age_mean = mean(age, na.rm = TRUE),
bmi_mean = mean(bmi, na.rm = TRUE)
)
```

As expected, the knee (TKA) cohort is both older and heavier on average (age ≈ 65 vs. 60; BMI ≈ 31 vs. 29), which fits the usual clinical profile of knee-arthroplasty candidates. Let's look at the full distributions with side-by-side boxplots.

```{r plot-age}
ggplot(demo_clean, aes(x = cohort, y = age, fill = cohort)) +
geom_boxplot(show.legend = FALSE) +
labs(title = "Age by cohort", x = NULL, y = "Age (years)") +
theme_minimal()
```

```{r plot-bmi}
demo_clean |>
filter(!is.na(bmi)) |>
ggplot(aes(x = cohort, y = bmi, fill = cohort)) +
geom_boxplot(show.legend = FALSE) +
labs(title = "BMI by cohort (quality-flagged values removed)",
x = NULL, y = expression(BMI~(kg/m^2))) +
theme_minimal()
```

The takeaway for planning is that the cohorts are not demographically interchangeable, so age and BMI are natural covariates to consider whenever we pool or compare them.

## Considerations While Working on the Project

### Data Generation

Demographic information is self-reported by participants at baseline through REDCap, following the A2CPS Manual of Procedures. Height and weight are used to derive BMI. The steps that turn the REDCap export into the reformatted file (recoding, BMI computation, and the plausibility and logic checks that populate the flag columns) are documented in the cleaning workflow that comes with the release.

### Other

- **Coded fields.** Most categorical fields are stored as numbers, and the coding differs from field to field, so it helps keep `demo_dict_updated.csv` handy.
- **Sparse Data.** A few fields have very sparse levels (for example, the handful of Unknown and Intersex responses in `sex`), which can make those subgroups unstable.
- **Self-report.** These fields are all self-reported at baseline, so the usual survey caveats apply, such as some missingness, rounding in height, weight, and pain duration, and bias in items like income.
- **Quality control.** The cleaning pipeline has been reviewed, and the flag columns serve as its record, showing which values were questioned and why. The checks can be re-derived from the accompanying cleaning workflow.

### Citations

{{< include _snippets/a2cps_citations.qmd >}}
15 changes: 15 additions & 0 deletions _freeze/Demographics_starter_kit/execute-results/html.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
{
"hash": "3b91bdd05cdfe8675527f86338e8c8c3",
"result": {
"engine": "knitr",
"markdown": "---\nauthors:\n - name: Briha Ansari\n - name: Patrick Sadil\n---\n\n# Demographics {#sec-demographics}\n\nThe demographics form captures the descriptive and background characteristics of each A2CPS participant: age, sex, gender identity, race and ethnicity, education, employment, marital status, household income, disability status, dominant hand, height and weight, pain duration, and a few COVID-related items. Almost every A2CPS analysis draws on these fields so demographics is often the first modality a project touches. Participants also belong to one of two surgical cohorts (knee arthroplasty and thoracic surgery), and since these cohorts differ systematically in age and body composition, the demographic variables matter when we interpret cohort comparisons.\n\nMany fields are coded categorical values (for example, `sex` is `1 = Male`, `2 = Female`), with the codes defined in the accompanying data dictionary. In addition to the raw fields, the cleaning workflow computes a validated **body mass index (BMI)** from height and weight, and attaches a set of **data-quality flags** that mark implausible or logically inconsistent values so that we can decide how to handle them.\n\n## Starting Project\n\n### Locate Data\n\nWhere are the relevant files?\n\n```bash\n$ {{< var release.root >}}/{{< var release.dir >}}/demographics/reformatted\n```\n\nThe reformatted data are in a single file, `reformatted_demo.csv` (one row per participant), together with `demo_dict_updated.csv`. The dictionary is essential here, since it lists the coded values for every categorical field (`sex`, `genident`, `ethnic`, `edulevel`, `empstat`, `maristat`, `incmlvl`, and so on) alongside the fields added during cleaning.\n\n### Extract Data\n\nThe file is a plain CSV, and here we use the [`tidyverse`](https://www.tidyverse.org/).\n\n\n::: {.cell}\n\n```{.r .cell-code}\nlibrary(tidyverse)\n```\n:::\n\n\n\n::: {.cell}\n\n```{.r .cell-code}\ndemo <- read_csv(\"data/pre-surgery/demographics/reformatted/reformatted_demo.csv\")\n```\n:::\n\n\nSince the categorical fields are stored as numeric codes, using the data dictionary as a reference is highly recommended (`demo_dict_updated.csv`).\n\n\n::: {.cell}\n\n```{.r .cell-code}\ndemo_dict <- read_csv(\"data/pre-surgery/demographics/reformatted/demo_dict_updated.csv\")\n\ndemo_dict |>\n filter(field_name == \"sex\") |>\n select(field_name, select_choices_or_calculations)\n```\n\n::: {.cell-output-display}\n<div class=\"kable-table\">\n\n|field_name |select_choices_or_calculations |\n|:----------|:-------------------------------------------------------------|\n|sex |1, Male &#124; 2, Female &#124; 3, Unknown &#124; 4, Intersex |\n\n</div>\n:::\n:::\n\n\nNext, a quick look at the continuous fields and the cohort label:\n\n\n::: {.cell}\n\n```{.r .cell-code}\ndemo |>\n select(record_id, cohort, age, sex, bmi_kg_m2, paindur) |>\n head()\n```\n\n::: {.cell-output-display}\n<div class=\"kable-table\">\n\n| record_id|cohort | age| sex| bmi_kg_m2| paindur|\n|---------:|:------|---:|---:|---------:|-------:|\n| 10001|TKA | 68| 2| 32.99943| 2|\n| 10003|TKA | 73| 2| 22.13865| 12|\n| 10004|TKA | 64| 2| 35.17960| 9|\n| 10005|TKA | 58| 1| 28.75630| 144|\n| 10006|TKA | 55| 1| 44.04175| 12|\n| 10007|TKA | 63| NA| 35.95317| 6|\n\n</div>\n:::\n:::\n\n\n`glimpse()` confirms the data types. Note that the coded categoricals come in as numbers, so we recode them to labelled factors before summarizing or plotting.\n\n\n::: {.cell}\n\n```{.r .cell-code}\ndemo |>\n select(record_id, cohort, age, sex, edulevel, bmi_kg_m2, paindur) |>\n glimpse()\n```\n\n::: {.cell-output .cell-output-stdout}\n\n```\nRows: 1,933\nColumns: 7\n$ record_id <dbl> 10001, 10003, 10004, 10005, 10006, 10007, 10008, 10010, 1001…\n$ cohort <chr> \"TKA\", \"TKA\", \"TKA\", \"TKA\", \"TKA\", \"TKA\", \"TKA\", \"TKA\", \"TKA…\n$ age <dbl> 68, 73, 64, 58, 55, 63, 73, 64, 54, 53, 71, 45, 65, 59, 56, …\n$ sex <dbl> 2, 2, 2, 1, 1, NA, 2, 1, 1, 1, 2, NA, 2, 2, NA, 1, 2, 1, 2, …\n$ edulevel <dbl> 5, 5, 4, 6, 4, 5, 3, 4, 5, 4, 6, 6, 4, 6, 3, 3, 5, 6, 6, 5, …\n$ bmi_kg_m2 <dbl> 32.99943, 22.13865, 35.17960, 28.75630, 44.04175, 35.95317, …\n$ paindur <dbl> 2, 12, 9, 144, 12, 6, 660, 64, 84, 192, 60, 60, 96, 36, 36, …\n```\n\n\n:::\n:::\n\n\n### Data Quality\n\nThe reformatted file has already been cleaned. It keeps the implausible values and flags them instead of dropping them, which lets the end user decide how to handle each flagged value. The flag columns are:\n\n- **`bmi_measurement_issue`** and **`height_weight_check`**: these mark BMI values built from missing or implausible height and weight. Most records are `\"Valid Range\"`, but a small number are flagged. The raw `bmi_kg_m2` column still contains extreme values (as low as 2 and as high as 64).\n- **`paindur_outlier_flag`**: flags pain-duration values that are zero or negative (~200 records), extreme statistical outliers (~80), or logically impossible (pain lasting longer than the participant has been alive).\n- **`age_outlier_flag`** and **`age_pain_logic_error`**: these mark missing ages and cases where the reported pain duration exceeds the participant's total months of life.\n- **`covid_logic_error`**: flags internally inconsistent COVID responses (e.g., a \"No\" with a follow-up detail still answered).\n- **`dom_hand_status`** and **`incmlvl_clean`**: cleaned, human-readable versions of dominant hand and income level.\n\nIt helps to tally a flag first and see how many records get affected:\n\n\n::: {.cell}\n\n```{.r .cell-code}\ndemo |>\n count(bmi_measurement_issue)\n```\n\n::: {.cell-output-display}\n<div class=\"kable-table\">\n\n|bmi_measurement_issue | n|\n|:-----------------------------------------------|----:|\n|Check Raw Data: Flag: Height Ft Implausible | 3|\n|Check Raw Data: Flag: Height Inches Implausible | 2|\n|Check Raw Data: Flag: Weight Implausible | 1|\n|Check Raw Data: Missing Raw Data | 68|\n|Flag: Implausibly High BMI | 3|\n|Valid Range | 1856|\n\n</div>\n:::\n:::\n\n\n### Cross-Modality Links\n\nEvery file includes `record_id` and `guid`, which uniquely identify a participant and let us link demographics to other A2CPS modalities. Since demographics is almost always used as covariates for another modality, the typical join attaches these background fields to a primary analysis table:\n\n\n::: {.cell}\n\n```{.r .cell-code}\ninner_join(other_modality, demo, by = intersect(names(other_modality), names(demo)))\n```\n:::\n\n\n## Exploratory data analysis\n\nSince the two surgical cohorts are recruited separately, a useful way to get oriented is to compare their basic demographics, because the differences here shape the interpretation of cohort comparisons. We look at age and BMI, and we use the quality flags to keep only the analysis-ready BMI values.\n\n\n::: {.cell}\n\n```{.r .cell-code}\ndemo_clean <- demo |>\n mutate(\n cohort = factor(cohort),\n bmi = if_else(bmi_measurement_issue == \"Valid Range\", bmi_kg_m2, NA_real_)\n )\n\ndemo_clean |>\n group_by(cohort) |>\n summarise(\n n = n(),\n age_mean = mean(age, na.rm = TRUE),\n bmi_mean = mean(bmi, na.rm = TRUE)\n )\n```\n\n::: {.cell-output-display}\n<div class=\"kable-table\">\n\n|cohort | n| age_mean| bmi_mean|\n|:--------|----:|--------:|--------:|\n|Thoracic | 527| 60.24658| 29.37161|\n|TKA | 1406| 65.31789| 31.34604|\n\n</div>\n:::\n:::\n\n\nAs expected, the knee (TKA) cohort is both older and heavier on average (age ≈ 65 vs. 60; BMI ≈ 31 vs. 29), which fits the usual clinical profile of knee-arthroplasty candidates. Let's look at the full distributions with side-by-side boxplots.\n\n\n::: {.cell}\n\n```{.r .cell-code}\nggplot(demo_clean, aes(x = cohort, y = age, fill = cohort)) +\n geom_boxplot(show.legend = FALSE) +\n labs(title = \"Age by cohort\", x = NULL, y = \"Age (years)\") +\n theme_minimal()\n```\n\n::: {.cell-output-display}\n![](Demographics_starter_kit_files/figure-html/plot-age-1.png){width=672}\n:::\n:::\n\n\n\n::: {.cell}\n\n```{.r .cell-code}\ndemo_clean |>\n filter(!is.na(bmi)) |>\n ggplot(aes(x = cohort, y = bmi, fill = cohort)) +\n geom_boxplot(show.legend = FALSE) +\n labs(title = \"BMI by cohort (quality-flagged values removed)\",\n x = NULL, y = expression(BMI~(kg/m^2))) +\n theme_minimal()\n```\n\n::: {.cell-output-display}\n![](Demographics_starter_kit_files/figure-html/plot-bmi-1.png){width=672}\n:::\n:::\n\n\nThe takeaway for planning is that the cohorts are not demographically interchangeable, so age and BMI are natural covariates to consider whenever we pool or compare them.\n\n## Considerations While Working on the Project\n\n### Data Generation\n\nDemographic information is self-reported by participants at baseline through REDCap, following the A2CPS Manual of Procedures. Height and weight are used to derive BMI. The steps that turn the REDCap export into the reformatted file (recoding, BMI computation, and the plausibility and logic checks that populate the flag columns) are documented in the cleaning workflow that comes with the release.\n\n### Other\n\n- **Coded fields.** Most categorical fields are stored as numbers, and the coding differs from field to field, so it helps keep `demo_dict_updated.csv` handy.\n- **Sparse Data.** A few fields have very sparse levels (for example, the handful of Unknown and Intersex responses in `sex`), which can make those subgroups unstable.\n- **Self-report.** These fields are all self-reported at baseline, so the usual survey caveats apply, such as some missingness, rounding in height, weight, and pain duration, and bias in items like income.\n- **Quality control.** The cleaning pipeline has been reviewed, and the flag columns serve as its record, showing which values were questioned and why. The checks can be re-derived from the accompanying cleaning workflow.\n\n### Citations\n\nIn publications or presentations including data from A2CPS, please include the following statement as attribution:\n\n> Data were provided (in part) by the A2CPS Consortium funded by the National Institutes of Health (NIH) Common Fund, which is managed by the Office of the Director (OD)/Office of Strategic Coordination (OSC). Consortium components and their associated funding sources include Clinical Coordinating Center (U24NS112873), Data Integration and Resource Center (U54DA049110), Omics Data Generation Centers (U54DA049116, U54DA049115, U54DA049113), Multi-site Clinical Center 1 (MCC1) (UM1NS112874), and Multi-site Clinical Center 2 (MCC2) (UM1NS118922).\n\n::: {.callout-note}\nThe following published papers should be cited when referring to A2CPS Protocol and Biomarkers: @sluka_predicting_2023 @berardi_multi_2022\n:::\n\n",
"supporting": [],
"filters": [
"rmarkdown/pagebreak.lua"
],
"includes": {},
"engineDependencies": {},
"preserve": {},
"postProcess": true
}
}
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
1 change: 1 addition & 0 deletions _quarto.yml
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,7 @@ book:
chapters:
- index.qmd
- Psychosocial_starter_kit.qmd
- Demographics_starter_kit.qmd
- genetic_variants.qmd
- genomic_imputation.qmd
- mri-idps.qmd
Expand Down