-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathsurvey-data.Rmd
More file actions
354 lines (278 loc) · 10.4 KB
/
Copy pathsurvey-data.Rmd
File metadata and controls
354 lines (278 loc) · 10.4 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
---
title: "PLAY Survey Data"
author: "Rick Gilmore"
date: "`r Sys.time()`"
output:
html_document:
self_contained: yes
code_folding: show
toc: true
toc_depth: 2
toc_float: true
params:
db_login: email@provider.com
play_volume_id: 1280
play_survey_session_id: 51539
---
# Purpose
This document shows how to access and visualize survey response data from the PLAY dataset.
# Rendering
To render the document, issue this command from the R console
`rmarkdown::render('survey-data.Rmd', params = list(db_login='<YOUR_DATABRARY_LOGIN>'))`
substituting your actual Databrary login (email account) for `<YOUR_DATABRARY_LOGIN>`.
If you try to 'knit' the document another way, it won't work.
# Setup
There are several R packages we need to run this code.
The following will test to see if the packages are installed on your system.
If not, the packages will be installed for you.
```{r setup, include=TRUE}
knitr::opts_chunk$set(echo = TRUE)
if (!require(tidyverse)) {
install.packages('tidyverse')
}
if (!require(devtools)) {
install.packages('devtools')
}
if (!require(databraryapi)) {
devtools::install_github('PLAY-behaviorome/databraryapi')
}
library(tidyverse) # for pipe `%>%` operator
```
You will need to login to Databrary in order to access the PLAY data.
The next chunk asks for your Databrary credentials.
It returns `TRUE` if your login is successful.
```{r login-databrary}
databraryapi::login_db(params$db_login)
```
# List files
Databrary 1.0 has a volume/session structure.
A volume is a dataset or collection.
A session is a collection of related files.
You can think of a session as a folder.
The survey data are in a session folder called 'KoBoToolbox Survey Data'.
Volumes and sessions have ID numbers that Databrary uses to access each.
The PLAY data release volume has an id of `r params$play_volume_id`.
The survey data session has a session id of `r params$play_survey_session_id`.
So, to go to the PLAY volume in your browser, you can visit <https://nyu.databrary.org/volume/1280>.
The following code, however, does all of this within an R session.
```{r load-play-survey-files}
play_survey_files <-
databraryapi::list_assets_in_session(session_id = params$play_survey_session_id)
# Print the files data in a 'pretty' way
play_survey_files %>%
kableExtra::kbl(.) %>%
kableExtra::kable_styling(., bootstrap_options = c("striped", "hover", "condensed", "responsive")) %>%
kableExtra::scroll_box(width = "800px", height = "300px")
```
Note that each file has an `asset_id`.
We'll use this to access the data file(s) we want.
# Import by measure
## Pets
```{r load-pets}
# Grab the asset_id for the pets data based on the name
pets_in_name <- stringr::str_detect(play_survey_files$name, 'pets')
pets_asset_id <- play_survey_files$asset_id[pets_in_name]
# Select and download the pets data
pets <-
databraryapi::read_csv_data_as_df(params$play_survey_session_id, as.numeric(pets_asset_id))
# Print the pets data
pets %>%
kableExtra::kbl(.) %>%
kableExtra::kable_styling(., bootstrap_options = c("striped", "hover", "condensed", "responsive")) %>%
kableExtra::scroll_box(width = "800px", height = "300px")
```
We can do a bit of clean-up here to see what families have pets and of what type.
```{r clean-pets}
pets_cleaned <- pets %>%
dplyr::mutate(., pets_at_home = recode(pets_at_home, 'Sí' = 'Yes'),
has_dog = stringr::str_detect(type_number, '[dD]og'),
has_cat = stringr::str_detect(type_number, '[cC]at'),
where_live = recode(where_live, 'Adentro' = 'Indoors')) %>%
dplyr::arrange(., child_age_grp)
```
### Homes with dogs
```{r homes-w-dogs}
pets_cleaned %>%
dplyr::filter(., has_dog == TRUE) %>%
kableExtra::kbl(.) %>%
kableExtra::kable_styling(., bootstrap_options = c("striped", "hover", "condensed", "responsive")) %>%
kableExtra::scroll_box(width = "800px", height = "300px")
```
It looks like there are more homes with dogs.
### Homes with cats
```{r homes-w-cats}
pets_cleaned %>%
dplyr::filter(., has_cat == TRUE) %>%
kableExtra::kbl(.) %>%
kableExtra::kable_styling(., bootstrap_options = c("striped", "hover", "condensed", "responsive")) %>%
kableExtra::scroll_box(width = "800px", height = "300px")
```
Most cats and dogs live indoors.
This must be in NYC!
Also, there seem to be dog families and cat families, but not both.
## Typical day
```{r typical-day}
# Grab the asset_id for the data based on the name
typical_in_name <-
stringr::str_detect(play_survey_files$name, 'typical')
typical_asset_id <- play_survey_files$asset_id[typical_in_name]
# Select and download the data
typical_day_data <-
databraryapi::read_csv_data_as_df(params$play_survey_session_id, as.numeric(typical_asset_id))
# Print the data
typical_day_data %>%
kableExtra::kbl(.) %>%
kableExtra::kable_styling(.,
bootstrap_options = c("striped", "hover", "condensed", "responsive")) %>%
kableExtra::scroll_box(width = "800px", height = "300px")
```
We'll do a little clean-up and select a smaller number of columns.
```{r clean-typical-day}
typical_cleaned <- typical_day_data %>%
dplyr::mutate(
.,
typical_day = recode(typical_day, 'Sí' = 'Yes'),
activities_similar = recode(activities_similar, 'Sí' = 'Yes'),
typical_night_morning = recode(typical_night_morning, 'Sí' = 'Yes'),
unusual_feelings = recode(unusual_feelings, 'Sí' = 'Yes')
) %>%
dplyr::select(
.,
play_site_id,
child_age_grp,
typical_day,
activities_similar,
typical_night_morning,
unusual_feelings
) %>%
dplyr::arrange(., child_age_grp)
# Print the data
typical_cleaned %>%
kableExtra::kbl(.) %>%
kableExtra::kable_styling(.,
bootstrap_options = c("striped", "hover", "condensed", "responsive")) %>%
kableExtra::scroll_box(width = "800px", height = "300px")
```
Note that the `unusual_feelings` variable might need to be examined more closely.
This short name might not reflect the actual question properly.
```{r}
typical_day %>%
dplyr::select(., child_age_grp, typical_day) %>%
dplyr::filter(., !is.na(typical_day)) %>%
dplyr::mutate(., typical_day = recode(typical_day, 'Sí' = 'Yes')) %>%
dplyr::group_by(child_age_grp) %>%
dplyr::summarise(n_families = n(),
n_rating_typical = sum(typical_day == 'Yes')) %>%
kableExtra::kbl(.) %>%
kableExtra::kable_styling(.,
bootstrap_options = c("striped", "hover", "condensed", "responsive")) %>%
kableExtra::scroll_box(width = "800px", height = "250px")
```
## Household labor
```{r household-labor}
# Grab the asset_id for the data based on the name
labor_in_name <-
stringr::str_detect(play_survey_files$name, 'labor')
household_labor_asset_id <- play_survey_files$asset_id[labor_in_name]
# Select and download the data
household_labor <-
databraryapi::read_csv_data_as_df(params$play_survey_session_id, as.numeric(household_labor_asset_id))
# Print the data
household_labor %>%
kableExtra::kbl(.) %>%
kableExtra::kable_styling(.,
bootstrap_options = c("striped", "hover", "condensed", "responsive")) %>%
kableExtra::scroll_box(width = "800px", height = "300px")
```
## Locomotor milestones
```{r loco-milestones}
# Grab the asset_id for the data based on the name
loco_in_name <-
stringr::str_detect(play_survey_files$name, 'loco')
loco_milestones_asset_id <- play_survey_files$asset_id[loco_in_name]
# Select and download the data
loco_milestones <-
databraryapi::read_csv_data_as_df(params$play_survey_session_id, as.numeric(loco_milestones_asset_id))
# Print the data
loco_milestones %>%
# Drop comments here for privacy...
dplyr::select(., -c('WHO_walk_comments', 'walk_onset_comments', 'crawl_onset_comments')) %>%
kableExtra::kbl(.) %>%
kableExtra::kable_styling(.,
bootstrap_options = c("striped", "hover", "condensed", "responsive")) %>%
kableExtra::scroll_box(width = "800px", height = "300px")
```
Clean-up to focus on milestone ages.
```{r}
loco_cleaned <- loco_milestones %>%
dplyr::select(.,
child_age_grp,
child_sex,
WHO_walk_month,
walk_onset_month,
crawl_onset_month) %>%
tidyr::pivot_longer(
.,
cols = c('WHO_walk_month', 'walk_onset_month', 'crawl_onset_month'),
names_to = "milestone",
values_to = "age_mos"
)
loco_cleaned %>%
ggplot(.) +
aes(x = age_mos, fill = child_sex) +
geom_histogram() +
facet_grid(milestone ~ .)
```
## Health
```{r health}
# # Grab the asset_id for the data based on the name
# health_in_name <-
# stringr::str_detect(play_survey_files$name, 'health')
# health_asset_id <- play_survey_files$asset_id[health_in_name]
#
# # Select and download the data
# health <-
# databraryapi::read_csv_data_as_df(params$play_survey_session_id, as.numeric(health_asset_id))
#
# # Print the data
# health %>%
# kableExtra::kbl(.) %>%
# kableExtra::kable_styling(.,
# bootstrap_options = c("striped", "hover", "condensed", "responsive")) %>%
# kableExtra::scroll_box(width = "800px", height = "300px")
```
## Media
```{r media}
# Grab the asset_id for the data based on the name
media_in_name <-
stringr::str_detect(play_survey_files$name, 'media')
media_asset_id <- play_survey_files$asset_id[media_in_name]
# Select and download the data
media <-
databraryapi::read_csv_data_as_df(params$play_survey_session_id, as.numeric(media_asset_id))
# Print the data
media %>%
kableExtra::kbl(.) %>%
kableExtra::kable_styling(.,
bootstrap_options = c("striped", "hover", "condensed", "responsive")) %>%
kableExtra::scroll_box(width = "800px", height = "300px")
```
Do some clean-up to summarize some of the findings.
```{r media-summarize}
media %>%
dplyr::select(., child_age_grp, have_tv, child_used_tv) %>%
dplyr::filter(., !is.na(child_used_tv)) %>%
dplyr::mutate(., child_used_tv = recode(child_used_tv, 'Sí' = 'Yes')) %>%
dplyr::group_by(child_age_grp) %>%
dplyr::summarise(homes_w_tv = sum(have_tv),
children_use_tv = sum(child_used_tv == "Yes")) %>%
kableExtra::kbl(.) %>%
kableExtra::kable_styling(.,
bootstrap_options = c("striped", "hover", "condensed", "responsive")) %>%
kableExtra::scroll_box(width = "800px", height = "250px")
```
# Clean-up
Log out of Databrary.
```{r logout-databrary}
databraryapi::logout_db()
```