Writing functions for your data, not just your vectors
new_mean(x), prop(x), and your new_sd(x) all take a vector of numbers.
But most of what you actually do all day takes a dataset and a variable:
This is the “copied it more than twice” signal. Silly example, as you could do it all in one summarize() call, but let’s write a function.
Error in `summarize()`:
ℹ In argument: `mean = mean(variable, na.rm =
TRUE)`.
Caused by error:
! object 'income' not found
Well, that didn’t work.
Here’s what we know works:
But income isn’t an object in your environment – if you typed income in the console you’d get an error. It only means something inside nlsy.
{dplyr} does this on purpose. It’s called data masking, and it’s why you don’t have to write nlsy$income in the mean() call.{dplyr} looks for a column literally named variable (that’s what you have in your function body!), doesn’t find one, and gives up.{{ }}Wrap the argument in curly-curly (I didn’t make up this name, it’s actually what they call it!) to say “don’t look for a column called variable, look for whatever the user passed in”:
And because it takes the data first, it pipes:
{{ }} works anywhere {dplyr} doesIncluding .by =:
Use "{{ variable }}" inside a string, with := instead of =:
ggplot() too{gtsummary}| Characteristic | Overall N = 12,6861 |
Male N = 6,4031 |
Female N = 6,2831 |
|---|---|---|---|
| race_eth_cat | |||
| Hispanic | 2,002 (16%) | 1,000 (16%) | 1,002 (16%) |
| Black | 3,174 (25%) | 1,613 (25%) | 1,561 (25%) |
| Non-Black, Non-Hispanic | 7,510 (59%) | 3,790 (59%) | 3,720 (59%) |
| eyesight_cat | |||
| Excellent | 2,916 (35%) | 1,582 (38%) | 1,334 (31%) |
| Very good | 2,970 (35%) | 1,470 (35%) | 1,500 (35%) |
| Good | 1,794 (21%) | 792 (19%) | 1,002 (23%) |
| Fair | 632 (7.5%) | 267 (6.4%) | 365 (8.5%) |
| Poor | 132 (1.6%) | 47 (1.1%) | 85 (2.0%) |
| Unknown | 4,242 | 2,245 | 1,997 |
| age_bir | 23 (20, 28) | 25 (21, 29) | 22 (19, 27) |
| Unknown | 6,743 | 3,652 | 3,091 |
| 1 n (%); Median (Q1, Q3) | |||
| Characteristic | Overall N = 12,4471 |
Northeast N = 2,5501 |
North Central N = 2,9341 |
South N = 4,5681 |
West N = 2,3951 |
|---|---|---|---|---|---|
| race_eth_cat | |||||
| Hispanic | 1,968 (16%) | 388 (15%) | 153 (5.2%) | 529 (12%) | 898 (37%) |
| Black | 3,128 (25%) | 552 (22%) | 561 (19%) | 1,764 (39%) | 251 (10%) |
| Non-Black, Non-Hispanic | 7,351 (59%) | 1,610 (63%) | 2,220 (76%) | 2,275 (50%) | 1,246 (52%) |
| eyesight_cat | |||||
| Excellent | 2,870 (35%) | 586 (38%) | 738 (36%) | 998 (32%) | 548 (35%) |
| Very good | 2,929 (35%) | 542 (35%) | 749 (36%) | 1,103 (35%) | 535 (34%) |
| Good | 1,762 (21%) | 298 (19%) | 426 (21%) | 691 (22%) | 347 (22%) |
| Fair | 621 (7.5%) | 101 (6.5%) | 131 (6.3%) | 270 (8.7%) | 119 (7.6%) |
| Poor | 130 (1.6%) | 30 (1.9%) | 29 (1.4%) | 50 (1.6%) | 21 (1.3%) |
| Unknown | 4,135 | 993 | 861 | 1,456 | 825 |
| age_bir | 23 (20, 28) | 25 (21, 30) | 24 (20, 28) | 22 (19, 27) | 24 (20, 28) |
| Unknown | 6,593 | 1,502 | 1,448 | 2,372 | 1,271 |
| 1 n (%); Median (Q1, Q3) | |||||
You only have to write the formatting once! Changing the stratifying variable is now just one word, not a copy-paste.
{{ }} is a tidyverse featureIt works because dplyr, ggplot2, and gtsummary are built to understand it. Base R functions are not.
lm(), pass a string and build the formulareformulate() takes character predictors and a character response:
lm() is easy tooIt takes a vector, so you get multivariable models for free:
This is a corner of R that changed a lot, and models are frequently out of date here.
If you see aes_string(), !!sym(), .dots =, or standardise_call() in a suggestion, it’s giving you a solution from before {{ }} existed (2019).
Warning
Those approaches mostly still run. But aes_string() has been formally deprecated in ggplot2 since 3.4, and the “curly-curly” approach was created due to frustrations with those older approaches, so learn the easier way to do it!
The website has a script to download with these examples, plus some additional exercises.