10.3 Add Columns

20200814 Adding columns to a dataset is accomplished with dplyr::mutate(). Within a pipeline dplyr::mutate() will modify the data as it passes through.

In the example below we add two new columns to the dataset. A tee pipe is used to print a sample of the resulting dataset using dplyr::select() and dplyr::sample_frac() within the curly braces. The ongoing pipe assigns the result into a new variable.

ds %>%
  mutate(range_temp=max_temp-min_temp,
         describe_temp=case_when(max_temp > 30 ~ "hot",
                                 max_temp > 20 ~ "mild",
                                 max_temp >  0 ~ "cold",
                                 TRUE          ~ "freezing")) %T>%
  {
    select(., date, location, ends_with("_temp")) %>%
    sample_frac() %>%
    print()
  } ->
newds
## # A tibble: 275,410 × 6
##    date       location     min_temp max_temp range_temp describe_temp
##    <date>     <chr>           <dbl>    <dbl>      <dbl> <chr>        
##  1 2020-06-13 Albury            3.6     16.4       12.8 cold         
##  2 2023-02-24 Katherine        NA       NA         NA   freezing     
##  3 2019-01-30 Adelaide         17.3     40.2       22.9 hot          
##  4 2018-12-26 Perth            16.8     28.2       11.4 mild         
##  5 2020-02-22 Canberra         13.9     26.3       12.4 mild         
##  6 2010-06-17 CoffsHarbour     10.1     19          8.9 cold         
##  7 2023-01-19 NorahHead        20.1     21.1        1   mild         
##  8 2025-09-13 Albury            3.1     18.8       15.7 cold         
##  9 2026-01-24 WaggaWagga       17.8     37.8       20   hot          
## 10 2011-09-29 WaggaWagga       10       13.7        3.7 cold         
## # ℹ 275,400 more rows

To overwrite the original dataset instead of saving it as a new dataset, replace the first pipe with an assignment pipe magrittr::%<>%.



Your donation will support ongoing availability and give you access to the PDF version of this book. Desktop Survival Guides include Data Science, GNU/Linux, and MLHub. Books available on Amazon include Data Mining with Rattle and Essentials of Data Science. Popular open source software includes rattle, wajig, and mlhub. Hosted by Togaware, a pioneer of free and open source software since 1984. Copyright © 1995-2022 Graham.Williams@togaware.com Creative Commons Attribution-ShareAlike 4.0