10.50 Wind Directions

20180723 The three wind direction variables (wind_gust_dir, wind_dir_9am, wind_dir_3pm) are also identified as character. We review the distribution of values here with dplyr::select() identifying any variable that tidyselect::contains() the string _dir and then build a base::table() over those variables.

# Review the distribution of observations across levels.

ds %>%
  select(contains("_dir")) %>%
  sapply(table)
##     wind_gust_dir wind_dir_9am wind_dir_3pm
## N           16541        21256        16216
## NNE         12612        15373        12754
## NE          13743        14276        15672
## ENE         15575        14991        14845
## E           17553        17767        15694
## ESE         14533        15250        16360
## SE          17929        17753        19749
## SSE         16995        17509        17123
## S           17588        16357        18466
## SSW         17296        14577        16147
## SW          16601        15771        16969
## WSW         16628        12998        17570
## W           18393        15463        18604
## WNW         15387        14196        16699
## NW          15161        15689        15663
## NNW         12409        14730        14473

Observe all 16 compass directions are represented and it would make sense to convert this into a factor. Notice that the directions are in alphabetic order and conversion to factor will retain that. Instead we can construct an ordered factor to capture the compass order (from N, NNE, to NW and NNW). We note the ordering of the directions here.

# Levels of wind direction are ordered compas directions.

compass <- c("N", "NNE", "NE", "ENE",
             "E", "ESE", "SE", "SSE",
             "S", "SSW", "SW", "WSW",
             "W", "WNW", "NW", "NNW")


Your donation will support ongoing availability and give you access to the PDF version of this book. Desktop Survival Guides include Data Science, GNU/Linux, and MLHub. Books available on Amazon include Data Mining with Rattle and Essentials of Data Science. Popular open source software includes rattle, wajig, and mlhub. Hosted by Togaware, a pioneer of free and open source software since 1984. Copyright © 1995-2022 Graham.Williams@togaware.com Creative Commons Attribution-ShareAlike 4.0