tidier package provides Apache Spark style window aggregation for R dataframes via mutate in dplyr flavor.
Example
Create a new column with average temp over last seven days in the same month.
set.seed(101)
airquality |>
# create date column
dplyr::mutate(date_col = lubridate::make_date(1973, Month, Day)) |>
# create gaps by removing some days
dplyr::slice_sample(prop = 0.8) |>
# compute mean temperature over last seven days in the same month
tidier::mutate(avg_temp_over_last_week = mean(Temp, na.rm = TRUE),
.order_by = date_col,
.by = Month,
.frame = range_between(lubridate::days(7), # 7 days before current row
lubridate::days(-1) # do not include current row
)
)
#> # A tibble: 122 × 8
#> Ozone Solar.R Wind Temp Month Day date_col avg_temp_over_last_week
#> <int> <int> <dbl> <int> <int> <int> <date> <dbl>
#> 1 10 264 14.3 73 7 12 1973-07-12 85.5
#> 2 NA 127 8 78 6 26 1973-06-26 75.4
#> 3 16 77 7.4 82 8 3 1973-08-03 81
#> 4 14 191 14.3 75 9 28 1973-09-28 71.8
#> 5 NA 138 8 83 6 30 1973-06-30 76.6
#> 6 NA 98 11.5 80 6 28 1973-06-28 75.8
#> 7 122 255 4 89 8 7 1973-08-07 83.7
#> 8 47 95 7.4 87 9 5 1973-09-05 92.5
#> 9 23 220 10.3 78 9 8 1973-09-08 90.7
#> 10 NA 286 8.6 78 6 1 1973-06-01 NaN
#> # ℹ 112 more rowsFeatures
-
mutatesupports-
.by(group by), -
.order_by(order by), -
.frame(window frame defined byrows_betweenorrange_between), -
.complete(whether to compute over incomplete window).
-
-
tidier::mutateis single-threaded. For heavy parallelization across many groups, users can combinetidyr::nest()with parallel map (e.g.furrr::future_map()) andtidyr::unnest().
Motivation
This implementation is inspired by Apache Spark’s windowspec, rows between and range between.