How to find an average for a Spark RDD?

I have read that reduce function must be commutative and associative. How should I write a function to find the average so it conforms with this requirement? If I apply the following function to count an average for an RDD it will not count the average correctly. Could anyone explain what is wrong with my function?

I guess that it takes two elements say 1, 2 and applies the function to them like (1+2)/2. Then sums up the result with the next element, 3 and divides it by 2 etc.

val rdd = sc.parallelize(1 to 100)

rdd.reduce((_ + _) / 2)

Solution

you can also use PairRDD to keep track of sum of all elements together with counts of elements.

val pair = sc.parallelize(1 to 100)
.map(x => (x, 1))
.reduce((x, y) => (x._1 + y._1, x._2 + y._2))

val mean = pair._1 / pair._2

Math.Sin() gives incorrect value
How to run my python script when the sunOS is start booting
Express-session: not resetting cookie expiration on each request
Getting a stack overflow exception when normalizing a vector
Edit default summary function in R gives error for multiple variables
What was a For loop? Why isn't it needed in R?
How to use download button in shiny and save results in various formats (csv, texte, pdf, spss...)?
Why are there two assignment operators, `<-` and `->` in R?
lm()$assign: what is it?
How to get the value of list(...) in R and S functions
Design matrix for MLM from library(lme4) with fixed and random effects
how to generate elements not included in my sample
Create a matrix with gradually changing values without a for loop
Emacs ESS and S-plus ( S+ ) 8.1 compatability
How to lag date-index in a time-series in R?
Nonlinear regression in R / S
Calling R from S-Plus?