How to find weighted sum on top of groupby in pyspark dataframe?

I have a dataframe where i need to first apply dataframe and then get weighted average as shown in the output calculation below. What is an efficient way in pyspark to do that?

data = sc.parallelize([
[111,3,0.4],
[111,4,0.3],
[222,2,0.2],
[222,3,0.2],
[222,4,0.5]]
).toDF(['id', 'val','weight'])
data.show()


+---+---+------+
| id|val|weight|
+---+---+------+
|111|  3|   0.4|
|111|  4|   0.3|
|222|  2|   0.2|
|222|  3|   0.2|
|222|  4|   0.5|
+---+---+------+

Output:

id  weigthed_val
111 (3*0.4 + 4*0.3)/(0.4 + 0.3)
222 (2*0.2 + 3*0.2+4*0.5)/(0.2+0.2+0.5)

Solution

You can multiply columns weight and val, then aggregate:

import pyspark.sql.functions as F
data.groupBy("id").agg((F.sum(data.val * data.weight)/F.sum(data.weight)).alias("weighted_val")).show()

+---+------------------+
| id|      weighted_val|
+---+------------------+
|222|3.3333333333333335|
|111|3.4285714285714293|
+---+------------------+

Logging using Logback on Spark StandAlone
How to properly checkpoint a dataframe in PySpark
How to construct Dataframe from a Excel (xls,xlsx) file in Scala Spark?
Is there any preference on the order of select and filter in spark?
Spark: What is the difference between repartition and repartitionByRange?
Updating values in apache parquet file
Difference between ReduceByKey and CombineByKey in Spark
Task not serializable exception while running apache spark job
Apache Spark with Spring boot - failed to start exception Factory method 'javaSparkContext' threw exception with message: javax/servlet/Servlet
which is the best way to convert json into a dataframe?
How to read a file stored in adls gen 2 using pandas?
How to define partitioning of DataFrame?
Spark Shell: spark.executor.extraJavaOptions is not allowed to set Spark options
Read CSV with "§" as delimiter using Databricks autoloader
What are the benefits of Apache Beam over Spark/Flink for batch processing?
How do I update a Spark setting in SparkR?
How to handle accented letter in Pyspark
difference between spark.kubernetes.driver.request.cores, spark.kubernetes.driver.limit.cores and spark.driver.cores
Pyspark Streaming data to Elastic search index from Kafka topic , running in Jupyter notebook, causing failure
Spark Send DataFrame as body of HTTP Post request
How to handle an AnalysisException on Spark SQL?
How convert a list into multiple columns and a dataframe?
PySpark Window functions: Aggregation differs if WindowSpec has sorting
Using rangeBetween considering months rather than days in PySpark
Pyspark replace strings in Spark dataframe column
Spark read from MongoDB and filter by objectId indexed field
How to specify file size using repartition() in spark
BloomFilter mergeInPlace() producing unexpected behavior
Spark reading from mutiple SQL databases in parallel
Spark partition size greater than the executor memory