scala apache-spark hiveql sqlperformance

Which is more efficient, max or order by desc limit 1 in HIVE using spark version 2

As Hive keeps the data in distributed, Which query will be more efficient out of below two, when we have not consider that column in partition by or in bucketing.

select max(stat_id) from stats_tbl ;
select stat_id from stats_tbl order by stat_id desc limit 1;

Solution

Definitely select max(stat_id) from stats_tbl because order by requires gathering (read "lots of shuffle") all the data into a single reducer (and that's why you have to supply a limit clause with it) which will be inefficient compared to an aggregate function that can be computed distributedly.

Math.Sin() gives incorrect value
How to run my python script when the sunOS is start booting
Express-session: not resetting cookie expiration on each request
Getting a stack overflow exception when normalizing a vector
Edit default summary function in R gives error for multiple variables
What was a For loop? Why isn't it needed in R?
How to use download button in shiny and save results in various formats (csv, texte, pdf, spss...)?
Why are there two assignment operators, `<-` and `->` in R?
lm()$assign: what is it?
How to get the value of list(...) in R and S functions
Design matrix for MLM from library(lme4) with fixed and random effects
how to generate elements not included in my sample
Create a matrix with gradually changing values without a for loop
Emacs ESS and S-plus ( S+ ) 8.1 compatability
How to lag date-index in a time-series in R?
Nonlinear regression in R / S
Calling R from S-Plus?