Ignoring corrupted Orc files when reading via Spark

I have multiple Orc files in HDFS with the following directory structure:

orc/
├─ data1/
│  ├─ 00.orc
│  ├─ 11.orc
├─ data2/
│  ├─ 22.orc
│  ├─ 33.orc

I am reading these files using Spark:

spark.sqlContext.read.format("orc").load("/orc/data*/")

The problem is one of the files is corrupted so I want to skip/ignore that file.

The only way I see is to get all the Orc files and validate(By reading them) one by one before passing it to Spark. But this way I will be reading the same files twice.

Is there any way I can avoid reading the files twice? Does Spark provide anything regarding this?

Solution

This will help you:

spark.sql("set spark.sql.files.ignoreCorruptFiles=true")

Math.Sin() gives incorrect value
How to run my python script when the sunOS is start booting
Express-session: not resetting cookie expiration on each request
Getting a stack overflow exception when normalizing a vector
Edit default summary function in R gives error for multiple variables
What was a For loop? Why isn't it needed in R?
How to use download button in shiny and save results in various formats (csv, texte, pdf, spss...)?
Why are there two assignment operators, `<-` and `->` in R?
lm()$assign: what is it?
How to get the value of list(...) in R and S functions
Design matrix for MLM from library(lme4) with fixed and random effects
how to generate elements not included in my sample
Create a matrix with gradually changing values without a for loop
Emacs ESS and S-plus ( S+ ) 8.1 compatability
How to lag date-index in a time-series in R?
Nonlinear regression in R / S
Calling R from S-Plus?