sql amazon-web-services aws-lambda etl aws-glue

Prevent files from being processed multiple times in AWS Glue

We are using glue for computing purposes. The data flow is happening like this landing->raw->stage->curated->Redshift.

However, when the everyday the data flows right -> the data is exactly getting doubled.

For example:

Aug 1: I have 100 records
Aug 2: I have 20 records

In Redshift, I would like to see 120 records at end of August 2. Instead of that, it is getting 220 records. Please refer me to a way to avoid this scenario.

Would like to retain partition based on the run date in both raw and stage.

Solution

It seems that you want to track files that have already been processed. You can prevent that by using the job bookmarking feature of Glue.

Math.Sin() gives incorrect value
How to run my python script when the sunOS is start booting
Express-session: not resetting cookie expiration on each request
Getting a stack overflow exception when normalizing a vector
Edit default summary function in R gives error for multiple variables
What was a For loop? Why isn't it needed in R?
How to use download button in shiny and save results in various formats (csv, texte, pdf, spss...)?
Why are there two assignment operators, `<-` and `->` in R?
lm()$assign: what is it?
How to get the value of list(...) in R and S functions
Design matrix for MLM from library(lme4) with fixed and random effects
how to generate elements not included in my sample
Create a matrix with gradually changing values without a for loop
Emacs ESS and S-plus ( S+ ) 8.1 compatability
How to lag date-index in a time-series in R?
Nonlinear regression in R / S
Calling R from S-Plus?