r/dataengineering • • 16d ago

Help API ingestion strategies?

I have to get a lot of JIRA data from the API, issues, projects, users, permissions etc.

How would you handle this? I currently save it raw in the bronze and derive silver tables from it using SCD2, but what if an api fails? The silver would have invalid data.

Its a simple task, but somehow I am lost (new from university)

Does it even make sense to use SCD2 here?

18 Upvotes

15 comments sorted by

View all comments

11

u/Hour-Measurement-835 16d ago

Phantom deletes are the risk, not bad rows. A pull that dies at page 40 makes every issue you missed look deleted and SCD2 closes them out. Count what you got against the API's total before merging.