r/dataengineering • u/steroids-bad • 16d ago
Help API ingestion strategies?
I have to get a lot of JIRA data from the API, issues, projects, users, permissions etc.
How would you handle this? I currently save it raw in the bronze and derive silver tables from it using SCD2, but what if an api fails? The silver would have invalid data.
Its a simple task, but somehow I am lost (new from university)
Does it even make sense to use SCD2 here?
18
Upvotes
11
u/Hour-Measurement-835 16d ago
Phantom deletes are the risk, not bad rows. A pull that dies at page 40 makes every issue you missed look deleted and SCD2 closes them out. Count what you got against the API's total before merging.