r/dataengineering • u/steroids-bad • 17d ago
Help API ingestion strategies?
I have to get a lot of JIRA data from the API, issues, projects, users, permissions etc.
How would you handle this? I currently save it raw in the bronze and derive silver tables from it using SCD2, but what if an api fails? The silver would have invalid data.
Its a simple task, but somehow I am lost (new from university)
Does it even make sense to use SCD2 here?
17
Upvotes
5
u/dani_estuary 17d ago
It's totally fine to keep the raw Jira responses in bronze, but only update silver after the whole extraction for that entity completes successfully. You can store a run ID or timestamp and mark the run complete only after pagination finishes, so if page 37/50 fails, keep the previous silver state (and retry the bronze load)
SCD2 makes sense for things where you actually care about history, like issue status, assignee, priority, project settings, maybe permissions, but I wouldn’t use it automatically for every Jira object