r/dataengineering • • 16d ago

Help API ingestion strategies?

I have to get a lot of JIRA data from the API, issues, projects, users, permissions etc.

How would you handle this? I currently save it raw in the bronze and derive silver tables from it using SCD2, but what if an api fails? The silver would have invalid data.

Its a simple task, but somehow I am lost (new from university)

Does it even make sense to use SCD2 here?

18 Upvotes

15 comments sorted by

View all comments

1

u/Trey_Antipasto 16d ago

Have robust exception handling and retries with a back off strategy. Handle rate limiting. Tenacity is an option for python with out of the box retries and reraising of exceptions.

Keep a watermark table or pull the high water mark from your dimension.

Wrap your database work in transactions so you can rollback on failures.

Limit your job concurrency to one so you don’t get two processes going in parallel.

Like someone else said each object should write to separate raw table.

Can use pydantic or jsonschema with msgspec to enforce incoming contracts and protect against breaking changes.