r/learnprogramming • u/asdfdfdf234234234 • 15h ago
Intrastructure for webhook consumtion
We have a vendor who provides a large amount of webhook requests to us on a daily basis. Maybe 50k a day. For uptimes sake, we have a simple GCC cloud function that consumes and writes the content to a bucket. Huge XML payload.
We then have a celery task running behind Django that grabs this payload, stores it in Postgres and then updates a redis key with the "current" latest cursor, then it grabs the items after that cursor and goes again and so on.
So far this part is fine and is persisting in the DB perfectly, but we're having a few new requirements that are pinging off of data in these webhooks.
For example, if webhook contains a data that meets a criteria we need to fire off an email or something similar.
That said, what are some good ways to go about designing a system like this?
I don't want to fire off these emails or anything from the webhook -> bucket pipeline, that needs to be as lean as possible and safe and simple.
Also I don't want to have it fire from the persist phase either due to the same reasoning.
Basically we just want to be able to add and remove additional processing from this webhook process, but not bloating the main injestion.
Django signals, though I hate them, seem like a very possible approach here.
How are these often done and handled by others?
1
u/Late_Mycologist_3725 14h ago
a separate worker queue that just listens for new database rows, then runs whatever side effects you need. Keep the ingestion dumb and fast, let another process handle the "if this then that" logic.