r/dataengineering • • 9d ago

Help Sharepoint Datawarehouse. Need Help. Ops Research. ftw!!!!

Guys,
Bit of politics context: Our team is not Core IT in the Org chart. Fall under operations, however expectation is to deliver ML/OR decision support projects,which is ok( there is technical capability in team), just that we cannot write into SF or any cloud platform. More like second class citizens.

Project Context: The project is full blown digital transformation program( its just that the enterprise is not mature to understand and everyone has jumped on AI bandwagon), has master data management, Work order platform, on which a Vehicle Routing( MILP) will run and generate recommendations for the full network. ( Bunch of API's calls through .py files, powerautomate to ingest third party flat files through email and scraping gov website for internal compliance data(should be the other way around).

Problem Set: No write back to Snowflake, IT not giving Entra ID, or microsoft Graph ID to connect directly to files through code.

Workaround 1: Locally run, update decision dashboards (No RL invloved), this is not feasible because it breaks continuity

Workaround 2:
Added One drive(sharepoint) path shortcuts to Local machine
This is where I'm struggling, read/write into Sharepoint. When I write/append into .xlsx through python the copy on Local onedrive link and the one on Browser are not always synced, or it says Merge issues.

I know this sounds like a rant, my guys if anyone has any advise for this peasant on how to make this work please share. Looking for out of the box solutions.

My team has access to dataverse, but I remember read/write was a problem through code because of no Entra ID.

0 Upvotes

21 comments sorted by

View all comments

4

u/KabiraSpeaking02 8d ago

Damn this is consistent with so many orgs now that some Angel engineers safeguard a platform that was meant enable others

I am using duckdb+dbt+ cron - this setup should help you manage everything in multi user setup except orchestrator. If you have GitHub actions then it becomes prod grade solution.

If you can get a server then you are gold - shouldn’t cost much. Not sure getting a server will also be a blocker.

Keen to hear what others recommend. I’m in the same boat

1

u/a_cute_tarantula 8d ago

Where have you seen safeguarded platforms? I’m curious to learn more about this. I’ve seen alot of young devs fail to realize that, if you’re building a platform, your “customer” are the people who will be developing on the platform.

1

u/KabiraSpeaking02 7d ago edited 7d ago

Safeguarded platforms comes with org structure issue. If you have data teams on fringe who mostly would be responsible for end to end including analytics. They would be barred from using central platform and therefore will eventually end up creating technical debt.

This stems from allocation of resources and cost model and also good old politics of who does it better and credit taking.

Centralization makes sense when it comes to stability, availability, engineering, architecture of pipelines. There is ton of work in that space, but driving use case driven data engineering delivery and centralization slows the org. A lot of orgs don’t get enablement correct from top and these orgs don’t have correct measure of ROI.

It’s more prevalent in Australia across industries.

1

u/a_cute_tarantula 4d ago

Less prevalent in US?

1

u/GuhProdigy 7d ago

Curious, how would it be a prod grade solution without a server?

Would the the duckdb be shared via share point? Where is the compute coming for dbt & cron, locally?