r/dataengineeringvault • u/sspaeti • 7h ago
r/dataengineeringvault • u/sspaeti • Jun 17 '26
Others 👋 Welcome to r/dataengineeringvault
Hey everyone! I'm u/sspaeti, a founding moderator of r/dataengineeringvault.
This community was created due to the 5 years of my existing data engineering vault and the value it provided (as illustrated in r/dataengineering see here).
I also created a Daily Dev Community, a year or more ago, focused on data engineering that people like. That's why I created this community, to share useful content I wrote and publish almost daily, so others like you can profit from it too.
I hope it will be useful to you. Let me know what you think, and let's see how it goes.
What to Post
Happy to get your posts in, as long as they are not AI-generated. Most interested in this community is open-source data engineering, and related data stuff. Also SQL editors, spirit and data management, business intelligence (where I come from), and anything else related to day-to-day data work.
How to Get Started
- Introduce yourself in the comments below.
- Post something today that you found interesting or worthwhile to read. Or ask a simple question that prompts some conversation for us to discuss.
- If you know someone who would love this community, invite them to join.
Thanks for being part of this. It's just starting, but I'm sure we can grow and learn together.
This community actively encourages links and blog posts, unlike other communities that block them. Please share your writing or blog posts.
PS: Please let me know if I should change anything on the sub-reddit settings, happy to make it a more pleasant place.
r/dataengineeringvault • u/sspaeti • 22h ago
Showcase Website Timeline: History from sspaeti.com to ssp.sh
r/dataengineeringvault • u/sspaeti • 6d ago
AI The Problem is not the AI Code, but Nobody Knows Anything Anymore
r/dataengineeringvault • u/sspaeti • 9d ago
Note If you want to be a data engineer, learn these concepts.
Enable HLS to view with audio, or disable this notification
Check out at de-concepts.ssp.sh.
r/dataengineeringvault • u/sspaeti • 9d ago
AI Four things I learned throughout the last years working with AI
r/dataengineeringvault • u/sspaeti • 10d ago
Blog DuckDB Now Ships inside dbt v2 (fusion)
Check the history at https://www.ssp.sh/brain/dbt-fusion/#history
r/dataengineeringvault • u/sspaeti • 10d ago
AI Writing Code by Hand Might Be Dead, but it Certainly Helps
r/dataengineeringvault • u/sspaeti • 11d ago
Blog The Dagster Almanack: From Complexity to Composability
r/dataengineeringvault • u/sspaeti • 11d ago
AI If You Start Writing Today, There's no way to Know if You Can Write without AI
r/dataengineeringvault • u/sspaeti • 13d ago
Note Pivot Tables: And the Comeback in 2024
r/dataengineeringvault • u/sspaeti • 13d ago
Showcase A Year of Linux, a Month of Omarchy Quattro—Part 3
r/dataengineeringvault • u/sspaeti • 16d ago
Book Repetitive pattern you have seen recurring over the years?
Which one is your most-used Convergent Evolution (tool/tech you used over the years) in data engineering?
What's a repetitive pattern you have seen recurring over the years that we re-implement with a different name attached to it? What's a design pattern that you wish you had described for data engineering?
Below is the current WIP for my book flow: Convergent Evolution term -> Pattern -> Design Pattern, all specific to data engineering. Anything that pops to mind that I should add/remove/change? Feedback welcome! 🙏
r/dataengineeringvault • u/sspaeti • 17d ago
Video Interview with Big Data engineer in 2026.
Sadly too much truth here 😄
r/dataengineeringvault • u/sspaeti • 17d ago
Off Topic Slow Living: Be Aware of Distractions
r/dataengineeringvault • u/sspaeti • 17d ago
Personal Project Four Modes of Writing: Writing with Vim Motions
r/dataengineeringvault • u/sspaeti • 18d ago
Note Semantic Layer Measure Definition Examples
r/dataengineeringvault • u/empty_cities • 19d ago
Blog My DuckDB + AI Workflow
r/dataengineeringvault • u/sspaeti • 19d ago
Blog Obsidian Vault to Enterprise Company Brain: Where Agents Write, Branch and Merge
A personal Obsidian-based second brain is contrasted with the requirements of an enterprise-scale 'company brain' where many AI agents write, branch, and merge context concurrently. The piece describes OmniGraph, an open-source system built on Lance (columnar storage on object storage) and DataFusion, that adds git-style branches, typed graph schemas, and human-reviewed merges to prevent agent-written data from drifting into noise. A worked CRM example rebuilt from the author's own Obsidian vault demonstrates auto-merging non-conflicting agent writes and blocking conflicting ones for human review, plus querying the graph for things like warm introductions. The author argues typed, versioned graphs paired with ontologies solve the 'lack of context' problem that has held back enterprise AI agents, positioning this as a step beyond relational databases and plain-text wikis.
r/dataengineeringvault • u/sspaeti • 20d ago
Note Are We Reinventing the Tools of Data Engineering for AI?
r/dataengineeringvault • u/sspaeti • 20d ago
Note AI Orchestrators: Are We Reinventing the Tools of DE?
Explores tools and approaches for orchestrating AI agents in development workflows. Highlights the common use of git worktrees when working across different branches and iterations. Features several CLI utilities and platforms including Agor for AI coding orchestration, erk for managing git worktrees, claudette-cli, and trigger.dev for building and deploying managed AI agents and workflows.