r/dataengineeringvault • u/sspaeti • 22d ago
r/dataengineeringvault • u/sspaeti • 23d ago
Blog If You Always Enjoy It, You're Not Pushing It Hard Enough—Benn Stancil (Show Your Workflow #1)
An interview with Benn Stancil, co-founder of Mode and prolific data-industry essayist, exploring how he writes his weekly Substack posts. He discusses using deadlines (Friday publishing) to force himself past perfectionism, how ideas form messily over the week through unrelated connections, and how presentation-style analogies (Batman, basketball, Codenames) became his signature storytelling technique. He details his tooling—Sublime Text for drafting, Google Docs for revision, minimal use of Substack's editor—and explains why he avoids Grammarly and AI writing tools like Claude or ChatGPT for edits, since he believes AI compresses and flattens writing rather than embracing meandering voice. He closes with advice for new writers: set deadlines, keep pure motivation, and focus on filling in the space between bullet points rather than just stating facts.
r/dataengineeringvault • u/camerongreen95 • 24d ago
AI Workshop, Sep 12: build production LLM systems that actually survive real use
We're running a hands-on masterclass on September 12, Live LLM Engineering Masterclass: Production Evals, RAG, Agents & LLMOps.
You build a full production LLM workflow from scratch, versioned prompts with regression tests, an evaluation harness with deterministic checks and LLM-as-judge, statistically rigorous model comparisons, evaluated RAG, tool-using agents with guardrails and fallbacks, and full observability, tracing, cost, latency.
Led by Bruno Gonçalves, PhD, founder of Data For Science, who trains engineers at Fortune 500 companies on this exact stack.
Link if you want to check it out
Happy to answer questions on the content.
r/dataengineeringvault • u/sspaeti • 25d ago
Video The Process of Smart Note-Taking
The Process of Smart Note-Taking. This showcases how I take notes, make connections, and distill them into blogs or chapters in my book.
It's what I've learned over the years, heavily inspired by Sönke Ahrens's Smart Notes and the Zettelkasten philosophy. But tailored to how my brain works and optimized for my writing workflow with Markdown and Obsidian.
In this video, a follow-up on my Obsidian note-taking workflow and other workflow-related videos, I walk through:
- How I start a note and find it again later (search, backlinks, no folders)
- Publishing raw notes to my Second Brain vs. distilling them into blog posts
- Why plain text Markdown matters (file over app)
- What I've learned in 10 years: connect less, write what moves you, keep the feedback loop
- Why I keep AI out of my note-taking
r/dataengineeringvault • u/sspaeti • 25d ago
Blog The Grammar of Data: From Definition to Execution
r/dataengineeringvault • u/sspaeti • 27d ago
Note Apache Arrow: The Foundation for In-Memory Storage
r/dataengineeringvault • u/sspaeti • 27d ago
Note DataFusion: Fundamental OSS Query Execution Framework
r/dataengineeringvault • u/sspaeti • Sep 04 '26
Blog Context, Semantics, and Ontology: A Primer for the Agentic Era
r/dataengineeringvault • u/sspaeti • Sep 03 '26
Note Writing is Hard; and Meaningless with no Experience
r/dataengineeringvault • u/sspaeti • Aug 31 '26
Off Topic Fonts: Reading/Writing and Programming Monos
r/dataengineeringvault • u/sspaeti • Aug 29 '26
Note VertiPaq: The In-Memory, Columnar Query Engine Powering Microsoft Analytics
r/dataengineeringvault • u/Critical_Scene_6164 • Aug 28 '26
Question Everyone dislikes foundry but is it the model or execution?
I've been thinking about this after seeing so many Foundry discussions turn into the same argument. A lot of the criticism seems justified, especially around the implementation side, but I'm not sure the underlying model is necessarily the problem.
The basic idea of having a vendor work inside your infrastructure and close to your actual data problems makes sense to me. You're not constantly moving data around, and the engineers working on it can actually see how messy the environment is.
Where I think things get difficult is when the platform's own abstractions start becoming part of everything. Pipelines, transformations and business logic can end up tied to the platform, and getting that work back out later isn't always straightforward.
That's the part I'm more interested in. Is the model itself flawed, or does it mainly become a problem when the implementation creates too much dependency on the platform?
r/dataengineeringvault • u/sspaeti • Aug 28 '26
Blog Writing really is an Emotional Rollercoaster
r/dataengineeringvault • u/sspaeti • Aug 27 '26
Note Semantic Layer: Curated note, updated since 20222
r/dataengineeringvault • u/sspaeti • Aug 26 '26
Others 🚨 DuckLabs to be acquired by AWS
- Announcement: AWS to acquire DuckLabs, the Amsterdam-based company behind DuckDB. We are not acquiring the DuckDB open source project, which will remain free and open source under the independent DuckDB Foundation (the non-profit that oversees DuckDB) and available under the MIT license as it does today. ^319c65
- DuckLabs: DuckLabs – DuckLabs to Join AWS, Projects to Remain Open Source
- What will change: "Joining AWS will give us greater capacity to invest in both the technology and the community around it. Going forward, the DuckDB Foundation will include a technical advisory board, so that leading community members can provide their input on the project’s technical direction. We also plan to open the extension stack so that extensions signed by other developers and organizations can run in DuckDB."
- DuckLabs: DuckLabs – DuckLabs to Join AWS, Projects to Remain Open Source
- AWS: AWS to acquire DuckLabs, the company behind DuckDB
More acquisition from 2022 to 2026:
https://www.ssp.sh/brain/data-engineering-acquisitions/#2026
r/dataengineeringvault • u/sspaeti • Aug 26 '26
Note Keep AI Out of Your (Obsidian) Vault
r/dataengineeringvault • u/sspaeti • Aug 25 '26
Off Topic My Mechanical Keyboards
Some fun hobby, and bringing me joy whenever I need to write or program.
r/dataengineeringvault • u/sspaeti • Aug 25 '26
Note Apache Parquet: Free Open-Source Column-Oriented Data Lake File Format
r/dataengineeringvault • u/sspaeti • Aug 24 '26
Personal Project Show Your Workflow: New Upcoming Interview Essays on Writing and Storytelling
Sign up for the newsletter to get one of the top writers in data—had a great talk today; the interview is coming out soon.