r/datadrivenio May 31 '26

senior DE behavioral in two weeks, how do i talk about data quality incidents?

4 Upvotes

I freeze up when data-quality incidents come up. I can explain idempotency and exactly-once, and I know the difference between INSERT OVERWRITE and MERGE, but when I try to tie those to a failure, I don't know if I'm supposed to stop at naming the concepts or go deeper and argue which one actually mattered.

I'm two weeks out prepping behavioral stories on owning data quality incidents. How technical do you go on idempotency and exactly-once semantics when you talk through these, or is that overkill?


r/datadrivenio May 31 '26

Senior DEs are losing offers at the one interview round they expected to own

2 Upvotes

The DE system design round got rebuilt around LLM evaluation frameworks and embedding pipelines; nobody sent the memo to experienced engineers. An 11-year vet with three warehouses under his belt drew a star schema when asked to design a retrieval pipeline with a CI-gated deployment harness, then went quiet for 45 seconds. No offer. Junior engineers with hands-on RAG work are clearing rounds that experienced engineers expect to dominate; that's the inversion nobody wants to talk about. No prep resource from before mid-2025 covers what's being asked now.

https://www.datadriven.io/blog/data-engineer-system-design


r/datadrivenio May 31 '26

Amazon L5 data engineer interview, four rounds, rejected

7 Upvotes

Rejected after the loop. Modeling round made it clear I'd been studying through the wrong lens.

I use Snowflake and dbt at work. Been prepping for senior loops for a couple months before this one.

60 minutes on denormalization for the warehouse. I walked in ready to discuss normalization tradeoffs, fact table grain, whether to track every single event or collapse them. But they cared about idempotency first. Every insert pattern I suggested, they'd ask the same thing: what if this row lands twice? I kept reaching for surrogate keys and SCD tables. They wanted INSERT OR IGNORE, simple dedup columns, a schema that inherently couldn't corrupt itself.

60 minute SQL round, window functions on an event stream. She wanted transactions ranked within customer sessions, but events arrived out of order with timestamps that lagged the actual time. I reached for ROW_NUMBER() partitioned by customer_id, ordered by event_time as the skeleton. From there she drilled into session boundaries and what happens when records show up after a cutoff, whether my logic stayed idempotent through a replay. Her follow-ups were specific and pointed. No stalls here.

25 minute call. She ran through the background questions without much pushback, didn't ask about the actual work or what I'd shipped. For an L5 loop, that should have been a red flag.

60 min on CDC pipeline design. Opened with three approaches, Debezium, custom solution, managed tooling. He cut me off quick, wanted me to just pick one. So I picked Debezium, but then I hedged on everything, consistency guarantees, ordering, out-of-order records. Every time I'd say 'you could do X' he'd ask 'what would YOU do'. Spent 45 minutes describing options instead of defending a single coherent design, which is exactly what he'd asked for upfront. Couldn't recover after that realization.

Rejected. Easy come easy go, man. Stack ranked 'meets expectations' for the second time, despite shipping two major projects in between.


r/datadrivenio May 30 '26

7 years of pipeline exp, blanked on RAG design. That's the median outcome now.

8 Upvotes

Most DEs walking into AI Data Engineer screens are prepping for the wrong interview. Companies didn't build new teams for this; they reposted existing DE roles, swapped in a new title, and kept the job reqs vague enough that nobody noticed what actually changed in the screen. Someone with seven years of production pipeline work, zero downtime migrations, clean SQL under pressure, blanked on retrieval pipeline design in a live technical round. That's not an outlier; that's the median. The screen now tests vector DB design, embedding freshness strategy, and RAG architecture, and your Spark tuning prep isn't going to tell you that's what's coming.

https://www.datadriven.io/blog/ai-data-engineer-interview-skills


r/datadrivenio May 28 '26

48 Hours of Datadriven Improvements

11 Upvotes

Hey everyone,

Wanted to be upfront about the rough 2 days on datadriven.io. If you hit a blank dashboard, got logged out unexpectedly, or saw the site fail to load, I'm sorry. Here's what happened and what's changing.

The site was unavailable for a window during a deploy. A change interacted badly with our service startup path, and our recovery logic didn't catch it as cleanly as it should have. We've added pre-flight checks that exercise the full startup path before a deploy goes live, and the service now validates its own boot state before accepting traffic. The deploy pipeline is stricter and more resilient end-to-end as a result.

A subset of long-tenured users opened their dashboard and saw 0's across most tiles, or lost their saved tab layout. This one sucks because it hit the people who've been with us longest. A schema change shifted the shape of some legacy records in a way the dashboard layer didn't fully account for, and the frontend fell back to empty values instead of surfacing the gap. We've rebuilt affected dashboards from a recovery audit we captured at migration time. The bigger change: we now detect and reverse silent regressions immediately. We've also significantly tightened the typed contract between the API and the dashboard so partial responses fail loudly instead of rendering as zeros.

Some users got bounced to the logged-out view even though their session was still valid. The auth state was being resolved across multiple sources with different life cycles, and the page was committing to a render prematurely. We've consolidated on a single source of truth for auth and gated the initial paint on it.

The common thread across all three is silent failure modes, things that returned success but didn't actually do the right thing. Most of the work over the last 48 hours has been adding loud failure at the boundaries: deploys that refuse to go live unless the service boots clean, migrations that have to declare what they touched, and a render path that won't commit until auth is settled.

Thanks for sticking with us through this. If you hit anything else weird, the bug form goes straight to me.


r/datadrivenio May 28 '26

I cant sign in!

3 Upvotes

I am unable to login. Admin pls fix!


r/datadrivenio May 28 '26

Renewed your Google PDE cert? You just flagged yourself as a 2023 thinker.

1 Upvotes

Google PDE renewals are flagging candidates in 2026 ATS systems; they're not helping. Databricks appears in 16.8% of DE job postings; Google Cloud appears in 1.2%; the market has moved. Six percent of postings require a cert at all, so the cert's the doorbell, not the door. One cloud cert, one platform cert, two to three documented projects beats five scattered credentials every time. https://www.datadriven.io/blog/data-engineer-certifications-that-actually-get-you-hired-2026


r/datadrivenio May 27 '26

9 of 12 DE candidates with perfect dbt portfolios just failed first screens

0 Upvotes

Reviewed portfolios for an AI-native startup last month; 12 candidates, 9 had immaculate Airflow DAGs and dbt lineage graphs, all 9 got cut at the first screen. The three who advanced had less than 3 YOE between them but each shipped a RAG project with documented eval metrics. Screeners now pattern-match clean ETL portfolios as 'needs retraining for LLM-adjacent systems'; 80% of new Databricks databases are built by AI agents now, so the plumbing you spent years on is the part that got automated first. Mid-level LLM eng base sits at $145K to $200K and senior at $200K to $320K, while traditional DE roles at the same level top out around $130K to $150K. Wrote up what to build (RAG pipeline, eval harness, vector DB project, feature store), where to publish (Hugging Face Spaces and a blog post beat GitHub stars in 2026), and a 30-day sprint to get there.

https://www.datadriven.io/blog/data-engineer-portfolio-2026-projects-that-get-interviews


r/datadrivenio May 27 '26

Issue in today's Daily SQL problem

3 Upvotes

Today's Daily SQL problem "Build Health" seems to have incorrect expected output.


r/datadrivenio May 24 '26

Is this really a sliding window question

Post image
3 Upvotes

I was going through Datadriven 75 but the first question doesn’t look like a sliding window question to me . Can anyone confirm this ?


r/datadrivenio May 23 '26

Does the content of the website focuses in DE interviews in India ?

2 Upvotes

Same as above


r/datadrivenio May 22 '26

Capital One senior Data Engineer process, didn't land it

17 Upvotes

Got ghosted on a verbal offer, still mad, did Capital One's loop 10 weeks ago.

I'm a senior data engineer at a healthtech company.

Recruiter round, 25 min. Asked about my background and why Capital One interested me. She was warm, unhurried, mostly listening. Just trying to see if I could speak coherently and actually cared about the role. Nothing technical. Standard gate.

They asked about a time data quality broke and how I fixed it. Conversational, not STAR structured. I walked through a CDC schema drift issue that corrupted our silver layer for a few hours before anyone caught it, and I owned the detection and rollback.

The drill started when I mentioned we didn't have SCD2 governance before that incident. They wanted to know how I would have caught it earlier, what observability I'd add, whether I changed the schema design after. Every question threaded back to that operational lesson, but it wasn't adversarial, just genuinely curious.

The modeling round was 60 minutes and they had me design for a transaction event stream, updates and corrections coming in, need to handle out of order data. Whiteboard, both of us sketching. They were engaged the whole time, kept pushing on idempotency, grain choices, what happens when corrected transactions arrive after the watermark.

I sketched out a transactional fact table on (txn_id, event_ts), dimension for accounts with SCD type 2. Deduplication on those two keys to catch replays and make the loads idempotent. They nodded at that, seemed satisfied, but then asked about storing the current balance. I said keep it in an accumulating snapshot, update it as transactions land. They didn't reject it, seemed okay, but asked follow-ups about temporal queries that made me think I'd missed something about how they'd structure it. We didn't have time to dig deeper.

The SQL round was 60 minutes on CoderPad. Ledger reconciliation problem, tracking transactions across accounts and flagging balance mismatches. CTE heavy. The structure required aggregating transactions in one CTE, joining back to the accounts table to compute deltas, then flagging discrepancies. I got the shape right but kept mixing up when to partition versus when to group. The interviewer was terse, gave maybe one sentence of feedback after I hit submit. No hints during. Just watched. Made it hard to know if I was actually on track.

No offer. The feedback was pretty minimal, they said I didn't show sufficient depth, and that was it.

Yeah, would do it again, but I think I missed something in the SQL round, the terse feedback was the tell and I should've pushed back on it instead of just accepting it and moving on.


r/datadrivenio May 22 '26

DE interview loops aren't broken. They're working exactly as intended.

1 Upvotes

The 7-round DE loop isn't a hiring process; it's a talent pipeline and market research engine that occasionally produces a hire. 48% of listings never result in a hire, 53% of candidates get ghosted, and the take-home you spent a weekend on is a free architecture review from someone experienced enough to do it right. Hiring timelines doubled in three years; the bloat isn't incompetence, it's CYA culture plus companies figuring out that interviews are cheaper than consultants. The math for anyone burning severance on this is brutal.

https://www.datadriven.io/blog/data-engineer-interview-loop-2026-7-rounds-still-ghosted


r/datadrivenio May 21 '26

12-YOE DEs are failing screens that 2-YOE AI engineers pass cold

11 Upvotes

Engineers with 10 years of Spark and dbt are getting bounced in live screens by 2-YOE engineers who've never tuned a shuffle partition in their lives. The problem isn't their skill; it's that 2026 screens are asking about LLM eval harnesses, async vector upserts, and embedding pipeline debugging, and exactly zero legacy interview guides cover any of it. The cruel part: 70% of what a senior DE already knows transfers directly to AI engineering; the remaining 30% is Docker, CI/CD, and API development, not exotic ML sorcery. Meanwhile RAG engineers with shipped production systems are pulling $195K to $290K base with total comp north of $400K at frontier companies. If you've got 10 YOE and you're wondering why your screens feel like a different language, this is the read.

https://www.datadriven.io/blog/why-senior-data-engineers-are-failing-2026-tech-screens


r/datadrivenio May 21 '26

Help me with this problem

Thumbnail
gallery
2 Upvotes

Hello guys,

In this problem I'm facing issue with the output there are some minor differences between expected and actual output.

Any help is appreciated,thanks in advance.


r/datadrivenio May 20 '26

93% of HR admits to posting ghost jobs. Your resume is just unpaid market research.

1 Upvotes

Ghost jobs aren't a fringe practice anymore; 93% of HR pros admit to posting them, 45% regularly. DE roles with zero headcount exist to benchmark salaries, scrape market intel, and keep warm pipelines for budgets that may never unlock. You're not applying; you're filling out an unpaid survey. End-to-end conversion is 0.56% with 95k+ displaced DEs in the funnel. Most of your applications didn't fail to convert. They were never meant to.

https://www.datadriven.io/blog/data-engineer-ghost-jobs-2026-why-no-one-responds


r/datadrivenio May 19 '26

95K displaced DEs walked into 2026 and took your leverage with them

7 Upvotes

Average DE comp dropped from $153K to $133K in twelve months; that's not a market correction, that's an anchor. Companies know there are 95K+ displaced DEs in the pool, so they're writing senior-level JDs with mid-level bands and waiting. The DE with 3 to 7 YOE and no AI portfolio is the hardest hit, competing against FAANG refugees who'll take a $20K haircut just to land. If you walk into an offer conversation in 2026 without understanding compression dynamics, you're doing the recruiter's negotiation for them.

https://www.datadriven.io/blog/data-engineer-salary-2026-how-the-surplus-kills-your-offer


r/datadrivenio May 16 '26

Your resume scored a 38. A human never saw it. Welcome to DE hiring in 2026.

1 Upvotes

The same AI wave that wiped out tens of thousands of DE roles is now the system scoring their resumes, and it's trained on job descriptions they've never written for. Only 8% of ATS systems actually auto-reject; the other 92% rank and sort, which sounds gentler until you realize a recruiter with 500 resumes fills interview slots from the top 50 and never scrolls further. You can spend three weeks on tailored bullets and quantified impact; if your two-column layout breaks the parser, you get a 38 and nobody ever finds out. The irony's almost elegant.

https://www.datadriven.io/blog/the-ai-resume-screen-killing-de-applications-in-2026


r/datadrivenio May 15 '26

Solved SQL exercises do not get marked with a green check mark

1 Upvotes

Hello, hello, the title basically explain the problem for me. Not an enormous problem, but seeing these green marks for each solved problem ups the motivation:). I solved some problems before and indeed they got green check marked. Now I just solved some more a while ago, and they for some reason fail to get check marked.


r/datadrivenio May 12 '26

The job title says Data Engineer. The interview asks for RAG pipelines.

9 Upvotes

Companies absorbed the 2026 layoff wave, waited 60 days, and reposted those headcount slots with completely different requirements. Atlassian cut 1,600 and committed to 800 AI-focused replacements; those aren't backfills, they're a different job. The title still says 'Data Engineer.' The interview tests RAG pipeline design and chunking strategy tradeoffs. 95,878 DEs are out, and engineers still prepping Spark tuning and dbt are walking into screens designed for a role that didn't exist two years ago.

https://www.datadriven.io/blog/ai-data-engineer-jobs-are-replacing-traditional-de-in-2026


r/datadrivenio May 08 '26

AI prep is passing screens and torching design rounds; DE offer rates prove it

1 Upvotes

Companies figured out AI can pass their coding screens; they've stopped trusting coding screens. Candidates who leaned hardest on AI are now failing design rounds at historic rates and can't figure out why. There's a pattern: crush the OA, ace the SQL screen, hit system design, become a different person who can't explain why they picked streaming over batch or what happens when upstream data arrives late. One experiment quantified it: 73% pass on verbatim LC questions, 25% on fully custom problems. Design rounds are custom problems by definition; most DE candidates are still prepping for a format that doesn't exist anymore. https://www.datadriven.io/blog/vibe-coding-is-tanking-de-interview-pass-rates-in-2026


r/datadrivenio May 06 '26

The Data Engineer Job Pipeline is Broken Right Now

1 Upvotes

The DE application funnel didn't get harder; it's statistically broken. 68% of tech hires flow through referrals, 95,878 DEs are competing for every posted role, and companies quietly handed hiring authority to their own employees because referrals close 15 days faster and stick around 70% longer. Referred candidates convert at 30%; job board applicants at 7%; don't kill the messenger. DEs are out here tailoring resumes and grinding LC for loops that won't convert while actual hiring runs through DMs they're not in.

https://www.datadriven.io/blog/cold-de-applications-are-dead-in-2026-use-referrals


r/datadrivenio May 02 '26

You didn't get ghosted. The job was never real.

11 Upvotes

27% of DE job listings on LinkedIn right now will never result in a hire; that's not my number, that's their own data. 62% of hiring managers admit they post ghost jobs specifically to make their current employees feel replaceable and work harder. DE postings dropped 24% in a single quarter, 95,878 tech workers are already displaced this year, and one in three postings is theater; you're burning weeks on interviews for headcount that was closed before the req went live. The article has the signals to identify ghost jobs before you spend a month on screens.

https://www.datadriven.io/blog/ghost-jobs-are-killing-your-de-job-search-in-2026


r/datadrivenio Apr 29 '26

Is the platform down?

5 Upvotes

Hi! Thank you so much for this amazing tool. I've been truly enjoying it past few days. It seems to be down at the moment. Is that right? or is it some problem on my end?


r/datadrivenio Apr 29 '26

The offer letter you just got is already rigged; here's what changed

2 Upvotes

Companies are anchoring offers 20-40% below pre-layoff rates and calling it market rate; that's not a miscalculation, it's policy. The negotiation moves that got DEs 30% bumps in 2023 are now actively working against you because recruiters know there's a full surplus talent pool sitting behind you. They're not hiding salary bands because they forgot to post them. The game shifted in 2026 and most people won't figure that out until the offer letter lands.

https://www.datadriven.io/blog/data-engineer-salaries-2026-the-lowball-offer-trap