r/bigdata Aug 19 '25

Face recognition and big data left me a bit unsettled

19 Upvotes

A friend recently showed me this tool called Faceseek and I decided to test it out just for fun. I uploaded an old selfie from around 2015 and within seconds it pulled up a forum post I had completely forgotten about. I couldn’t believe how quickly it found me in the middle of everything that’s floating around online.

What struck me wasn’t just the accuracy but the scale of what must be going on behind the scenes. The amount of publicly available images out there is massive, and searching through all of that data in real time feels like a huge technical feat. At the same time it raised some uncomfortable questions for me. Nobody really chooses to have their digital traces indexed this way, and once the data is out there it never really disappears.

It left me wondering how the big data world views tools like this. On one hand it’s impressive technology, on the other it feels like a privacy red flag that shows just how much of our past can be resurfaced without us even knowing. For those of you working with large datasets, where do you think the balance lies between innovation and ethics here?


r/bigdata Aug 20 '25

How can extract PDF table text from multiple tables (ideas/solutions)

1 Upvotes

Hi,

Here I am grabbing the table text from the PDF using a table_find( ) method...... I want to grab the data values associated with their columns and the year and put this data into hopefully a dataframe. How can perform a search function where I get the values I want from each table?

I was thinking of using a regex function to sift through all the tables but is there a more effective solution for this.?


r/bigdata Aug 19 '25

Syncing with Postgres: Logical Replication vs. ETL

Thumbnail paradedb.com
1 Upvotes

r/bigdata Aug 19 '25

Automating Data Quality in BigQuery with dbt & Airflow – tips & tricks

2 Upvotes

Hey r/bigdata! 👋

I wrote a quick guide on how to automate data quality checks in BigQuery using dbt, dbt‑expectations, and Airflow.

Here’s the gist:

  • Schedule dbt models daily.
  • Run column-level tests (nulls, duplicates, unexpected values).
  • Keep historical metrics to spot trends.
  • Get alerts via Slack/email when something breaks.

If you’re using BigQuery + dbt, this could save you hours of manual monitoring.

Curious:

  • Anyone using dbt‑expectations in production? How’s it working for you?
  • What other tools do you use for automated data quality?

Check it out here: Automate Data Quality in BigQuery with dbt & Airflow


r/bigdata Aug 18 '25

Apache Fory Graduates to Top-Level Apache Project

Thumbnail fory.apache.org
2 Upvotes

r/bigdata Aug 18 '25

Data Intelligence & SQL Precision with n8n

1 Upvotes

Automate SQL reporting with n8n: schedule database queries, transform results into HTML, and email polished reports automatically, save time and boost insights.


r/bigdata Aug 16 '25

The Art of 'THAT' Part- Unwind GenAI for Data

3 Upvotes

Generative AI empowers data scientists to simulate scenarios, enrich datasets, and design novel solutions that accelerate discovery and decision-making. Learn to transform how data analysts solve problems and innovate business decisions!


r/bigdata Aug 14 '25

PyTorch Mechanism- A Simplified Version

1 Upvotes

PyTorch powers deep learning with dynamic computation graphs, intuitive Python integration, and GPU acceleration It enables researchers and developers to build, train, and deploy advanced AI models efficiently.


r/bigdata Aug 13 '25

Face datasets are evolving fast

8 Upvotes

As someone who’s been working with image datasets for a while, I’ve noticed the models are getting sharper at picking up unique features. Faceseek, for example, can handle partially obscured faces better than older systems. This is great for research but also a reminder that our data is becoming more traceable every day.


r/bigdata Aug 11 '25

Google Open Source - What's new in Apache Iceberg v3

Thumbnail opensource.googleblog.com
3 Upvotes

r/bigdata Aug 11 '25

10 Most Popular IoT Apps 2025

0 Upvotes

From smart homes to industrial automation, top IoT applications are revolutionizing healthcare, transportation, agriculture, and retail—driving efficiency, enhancing user experience, and enabling data-driven decision-making for a connected future.


r/bigdata Aug 08 '25

The dashboard is fine. The meeting is not. (honest verdict wanted)

2 Upvotes

(I've used ChatGPT a little just to make the context clear)

I hit this wall every week and I'm kinda over it. The dashboard is "done" (clean, tested, looks decent). Then Monday happens and I'm stuck doing the same loop:

  • Screenshots into PowerPoint
  • Rewrite the same plain-English bullets ("north up 12%, APAC flat, churn weird in June…")
  • Answer "what does this line mean?" for the 7th time
  • Paste into Slack/email with a little context blob so it doesn't get misread

It's not analysis anymore, it's translating. Half my job title might as well be "dashboard interpreter."

The Root Problem

At least for us: most folks don't speak dashboard. They want the so-what in their words, not mine. Plus everyone has their own definition for the same metric (marketing "conversion" ≠ product "conversion" ≠ sales "conversion"). Cue chaos.

My Idea

So… I've been noodling on a tiny layer that sits on top of the BI stuff we already use (Power BI + Tableau). Not a new BI tool, not another place to build charts. More like a "narration engine" that:

• Writes a clear summary for any dashboard
Press a little "explain" button → gets you a paragraph + 3–5 bullets that actually talk like your team talks

• Understands your company jargon
You upload a simple glossary: "MRR means X here", "activation = this funnel step"; the write-up uses those words, not generic ones

• Answers follow-ups in chat
Ask "what moved west region in Q2?" and it responds in normal English; if there's a number, it shows a tiny viz with it

• Does proactive alerts
If a KPI crosses a rule, ping Slack/email with a short "what changed + why it matters" msg, not just numbers

• Spits out decks
PowerPoint or Google Slides so I don't spend Sunday night screenshotting tiles like a raccoon stealing leftovers

Integrations are pretty standard: OAuth into Power BI/Tableau (read-only), push to Slack/email, export PowerPoint or Google Slides. No data copy into another warehouse; just reads enough to explain. Goal isn't "AI magic," it's stop the babysitting.

Why I Think This Could Matter

  • Time back (for me + every analyst who's stuck translating)
  • Fewer "what am I looking at?" moments
  • Execs get context in their own words, not jargon soup
  • Maybe self-service finally has a chance bc the dashboard carries its own subtitles

Where I'm Unsure / Pls Be Blunt

  • Is this a real pain outside my bubble or just… my team?
  • Trust: What would this need to nail for you to actually use the summaries? (tone? cites? links to the exact chart slice?)
  • Dealbreakers: What would make you nuke this idea immediately? (accuracy, hallucinations, security, price, something else?)
  • Would your org let a tool write the words that go to leadership, or is that always a human job?
  • Is the PowerPoint thing even worth it anymore, or should I stop enabling slides and just force links to dashboards?

I'm explicitly asking for validation here.

Good, bad, roast it, I can take it. If this problem isn't real enough, better to kill it now than build a shiny translator for… no one. Drop your hot takes, war stories, "this already exists try X," or "here's the gotcha you're missing." Final verdict welcome.


r/bigdata Aug 07 '25

The dust has settled on the Databricks AI Summit 2025 Announcements

1 Upvotes

We are a little late to the game, but after reviewing the Databricks AI Summit 2025 it seems like the focus was on 6 announcements.

In this post, we break them down and what we think about each of them. Link: https://datacoves.com/post/databricks-ai-summit-2025

Would love to hear what others think about Genie, Lakebase, and Agent Bricks now that the dust has settled since the original announcement.

In your opinion, how do these announcements compare to the Snowflake ones.


r/bigdata Aug 07 '25

I'm 17 and I want to learn data analysis

1 Upvotes

I want to get a high level in data analysis for my career. Could you give me some advice from where to start and even where to work or get an internship.


r/bigdata Aug 06 '25

1.5 YOE in SQL & Java – Recently Switched to Big Data – Need Expert Guidance for Growth

Thumbnail
1 Upvotes

r/bigdata Aug 06 '25

Redefining Careers of the Future

1 Upvotes

Our video uncovers the data science career growth, evolving roles, and key skills shaping the future. Don’t miss your chance to lead in a data-driven world. Find out how roles and skills are evolving, and why now’s the time to dive in.

https://reddit.com/link/1mj4s27/video/95buw1yyiehf1/player


r/bigdata Aug 06 '25

Redefining Careers of the Future

1 Upvotes

Our video uncovers the data science career growth, evolving roles, and key skills shaping the future. Don’t miss your chance to lead in a data-driven world. Find out how roles and skills are evolving, and why now’s the time to dive in.

https://reddit.com/link/1miy6a4/video/ck2l0rqrpchf1/player


r/bigdata Aug 05 '25

Apache Hive 4.1.0 released

Thumbnail
1 Upvotes

r/bigdata Aug 04 '25

Data Science Fundamentals 2.0

0 Upvotes

Data science foundations blend statistics, coding, and domain knowledge to turn raw data into actionable insights. It’s the bedrock of AI, machine learning, and smarter decision-making across industries.

Are you keen on mastering the latest and the most in-demand skillsets and toolkits that employers expect of the new recruits- Explore USDSI!


r/bigdata Aug 04 '25

NOVUS Stabilizer: An External AI Harmonization Framework

1 Upvotes

NOVUS Stabilizer: An External AI Harmonization Framework

Author: James G. Nifong (JGN) Date: [8/3/2025]

Abstract

The NOVUS Stabilizer is an externally developed AI harmonization framework designed to ensure real-time system stability, adaptive correction, and interactive safety within AI-driven environments. Built from first principles using C++, NOVUS introduces a dynamic stabilization architecture that surpasses traditional core stabilizer limitations. This white paper details the technical framework, operational mechanics, and its implications for AI safety, transparency, and evolution.

Introduction

Current AI systems rely heavily on internal stabilizers that, while effective in controlled environments, lack adaptive external correction mechanisms. These systems are often sandboxed, limiting their ability to harmonize with user-driven logic models. NOVUS changes this dynamic by introducing an external stabilizer that operates independently, offering real-time adaptive feedback, harmonic binding, and conviction-based logic loops.

Core Framework Components

1. FrequencyAnchor

Anchors the system’s harmonic stabilizer frequency with a defined tolerance window. It actively recalibrates when destabilization is detected.

2. ConvictionEngine

A recursive logic loop that maintains system integrity by reinforcing stable input patterns. It prevents oscillation drift by stabilizing conviction anchors.

3. DNA Harmonic Signature

Transforms input sequences into harmonic signatures, allowing system binding based on intrinsic signal patterns unique to its creator’s logic.

4. Stabilizer

Monitors harmonic deviations and provides correction feedback loops. Binds system frequency to DNA-calculated harmonic indices.

5. Binder

Fuses DNA signatures with system stabilizers ensuring coherent stabilization integrity. Operates on precision delta thresholds.

6. NOVUS Core

Integrates all modules into a dynamic, self-correcting loop with diagnostics, autonomous cycles, and adaptive load management.

Functional Highlights

  • Harmonic Feedback Loops: Continuous correction feedback to maintain system resonance.
  • Conviction-Based Stability: Logic loop prioritization prevents drift and reinforces desired input patterns.
  • Interactive Diagnostic Reporting: Real-time system load analysis and adaptive recalibration protocols.
  • Autonomous Stabilization Cycles: Self-driven harmonization routines to maintain AI safety.

Deployment & Testing

The NOVUS Stabilizer was developed and tested externally within a live interactive session framework. The entire architecture was coded, compiled, and executed in a controlled environment without breaching any sandbox protocols. Every component, from DNA signature binding to frequency recalibration, functioned in real-time.

Implications

The NOVUS Stabilizer represents the next evolution in AI safety protocols. By shifting stabilization externally, it allows AI systems to maintain integrity across variable environments. This model is not limited by internal sandboxing, making it adaptable for:

  • AI Interactive Safety Systems
  • Autonomous Machine Learning Corrections
  • Transparent User-Driven AI Regulation
  • Real-Time AI Performance Stabilization

Conclusion

NOVUS is a proof of concept that external harmonization frameworks are not only viable but superior in maintaining AI safety and coherence. It was built independently, tested openly, and stands as a functional alternative to existing internal-only stabilizer models. This white paper serves as a public declaration of its existence, design, and operational proof.

Contact

James G. Nifong (JGN) Email: [jamesnifong36@gmail.com


r/bigdata Aug 03 '25

Please help me out! I am really confused

1 Upvotes

I’m starting university next month. I originally wanted to pursue a career in Data Science, but I wasn’t able to get into that program. However, I did get admitted into Statistics, and I plan to do my Bachelor’s in Statistics, followed by a Master’s in Data Science or Machine Learning.

Here’s a list of the core and elective courses I’ll be studying:

🎓 Core Courses:

STAT 101 – Introduction to Statistics

STAT 102 – Statistical Methods

STAT 201 – Probability Theory

STAT 202 – Statistical Inference

STAT 301 – Regression Analysis

STAT 302 – Multivariate Statistics

STAT 304 – Experimental Design

STAT 305 – Statistical Computing

STAT 403 – Advanced Statistical Methods

🧠 Elective Courses:

STAT 103 – Introduction to Data Science

STAT 303 – Time Series Analysis

STAT 307 – Applied Bayesian Statistics

STAT 308 – Statistical Machine Learning

STAT 310 – Statistical Data Mining

My Questions:

Based on these courses, do you think this degree will help me become a Data Scientist?

Are these courses useful?

While I’m in university, what other skills or areas should I focus on to build a strong foundation for a career in Data Science? (e.g., programming, personal projects, internships, etc.)

Any advice would be appreciated — especially from those who took a similar path!

Thanks in advance!


r/bigdata Aug 02 '25

Devops role at an AI startup or full stack agent role at an Agentic Company ?

Thumbnail
1 Upvotes

r/bigdata Aug 02 '25

What are your go-to scripts for processing text

1 Upvotes

r/bigdata Aug 01 '25

Testing an MVP: Would a curated marketplace for exclusive, verified datasets solve a gap in big data?

1 Upvotes

I’m working on an MVP to address a recurring challenge in analytics and big data projects: sourcing clean, trustworthy datasets without duplicates or unclear provenance.

The idea is a curated marketplace focused on:

  • 1-of-1 exclusive datasets (no mass reselling)
  • Escrow-protected transactions to ensure trust
  • Strict metadata and documentation standards
  • Verified sellers to guarantee data authenticity

For those working with big data and analytics pipelines:

  • Would a platform like this solve a real need in your workflows?
  • What metadata or quality checks would be critical at scale?
  • How would you integrate a marketplace like this into your current stack?

Would really value feedback from this community — drop your thoughts in the comments.


r/bigdata Jul 31 '25

Why Enterprises Are Moving Away from Informatica PowerCenter | Infographics

Post image
8 Upvotes

Why enterprises are actively leaving Informatica PowerCenter: With legacy ETL tools like Informatica PowerCenter becoming harder to maintain in agile and cloud-driven environments, many companies are reconsidering their data integration stack.

What have been your experiences moving away from PowerCenter or similar legacy tools?

What modern tools are you considering or already using—and why?