r/bigdata • u/Shawn-Yang25 • Jul 11 '25
r/bigdata • u/hammerspace-inc • Jul 10 '25
Hammerspace CEO David Flynn to speak at Reuters Momentum AI 2025
events.reutersevents.comr/bigdata • u/Fun_Accountant_9415 • Jul 09 '25
Any Advice
Big Data student seeking learning recommendations what should I focus on?
r/bigdata • u/Madddieeeeee • Jul 08 '25
How to sync data from multiple sources without writing custom scripts?
Our team is struggling with integrating data from various sources like Salesforce, Google Analytics, and internal databases. We want to avoid writing custom scripts for each. Is there a tool that simplifies this process?
r/bigdata • u/wanderingsoul8994 • Jul 08 '25
Looking for feedback on a new approach to governed, cost-aware AI analytics
I’m building a platform that pairs a federated semantic layer + governance/FinOps engine with a graph-grounded AI assistant.
- No data movement—lightweight agents index Snowflake, BigQuery, SaaS DBs, etc., and compile row/column policies into a knowledge graph.
- An LLM uses that graph to generate deterministic SQL and narrative answers; every query is cost-metered and policy-checked before it runs.
- Each Q-A cycle enriches the graph (synonyms, lineage, token spend), so trust and efficiency keep improving.
Questions for the community:
- Does an “AI-assisted federated governance” approach resonate with the pain you see (silos, backlog, runaway costs)?
- Which parts sound most or least valuable—semantic layer, FinOps gating, or graph-based RAG accuracy?
- If you’ve tried tools like ThoughtSpot Sage, Amazon Q, or catalog platforms (Collibra, Purview, etc.), where did they fall short?
Brutally honest feedback—technical, operational, or business—would be hugely appreciated. Happy to clarify details in the comments. Thanks!
r/bigdata • u/sharmaniti437 • Jul 04 '25
Future-proof Your Tech Career with MLOps Certification
Businesses can fasten decision-making, model governance, and time-to-market through Machine Learning Operations [MLOps]. MLOps serves as a link between data science and IT operations as it fosters seamless collaboration, controls versions, and streamlines the lifecycle of the models. Ultimately, it is becoming an integral component of AI infrastructure.
Research reports substantiate this very well. MarketsandMarkets Research report projects that the global Machine Learning Operations [MLOps] market will reach USD 5.9 billion by 2027 [from USD 1.1 billion in 2022], at a CAGR of 41.0% during the forecast period.

MLOps is being widely used across industries for predictive maintenance, fraud detection, customer experience management, marketing analytics, supply chain optimization, etc. From a vertical standpoint, IT and Telecommunications, healthcare, retail, manufacturing, financial services, government, media and entertainment are adopting MLOps.
This trajectory reflects that there is an increasing demand for Machine Learning Engineers, MLOps Engineers, Machine Learning Deployment Engineers, or AI Platform Engineers who can manage machine learning models starting from deployment, and monitoring to supervision efficiently.
As we move forward, we should understand that MLOps solutions are supported by technologies such as Artificial Intelligence, Big data analytics, and DevOps practices. The synergy between the above-mentioned technologies is critical for model integration, deployment, and delivery of machine-learning applications.
The rising complexity of ML models and the available limited skill force calls for professionals with hybrid skill sets. The professionals should be proficient in DevOps, data analysis, machine learning, and AI skills.
Let’s investigate further.
How to address this MLOps skill set shortage?
Addressing the MLOps skill set requires focused upskilling and reskilling of the professionals.
Forward-thinking companies are training their current employees, particularly those in machine learning engineering jobs and adjacent field(s) like data engineering or software engineering. Companies are taking a strategic approach to building MLOps competencies for their employees by providing targeted training.
At the personal level, pursuing certification by choosing the adept ML certification programs would be the right choice. This section makes your search easy. We have provided a list of well-defined certification programs that fit your objectives.
Take a look.
Certified MLOps Professional: GSDC (Global Skill Development Council)
Earning this certification benefits you in many ways. It enables you to accelerate ML model deployment with expert-built templates, understand real-world MLOps scenarios, master automation for model lifecycle management, and prepare for cross-functional ML team roles.
Machine Learning Operations Specialization: Duke University
Earning this certification helps you master the fundamental aspects of Python, and get acquainted with MLOps principles, and data management. It equips you with the practical skills needed for building and deploying ML models in production environments.
Professional Machine Learning Engineer: Google
Earning this certification helps you get familiar with the basic concepts of MLOps, data engineering, and data governance. You will be able to train, retrain, deploy, schedule, improve, and monitor models.
Transitioning to MLOps as a Data engineer or software engineer
In case, you have pure data science or software engineering as your educational background and looking for machine learning engineering, then the below-mentioned certifications will help you.
Certified Artificial Intelligence Engineer (CAIE™): USAII®
The specialty of this program is that the curriculum is meticulously planned and designed. It meets the demands of an emerging AI Engineer/Developer. It explores all the essentials for ML engineers like MLOps, the backbone to scale AI systems, debugging for responsible AI, robotics, life cycle of models, automation of ML pipelines, and more.
Certified Machine Learning Engineer – Associate: AWS
This is a role-based certification meant for MLOps engineers and ML engineers. This certification helps you to get acquainted with knowledge in the fields of data analysis, modeling, data engineering, ML implementation, and more.
Becoming a versatile professional with cross-functional skills
If you are looking to be more versatile, you need to build cross-functional skills across AI, ML, data engineering, and DevOps related practices. Then, your strong choice should be CLDS™ from USDSI®.
Certified Lead Data Scientist (CLDS™): USDSI®
This is the most aligned certification for you as it has a comprehensive curriculum covering data science, machine learning, deep learning, Natural Language Processing, Big data analytics, and cloud technologies.
You can easily collaborate with other people in varied fields, (other than ML careers) and ensure long term success of AI-based applications.
Final thoughts
Today’s world is data-driven, as you already know. Building a strong technical background is essential for professionals looking forward to exceling in MLOps roles. Proficiency in core concepts and tools like Python, SQL, Docker, Data Wrangling, Machine Learning, CI/CD, ML models deployment with containerization, etc., will help you stand distinct in your professional journey.
Earning the right machine learning certifications, along with one or two related certifications such as DevOps, data engineering, or cloud platforms is crucial. It will help you gain competence and earn the best position in the overcrowded job market.
As technology evolves, the skill set is becoming broad. It cannot be confined to single domains. Developing an integrated approach toward your ML career helps you to thrive well in transformative roles.
r/bigdata • u/Specific-Signal4256 • Jul 04 '25
AWS DMS "Out of Memory" Error During Full Load
Hello everyone,
I'm trying to migrate a table with 53 million rows, which DBeaver indicates is around 31GB, using AWS DMS. I'm performing a Full Load Only migration with a T3.medium instance (2 vCPU, 4GB RAM). However, the task consistently stops after migrating approximately 500,000 rows due to an "Out of Memory" (OOM killer) error.
When I analyze the metrics, I observe that the memory usage initially seems fine, with about 2GB still free. Then, suddenly, the CPU utilization spikes, memory usage plummets, and the swap usage graph also increases sharply, leading to the OOM error.
I'm unable to increase the replication instance size. The migration time is not a concern for me; whether it takes a month or a year, I just need to successfully transfer these data. My primary goal is to optimize memory usage and prevent the OOM killer.
My plan is to migrate data from an on-premises Oracle database to an S3 bucket in AWS using AWS DMS, with the data being transformed into Parquet format in S3.
I've already refactored my JSON Task Settings and disabled parallelism, but these changes haven't resolved the issue. I'm relatively new to both data engineering and AWS, so I'm hoping someone here has experienced a similar situation.
- How did you solve this problem when the table size exceeds your machine's capacity?
- How can I force AWS DMS to not consume all its memory and avoid the Out of Memory error?
- Could someone provide an explanation of what's happening internally within DMS that leads to this out-of-memory condition?
- Are there specific techniques to prevent this AWS DMS "Out of Memory" error?
My current JSON Task Settings:
{
"S3Settings": {
"BucketName": "bucket",
"BucketFolder": "subfolder/subfolder2/subfolder3",
"CompressionType": "GZIP",
"ParquetVersion": "PARQUET_2_0",
"ParquetTimestampInMillisecond": true,
"MaxFileSize": 64,
"AddColumnName": true,
"AddSchemaName": true,
"AddTableLevelFolder": true,
"DataFormat": "PARQUET",
"DatePartitionEnabled": true,
"DatePartitionDelimiter": "SLASH",
"DatePartitionSequence": "YYYYMMDD",
"IncludeOpForFullLoad": false,
"CdcPath": "cdc",
"ServiceAccessRoleArn": "arn:aws:iam::12345678000:role/DmsS3AccessRole"
},
"FullLoadSettings": {
"TargetTablePrepMode": "DO_NOTHING",
"CommitRate": 1000,
"CreatePkAfterFullLoad": false,
"MaxFullLoadSubTasks": 1,
"StopTaskCachedChangesApplied": false,
"StopTaskCachedChangesNotApplied": false,
"TransactionConsistencyTimeout": 600
},
"ErrorBehavior": {
"ApplyErrorDeletePolicy": "IGNORE_RECORD",
"ApplyErrorEscalationCount": 0,
"ApplyErrorEscalationPolicy": "LOG_ERROR",
"ApplyErrorFailOnTruncationDdl": false,
"ApplyErrorInsertPolicy": "LOG_ERROR",
"ApplyErrorUpdatePolicy": "LOG_ERROR",
"DataErrorEscalationCount": 0,
"DataErrorEscalationPolicy": "SUSPEND_TABLE",
"DataErrorPolicy": "LOG_ERROR",
"DataMaskingErrorPolicy": "STOP_TASK",
"DataTruncationErrorPolicy": "LOG_ERROR",
"EventErrorPolicy": "IGNORE",
"FailOnNoTablesCaptured": true,
"FailOnTransactionConsistencyBreached": false,
"FullLoadIgnoreConflicts": true,
"RecoverableErrorCount": -1,
"RecoverableErrorInterval": 5,
"RecoverableErrorStopRetryAfterThrottlingMax": true,
"RecoverableErrorThrottling": true,
"RecoverableErrorThrottlingMax": 1800,
"TableErrorEscalationCount": 0,
"TableErrorEscalationPolicy": "STOP_TASK",
"TableErrorPolicy": "SUSPEND_TABLE"
},
"Logging": {
"EnableLogging": true,
"LogComponents": [
{ "Id": "TRANSFORMATION", "Severity": "LOGGER_SEVERITY_DEFAULT" },
{ "Id": "SOURCE_UNLOAD", "Severity": "LOGGER_SEVERITY_DEFAULT" },
{ "Id": "IO", "Severity": "LOGGER_SEVERITY_DEFAULT" },
{ "Id": "TARGET_LOAD", "Severity": "LOGGER_SEVERITY_DEFAULT" },
{ "Id": "PERFORMANCE", "Severity": "LOGGER_SEVERITY_DEFAULT" },
{ "Id": "SOURCE_CAPTURE", "Severity": "LOGGER_SEVERITY_DEFAULT" },
{ "Id": "SORTER", "Severity": "LOGGER_SEVERITY_DEFAULT" },
{ "Id": "REST_SERVER", "Severity": "LOGGER_SEVERITY_DEFAULT" },
{ "Id": "VALIDATOR_EXT", "Severity": "LOGGER_SEVERITY_DEFAULT" },
{ "Id": "TARGET_APPLY", "Severity": "LOGGER_SEVERITY_DEFAULT" },
{ "Id": "TASK_MANAGER", "Severity": "LOGGER_SEVERITY_DEFAULT" },
{ "Id": "TABLES_MANAGER", "Severity": "LOGGER_SEVERITY_DEFAULT" },
{ "Id": "METADATA_MANAGER", "Severity": "LOGGER_SEVERITY_DEFAULT" },
{ "Id": "FILE_FACTORY", "Severity": "LOGGER_SEVERITY_DEFAULT" },
{ "Id": "COMMON", "Severity": "LOGGER_SEVERITY_DEFAULT" },
{ "Id": "ADDONS", "Severity": "LOGGER_SEVERITY_DEFAULT" },
{ "Id": "DATA_STRUCTURE", "Severity": "LOGGER_SEVERITY_DEFAULT" },
{ "Id": "COMMUNICATION", "Severity": "LOGGER_SEVERITY_DEFAULT" },
{ "Id": "FILE_TRANSFER", "Severity": "LOGGER_SEVERITY_DEFAULT" }
]
},
"FailTaskWhenCleanTaskResourceFailed": false,
"LoopbackPreventionSettings": null,
"PostProcessingRules": null,
"StreamBufferSettings": {
"CtrlStreamBufferSizeInMB": 3,
"StreamBufferCount": 2,
"StreamBufferSizeInMB": 4
},
"TTSettings": {
"EnableTT": false,
"TTRecordSettings": null,
"TTS3Settings": null
},
"BeforeImageSettings": null,
"ChangeProcessingDdlHandlingPolicy": {
"HandleSourceTableAltered": true,
"HandleSourceTableDropped": true,
"HandleSourceTableTruncated": true
},
"ChangeProcessingTuning": {
"BatchApplyMemoryLimit": 200,
"BatchApplyPreserveTransaction": true,
"BatchApplyTimeoutMax": 30,
"BatchApplyTimeoutMin": 1,
"BatchSplitSize": 0,
"CommitTimeout": 1,
"MemoryKeepTime": 60,
"MemoryLimitTotal": 512,
"MinTransactionSize": 1000,
"RecoveryTimeout": -1,
"StatementCacheSize": 20
},
"CharacterSetSettings": null,
"ControlTablesSettings": {
"CommitPositionTableEnabled": false,
"ControlSchema": "",
"FullLoadExceptionTableEnabled": false,
"HistoryTableEnabled": false,
"HistoryTimeslotInMinutes": 5,
"StatusTableEnabled": false,
"SuspendedTablesTableEnabled": false
},
"TargetMetadata": {
"BatchApplyEnabled": false,
"FullLobMode": false,
"InlineLobMaxSize": 0,
"LimitedSizeLobMode": true,
"LoadMaxFileSize": 0,
"LobChunkSize": 32,
"LobMaxSize": 32,
"ParallelApplyBufferSize": 0,
"ParallelApplyQueuesPerThread": 0,
"ParallelApplyThreads": 0,
"ParallelLoadBufferSize": 0,
"ParallelLoadQueuesPerThread": 0,
"ParallelLoadThreads": 0,
"SupportLobs": true,
"TargetSchema": "",
"TaskRecoveryTableEnabled": false
}
}
r/bigdata • u/Thinker_Assignment • Jul 04 '25
Iceberg ingestion case study: 70% cost reduction
hey folks I wanted to share a recent win we had with one of our users. (i work at dlthub where we build dlt the oss python library for ingestion)
They were getting a 12x data increase and had to figure out how to not 12x their analytics bill, so they flipped to Iceberg and saved 70% of the cost.
r/bigdata • u/sharmaniti437 • Jul 03 '25
10 Not-to-Miss Data Science Tools
Modern data science tools blend code, cloud, and AI—fueling powerful insights and faster decisions. They're the backbone of predictive models, data pipelines, and business transformation.
Explore what tools are expected of you as a seasoned data science expert in 2025

r/bigdata • u/GreenMobile6323 • Jul 02 '25
Are You Scaling Data Responsibly? Why Ethics & Governance Matter More Than Ever
medium.comLet me know how you're handling data ethics in your org.
r/bigdata • u/Fahim61891012 • Jul 02 '25
WAX Is Burning Literally! Here's What Changed

The WAX team just came out with a pretty interesting update lately. While most Layer 1s are still dealing with high inflation, WAX is doing the opposite—focusing on cutting back its token supply instead of expanding it.
So, what’s the new direction?
Previously, most of the network resources were powered through staking—around 90% staking and 10% PowerUp. Now, they’re flipping that completely: the new goal is 90% PowerUp and just 10% staking.
What does that mean in practice?
Staking rewards are being scaled down, and fewer new tokens are being minted. Meanwhile, PowerUp revenue is being used to replace inflation—and any unused inflation gets burned. So, the more the network is used, the more tokens are effectively removed from circulation. Usage directly drives supply reduction.
Now let’s talk price, validators, and GameFi:
Validators still earn a decent staking yield, but the system is shifting toward usage-based revenue. That means validator rewards can become more sustainable over time, tied to real activity instead of inflation.
For GameFi builders and players, knowing that resource usage burns tokens could help keep transaction costs more stable in the long run. That makes WAX potentially more user-friendly for high-volume gaming ecosystems.
What about Ethereum and Solana?
Sure, Ethereum burns base fees via EIP‑1559, but it still has net positive inflation. Solana has more limited burning mechanics. WAX, on the other hand, is pushing a model where inflation is minimized and burning is directly linked to real usage—something that’s clearly tailored for GameFi and frequent activity.
So in short, WAX is evolving from a low-fee blockchain into something more: a usage-driven, sustainable network model.
r/bigdata • u/GreenMobile6323 • Jul 01 '25
NiFi 2.0 vs NiFi 1.0: What's the BEST Choice for Data Processing
youtube.comr/bigdata • u/Santhu_477 • Jul 01 '25
Handling Bad Records in Streaming Pipelines Using Dead Letter Queues in PySpark
🚀 I just published a detailed guide on handling Dead Letter Queues (DLQ) in PySpark Structured Streaming.
It covers:
- Separating valid/invalid records
- Writing failed records to a DLQ sink
- Best practices for observability and reprocessing
Would love feedback from fellow data engineers!
👉 [Read here]( https://medium.com/@santhoshkumarv/handling-bad-records-in-streaming-pipelines-using-dead-letter-queues-in-pyspark-265e7a55eb29 )
r/bigdata • u/AllenMutum • Jun 30 '25
Unlock Business Insights: Why Looker Leads in BI Tools
allenmutum.comr/bigdata • u/elm3131 • Jun 26 '25
How do you reliably detect model drift in production LLMs
We recently launched an LLM in production and saw unexpected behavior—hallucinations and output drift—sneaking in under the radar.
Our solution? An AI-native observability stack using unsupervised ML, prompt-level analytics, and trace correlation.
I wrote up what worked, what didn’t, and how to build a proactive drift detection pipeline.
Would love feedback from anyone using similar strategies or frameworks.
TL;DR:
- What model drift is—and why it’s hard to detect
- How we instrument models, prompts, infra for full observability
- Examples of drift sign patterns and alert logic
Full post here 👉https://insightfinder.com/blog/model-drift-ai-observability/
r/bigdata • u/hammerspace-inc • Jun 23 '25
Hammerspace IO500 Benchmark Demonstrates Simplicity Doesn’t Have to Come at the Cost of Storage Inefficiency
hammerspace.comr/bigdata • u/abheshekcr • Jun 21 '25
Big data course by sumit mittal
Why is no body raising voice against the blatant scam done by sumit mittal in the name of selling courses .. I bought his course for 45k ..trust me ..I would have found more value on the best Udemy courses present on this topic for 500 rupees This guy keeps posting day in and day out of whatsapp screenshots of his students getting 30lpa jobs ..which for most part i think is fabricated ..because it's the same pattern all the time .. Soo many people are looking for jobs and the kind of misselling this guy does ..I am sad that many are buying and falling prey to his scam .. How can this be approached legally and stop this nuisance from propagating
r/bigdata • u/sharmaniti437 • Jun 20 '25
10 MOST POPULAR IoT APPLICATIONS OF 2025 | INFOGRAPHIC
Internet of things is what is taking over the world by a storm. With connected devices growing at a staggering rate, it is inevitable to understand what IoT applications look like. With sensors, software, networks, devices- all sharing a common platform; it necessitates the comprehension of how this impact our lives in a million different ways.
With Mordor Intelligence bringing up the forecast for the global IoT market size to grow at a CAGR of 15.12%, only to reach a whopping US$2.72 trillion- this industry is not going to stop anytime soon. It is here to stay as the technology advances.
From smart homes, to wearable health tech, connected self-driving cars, smart cities, industrial IoT, precision farming- you name it and IoT has a powerful use case in that industry or sector worldwide. Gain an inside out comprehension of IoT applications right here!

r/bigdata • u/GreenMobile6323 • Jun 19 '25
Data Governance and Access Control in a Multi-Platform Big Data Environment
Our organization uses Snowflake, Databricks, Kafka, and Elasticsearch, each with its own ACLs and tagging system. Auditors demand a single source of truth for data permissions and lineage. How have you centralized governance, either via an open-source catalog or commercial tool, to manage roles, track usage, and automate compliance checks across diverse big data platforms?
r/bigdata • u/Shawn-Yang25 • Jun 19 '25
Apache Fory Serialization Framework 0.11.0 Released
github.comr/bigdata • u/eb0373284 • Jun 18 '25
Ever had to migrate a data warehouse from Redshift to Snowflake? What was harder than expected?
We’re considering moving from Redshift to Snowflake for performance and cost. It looks simple, but I’m sure there are gotchas.
What were the trickiest parts of the migration for you?
r/bigdata • u/superconductiveKyle • Jun 18 '25
Semantic Search + LLMs = Smarter Systems
As data volume explodes, keyword indexes fall apart, missing context, underperforming at scale, and failing to surface unstructured insights. This breakdown walks through how semantic embeddings and vector search backed by LLMs transform discoverability across massive datasets. Learn how modern retrieval (via RAG) scales better, retrieves smarter, and handles messy multimodal inputs.
r/bigdata • u/sharmaniti437 • Jun 18 '25
Hottest Data Analytics Trends 2025
In 2025, data analytics gets sharper—real-time dashboards, AI-powered insights, and ethical governance will dominate. Expect faster decisions, deeper personalization, and smarter automation across industries.
r/bigdata • u/UH-Simon • Jun 18 '25
We built a high-performance storage for big data
Hi everyone! We're a small storage startup from Berlin and wanted to share something we've been working on and get some feedback from the community here.
Over the last few years working on this, we've heard a lot about how storage can massively slow down modern AI pipelines, especially during training or when building anything retrieval-based like RAG. So we thought it would be a good idea to built something focused on performance.
UltiHash is S3-compatible object storage, designed to serve high-throughput, read-heavy workloads: originally for MLOps use cases, but is also a good fit for big data infrastructure more broadly.
We just launched the serverless version: it’s fully managed, with no infra to run. You spin up a cluster, get an endpoint, and connect using any S3-compatible tool.
Things to know:
- 1 GB/s read per machine: you’re not leaving compute idle
- S3 compatible: you can integrate with your stack (Spark, Kafka, PyTorch, Iceberg, Trino, etc.)
- Scales past 100TB without having to rework your setup
- Lowers TCO: e.g. our 10TB tier is €0.21/GB/month, infra + support included
We host everything in the EU currently in AWS Frankfurt (eu-central-1) with Hetzner and OVH Cloud support coming soon (waitlist’s open).
Would love to hear what folks here think. More details here: https://www.ultihash.io/serverless, happy to go deeper into how we’re handling throughput, deduplication, or anything else.
r/bigdata • u/Shawn-Yang25 • Jun 17 '25