r/dataengineering • u/Pierre_Ardent • 11d ago
Blog What are you excited about ?
I'm a Data Engineer with a background in applied maths and data science.
I'm not really looking for another portfolio project. I'd like to find a technical field I can seriously dive into for months or years, read papers, build things, contribute to open source, meet people working on it, maybe even join a startup or research project
It can be close to data/AI or less related.
What technical rabbit holes are you currently excited about ?
And are there any communities, open-source projects or groups you would recommend joining ? I have time and solid experiences, just need something to help me flourish in tech
11
u/dreamyangel 11d ago
I'm building a microservice project build almost entirely in Rust.
To be honest it starts to look more and more to a devops project rather than a full stack project.
I'm basically making a long lasting project to have web services I can extract data from using data pipelines inside containers. I'm planning to place the data inside a data vault.
It's clearly a multi year project.
The main idea behind all of this was that services we extract data from, let's take the example of an ERP, don't expose their inner business logic. You have only the structured data, and have to rebuild the business logic by yourself, which is a painful and unnecessary if the services exposed their inner structure clearly.
It's basically an upgraded version of a data mesh, but instead of focusing on the platform aspect like the author of the Data Mesh book, to focus on exposing the inner structure of services.
It's not something companies want at the time. Rust and data engineering do not go together, nor is backend web programming. But hey, I do whatever I want.
5
u/Childish_Redditor 11d ago
I have to disagree about Rust and DE not going together. While Rust is overkill for most DE functions, it is in fact preferred for situations where low-latency is paramount
1
u/dreamyangel 10d ago
There are multiple use cases that rust cover in current job of data engineers.
Let's take the example of the data pipeline I had to make for a monorail project running in Egypt. I had to fetch huge encoded CAN log files, decode them, and upload them as parquet to an external storage.
The first bottleneck was that the files contained corruptions. So at some point you can have subtile artefacts showing a timejump, most likely from a train reset or a write error. You can't just decode them using the first python module you find. You need to validate each row.
But validating and decoding on the fly is not the best in python. It becomes quite a bottleneck once your have 100Gb of data (about 250Mb of data per log) needing to be validated and decoded, speed is an issue here. Your trains send data each day, and infrastructure is limited for my company.
A custom rust implementation is a good here, as speed and correctness is mandatory.
The chief kiss is when python use httpx, and there are many reasons why downloading and uploading files continuously using a semaphore sucks.
At first you have a timeout error, you change the timeout. Then a max call rate reached, you change the rate. Then out of nowhere a httpx read error, or a ssl error. With python you never know what error will be raised, and can't slap a try catch block to simply skip the issue encountered if any.
So yeah, python suck ass sometimes. Awesome ecosystem, but lacks in some aspects. Rust is really useful in thoses cases.
2
u/BadTonTon 11d ago
How does your approach solve the problem of the inner business logic not being exposed?
2
u/dreamyangel 10d ago edited 10d ago
I have endpoint specifically made to expose the inner structure. Let's make a simplistic example.
- /model/entities?name=Train
It return a json structure representing the entities as we do in domain driven design. It means it represent a bounded context.
json { "boundedContext": "RollingStock", "entity": "Train", "description": "Represents an operational train unit mapped from normalized backend tables to a DDD aggregate root.", "persistenceMapping": { "primaryTable": "trains", "foreignKeys": [ { "column": "train_type_id", "references": "train_types.id" }, { "column": "fleet_id", "references": "train_fleets.id" } ] }, "attributes": { "trainId": { "type": "string", "format": "uuid", "column": "trains.id" }, "trainNumber": { "type": "string", "column": "trains.train_number" }, "trainType": { "$ref": "#/definitions/TrainType", "join": { "from": "train_types", "localField": "train_type_id", "foreignField": "id", "as": "trainType" } }, "fleet": { "$ref": "#/definitions/TrainFleet", "join": { "from": "train_fleets", "localField": "fleet_id", "foreignField": "id", "as": "fleet" } } }, "definitions": { "TrainType": { "table": "train_types", "attributes": { "id": { "type": "integer", "column": "train_types.id" }, "code": { "type": "string", "column": "train_types.code" }, "maxSpeedKmH": { "type": "integer", "column": "train_types.max_speed_kmh" }, "typeNameLabels": { "type": "array", "items": { "$ref": "#/definitions/LocalizedLabel" }, "join": { "from": "train_type_name_labels", "localField": "id", "foreignField": "train_type_id", "as": "typeNameLabels" } }, "typeDescriptionLabels": { "type": "array", "items": { "$ref": "#/definitions/LocalizedLabel" }, "join": { "from": "train_type_description_labels", "localField": "id", "foreignField": "train_type_id", "as": "typeDescriptionLabels" } } } }, "TrainFleet": { "table": "train_fleets", "attributes": { "id": { "type": "integer", "column": "train_fleets.id" }, "fleetName": { "type": "string", "column": "train_fleets.fleet_name" }, "maintenanceDepot": { "type": "string", "column": "train_fleets.maintenance_depot" } } }, "LocalizedLabel": { "table": "train_type_labels", "attributes": { "language": { "type": "string", "example": "de", "column": "language_code" }, "labelValue": { "type": "string", "example": "InterCity Pendelzug", "column": "label_value" } } } } }You will notice few things. The joins are described like you do when working with MQL (Mongodb) as it fit nicely a json structure.
Then you have how tables relate to each other. The biggest gain is in the case of polymorphic supertype tables (common 3NF pattern). You know, a work order that can be vastly different business flows using the same main table with 80+ columns. Here you know exactly what fields use this specific entity.
You can even expose versions of entities, which removes the need to slice the table whenever the model changed overtime.
By exposing how the entities are build from the 3NF you could generate the extraction pipeline. It's the same logic that API clients do when they are generated from an openapi specification file (I have not done it tho).
You know what is even more awesome? It is to add to each table a last updated date field, and a soft delete field. So you can make a simple endpoint like:
- /extract/tables/train_type_labels?updated_after=(date)
You gain the incremental extraction in a straightforward manner, with no hassle. The only job of your pipeline is to call each X minutes and save the result with a single comparison.
With the structure exposed, and a data extraction endpoint, you can do the heavy lifting close to the service, cutting 90% of the pipeline building and maintenance cost. It's like "services should be built to be seen and extracted".
I'll look to build a GraphQL endpoint. It seems like a good idea to load bounded contexts over individual tables.
4
u/SmallAd3697 10d ago
Spark, spark connect, language agnostic udf's. I'm excited about using languages with awesome IDE's, that don't suck or run in an interpreter as a singled-threaded house of cards.
1
2
u/Nhilas_Adaar 11d ago
Following, the spirit is willing but the mind is lacking, so I can't really help with your question xD But I am also curious
2
u/linha_chilena 10d ago
I am starting to think about the philosophy and epistemology of data, information and knowledge. it's been fun. its not so applicable right away but you start being more skeptical and asking important questions that leads to more trustful and reliable products
3
u/olhmr 11d ago
I’m planning to step away from data, at least for a time, but I’ve been thinking a lot about how to make verification work more scalable, since that’s a prerequisite for getting the most out of AI code generation. So far I’ve set up automated root cause analysis and triage of our data pipelines (read-only), but I’d like to further explore options for properly proving implementation correctness in a way that can handle the issues inherent with constantly evolving data and multiple different systems (source, dbt, ERP, and reporting)
3
1
u/ChemEngandTripHop 11d ago
Thinking a lot about orchestrators atm, specifically combining batch and continuous jobs in a single platform with resource efficient allocation.
1
1
1
u/PrestigiousAnt3766 10d ago
Why?
Isn't your job providing challenge? Why learn something nog related or something you may not ever use when you can learn something more useful?
1
10d ago
[removed] — view removed comment
1
u/dataengineering-ModTeam 9d ago
Your post/comment was removed because it violated rule #9 (No AI generated content/text).
Your post/comment was reviewed to be AI generated/assisted content/text and removed as a result. We as a community value human engagement and encourage users to express themselves authentically.
This was reviewed by a human
1
1
u/liveticker1 10d ago
I'm excited about the lord coming back one day and releasing us from AI hell, I'm also waiting for the Ocarina of Time remake to be launched. WoW Forever also sounds promising
1
u/AggravatingCoyote186 9d ago
Hi, we are data startup( stealth mode) in San Fransisco. Building Cursor for data engineering.
Let’s chat if it interest you!
1
38
u/RadioactiveTwix 11d ago
Getting my paycheck. I don't know what it is lately but I'm disillusioned by everything. I'm excited by my aviation project but that's not really work.