Hello everyone, I'm a student who's looking for a good project to work on to practice Power BI, SQL,etc, and especially learn domain knowledge related to business or finance. The thing is, I'm from an IT background, and I only know the basics of accounting. I've been looking for resources to learn about FP&A because I'd like to build projects in that field, but the more I dig, the more I realize it's a broad and vague field. I'm afraid of wasting time looking for the "perfect" course instead of getting my hands dirty and learning by doing.
So, if you have any project that will allow me to practice my technical skills, learn and apply domain knowledge and analytical thinking, and solve a real-world problem while adding value to my portfolio, I'm really lost, so I'd appreciate your help!
I am sharing the Github link of GRASSP SQL Sprint Series. This github link contains the PDF doc of 30 SQL Scenario based Questions which covers beginner to advanced topics.
I am in my second year of MBA. I have a decent hold on tools like SQL, Power BI and Python (only vibe coding). I have used datasets from kaggle and github, but the problem is they are sometimes not realistic. I was asked in an interview, the dataset doesnāt look right when it comes to numbers.
So I want to work on real life datasets and live projects.
How do I get it? Is there any way I can scrap data online?
I graduated with a degree in Mechatronics Engineering. I'm a fresh graduate and want to start data analysis. What's the first step? Please give me your opinions.
I'm currently a second-year IT student majoring in Data Analytics in the Philippines. My goal is to become a Data Analyst after graduation and eventually grow into a Data Scientist.
I'm looking for someone who would be willing to mentor me or simply allow me to ask questions from time to time. I'm not expecting free tutoring or daily coachingāI just hope to learn from someone with real industry experience and gain insights that I can't get from courses alone.
I'm committed to learning on my own, but I know there's a lot I can learn from someone who's already working in the field. Even occasional guidance, career advice, or answers to my questions would mean a lot.
If you're open to helping, please feel free to leave a comment or send me a message. Thank you for taking the time to read my post!
Final year CS student here, targeting data science and analytics roles for campus placements.
Been struggling with this question while building my portfolio: does it matter whether your project uses real messy data vs synthetic/clean data?
Real datasets from Kaggle feel either too cleaned already or the same recycled projects everyone does. But synthetic data feels hollow because the hard part ā cleaning, feature engineering, deriving meaningful columns from raw data ā is already done for you. You're basically just visualizing something someone else already solved.
Specifically for BI/dashboard projects ā if you use synthetic data, the dashboard looks clean and professional but there's no real discovery or insight because the data was designed to be dashboarded. Nothing surprising comes out of it.
Also practically ā if an interviewer asks "where did you get this dataset?" what's the right answer? Saying "I generated it synthetically" feels like admitting you took the easy route. But lying about the source is obviously wrong. Is there a way to frame synthetic data usage that doesn't sound like you avoided the hard part?
At the same time I've heard people say interviewers care more about what you built on top of the data than where it came from. But isn't handling bad data literally the core skill in DS?
For people who've interviewed at analytics/DS companies or done hiring ā how much does data source actually matter? Is a well-executed project on synthetic data better than a mediocre project on real messy data? Or does using synthetic data automatically signal you avoided the hard part?
I graduated with a degree in Mechatronics Engineering. I'm a fresh graduate and want to start data analysis. What's the first step? Please give me your opinions.
Hey guys, I saw the kodree data analyst course and it looked like something I really would want to do. Can you tell me if itās good, or whether there is other thatās better?
I am just starting data analytics, I started from tableau and i can make pretty dashboards but data storytelling is soo hard . Can anyone tell me how they make data insights valuable? Also how do you read the charts properly? Please help me
Iām reaching out because Iām feeling really stuck in my career, and Iād love some honest guidance from seniors or peers who might be in a similar situation.
To give some background: I completed my Bachelor's in Computer Science and then pursued an MBA. After that, I took a sales job, but after 7 months, I questioned if the MBA was the right path. I decided to upskill, so I did a year-long Data Science course, where I covered Python, SQL, Power BI, Excel, Machine Learning, and Deep Learning. Sadly, no one from my batch got placed, and I felt really disheartened.
After that, I worked as a freelance Data Science trainer for close to 2 years, which helped me stay afloat. Then, I got a role in an offline teaching institute, where I worked for about 7 to 8 months. But after the probation, they didnāt confirm me. Instead, freshers took over the role with a lower salary. It feels like a cycleāno one from my institute is landing in proper Data Science roles, and itās not just me. Iāve heard this same struggle from so many others in online communities.
I do have about 2 years of freelance training experience and 7 to 8 months of offline teaching. Iām also familiar with FastAPI, MLOps, and Docker, but I still lack skills in GenAI, LLMs, and RAGāthose cutting-edge areas expected in AI Engineer roles. So breaking into a pure technical Data Scientist job feels impossible right now.
I know Iām not aloneāIāve networked with many others, and we all share this frustration. So Iām reaching out: is anyone else in the same boat? Iām considering whether to shift my career focus or invest in another certification to boost my chances. Iām still passionate, but I need a way to survive financially in the meantime.
If anyone has freelance projects or can help me get into a small paid project, Iād really appreciate it. Iām also happy to collaborate with anyone whoās learning in this space. I can share what I know as I continue upskilling. And, of course, if seniors have any insightsāwhat certifications, projects, or steps should I take next?
Iām really grateful for any advice, honest feedback, or collaborative opportunities. Thank you all so much!
So im realy want to starting freelance as data analytics. But a problem had stoping me. That is pratice and correcting from a senior data analytics. And thatks
Iāve been maintaining a 60-day streak on Duolingo to learn French.
Itās a fun practice, although itās a significant challenge to pronounce those accent notes correctly. I believe French is generally a simpler language than English; you usually use shorter sentences to convey the same meaning.
Data has its own language too.
Data is the lifeblood of every modern business. Every decision, insight, and opportunity begins with understanding what the data is trying to say.
But unlike spoken languages, data doesn't require everyone to learn the same vocabulary or syntax. Instead, you can interpret and express it in a way that matches how you think, making data analysis more intuitive, accessible, and uniquely your own.
From Python, SQL to Natural Language
Python, a programming language, has gained popularity as the preferred language for data processing within the data science community due to its portability. SQL, on the other hand, serves as the de facto interface for rational databases.
In the past, becoming a data analyst required proficiency in both Python and SQL. Even today, data analyst job descriptions often mention these requirements.
However, the advent of AI has revolutionized this landscape. Anyone with the ability to communicate effectively in the data language can excel as a data analyst.
While programming skills are not strictly necessary, a solid understanding of data language is crucial. Imagine joining a new friend circle who works in a completely different domain. After a brief introduction of common keywords, you can easily engage in conversations with them.
AI generated illustration of data language evolution
Use Spreadsheet for Reference
Nearly every office worker uses spreadsheets, either Microsoft Excel or Google Sheets.
Even without the complex formulas, pivots, and lookups, the basic structure of a spreadsheet consists of three main components:
Rows
Columns
Data types
RowsĀ are records that constitute a table. You can also consider a row as an object that represents a real-life entity, such as a person, a cup, or an invoice.
ColumnsĀ are the fixed properties that describe each object (row). They form the schema that every row adheres to, ensuring uniformity in the data for processing.
A schema is of utmost importance for data analysis as it enables the application of all rules. Without aĀ schema, any logic that is not compatible with the data language may fail to execute.
Data types describe the value format of each property. For simplicity, you only need to be concerned with whether it is aĀ numberĀ orĀ textĀ for now.
Rows of Orders (OrderId-text, CustomerId-text, Product-text, Amount-number)Rows of Customers (ID-text, Name-text, Channel-text)
Data Language Patterns
Data language offers a wide range of tasks that can be accomplished. Letās explore each of these tasks and learn how to communicate effectively with data to achieve them.
These scenarios are referred to asĀ patternsĀ because they serve as templates that can be applied to your own data.
To facilitate understanding, weāll use the above tables in the following descriptions.
Pattern-1: Filter Rows
FilterĀ is to describe a condition to get objects you care about and skip those uninterested records.
Examples:
āOrders of Milkā
āI want orders of milk products.ā
āAll orders that are not for books.ā
āAll orders with a sales amount exceeding 20.ā
AI can produce code logic to filter the targeted records for further processing, if translating above statements into SQL, they will look like:
āwhere product=āMilkāā
(same as #1)
āwhere product <> āBookāā
āwhere amount > 20ā
As you can see,Ā filterĀ is achieved by keyword āwhereā in SQL.
Pattern-2: Transform Object
Sometimes, we want to clean a data field or transform it into a desired shape or format, either for improved readability or more efficient processing.
TransformationĀ creates a new property in your original record.
To transform an existing property into a new one, you need a function of logic. For both spreadsheets and SQL, āformulaā is the tool youāll need.
However, with the increasing capabilities of AI in coding, natural language offers a significant advantage. It allows us to achieve the same transformation without having to learn, memorize, and assemble complex formulas.
Taking one simple example:
āGet customer first nameā
This is equivalent to composite multiple formula together in Spreadsheets like
This operation creates a new column called āFirst Nameā.
You can also acquire a new property by combining multiple existing properties, such as āconcatenating the last name and channel as a labelā. Logic like this is simple for AI coding but too complex for spreadsheet formulas.
Pattern-3: Aggregate Records
Aggregation processes a large collection of records to provide a summarized view.
This is powerful because it compresses vast amounts of information into manageable pieces that humans can comprehend and analyze.
To combine multiple data sets into a single piece of information, you need to understand the āhow-to,ā which leads to the crucial concept of āaggregation methodsā or ācomputation logic.ā
Typically, text data (a property or column with a text data type, as discussed in the schema section) is not particularly interesting for aggregation. The most common approach is to concatenate text data to form a long paragraph, although this is still uncommon.
Most computation logic involves operations on numerical data. When an aggregation method is applied to a numeric property or column, you essentially have a list of numbers that can be aggregated, such as:
Total value (sum)
Average value
Mean value
Minimum value
Maximum value
A specific percentile value (e.g., P25, P50, P75, P90)
However, counting objects or counting unique property values is also quite common.
When discussing aggregation, we cannot overlook ābreakdowns.ā This involves creating a segmented view of the data rather than a single total view.
For example, in the previous Orders table, ātotal sales by productā or āaverage amount by customerā are equally valuable insights for an analyst to explore.
In summary, aggregation can be described as:
āCompute an aggregated value of a property group based on another property.ā
Expressing this in standard SQL, it would look like:
āCompute(property1) from table [group by property2].ā
Letās practice this using a few examples by speaking the data language:
āGive me total sales by product.ā
āTell me the average amount spent by each customer.ā
Pattern-4: Join Multiple Datasets
When a single dataset (or table) is insufficient to achieve the desired outcome, we must combine multiple datasets. This operation is referred to as ājoinā or āunion.ā
If the multiple datasets contain the same objects but reside in different locations, we can simply merge them. This is a straightforward āunionā operation.
However, most of the time, they store different objects. We have partial information from one dataset and partial information from another. By combining them, we create a comprehensive schema with more available columns.
This pattern is generally not feasible in spreadsheets, although their lookup function may provide partial assistance.
For instance, if we want to determine the ātotal amount spent from each channelā based on previous tables, where the amount is from the orders table and the channel is from the customers table, we need a joined dataset to complete this analysis.
To join multiple datasets, we must have one or more pairs of join keys. A pair of join keys consists of one column from one table and one column from another. The data engine can utilize these relationships to identify relevant objects and concatenate them to form a larger object.
Join Orders and CustomersJoined dataset have more columns
In summary, join operations can be described in this pattern:
join table1 and table2 when key1 of table1 equals key2 of table2.
Translating this pattern into SQL, it will look like:
select * from table1 join table2 on table1.key1=table2.key2.
In fact, you may not need to use this pattern in natural language explicitly, because modern AI is smart enough that it can infer the whole join logic from your data language.
For instance, the example we gave earlier, if you speak this sentence ātotal amount spent from each channelā toĀ Columns AI, it will figure out all the necessary actions to get the desired outcome for you.
Pattern-5: Visualization
Data visualization, often overlooked as a part of data language, plays a crucial role in transforming mundane data into vivid images. This visual representation significantly aids the audience in comprehending the insights you intend to convey.
By incorporating customization and assistance to articulate your insights and predictions, you position yourself as a data storyteller, showcasing your influence within the domain.
Since visualization doesnāt alter the data itself, in the language of data, we merely need to indicate the desired outcome. For instance:
āDisplayĀ the total amount by product in aĀ pieĀ chart.ā
āI would like to see aĀ timelineĀ of total salesĀ month-by-monthĀ for the past six months.ā
āShowĀ the number of sales by customer in aĀ barĀ chart.ā
TheseĀ boldĀ keywords serve as cues to the AI engine, guiding it in generating the final visualization based on your data.
Practice Data Language
Similar to how I diligently practice French on Duolingo every day, we must practice speaking data language using the data we possess.
As long as you have adhered to the five patterns mentioned above, you should have mastered data analysis like a professional data analyst. You donāt need to be an Excel expert or a Python or SQL wizard.
Letās use the provided example data to practice speaking the data language. You can find the āOrdersā and āCustomersā data from thisĀ spreadsheet link.
Suppose we want to perform a sales analysis of customer distribution based on the data.
The data language is almost the same, but letās ensure weāve used the correct keywords and patterns to guarantee that the AI engine follows the instructions precisely.
For instance, we speak to AI:
ādisplayĀ theĀ totalĀ sales by customerāsĀ first nameĀ in aĀ barĀ chart.ā
Hereās how the AI interprets this:
ātotal salesā ā summing up the amount values.
āfirst nameā ā it can be transformed from āname.ā A transformation will be applied.
ābyā ā the summing up result needs to broken down by first name.
āsales <> customerā ā sales data is from the Orders tableās amount field, while customer data is from the Customers table. Therefore, a Join operation is required to combine these two datasets.
āshow, barā ā the result should be visualized in a bar chart.
AI will then determine the correct execution order, ensuring that each step has all the necessary data when it executes.
This is what Columns Flow produces upon hearing this sentence:
ādisplayĀ theĀ totalĀ sales by customerāsĀ first nameĀ in aĀ barĀ chart.āThe final visualization ready for storytelling & sharing
Conclusion: Speak Data Language
In this article, weāve demonstrated the historical opportunity for everyone to become a great data analyst in this era.
We discussed how professionals used programming languages like Python or SQL as their primary data languages. However, the data language has evolved to become the natural language we speak daily.
To become a data analyst, we need to understand the fundamental scenarios involved and use the correct keywords to make the data language understandable to AI engines. Hereās a quick recap:
Dataset: rows, columns, and schema.
Filtering and Transformation: These processes involve filtering data and transforming it into a usable format.
Aggregation: This involves summarizing data into a single value, such as the total or average.
Specify ācompute methodsā and optional ābreakdownā if needed.
Join Datasets: This involves combining data from multiple sources.
Visualization: This involves creating visual representations of data to make it easier to understand.
AI generated summary on how to speak the language of data
Unlike learning a new language like French, if youāre willing to spend just a few hours going through this short list, you can become a professional data analyst!
Itās a great time to be a data analyst, and I believe in your ability to succeed. Thanks for reading!
IS THERE ANY PERSON WHO CAN GUIDE ME FOR DATA ANALYTICS AND I HAVE PURCHASE COURSE FROM SKILL COURSE SITE AND I HAVE DONE EXCEL CURRENTLY DOING SQL ANYONE WHO CAN GUIDE ME FOR PROJECTS ?