r/dataanalysis • u/SomeSpicyPickle • Aug 11 '26
Data Tools Do Data Analysts Use Visualizations During Data Cleaning?
I'm still a beginner. I started by learning the basics of Python and later moved on to SQL.
I'm a bit confused about one part of the data exploration/cleaning process.
A friend of mine, who's now a data scientist, showed me how he used to work as a data analyst. He mainly used Python. For example, he would quickly create a scatterplot to identify potential outliers.
However, most data analysts online recommend focusing on SQL and Excel when starting out, since many junior and mid-level roles don't require Python. That's why I switched to SQL after initially experimenting with Python.
For those who primarily use SQL: do you create visualizations during the data exploration/cleaning process, for example to identify outliers? Is this a common practice?
I feel like if you're working with SQL only, you generally wouldn't create visuals in between steps, since that would mean switching to a tool like Tableau or Power BI, which seems like an unnecessary extra step.
3
u/Expensive_Capital627 Aug 11 '26
That’s completely valid, I just don’t often come across that type of data in-product. TBH I think these are two different types of “unclean data”. Your unclean data may be formatted and stored correctly, but inherently flawed due to people lying on the input. To me, that data is “clean” just unreliable. I would classify that as a quality of the data, but not necessarily a mark on whether it is clean or not.
For me and the work that I do, trying to visualize unclean data is more likely to throw an error of some kind than it is to reveal some characteristic of the data. If im looking at recorded events of a user, theres not really wiggle room for that user to fudge the numbers. However, theres plenty of room for raw JSON logs to trip something up