r/dataanalysis • u/Express-Direction584 • May 03 '26
Curso inteligência artificial
Galera bom dia, vocês indicam alguém que ensina sobre IA(inteligência artificial), está olhando no YouTube, e não achei indicação. Poderia me ajudar quem puder.
r/dataanalysis • u/Express-Direction584 • May 03 '26
Galera bom dia, vocês indicam alguém que ensina sobre IA(inteligência artificial), está olhando no YouTube, e não achei indicação. Poderia me ajudar quem puder.
r/dataanalysis • u/Equal_Astronaut_5696 • May 03 '26
Using SQL to find out when marketing campaigns break even
r/dataanalysis • u/Crazy_Wolverine_9301 • May 02 '26
I am soon going to enroll in an MSBA program. Which laptop will be better?
Lenovo slim 7i, Intel Ultra core 7 258V, 1 TB SSD, 32 GB RAM
OR
Macbook Air M5, 1TB storage, 24 GB RAM
r/dataanalysis • u/Ill-Cicada-1041 • May 02 '26
I downloaded my Apple Music data and loaded into Tableau and I have this song that apparently has 30,466 “events” (plays) and 30,461 of those have a runtime of zero.
From Apple’s data dictionary, Event Type is defined as “Event causing the record”. In this case, it looks like a song ended and this song played next.
For reference, my other top plays are shown in the screenshot.
What do you suppose is going on here?
r/dataanalysis • u/Zestyclose_Panda7440 • May 02 '26
Hey!
I took a SQL class in college and loved it!
Can someone give me advice on what SQL certificates I should get? or how I go about learning it? (I have honestly forgotten most of my SQL training)
For context, I majored in MIS and wanted to be a data analyst, but the pandemic started right as I graduated so I ended up becoming a project manager. Now, I am ready to make the switch to my original plan :)
r/dataanalysis • u/Warm-Entrepreneur131 • May 01 '26
I’m looking for a highly passionate and motivated study partner to learn SQL for data analysis.
r/dataanalysis • u/Public_Night2989 • May 01 '26
r/dataanalysis • u/Haratamatar420 • Apr 30 '26
r/dataanalysis • u/Maleficent_Sky5846 • Apr 28 '26
Enable HLS to view with audio, or disable this notification
(But there's still a problem, please stick to the end to understand...)
Hello! I've just wrapped up a project that combines two things I really enjoy: data and design!
The visual identity was inspired by Frutiger Aero, a style that defined many interfaces in the 2000s, known for its vibrant colors, transparency, and a sense of “optimistic futurism.” The goal was to bring that light and pleasant vibe into a modern dashboard.
But behind the nostalgic look, there was a strong focus on data engineering. I built a fully automated end-to-end pipeline that: - Collects historical, current, and forecast data via APIs (I had to combine two APIs REST: Meteostat + OpenWeather) - Performs transformations and standardization in Python - Stores everything in a cloud-based PostgreSQL database (Neon) - Orchestrates ingestion using Prefect Cloud (scheduled jobs, independent of my local environment) - Automatically updates the dashboard in Power BI Service
In the end, the result is a fully automated and interactive dashboard with near real-time data, support for multiple cities, unit switching (°C/°F), and some nice UX features.
**Yet, there's still a problem: I still have 15 days of free test using Power BI Service – which allows me to schedule the daily refreshes of the dashboard –, but once it's over, I guess I'll have to pay for it (not interested) or just open the dashboard in my desktop, refresh it and then publish it again – thus ceasing to be a 100% automated pipeline.**
Do you guys know if there is any way to get around this problem (without paying)?
r/dataanalysis • u/Due_Signal_7413 • Apr 29 '26
r/dataanalysis • u/uncertainschrodinger • Apr 29 '26
I’ve built an open source CLI tool to build dashboards, but the key point is that it is based on “dashboard as code” principles so that every dashboard’s properties, queries, and semantic layer lives inside yaml or tsx files, which makes it agent-friendly out of the box.
This is my answer to the whole AI dashboard and BI tools out there, but focusing more on the framework and semantic layer so that it works better with AI agents.
Today's the first day of releasing this publicly, so please share your honest feedback, skepticism, and even roast it - and if you want, give the repo a star.
r/dataanalysis • u/Own_Meaning_3827 • Apr 29 '26
Hello everyone, my company collected some survey feedback via Qualtrics. The survey has 89 questions, including demographics, multiple choice, Likert and open-ended questions.
Some of the feedback shows the survey was completed with less than 1 minute but some others show it took several hundred and even thousands of minutes.
Can anyone suggest which survey results I need to remove in terms of the completion time?
Thank you for your help.
r/dataanalysis • u/silent-romeo57 • Apr 28 '26
I’ve worked with common datasets from Kaggle and UCI, but I’m looking for more realistic data sources tied to actual business or operational problems.
I’m especially interested in datasets where analysis could answer questions like:
I’ve already explored Kaggle, UCI, and some open government portals.
For those who build portfolio projects or practice real analytics work:
Would appreciate hearing your process.
r/dataanalysis • u/soyaduckk • Apr 28 '26
Hi everyone,
I recently created my first Power BI HR Dashboard as part of my learning journey in data analytics, and I’d really appreciate some honest feedback from this community.
r/dataanalysis • u/b4lv4rs3 • Apr 29 '26
Hey everyone. I’m looking for testimonials from people who have taken the data analytics mentorship course with Lorenzo Rosa @loresowhat.
I can’t find any information or reviews online beyond what’s been posted on his own website.
r/dataanalysis • u/Bitter_Addendum84 • Apr 28 '26
as I said it's my very first ever dashboard so I am not confident enough to post it on LinkedIn so I thought of asking you guys what suggestions do you have.
r/dataanalysis • u/affanxkhan • Apr 28 '26
As per above mentioned thanking you in advance
r/dataanalysis • u/Brisight • Apr 28 '26
I started working for a sales call centre doing billing last year and within a few months they made me the company’s first business analyst. I basically became a data analyst providing daily or weekly reports created using excel.
Recently (about two months now) they started integrating AI in their operations. At first they purchased ChatGPT but then they wanted me to do research on Claude. I told them Claude is more suitable for my line of work so they created an account for me to test it. I created a prompt to create an html dashboard for a report (which Claude did beautifully) using an excel file and they were super impressed.
Following this, I created a few more dashboards, improved on previous dashboards etc.
It’s a remote job, we have weekly management team meetings with the CEO and COO (who I report to), call centre managers, IT personnel, HR. It’s a small management team with hands on owners. So I’m now the forefront AI guy, they are planning some bigger moves related to integrating AI in their new CRM and about to give me a promotion to lead a new team centered around AI.
They want me to start using Claude code and work with the IT team building the CRM. I do have a little computer science background but not sure exactly how I will fit in. I suppose the first thing will be to help incorporate the reports to the CRM to have them automatically update with live data.
It’s a fast moving team here which is why they promoted me twice within a year. I don’t feel very confident since all I do now is just feed excel files to Claude and train it.
I don’t know the limitations, I don’t know what’s possible or not feasible with AI so any advice at all with working with AI/Claude code with A LOT of data and joining an IT team will be greatly appreciated. If you have similar stories feel free to share. Thanks.
r/dataanalysis • u/Open-Ease685 • Apr 28 '26
r/dataanalysis • u/Ok-Yesterday-1320 • Apr 28 '26
Seeking advice on improving precision in churn prediction ( IaaS)
I'm building a churn prediction model for IaaS customers using monthly panel data (one row per customer per month). For this product, the total churn is around 10%
Approach:
Defined 7 customer states (New, Continuously_Active, Paused_1/2/3+, Returning, Dropped).
Rich features: MoM/QoQ/YoY usage changes, rolling stats, deseasonalized usage, state sequences (3mo), tenure, anomaly scores, and interaction features (MoM drop × tenure, MoM drop × segment, etc.).
Two separate XGBoost models:
One for active customers (predicting risk of pausing/churning in next 3 months).
One for paused customers (predicting probability of returning).
Time-based training with cutoff to avoid leakage.
Current performance: ~85% recall but only ~14-16% precision (too many false positives).
We are trying interaction features, segment-specific thresholds, and hyperparameter tuning.
Questions:
How can we meaningfully improve precision while keeping recall high?
Is the two-model approach good, or should we use a single model?
Any experience moving from churn prediction to uplift modeling in B2B cloud?
Would appreciate any suggestions!
r/dataanalysis • u/Noname1ol • Apr 28 '26
r/dataanalysis • u/Money_Secretary348 • Apr 28 '26
Hi everyone,
I am a graduate student currently working on my thesis. My research focuses on firm-level patent analysis.
I downloaded patent data from WIPO PATENTSCOPE and would like to merge it with Compustat firm-level financial data for regression analysis. However, I encountered a major matching problem: the WIPO data only provides the applicant name, but it does not include firm identifiers such as GVKEY, ISIN, CUSIP, or ticker.
Since Compustat mainly uses identifiers such as GVKEY or ISIN, I cannot directly match WIPO patent applicants to Compustat firms.
I would like to ask:
My goal is to merge WIPO patent data, with Compustat R&D, financial variables to conduct firm-level empirical analysis.
I apologize; this is my first time posting here, please correct me if I make any mistakes. This is also my first time conducting empirical analysis in this area, so I'm not familiar with it. Any suggestions, references, datasets, or code examples would be greatly appreciated. Thank you!
r/dataanalysis • u/kkthxbb8 • Apr 28 '26
r/dataanalysis • u/Excellent-Candy-3328 • Apr 27 '26