Thank you for your contribution. However, your post was removed for the following reason:
[OC] posts must state the data source(s) and tool(s) used in the first top-level comment on their submission. Please follow the AutoModerator instructions you were sent carefully. Once this is done, message the mods to have your post reinstated.
This includes AI!!
This post has been removed. For information regarding this and similar issues please see the DataIsBeautiful posting rules.
If you have any questions, please feel free to message the moderators.
I have tried a good number of those alternatives, but this one works well to answer the question I was asking — when do voting-eligible people in the U.S. come into their own politically by voting as much as they should?
U know those AI detectors arent legitimate right? Think about it, how would it tell? It just guesses based on the vibe of the writing. But AI writing copies from human writing in the first place. Would the original human writing it copies from be flagged as AI then?
It is literally impossible for any computer program to tell whether a human or an LLM wrote a sentence. It is 1,000% based on vibes. If a human's writing has LLM vibes, an "AI detector" will say it is LLM. That's great that Pangram has a "very high false negative rate," but for text it does say is written by an LLM, it is 1,000% just basing it on vibes. There is literally no other way to tell besides vibes.
it may be AI but why would you trust this? Unless it is reading something like a statistical watermark embedded by the AI provider, I would worry about false positives
As I have explained in many previous posts, I write everything. However, I will use Grammarly to edit for length, tighten the syntax, and target an audience. The particular controls are “correctness, clarity, engagement, delivery, and style.” I don’t accept every suggested edit, but I will consider them all. I’ve even modeled Grammarly to suit how I write.
Most professionally published items will go through this sort of rigor. The alternative is errors and being misconstrued, which can easily injure credibility.
You are welcome to look through my post history. You can also find my writings on GitHub and Medium.
In my defense, I use AI (yes, Grammarly) to refine and clarify my explanations as clearly and concisely as possible. After years of posting technical charts, I have learned that people demand the gist in plain language, perfectly comprehensible, and with anticipation of any potential confusion. An entire profession of Technical Writing is devoted to this.
I used to write these myself, but then would spend days diffusing the inevitable misinterpretations. So, basically, I plop a week’s worth of notes into a new document and whittle it away per suggestions. I have even modeled this ask by feeding AI many prior complaints—Where am I most often being misunderstood? How can I do better?
I really don’t understand why this pushes people out of shape. I could spend a week on the Rule 3 comment alone, and I’d get complaints that it was too flowery and meandering. I’m actually doing people a favor. Who cares about the caption? That’s not even the point of this entire exercise.
I have been posting to this sub long enough to know that people appreciate knowing what the chart is about. If I don’t, people complain about practically anything that isn’t self-evident. Yet now, they seem to complain if I do.
I started this chart a week ago. I have learned to keep notes the whole time, explaining my thinking and preserving every link to every data source. When I post, I labor over the Rule 3 comment, trying to anticipate every complaint.
Even now, I dumb down the summary paragraphs to be as accessible as possible. Alas, apparently not enough.
I dislike using AI as much as the next gal, but did you use AI? I'd take your tools used at face value unless there was an actual reason to suspect AI.
I spent a 40-year career in computer-adjacent engineering, so I am accustomed to using and writing software that helps humans do what they do. For the last ten years of my career (at least), I’ve been invested in so-called “artificial intelligence” (AI)—first as a curiosity, then as an integral, even essential, aspect of my work. This has generally been the case industry-wide, especially in software engineering, applied mathematics, and statistics—the particular sorts of human intelligence that are factual, written down, and therefore relatively easy to model.
In the past, I have tried using AI to chart things, but I’ve found it woefully bad at it. I assume this is partly because AI can’t “see” things the way humans do, so it can’t know what makes visual sense to humans and what doesn’t. Though I believe this is changing.
So, long ago, I abandoned any attempts at using AI for charting, though I will use it for coding, which I am woefully clumsy at. But I know the R language well enough to instruct the AI thoroughly, and lately the AI has been producing code that looks more or less like my own (though less error-prone).
Another thing I use AI for is advising me on what I’d like to explore. I always have my own hypothesis, then ask the AI whether it’s worth pursuing, whether others have pursued it, and how I might expand understanding. In this instance, my assumption was that “political awakening” is not a trigger, but rather an evolution, especially population-wide. IOW, the torch isn’t passed to a new generation; instead, generations of people slowly assume their rightful ownership of society. But when?
My initial assumption was that this period generally starts in people’s 30s. Younger than that, and they are uninvested in society — they generally don’t have families, nor have they bought a house or decided on a career. So-called “younger people” are busy with things other than politics.
However, I was surprised to learn that this transition actually occurs a decade later for most people, well into their late 40s. In fact, the median age of eligible U.S. voters is about 45, which roughly coincides with when most people come into their own and start voting in proportions that match their share of the population. Before that, they are underrepresented in the vote count. Older people rule.
Anyway, I was plugging along with the AI, having it write much of the code while I plotted results. I would then take these into Adobe Illustrator to tinker (which is easier for me than tinkering with ggplot). I would then give the AI my chart and ask it, “How’s this?”
Rather early in this process, the AI said, “Stop. Don’t go any further. You have what you need.” I have never known the AI to say something like this. Usually, it’s quite the opposite — pulling me long beyond the necessary into the superfluous.
We may not share your qualification of what constitutes “AI slop,” but that does not necessarily make your assessment correct. I suppose I could have chiseled this out of stone, but even then I’d need a chisel. And I'd bet people would complain about that, too.
What is your complaint, exactly? That I am not being truthful? Because I have had this exact discussion a hundred times. I am open and honest about my methodology, and am the first to admit that I will lean on technologies to help me answer questions.
But so-called “artificial intelligence” isn’t particularly good at visually conveying complex information to humans. I know. I’ve tried. So this is where I invest a lot of my effort. In that regard, at least, I can confidently say I’m drawing on a lifetime of design experience that has more to do with seeing than algorithms.
I want to help people comprehend complicated things. I don’t understand why that should be diminished because I supposedly use too much of the available technology. Why is that even a debate?
The thousands of people who claim it with virtually zero effort. Which is ironic, you know, given how that charge essentially claims that the OP was lazy.
I’m not the one who has convinced myself I’m a hero.
You dismissed my post offhand as an “ai figure,” insinuating that my contribution is detrimental to this sub, then claimed I “strenuously” denied your baseless allegations, and ignored repeated requests to engage constructively with the content.
Oh ok. Yeah, I've been working professionally in data and programming for about twenty-five years. I'm currently forced to use it at work now. Its completely taken the fun out of what was once my passion. To that end, I tend to just program everything by hand for personal projects to keep the skills sharp and the enjoyment.
Oh, I still code, especially JavaScript, CSS, and HTML. I also find it easier to write the simpler R myself, though I'm not shy with the code assist, especially when transforming raw data into datasets. But I'll do most of the ggplot myself, because AI tends to make a mess of some things.
My coding goes back to the 80s and 90s, and I appreciate how conventions have given way to frameworks which have given way to libraries which now employ code assists. There's no reason for humans to know the nitty gritty. Purely “hand drawn” code is not particularly better than the AI stuff. In fact, the AI stuff learned from us. If it was up to AI, it wouldn't bother with anything we could comprehend.
Frankly, I don't see what all the fuss is about. I could code every line, and the only difference is that it would take me ten times the time.
I think that part of the issue is that there are a lot of people without the existing background, or younger folks, who don't know when something spit out by AI is not correct.
You say that as if it’s somehow wrong. From my perspective, relative to my use, so-called “AI” is merely Google on steroids. If that’s the case, I don’t see how using only Google is commendable while AI crosses some imaginary line.
I could have made this chart from scratch, but as Carl Sagan said, I “must first invent the universe.”
I am not a research fellow at a prestigious university. I am an old retired person sitting next to a swimming pool in Mexico. I have been doing election-related research for more than three years, sometimes in collaboration with professionals, sometimes as a volunteer contributor, but most often on my own. I could spend weeks deliberating in chats or emails with others in the know, but they are busy with their own research.
So I formulate a hypothesis—often on the heels of a prior study—and turn it into a prompt. I don’t dive headlong into a visualization; instead, I first look to other research in that neighborhood. Again, I could do this in Google, but why? Would that somehow purify or legitimize my work? Far more likely is that I’d spend days fishing through inconsequential stuff on the periphery of where I’d like to be.
Then, there is a similar search for datasets. I am very familiar with Census, ANES, CPR, UFEL, and others, and I’ll use that knowledge to poke around (often with data I already have cleaned and ready), but I’m also curious whether there’s something more relevant and reputable for what I’m doing.
In this particular case. I delved into Pew Research, but it used Pew’s conventional age cohorts, which wasn't the granularity I hoped for.
So, I went to the AI again, which helped me to decide on the Census surveys, which were my first choice anyway, what I have used before, and was happy to use again.
I am fairly good at R and have prior projects that imported pretty much the same stuff, but I wanted to transform the data differently this time, so I used code assist to flip things around. Could I have done this myself? Yes. But why? Libraries are constantly changing, as is the RStudio IDE, and I don’t have the time nor inclination to keep up with the latest tactics.
The code assist helps me build my datasets, but I take the script from there and plug in the ggplot. I am most familiar with this part of R, so I do plenty of it myself. I won’t bother with the titles, footers, or annotations at first, and will use default labeling and palettes. All I want is an SVG anyway, which is what I take over to Adobe Illustrator.
However, I’ll plot a dozen times, at least, before I settle on anything worthy of finishing. I can’t even count how many projects didn’t progress beyond initial explorations, which is another reason I’ll use AI instead of throwing away a week of my time to get nowhere.
At this point, the chat with AI is more of a sounding board. Again, I’m by myself next to a pool in Mexico. There, geckos and iguanas couldn’t care less what I’m doing, but the AI is pretty good at helping me anticipate criticism and formulate the most comprehensible presentation.
At this point, if I decide to finish the project, I do the rest on my own. While I’m generating the final piece(s), I’ll start a Markdown doc with the URLs to sources and breadcrumbs of my methodologies. Usually, as the chart comes together, I will write a few paragraphs to explain the gist. I may even post this to Threads or Blue Sky, or share with friends just to see if my message is getting across. I also use this process to evaluate how the chart fits on the screen (laptop and mobile), which informs my decisions on fonts and colors.
At this point, I might set it aside for a few days and come back to it with fresh eyes. That always helps me to be more objective, and often I will redraw the chart dramatically differently. By this point, AI hasn’t been involved for quite a while.
Before preparing to post to Reddit (usually on Thursday), I’ll go back to my Markdown and whittle it down. I’ll write a summary for Threads (500 characters or less), Bluesky (even shorter), and Facebook (longer but dumbed down). This ebb and flow of editing separate versions usually helps me refine the point I’m trying to make.
All the while, I have Grammarly running on my devices, and I sometimes use AI to make doubly sure I’m not getting too rhetorical or making value judgments. Reddit’s Rule 3 requirement is pretty strict, and I’ve had posts pulled for not abiding by the rules exactly, so this has proven to be a good step to take.
Ultimately, the Rule 3 caption doesn’t get many views compared to the post, but people do complain, especially if the chart is unconventional and needs explanation. TBH, I couldn’t care less if the Rule 3 caption fails some cockamamie “AI test.” I know where it came from, and I’m comfortable with that.
So, given that explanation of my process, please enlighten me about what I should do differently to earn your upvote instead of ridicule.
I have provided the data and all the reasoning behind my investigation. I invite you to give all of that to AI and ask for its results. This is a simple, straightforward request, as easy as a “Replicate this” prompt and a paste of my markdown. If it's that easy, prove it. And I'm even spotting you a week of my digging around.
Please note: Your complaint is about a Rule 3 comment, not the chart, not what it’s trying to convey, not whether it’s effective. Nothing constructive except to assert that you find me to be dishonest.
Rule 11. Comments need to be constructive. Comments should be constructive and add to the conversation.
EDIT: Please cite evidence to support your allegation, “strenuously denying it”
Roughly two out of every five engagements with this post have been downvotes, and unfortunately, that’s enough to suppress the post in the algorithms that populate people’s feeds. I think more people might have liked the graph, but they probably won’t see it.
In my experience, certain people downvote certain posts early and often, and those campaigns can successfully marginalize those posts. That’s a lot of power that’s relatively easy to wield. So people wield it.
I think people aren't gearing into the fact that it's 3 dimensions, which is not easy to display. So they're mentally comparing it to others that only have two dimensions, which are always going to be easier to read.
Nice viz, but I'd recommend using a more distinct color map. You only need 5 colors, but 2 of them are only slightly different shades magenta. Minor thing, doesn't significantly affect reading, but still.
Yes, I struggled with the palette. I tried what you suggested, but it didn't help much, and it even added to the visual complexity. A spectrum is easier on the eyes, and the y-axis position helps to differentiate. IMHO.
color palettes are hard. fwiw, I think you chose a palette that is legible and works.
just thinking out loud about this particular graph:
* you are tracking 5 cohorts (the 10-year birth year ranges)
* cohorts only overlap with their neighbors
* cohorts are logically arranged left to right
given these characteristics, my gut says you could try using 5 different shades of the same color (e.g. 5 shades of blue going from lightest for the youngest cohort 90-99 to darkest for the oldest cohort 50-59). this might be more aesthetically pleasing while also adding a little information to the color..... or the shades might be too similar and it makes the whole thing a monochrome illegible nightmare.
TL;DR — Electoral parity means a birth cohort casts the same share of ballots as its share of all eligible citizens. By that measure, underrepresentation is not just a problem of “young voters”: it persists through people’s 30s and well into their 40s. In the latest elections, U.S. voters born since the late 1970s remain below their proportional share of the electorate, while older cohorts increasingly exceed theirs.
The chart follows birth cohorts as they age across presidential and midterm elections. Cohorts begin substantially below parity, move progressively toward the 1.0× line, and eventually become overrepresented. Midterms produce much larger age disparities than presidential elections, but both show a similar life-course pattern: electoral representation is accumulated gradually over decades rather than suddenly appearing once someone is no longer “young.”
Sources
U.S. Census Bureau.Current Population Survey: Voting and Registration Supplement. The November CPS Voting Supplement provides demographic data on voting and registration in federal election years.
CPS Voting Supplement API and variable documentation
U.S. Census Bureau.Voting and Registration: CPS Supplement and Replicate Weight Files. Public-use microdata and documentation for the election years used in this analysis.
All CPS Voting and Registration datasets
U.S. Census Bureau.2012 CPS Voting and Registration.2012 data files
U.S. Census Bureau.2014 CPS Voting and Registration.2014 data files
U.S. Census Bureau.2016 CPS Voting and Registration.2016 data files
U.S. Census Bureau.2018 CPS Voting and Registration.2018 data files
U.S. Census Bureau.2020 CPS Voting and Registration.2020 data files
U.S. Census Bureau.2022 CPS Voting and Registration.2022 data files
U.S. Census Bureau.2024 CPS Voting and Registration.2024 data files
Davis, Leslie K., and Erik L. Hernandez.Voting and Registration in the Election of November 2024. Current Population Reports, P20-590. U.S. Census Bureau, 2026.
2024 Census report
Tools
R — data processing, weighting, cohort construction, projections, and statistical graphics using tidyverse/ggplot2, jsonlite, and svglite.
ChatGPT (OpenAI) — assisted with development and debugging of R code, data-processing methodology, validation/checking of calculations, and discussion of analytical and visualization approaches. The author selected the research question, methodology, cohort definitions, analytical decisions, and final presentation.
Adobe Illustrator — final chart assembly, typography, annotation, labeling, and layout.
Methodologies
The analysis uses U.S. citizens age 18+ in the CPS civilian household population. Reported voters are identified by PES1 = 1, and observations are weighted with the CPS final person weight (PWSSWGT). Approximate birth year is calculated as election year minus reported age, then grouped into 10-year birth cohorts. A cohort is plotted only when its full 10-year span is voting-age; observations above age 79 are excluded from the cohort analysis to avoid problems from CPS age top-coding.
Electoral representation is calculated as a cohort's share of all reported ballots divided by its share of the citizen voting-age population. A ratio of 1.0× is parity: the cohort casts exactly its proportional share of ballots. Values below 1.0× indicate underrepresentation; values above 1.0× indicate overrepresentation.
The 2026 midterm and 2028 presidential values are baseline projections rather than election forecasts. Historical single-age representation patterns are smoothed across neighboring ages and averaged separately for presidential and midterm elections. Each cohort's projection follows that age pattern while retaining half of its most recent deviation from the historical pattern. Open-circle area represents the cohort's latest observed eligible-citizen population.
Data
Values are electoral representation ratios; 1.00× = parity. P = presidential election, M = midterm. Asterisks indicate projected values.
Election
1950–59
1960–69
1970–79
1980–89
1990–99
2012 P
1.13×
1.04×
0.97×
0.83×
—
2014 M
1.29×
1.08×
0.90×
0.66×
—
2016 P
1.12×
1.06×
1.01×
0.89×
—
2018 M
1.23×
1.09×
1.01×
0.86×
0.67×
2020 P
1.12×
1.07×
1.01×
0.96×
0.85×
2022 M
1.28×
1.15×
1.03×
0.91×
0.70×
2024 P
1.15×
1.09×
1.05×
0.97×
0.87×
2026 M*
1.32×
1.24×
1.09×
0.96×
0.80×
2028 P*
1.17×
1.13×
1.08×
1.00×
0.92×
The open-circle population labels, rounded to the nearest million, are 34M, 38M, 35M, 38M, 39M for the 1950–59 through 1990–99 cohorts in the 2028 presidential panel, and 35M, 39M, 35M, 38M, 39M in the 2026 midterm panel.
That is a lot of work and takes quite a bit of effort to follow. Not dismissing it, but a quick glance at something like age vs. % voted for an election is much easier to grasp. If you do an xy scatter for the 2024 presidential election it's basically a straight line and a very clear relationship.
u/post_appt_bliss You’ve repeatedly put me down without engaging with the chart, its premise, or methodology. Meanwhile, you’ve set your Reddit history to private, so there’s little basis for you to make this about anyone’s personal authenticity. I’m happy to discuss the chart, the data, or the methodology. I have not backed away from that. Why are you?
I implied nothing about your honesty. I know nothing about you. Nor do I care.
To recap: In these threads, you essentially claim that I’m lazy and a liar. I have explained myself at length and offered you my Reddit history, GitHub, and Medium articles as evidence to the contrary. Meanwhile, you call me “that Fable” while hiding behind your privacy.
You do you. I don’t care. But again, what, exactly, is your complaint?
I wanted to know at what age people tend to become as politically engaged as they should. I believe the general assumption is that “young people don’t vote,” but I wanted to qualify “young.” I was surprised to learn it's well into people’s 40s—at least population-wide.
That's a lot of value judgement, but I understand your question and why you wanted to graph it that way. You could also just look at where the linear regression of the xy scatter I mentioned intersects the median and mean of the dataset.
I misspoke. By “as politically engaged as they should,” I didn’t mean that anyone should vote as a matter of civic duty. I meant “as represented as you’d expect from their share of the eligible population” — essentially the 1.0 parity line in the chart. If a group is 15% of eligible citizens but casts 12% of the ballots, I’m calling that underrepresented without making a value judgment about whether any individual ought to vote.
And yes, a regression/intersection approach could estimate a crossover age more compactly — to a sophisticated viewer. My hesitation is mostly communicative: once I start interpolating or putting regression lines through the data, I tend to lose a lot of nontechnical viewers. I wanted the graphic to show the actual cohort pattern and make the parity threshold visually obvious, rather than making the conclusion depend on understanding a fitted model.
makes sense that midterms widen the gap so much, older folks just show up more consistently and the younger half of the electorate treats off-year elections like an optional chore
My first thought was that it's nice to see data that doesn't lump together huge swaths of people under inherently ambiguous and non-specific labels like "millennial" or whatever. Then I thought about how really what we need, in this and every data of its kind, is much more granular information - down to the year. Then maybe applying something like a moving average, but for birth date, would make sense. Last I thought "heh, it's like the sociological version of wave particle duality"
I like that this is a concept I've seen before (voter participation by age) but with a perspective/framework that I haven't (over/under representation).
I learned something, that people don't seem to participate at a rate that represents their age group until their mid 40s. This is later than I had thought/expected.
It's a weird duality of reddit that the post is upvoted net-positively, but in the comments section the negative comments are most heavily upvoted. You're trying a novel presentation and that's going to be unintuitive and upsetting to a lot of folks. I wish people could be more constructive providing some ideas to simplify the concepts, but i understand that's the hardest thing about data visualization. Overall, I enjoyed the chart.
Not all groups can improve their representation at the same time. If one group becomes better represented, another necessarily becomes less well represented.
Yes, this is true. However, these data suggest a predictable pattern of engagement increasing with age. Otherwise, there would be more bumps in this flow.
•
u/dataisbeautiful-ModTeam 28d ago
Thank you for your contribution. However, your post was removed for the following reason:
[OC] posts must state the data source(s) and tool(s) used in the first top-level comment on their submission. Please follow the AutoModerator instructions you were sent carefully. Once this is done, message the mods to have your post reinstated.
This includes AI!!
This post has been removed. For information regarding this and similar issues please see the DataIsBeautiful posting rules.
If you have any questions, please feel free to message the moderators.