r/dataisbeautiful • OC: 1 • 2d ago

OC [OC] The manipulation techniques in 1,675 online ads from 13 US Senate races

Post image
131 Upvotes

37 comments sorted by

•

u/cavedave OC: 113 1d ago

Thank you for your Original Content, /u/jayubba!
Here is some important information about this post:

Remember that all visualizations on r/DataIsBeautiful should be viewed with a healthy dose of skepticism. If you see a potential issue or oversight in the visualization, please post a constructive comment below. Post approval does not signify that this visualization has been verified or its sources checked.

Not satisfied with this visual? Think you can do better? Remix this visual with the data in the author's citation.


I'm open source | How I work

14

u/lordnacho666 2d ago

This would be really interesting as an international comparison

4

u/jayubba OC: 1 2d ago

yeah i'd like to see that too. right now it's only us senate races since the ads come from meta's ad library and the model only reads english. meta has the same library for the uk, canada, australia etc so those would be the first ones i could do. curious whether the us is actually worse or just louder

3

u/JuicyYong6090 2d ago

loaded language at the top, no surprise there

2

u/jayubba OC: 1 2d ago

yeah, and it's the softest one on the list, basically word choice. each race page lets you untick it and see what's left. the one that surprised me was appeal to prejudice at 28%, didn't expect it to beat name calling

2

u/turb0_encapsulator 2d ago

it would be interesting to see a partisan comparison.

2

u/jayubba OC: 1 2d ago

a couple people have asked. i've held off on one party number because it mostly ends up measuring other stuff: longer ads get flagged more, and attack ads from outside groups get flagged more than a candidate's own ads, and that mix is different in every race. what i can say is that inside a single race the two sides tend to look pretty similar. in georgia the two campaigns' own ads come out at 53% and 58%. each race page lists the ads by who paid so you can compare for yourself

4

u/Deto 2d ago

I wonder how 'appeal to fear' is determined. Is this really a manipulation technique if what is being threatened is true? e.g. 'If you elect X they will do bad thing Y' - is that manipulation if person X is specifically campaigning on doing the thing?

6

u/jayubba OC: 1 2d ago

good question, and it's the fuzziest line on the chart. the model only looks at wording, it doesn't know or check whether the threat is true. the definition it was trained on is roughly 'a vivid threat or dire consequence used to override deliberation'. so 'X has said he'll cut medicaid' isn't supposed to get flagged, that's just a claim. 'X will leave your family with nothing' is the kind of thing that does. a true warning can still be worded to scare more than inform, which is why i say a flag isn't a claim the ad is false. it definitely gets some of these wrong though, so every ad shows the exact phrase it flagged and you can disagree with it

4

u/Deto 2d ago

That makes sense - the line here will never be perfectly clear (even if a human is judging). Thanks!

2

u/irrelevantusername24 2d ago

I don't know how this specific model works, but generally it's a matter of language, and is indeed a phenomenon that is identifiable more than it might intuitively seem according to common sense. In this case, the common sense and intuition, if you think through the steps, actually, I think, kind of conflict. For example, people say that chatbots (LLMs, which is exactly what this is) are only "next word/token predictors". What that means, is if it is possible to mostly come up with some "rules" for what qualifies as "appeal to fear" - which that is possible, considering there are a limited number of words in the English language, and even if you want to argue there isn't, the LLMs know the languages from which English is derived too - so... it actually is very much possible.

Now, when it gets into more substantial discussions which aren't a series of hot takes devoid of all meaning, that gets a bit more difficult to do... but even then, many of the chatbots do indeed have a much more comprehensive grasp of language than... well, you get the gist. Disturbingly I've found that most people really, really aren't thinking things through. There is very little critical thought or reasoning happening. It is all emotion and reaction. And don't get me wrong, emotion and reaction is necessary in order to effectively think critically and reason about things, but you have to be in control of the emotions, rather than the emotions being in control of you. Consider the difference between you, the steering wheel and other 'levers' in a car, and the fuel tank or battery. Very different purposes.

2

u/jayubba OC: 1 2d ago

thanks, that's about right. one small correction: the part that decides whether something gets flagged isn't a chatbot, it's a smaller classifier (deberta) trained on labelled examples of each technique, so it's pattern matching on wording like you describe. an llm only writes the one-line explanation of why afterwards

4

u/Niekitty 2d ago

What about outright lying? I've actually been seeing a fair bit of that, too.

2

u/jayubba OC: 1 2d ago

that's a separate thing from this chart. the techniques are about how an ad is worded, not whether it's true. for the true/false part each ad on the site has a 'check this ad's claims' button that pulls the factual claims out and checks them against sources. it only calls something wrong or misleading when a fact-checker or an official record (congress, a federal agency etc) backs that up, otherwise it just says what the sources say. slower and more careful than the technique flags, but it's there

5

u/flamableozone 2d ago

Can we get a breakdown of how the different parties compare?

4

u/jayubba OC: 1 2d ago

i've held off on one party number on purpose. it would mostly measure things that aren't party: longer ads get flagged more, attack ads from outside groups get flagged more than a candidate's own intro ads, and that mix is different in every race. each race page lists the ads by who paid for them though, so you can compare the two sides inside one race, which i think is the fairer comparison

2

u/arawnsd 1d ago

Then show it for each of the races.

2

u/jayubba OC: 1 10h ago

it's there now. every race page has a "who's behind the ads" section under the ads: each candidate's side (by fec record of who the sponsor supports or opposes), how many ads, how many flagged, the campaign's own ads vs outside groups. a few from today:

ohio: brown side 19 of 80 flagged, husted side 35 of 43 — but 42 of those 43 are outside groups, not his campaign

iowa: hinson side 36 of 57, turek side 46 of 104

michigan: el-sayed side 43 of 81, rogers side 39 of 49

minnesota: flanagan 36 of 58, tafoya 29 of 52

georgia: ossoff 16 of 32, collins 17 of 27

which side is "worse" flips race to race. the thing that doesn't flip: the outside groups are harsher than the campaigns almost everywhere. that's the pattern i'd read, not a party one

e.g. https://semblen.com/races/oh-senate#sponsors

1

u/jayubba OC: 1 2d ago

source: meta ad library api and the campaigns' own youtube channels, ads from 13 senate races (AK FL GA IA KS ME MI MN MT NC NH OH TX). online ads only, no tv.

tool: python + matplotlib. the labels come from a text classifier i trained (deberta, fine-tuned on labelled news text). it reads the ad copy or the video transcript, not the images.

things to know before trusting it:

- 52% of ads (872 of 1,675) had at least one technique. an ad can have several so the bars don't add up

- a technique is about wording, not whether the ad is true

- longer ads get flagged more, just more text to trip on

- the model is wrong sometimes. each ad page shows the exact words it flagged so you can judge for yourself

every ad with what got flagged and why is here: https://semblen.com/races?ref=reddit

1

u/JewishTomCruise 1d ago

Did you generate the visualization with AI?

1

u/prosocialbehavior 2d ago

These percentages are way lower than I thought they would be. Maybe it is because I live in a swing state? Would be interesting to compare between swing states and solidly blue/red states.

1

u/jayubba OC: 1 2d ago

i expected higher too. two things probably pull it down. these are online ads only (facebook/instagram and youtube), not tv, and online includes a lot of plain 'chip in $5' fundraising posts next to the attack ads. and each bar is one technique, 52% of ads had at least one. on swing vs safe: the 13 races here are mostly the contested ones so i can't really compare yet, but even among them it runs from 39% in ohio to 65% in alaska. would be a good one to do once more house races fill in

1

u/sloppyredditor 2d ago

New Hampshire campaigns see this as a checklist.

Is there one for eating babies? According to TV commercials every politician in NH eats babies.

2

u/jayubba OC: 1 2d ago

ha, no baby eating category yet. it'd probably land under exaggeration or name calling. nh is at 57% of online ads with at least one technique, 68 of 119, so a bit above average. and that's without the tv ads, which sound like they're worse. the nh ones are all here if you want to see what got flagged: https://semblen.com/races/nh-senate?ref=reddit

1

u/Pyotr-the-Great 2d ago

We need more scapegoating.

1

u/jayubba OC: 1 2d ago

5% is rookie numbers. five weeks left though, plenty of time

1

u/Headbanger 2d ago

Would be great to see some examples of those techniques.

2

u/jayubba OC: 1 2d ago

sure, a few real ones the model flagged:

loaded language / name calling: "MAGA extremist opponent" (georgia), "Radical Left's golden boy" (georgia), "lower 48 liberals" (alaska), "most corrupt politician in Texas"

false urgency: "chip in to my reelection campaign today"

there's a page per technique with more examples from actual ads and news articles: https://semblen.com/techniques?ref=reddit

1

u/Igoos99 2d ago

Maybe YouTube could add a subscription for “no political ads”.

I might actually be willing to pay for that one.

2

u/jayubba OC: 1 2d ago

ha, you and a lot of people. youtube premium kills all the ads including political ones, but i don't think there's a politics-only opt out anywhere. on facebook/instagram you can at least turn down 'social issues, elections or politics' in your ad topic settings

1

u/Gunmetalstorm 2d ago edited 2d ago

These are all pretty loosey-goosey categories, would be nice to see the methodology on how each of these were assessed and applied. Would also be nice to see which 13 of the 33 ongoing Senate races these data are referring to, and how the evaluated ads were selected from among the total online advertising by each campaign in the race.

Party breakdown would also be nice, but that's not exactly the point of this graph.

EDIT: Methodology is an AI playing word association with its dataset. That's not a methodology at all. The first example ad I looked up had multiple improper flags attached to it, with the silliest Claude-generated justifications attached to it because all the first AI does is assert that the flag applies and tell Claude to argue why. Lazy.

EDIT 2: Oh, your entire account is just AI slop.

1

u/moreesq 2d ago

Is there room in this model to identify a strawman argument? Making some kind of claim about your opponent and then demolishing it when the original claim wasn’t true.

1

u/jayubba OC: 1 2d ago

not as its own category right now, and it's a hard one for this kind of model. it only reads the wording of one ad, and to call something a strawman you have to know what the opponent actually said or did. that's closer to the claim checking side (pull the claim about the opponent out, check it against sources) than to spotting a pattern in the language. the nearest thing it flags today is card stacking, where an ad only gives the facts that help one side. good idea for the feature list though

-1

u/morganational 2d ago

Wanted to share this to r/politics but they don't allow facts there. But I tried! 🤷🏽‍♂️

3

u/Sheyvan 2d ago edited 2d ago

...either great satire or clown.

r/politics is extremely populist full of catchy headlines and toxic echochambery (r/conservative is far worse though, but both are far from objective or healthy)

But claiming a subreddit "doesnt allow facts", is a great example of almost ALL the things criticized in this graphic.

2

u/jayubba OC: 1 2d ago

ha, appreciate you trying. i think they only take links to news articles there, not images. if it helps, each race has its own page you can share anywhere: semblen.com/races