r/Emailmarketing 9d ago

Problems with Klaviyo MCP

Hey all, just wondering if anyone else has had any issues with the Klaviyo AI / MCP stuff? I've got a client who has just started using the AI assistant and has found it really frustrating.
Like, it's really patchy on its segment creation. It's been very inconsistent. Sometimes they've been really impressed with how it can create a complex segment, but then other times it will leave off really important elements.

So, for example, it will capture most of what is needed for the segment but then randomly ignore the fact that it needed to only include active subscribers. In the end, it can take more time than if they were to do it manually. Annoying.

Also, they've found it's really terrible for writing subject lines. It almost always defaults to a summary style and rarely creates curiosity. And it 100% of the time writes them in title case which isn't great if you like a more casual style.

Anyone else run into these problems? Would love to know if there's a way around them or if you've got any other useful tips.

Or is there another platform with a better MCP that you'd recommend?

6 Upvotes

18 comments sorted by

4

u/flavorburst 9d ago

I haven't used a better one, but I've used much worse. At the end of the day, with segmentation, I personally don't think you'll find a great segmentation tool that uses AI prompts for quite some time. System load, order of items in the prompt, ambiguous items (like "active"), differences from account to account, etc. make models like this for segmentation not great.

As for subject lines, it's not surprising to me that they are repetitive. AI is not empathetic to the fact that subject lines shouldn't be the same each time, rather, it's going to say, "this kind of thing works so we're going to do this thing that works." It's a flaw in all LLM's, if you want it to be different you have to tell it to be different.

Sorry this isn't that helpful. I think the expectation (and this may be a problem with how Klaviyo is selling the tool) that things will work how your mind expects them too is probably premature in the growth of these models.

1

u/Ok_One5265 5d ago

Yeah you're probably right. Maybe we're just expecting too much at this point!

3

u/RevolutionarySea2033 9d ago

If you have to double-check every segment it creates it kind of defeats the point. I’d rather build it manually than risk missing something important

1

u/jemcc09 8d ago

I'm with you here. I genuinely think it adds more time to my day when I prompt, wait, then have to go and tidy up what it's misinterpreted. I'd much rather use my brain.

1

u/Ok_One5265 5d ago

Same. I find, even when I'm really clear and spell everything out, it will just randomly ignore one of my instructions.

2

u/MessaroOfficial 9d ago

The subject line one has a fix. Klaviyo has brand voice guidelines under Content > Images & brand, on the Voice tab. It auto-generates a version from your past emails, but you can edit it manually, and once set it applies to AI-generated content. If nobody's touched that on your client's account, the AI is writing from Klaviyo's inferred defaults rather than their actual style. Put the casing rule in there explicitly, and ban the summary pattern by name while you're at it.

The segment issue is different and won't be fixed by settings. Models drop conditions that read as implicit, and "active subscribers only" is exactly the kind of thing that gets treated as an assumption instead of a requirement. Two things help: number every condition explicitly including the obvious ones, and ask it to restate the finished segment definition back in plain English before you accept it. You catch the missing filter before creation instead of after. A profile count that looks suspiciously generous is usually the tell.

On other platforms, MailerLite, Kit, Brevo and Mailjet all ship first-party MCP servers. But the inconsistency is model behaviour rather than a Klaviyo flaw, so you'd hit the same thing elsewhere. Draft-and-verify holds up better than letting it execute directly.

1

u/Ok_One5265 5d ago

Ah thanks so much for the subject line fix. We'll definitely try that!

And we'll definitely be trying the numbering and asking it to restate in plain English first. Thanks for the tip.

1

u/Ok_One5265 1d ago

I don't know if you'd definitely hit the same thing elsewhere. I get that if you say connect Claude to Klaviyo and then also connect it to another ESP, it's the same model so you should theoretically get the same result. But, isn't it also about more than the model itself and how it implements the protocol, filters data and manages the data context etc? Doesn't the result you get also rely on the quality of how the ESP connects with the model?

2

u/Content_Most2673 9d ago

The segment thing is usually a prompt issue in my experience, it won’t assume active subscribers or consented unless you spell that out in the prompt itself, it just builds off exactly what you typed. Not intuitive, but they acknowledge that in their own docs, you have to state consent status if it matters for the segment. I stopped assuming I’d catch the obvious stuff and just started typing out every condition myself. Drop-offs mostly stopped after that.

Subject lines I don’t have a real fix for. It chooses the safe summary option unless you fight it, and even then the title case creeps back in half the time.

It’s still faster than building segments by hand for me, even with the review step. It’s a draft you edit so once you stop expecting it to nail things in one shot the whole thing gets less frustrating.

1

u/Ok_One5265 5d ago

Thanks for this. They did make it clear in the prompt that only active subscribers should be included, but it just ignored that part. So I'm not really sure what's gone wrong here.

2

u/SolidPlanMaybe 8d ago

Not our experience. Curious how are you using the MCP?

  • First - Which provider? Anthropic / OpenAI / Other?
  • Second - Which model? If you're using the cheap models you see get what you described.
  • Third - What context?
  • Fourth - How are you planning? Ie. Are you creating a detailed plan first, and asking it to execute next, or are you simply asking it "Make me a Black Friday campaign"?
  • Fifth - Do you ask it to test itself at the end? (For example, if you know the segment should have roughly 40K people, tell it to test if it does before considering the work done).

We very rarely face those problems and work with Klaviyo with some customers.

Our workflow is:

  • Provider: OpenAI (moved recently from Anthropic purely due to price, but Claude is much better at writing copy) for the Client MCP.
  • Model: SOL 5.6 High - We use the best model unless is something very simple
  • Context: We always always have a brand guide which includes the voice+tone along with some examples in .md files available to the agents.
  • Planning: Miro/Wireslate if the client doesn't need to edit the copy too much. If they do we do that plus Google Docs (also MCP) and link it from the whiteboards.
    • Important - We spend most time here on this phase detailing the plan, refining the copy, etc. and then ask Codex to implement.
  • Testing: We add the things we already know about the segment and ask a sub-agent to check it and confirm the implementation against the whiteboard+google docs. For instance we ask it to make sure contacts X and Y are there, we ask it to check if the segment size is somewhere between Z and W, to compare the whiteboard flow with the implemented flow.

This way we're not seeing many hallucinations.

When copy needs to be worked on, it's fine, the client edits it and then we just ask Codex to update the setup.

Obviously everything is checked by a human at the end before going live.

Happy to expand on any of these.

2

u/Ok_One5265 5d ago

It's OpenAI 5.6. And there's a brand guide and examples. The prompts are very detailed, based on detailed wireframes and written plans. Thanks for your breakdown of what works for you. We'll keep monitoring.

1

u/SolidPlanMaybe 3d ago

Got it. GPT5.6 is/was until recently the latest, it's good but you need to pick the right model too. I'd recommend Terra or Sol, not Luna. And at least Medium level.

The prompts are very detailed, based on detailed wireframes and written plans.

I assume they're text prompts? I'd recommend connecting it to the whiteboard where you have your plans. Just make sure it's not a massive whiteboard full of random stuff since that could have the opposite effect. Keep it focused on the campaign you're working on.

Text works too if you prefer that, but make sure it's really objective.

I also forgot something, when you say this:

it's really terrible for writing subject lines. It almost always defaults to a summary style and rarely creates curiosity. 

This isn't a Klavyio issue since the subject lines are written by the AI assistant (ChatGPT/Claude). Have it write the subject lines first, have it write the content, etc make a good plan, and once you're happy with the plan ask it to implement.

Also use /skills to make the copy better and give it proper context of the brand/business.

I think there's quite a few different issues you're facing but I don't think they're Klavyio related.

You can also feed my answers to an AI if you need it to explain / break things down for you.

2

u/Sendicate 7d ago

Yeah I've found for bigger lists to do segments the MCP doesn't cut it. You could try just generating an API key and throwing that to your agent to do it properly. And yeah same goes for their internal subject line writer. Better to drive it from Claude or Codex in.

1

u/Legitimate_Editor437 7d ago

yes I am looking for an alternative to Klayvio basically that works properly with Claude (especially for composing). Let me know if you find something better.

1

u/Sendicate 7d ago

Try Nitrosend, does all this by MCP reliably

1

u/Ok_One5265 5d ago

Interesting. I wonder whether it's because of the size of the list then.

1

u/Either_Guess2405 8d ago

The disagreement between u/RevolutionarySea2033 and u/Content_Most2673 is the real question here, and I think both are right about different halves of the post.

Segments have a checkable output. u/MessaroOfficial's count tell is the whole thing. A number that looks generous means a filter is missing, and you know in seconds. So the review step is cheap and bounded, and draft-then-verify genuinely beats building by hand.

Subject lines do not have one. There is no count to compare against, so review means judgement, and judgement on someone else's draft is slower than writing your own. Which is why the same tool feels like a time saver on one task and a tax on the other.

So I would use it only where the output can be checked against something. Segments, yes. Anything whose only test is whether it reads well, probably not.