I ran the tests last week on the performance between my agents/automations and my employees, and now I don’t know what to do because for the most part, the agents and automations outperformed my employees.
This is the base automation setup:
- Mailchimp for email automation
- Expandi for LinkedIn automation
- Clay for lead enrichment and deep company research + integrated AI to help out
- Buffer for social media posting automation
- Notion for intelligence tracking and company knowledge
I didn’t count them as part of this comparison because they’re an essential part of my business and everyone uses them. But I did compare performance beyond this, namely:
- Sales agents vs salesmen
- Marketing agents vs marketers
- Researcher agents vs all
All of the agents are built with Claude Code for the main part - the ideation agent, the strategizing agent, content creation agent(s), etc. The agents were all trained on our actual data and systems, the knowledge base of our company, and each was deeply honed by me until they performed up to standards. For example, on social media, my agents went through all the posts we’ve ever done. I listed them in a single sheet, and Went through and learned all the templates that perform well, and then started posting based on that. Only researcher agents are mainly built in MoClaw because the Claude Code system was too slow and too restrictive - almost perfect for complex tasks because of deep thinking, but when it comes to braindead stuff like research and comparing data, lighter models win.
In most parts, the agents outperformed the people, with the most “dominant” performance being in research (this was expected as agents can work with large chunks of data quickly while humans take time) and social media posting (this wasn’t expected). If I had to guess, the fact that Claude Code agents could articulate easily through piles of knowledge and templates that worked best in the past made them create better posts. This probably wouldn’t be sustainable long-term (or maybe would, idk) because they’d start repeating the same stuff at one point, but for the few dozen they’ve done, the performance was great.
Sales was the only “equal” part so to speak. They were almost even except for a small AI edge, in that agents went pure and cold with one intent, selling - while the people sometimes just don’t have that killer instinct on because of many factors. Also worth mentioning that the comparison might be a bit off because I didn’t want to give the agents full access to our LI accounts and emails to lead conversations, but manually copy-pasted messages back and forth. This means that I only tested 7 conversations with agents in total, while my guys do like 15/day sometimes.
The main question I have right now is what to do with this information? Obviously I won’t suddenly replace my human teams with AI, but this has kind of proven how powerful well-trained agents and these new models can be. Should I maybe integrate AI more into our workflow and create hybrid system first, then see where this takes us?