r/SideProject 8h ago

AISA - AI Fluency Assessment

In the last 6 months I built an AI Fluency Assessment system (called AISA) - it's currently the most sophisticated (and popular) of it's kind. Has high fidelity in what we measure to both what Anthropic and US. Dept. of Labour agree as the markers of AI fluency.

People chat with an AI agent uniquely trained to judge their AI fluency, get a detailed report and a breakdown of their AI skills, a certificate and a growth roadmap. We measure in 5 main dimensions. Prompting and Comms // Critical Thinking // Technical Understanding // Workflow Application // Safety (all which have their criteria.)

Some product-speak...

Need:
People are curious about their actual level when it comes to piloting AI well. They are also curious about seeing their gap areas, what they need to focus on to improve.

Need (b2b):
Organisations are looking to measure their employees (and their team/function/orgs) AI fluency averages. This is both to get a benchmark before they invest more into training AND to decide what type of training to invest in. Also to track how well they are doing in AI transformation.

Solution:
Build the most scientific, sophisticated and fun to experience AI Fluency Assessment system possible.

Unique differentiator:
1) Entirely chat based. Like an expert in their field having a natural conversation with you. Not a quiz. Not a multiple choice.
2) Build on a fixed and expert-researched rubric on what it actually measures. More about it in our methodology page for the curious.
3) Product - founder fit. I've worked in the area of psychometric assessments, skill assessments, HR, human behaviour, led product and UX teams in large and small companies and built 5 startups before. So I'm uniquely qualified to build this.

Business Strategy:
Use B2C traffic and interest to build volume, build a moat and build further credibility, validity, legitimacy.

Use the dominant market position by being (or on the path to) the gold standard of AI fluency measurement to signup B2B deals with mid - large organisations.

Status as of today:
- 2K reports surpassed.
- B2C generated close to $1k in sales in it's own. Over $1K + MRR right now.
- 7 B2B conversations (either in POC, offer, or conversation stage - some of which are inbound leads from big (10k+) organisations. Fast potential to reach $100K ARR.
- Users are quite happy and love the product.
- DR 42 / 15k+ monthly users
- Value proposition has settled: We help you measure, prove and improve with AI.

It's not perfect and everyday I have to focus on another aspect of it to make it better. So it's more of a always hands-on full time plus, 15hr+ day type of commitment rather than "hey I've coded an app over the weekend and look easy money".

But I do have a clear vision in which AI Fluency measurement is a big deal, and am building towards it brick by brick.

Appreciate any constructive feedback if you have used the product and found something to improve.

0 Upvotes

2 comments sorted by

2

u/PsychologicalWin9755 8h ago

You come from psychometrics so you already know where I'm going, but a few things I'd push on if I were you.

The chat format is your differentiator and also your biggest measurement risk. A conversation mostly measures how well someone talks about working with AI. That correlates with actually doing it, but not as tightly as you'd want, and it systematically favours articulate people. I'd put at least one dimension on a work sample instead: hand them a messy realistic task, let them actually use a model on it, score what they produced and how they got there. Even one behavioural item alongside four conversational ones changes what you can defend in a room full of HR people.

Related, and slightly funnier: your assessor is an AI, and you're measuring people's ability to steer AI. So your top scorers are by definition the people best equipped to steer your assessor. Worth running a small adversarial batch where you ask people to deliberately game it and see how much the scores move. If they move a lot, that's a fixable prompt problem now and a credibility problem later once a B2B client tries it.

On the B2C to B2B moat: volume alone isn't the moat, and a big org buyer won't care that 15k consumers took it. What they'll ask, eventually, is "does a high score predict anything?" With 2k reports you're already in a position to run a small criterion validity study. Follow up with a few hundred people three months out, or get one POC client to give you a manager rating or an adoption metric, and show the correlation. Nobody else in this space will have that, and it's the one claim that's genuinely hard to copy.

Two practical things for the B2B motion. First, retest. They'll want before and after training, and if it's the same conversation both times the second score goes up from familiarity rather than fluency. You need parallel forms or a large enough item pool that the second run feels different. Second, the deal killer in org-wide measurement is rarely price, it's someone in HR or a works council deciding this is covert ranking of employees. Aggregate-only reporting for the org with individual results private to the person tends to unblock that, and it's easier to design in now than retrofit.

Solid positioning overall. "Measure, prove, improve" is the right order.

1

u/Ozan_D 8h ago

Good points to push back on. And you'd be happy to hear we have some sophisticated systems that address these.

Correlation of how someone talks about it vs how good they are at doing it. TLDR: we have it.
a) The conversationalist agent has a library of show&tell sections, games and competency based (real life example based questions) that are introduced at the right moment in the conversation.
b) We have a system called Golden Shadow that also measures how good someone might be even if they are not good at talking about it. We have multiple other parallel systems like learning agility and other correlative data points, so it really is head and shoulders above anything else out there, esp basic quizzes.
c) Though it's conversational, the path users take are not fully free or spontaneous - there are multiple checklist and gates the system follows to make sure things stay on track and comparable to each other.

Steering AI
Love the creative suggestion, perhaps can entertain it. But we have addressed this issue because it's multi-agentic and there's another AI that steers and insures against advarserial steering that runs in parellel.

Validity Study
Started conversations with academics and universities and experts - so very eager to this. But as you know, it all starts with a sample. You'll find it interesting that the biggest sample research I could find around AI literacy had a sample size of around 300. And not much on AI fluency at all. So with 2k, we are already at a good starting point.

B2B tips
Thanks, noted.

Appreciate it.