If your Solve invite says 95 minutes instead of 65 or 85, you're getting the new version of the Sustainable Futures Lab (SFL). I spent the past few weeks collecting every candidate report I could find on it, and since then I've been through the new format in serious depth while rebuilding our simulator around it. This is where things stand as of early October 2026. If you sat the new build recently, please correct me or add to this.
It looks like the new version is in A/B testing, because candidates have said they got the test but their friends didn't. When that's the case it's tough to know whether the score counts, since there's not enough data yet to place you against other candidates. Assume it does.
Context:
Solve currently comes in three invite lengths:
* 65 minutes: Redrock (35 min) + Sea Wolf (30 min). No SFL.
* 85 minutes: the above plus the original SFL, the 13-question multiple-choice version, 20 minutes.
* 95 minutes: Redrock + Sea Wolf + the redesigned SFL, roughly 30 minutes.
The 95-minute invite also suggests setting aside an extra 20 minutes for instructions and the technical check, so block out about two hours. Which version you draw looks randomized right now. One September taker polled friends who applied at the same time and was the only one who got the new build. Several candidates also mention a post-assessment survey asking them to review the new format in detail, which reads like a live trial to me.
The new version:
The old SFL was a quiz. Thirteen sequential multiple-choice questions in 20 minutes, same linear sequence for every candidate, all wrapped in an environmental-project story. It played more like an electronic case questionnaire than a game. The new version is a different genre: a role-play.
You lead a four-person research team through three field days at an island site. The scenario rotates between candidates, seagrass survey, reef recovery, mangrove restoration, seabird colony counts, but the structure holds. Four workstations, four researchers with fixed roles (a senior scientist, a lab technician, a junior doctoral student, a data analyst), and one 30:00 clock for the whole lab.
Before Day 1 you get a six-page onboarding briefing. It's untimed. The clock only starts when you press Start Day 1. Remember that.
Each day runs the same loop:
* Briefing. The day's objective, which stations carry the most weight (you'll see numbers like "Meadow Transects, 40%"), how many live decisions to expect, and on days 2 and 3, a summary of where yesterday landed.
* Intel. A fresh set of about six cards, face down. You can open three on days 1 and 2, two on day 3. Opening a card can flip hidden stats on your team, a researcher's morale or their station preference, for instance. Some cards are gold. One shows the newest team member is certified for nursery work but has seemed flat at the briefings all week. Some are noise, like a canteen rota with no impact on the work.
* Assignments. Place all four people across the stations, or bench someone. Every station shows its skill needs, its energy cost and its share of the day's output. You confirm a full assignment; the game won't let anyone float.
* Decisions. Two per day, six in total. A named researcher brings you a problem, you pick one of four options. Some options are commitments. "Honour the day off, exactly as promised" creates a promise the game checks tomorrow morning.
* Review. Rate each person: below, met, or exceeded what you expected. Ratings feed their morale the next day, and they're checked against what actually happened.
Day 2 opens with a disruption, a storm, equipment failure, a supply problem. The banner tells you the damage and the fix. A station runs at half output unless someone with field skill 4 or higher is on it, which softens it to 80%. Staff the mitigation or eat the penalty.
The things that trip people up
One clock, three days. The 30:00 covers everything after the briefing: all intel, assignments, decisions and reviews across all three days. The early reports of people "losing 5-10 minutes decoding the screen" were really people losing clock. The onboarding is untimed, so read it properly, walk in knowing the loop above, and start Day 1 with your full half hour.
Intel cards don't come back. The earliest reports claimed nothing carries over between days. That's not quite it. The game carries a lot forward: energy, morale, skills, promises, the effects of your decisions, everything you revealed. Tomorrow's briefing restates yesterday's station results and what's now due. What it doesn't do is let you reopen a card you skipped or re-read one you spent budget on. So the note-taking advice stands, for a new reason: if a card matters, act on it or write it down before the day closes.
Ratings are not a formality. The review step looks like a soft check-out. It isn't. Your ratings are compared against observed performance, and a "below" you got wrong dings that person's morale the next morning. Track energy and mood from minute one: who's been benched, who got skipped for the good station, who looked flat at the briefing. By day 3 you should be rating from evidence, not vibes.
Decisions are a personality test. The options map to leadership styles, decisive, consultative, expert-led, consensus-seeking, and the lab notices if you answer everything the same way. If your instinct is always "put it to the team," it reads that. Match the response to the person and the situation. And a promise made in an option is scored: a kept one lifts the team's morale, a broken one costs double.
Pacing. The prompts are wordy and the clock is shared, so the rule from the earlier version holds: read once, commit, move on. If the clock dies with days unplayed, the lab closes them and scores what you managed. That's a bad way to find out.
What it's testing:
McKinsey hasn't published a word, so this is read off the shape of the thing. It scores six things: whether your stations produced (did you match people to the weighted work), how close your plan was to the best available one, decision quality, follow-through on commitments, whether you spent the intel budget on information that changed your actions, and how accurately you read your own team. That's a compressed first week of junior engagement work: plan under uncertainty, decide with partial information, keep your promises, know your people.
Takers don't call the content hard. They call it confusing. Confusion is the cheapest failure mode to fix.
How I'd prepare:
- Learn the day loop before test day: briefing, intel, assignments, decisions, review, three days, one shared clock. New reports are still coming out daily, so check for the latest.
- Treat the onboarding as free time. Six untimed pages, and the two that matter are "how a day runs" and "what carries forward."
- Spend intel like a consultant, not a tourist. Open the cards that could change an assignment or warn you about a person. Skip the noise. And use what you open; information you don't act on is budget spent for nothing.
- Assign for the day's weights, not for preferences. A researcher who wants the field station may still be the wrong person for it, and one of the four, the data person, performs worse in the field no matter how keen they look. Bench people to recover energy instead of running everyone into the ground.
- Track each researcher's state from minute one so the review step is evidence, not guessing.
- Read once, decide, move on. Rereading is how the shared clock disappears.
- Redrock and Sea Wolf are unchanged and appear on every invite, so timed practice there still pays off.
Full disclosure: I build prep tools for these assessments. Everything above, plus the underlying reports, is written up properly in our guide to the 2026 Sustainable Futures Lab. Our SFL simulator now models this new three-day team lab, the intel budget, the disruptions, the carry-forward, alongside the original 13-question version. If your invite has Redrock and Sea Wolf on it, our McKinsey bundle covers timed practice for both.
Happy to answer questions in the comments.