Working With AI · Course 3

Pressure-Test My Strategy in the AI Era

Find the weak points before your stakeholders do.

FormatOnline, asynchronous
Learning time2 hours
You needAn AI assistant + one recommendation you must defend
Start here

How this course works

5 minutes

Good strategy is no longer enough. It has to hold up under pressure, and the pressure has gotten faster and sharper. This course teaches a repeatable method for stress-testing your own recommendations before anyone else gets the chance.

You will work on two tracks throughout: the safe case, a fictional recommendation in Module 2 with deliberate weaknesses planted in it, and your real case, a recommendation you will need to defend in the next 90 days.

Confidentiality note: the safe case exists so you can practice with AI tools freely. When you move to your real case, follow your organization's policy on what may be shared with AI tools. The method works on an abstracted version of your case (structure, logic, and rounded numbers with names removed) nearly as well as on the full version.
Before you start
Name your real case
Module 1

The new scrutiny: how AI armed your stakeholders

18 minutes

What changed in the room

Ten years ago, a well-built deck had a structural advantage: management held the information. Board members, investors, and executive stakeholders could probe, but real verification took weeks they did not have. A confident narrative, delivered well, often carried the day on delivery alone.

That advantage is gone. Governance research now documents the shift directly: directors are using AI to independently benchmark company disclosures against peers, test management's numbers against outside market data, and analyze the historical board materials management itself provided. PwC reports that 35 percent of directors say their boards have begun using AI in their oversight work, including preparing for meetings and pressure-testing strategy. Board platforms now generate director-specific probing questions from the pre-read materials automatically.

Translate that into what happens to your recommendation:

  • Your benchmark gets checked in real time. The peer comparison you selected can be re-run, with peers you did not select, before you finish presenting.
  • Your numbers get tested against your own history. "This projects 25 percent growth; the last three initiatives you presented projected similarly and delivered 8. What is different?" is now a question anyone can construct in minutes from your old decks.
  • Your pre-read gets red-teamed. The two-day pre-read window is enough time for a stakeholder to have AI generate the twenty hardest questions your document raises.
  • Alternatives you did not present get surfaced. "What options did you consider and reject?" now comes with the stakeholder having already generated the option list themselves.

The asymmetry, and how to flip it

Challenging a recommendation with AI takes minutes, while building one still takes weeks. The attacker's cost dropped faster than the builder's. That asymmetry is why polished-but-fragile narratives are failing at higher rates in rooms that used to accept them.

The flip is equally available: every tool your stakeholders can point at your recommendation, you can point at it first. The same AI that generates their twenty hardest questions can generate them for you, two weeks earlier, while you still have time to fix what the questions expose. Scrutiny is coming either way. The only variable you control is whether it happens before the meeting, run by you, or during the meeting, run by them.

What survives scrutiny

Recommendations that hold up under AI-enabled pressure share a structural property: their reasoning is visible and already stress-marked. They state their assumptions rather than burying them. They show the alternatives considered and the honest reason each lost. They name their risks with responses attached, and they define in advance what evidence would change the recommender's mind.

Notice what that list does not include: certainty. Defensibility does not come from having no weaknesses. It comes from knowing your weaknesses better than your challengers do, and having decided what to do about each one. A leader who says "this assumption is the load-bearing one, here is why I believe it, and here is what we will see within 90 days if I am wrong" is harder to rattle than one defending a claim of confidence.

Check your understanding
What makes a recommendation defensible under AI-enabled scrutiny?
Exercise 16 minutes
Name your exposure

Quick answers on your real case, no deep analysis yet. Do not fix anything: this is a baseline of where you already suspect the soft spots are.

Module 2

The seven weaknesses of executive narratives

22 minutes

Strategic recommendations fail under scrutiny in patterned ways. The patterns repeat because they are not lying, they are motivated reasoning: each one is a shortcut a smart, honest person takes when they already believe their conclusion. Knowing the seven gives you a checklist for finding them in your own work, where they are hardest to see.

  1. The buried assumption. The recommendation depends on something never stated, so it is never examined. "The rollout completes in Q2" quietly assumes the two key hires close in Q1 and vendor timelines hold.
  2. The chosen baseline. Every comparison requires choosing what to compare against, and the choice is doing silent work. Growth "versus last year" when last year dipped. Benchmarks against the top quartile when the honest peer set is the median. The number is accurate. The frame was selected.
  3. The hockey stick. History grows at 6 percent; the projection grows at 25, with the bend arriving conveniently after the decision. The tell is a projection whose driver is the initiative's assumed success rather than any mechanism you can name.
  4. The friendly pilot. Evidence from a test run under the conditions most likely to make it succeed: the flagship location, the strongest team, volunteer users. Real result, wrong denominator. The pilot proves the idea can work; the narrative treats it as proof it will work everywhere.
  5. The missing alternative. The case argues for the recommendation without arguing against the alternatives, including the strongest one: doing nothing, or a smaller reversible version first. When no alternative appears, stakeholders correctly suspect the comparison was not run, or was run and lost.
  6. The convenient source. Key evidence from a party with a stake in the conclusion: the vendor's ROI study, the consultant hired by the sponsor. Also in this family: the three-year-old market study still cited because updating it might change the answer.
  7. The all-or-nothing ask. Full commitment requested up front, no staged decision points, no defined conditions under which the initiative stops. The absence of kill criteria signals that the recommender has not seriously imagined being wrong.

The safe case

Read this fictional recommendation. It contains at least five of the seven weaknesses, planted deliberately. Find them before opening the reveal in Module 3.

Bluewater Coffee Roasters: Retail Expansion Recommendation (excerpt from CEO memo to the board)

Bluewater operates 38 cafes across the Pacific Northwest with $46M in annual revenue. I recommend the board approve an $18M investment to open 12 new locations across two new metro markets over 18 months, alongside launch of the Bluewater mobile app.

The specialty coffee retail segment continues to grow, with leading chains posting 22 percent annual revenue growth. Our own app pilot at the Portland flagship generated a 31 percent increase in repeat visits, demonstrating strong customer appetite for digital engagement across our footprint. A market study by Cascade Growth Partners, who supported our original expansion in 2023, identifies both target metros as high-opportunity markets.

We project the new locations reach $14M combined annual revenue by year three, reflecting 25 percent year-over-year growth after opening. Company revenue has grown steadily at 5 to 7 percent annually over the past four years, and this initiative moves us decisively beyond that plateau. Speed matters: delaying entry risks ceding both markets to competitors. I recommend we commit to all 12 locations now to secure favorable lease terms, with construction beginning in Q1.

Exercise 212 minutes
Two audits

The reveal is in Module 3, inside the claim stack teardown.

Module 3

The stress test: strip, rate, attack, compare, break

32 minutes

The stress test is five steps run in order, each with AI as your red team. It takes 60 to 90 minutes for a significant recommendation, which is why it belongs a week before the meeting, not the night before. This module compresses it into a guided pass on the safe case; the toolkit appendix carries the full prompts for your real work.

Step 1: Strip it to the claim stack

Reduce the narrative to its logical skeleton: the recommendation on top, the three to five key claims holding it up, and under each claim, the assumptions it rests on. Prose hides logic. Stacks expose it.

The Bluewater memo, stripped (this is the flaw reveal — open after Exercise 2)

Recommendation: invest $18M in 12 locations plus the app, all at once.

Claim 1: the market opportunity is large and growing. Rests on: segment growth applies to us (benchmark is "leading chains": pattern 2, the chosen baseline); the Cascade study is current and independent (commissioned by the sponsor of the last expansion: pattern 6, the convenient source; and its currency is unstated).

Claim 2: customer demand for our model is proven. Rests on: the flagship pilot generalizes to 12 new locations in metros where the brand is unknown (pattern 4, the friendly pilot).

Claim 3: the financial projection is achievable. Rests on: 25 percent growth from a company that has grown 5 to 7 percent for four years, with no stated mechanism for the bend (pattern 3, the hockey stick), plus unstated assumptions about hiring, construction, and same-store performance (pattern 1).

Claim 4: committing fully now is the right structure. Rests on: an urgency claim ("ceding markets") that carries no evidence, and the absence of any staged alternative or stopping condition (patterns 5 and 7).

If your audit caught five of those, you found the plants. The stack format is what makes them findable in your own work, where motivated reasoning hides them from the prose reader you become.

Step 2: Rate the assumptions

Not every assumption deserves attack. Rate each on two axes: how load-bearing (does the recommendation survive if this is wrong?) and how uncertain (how strong is the actual evidence?). Your attack list is the quadrant that is both load-bearing and uncertain.

For Bluewater, the pilot-generalization assumption and the 25 percent growth rate are both squarely in that quadrant. The lease-terms urgency claim is uncertain but less load-bearing. Spend your scrutiny where wrongness is fatal.

Step 3: Attack the evidence

For every piece of evidence supporting a load-bearing claim, run five questions:

  1. Source: who produced this, and what did they want to be true?
  2. Age: when is this from, and what has changed since?
  3. Denominator: what population does this actually describe, and is it the population I am projecting onto?
  4. Counterfactual: compared to what? What would this look like under the do-nothing case, or a different baseline?
  5. Survivors: am I looking at the winners only? (The "leading chains at 22 percent" benchmark excludes every chain that expanded and shrank.)

This is where AI earns its place. Hand it your evidence summary and have it run the five questions adversarially. It will not know your industry's ground truth, but it is relentless at spotting which claims rest on which sources, and it does not share your motivation to go easy.

Step 4: Compare against the steelman alternatives

Build the strongest honest version of at least two alternatives: the do-nothing case and the strongest different approach. For Bluewater: invest in same-store growth and the app across all 38 existing cafes, or stage it: open 3 locations in one metro with defined success gates before committing to the rest.

The test your recommendation must pass: state, in numbers where possible, why it beats each steelman. If the staged alternative is nearly as good with a fraction of the downside, the honest recommendation may be the staged one, and discovering that before the board does is the difference between leading the meeting and losing it.

Step 5: Break it, then set the tripwires

Run a pre-mortem: it is 18 months from now and the initiative failed. Write the three most plausible failure stories, specifically, with causes. Then convert each into two artifacts:

  • A tripwire: the earliest observable evidence that this failure story is beginning. For Bluewater's hockey stick: "if the first 3 locations are below 60 percent of projected revenue at month 6."
  • A response: what happens when the tripwire trips. Pause openings, revisit the model, stop.

Tripwires plus responses are your kill criteria, and they repair pattern 7. Presenting them does not weaken your ask. It is usually the single strongest credibility move in the room, because it demonstrates you have imagined being wrong and priced it.

Rebuild: the stress-marked narrative

The output of the five steps is not a pile of doubts. It is a stronger document: claims stated with their assumptions visible and the load-bearing one flagged; evidence that survived the five-question attack; alternatives shown, with the numeric reason each lost; risks named with tripwires and responses attached; and an ask restructured where the stress test demanded it (for Bluewater, almost certainly staged).

Check your understanding
Which assumptions belong on your attack list in Step 2?
Exercise 314 minutes, first pass
Stress-test your real case
Module 4

Preparing for sharper questions

28 minutes

The question bank: generate their prep before they do

Stakeholders can have AI generate their hardest questions from your pre-read. Your move is to generate that question bank first, from personas that match how real scrutiny arrives. Four personas cover most rooms:

  • The verifier: independently checks your numbers against public benchmarks and your own history. "Your last two initiatives projected above 20 percent and delivered under 10. Walk me through why this projection is different."
  • The capital allocator: treats your ask as competing with every other use of the money. "Why is this the best $18M we can spend, versus the alternatives you have not shown me?"
  • The operator: attacks execution, not strategy. "Twelve locations in 18 months means an opening every six weeks. Show me the hiring and construction plan that supports that cadence."
  • The historian: remembers everything the organization has tried. "How is this different from the 2021 expansion, and what specifically did we learn from it that changed this plan?"

Run your rebuilt document through all four (prompts in the toolkit). Merge the output into a single ranked bank: the twenty hardest, ordered by how much damage an unprepared answer would do.

The answer architecture

For each question, prepare the answer in a fixed shape: the direct answer in the first sentence, the evidence in the second, the implication for the decision in the third. The shape matters under pressure because pressure produces rambling, and rambling reads as evasion even when it is not.

Weak: "That is a great question, and there are several factors to consider around the growth rate..."

Strong: "The projection assumes 25 percent; our four-year history is 6. The difference is the app-driven repeat rate, which is the load-bearing assumption in this case, and it is exactly what the month-6 tripwire tests. If the first locations track below 60 percent of the projection, the staged structure stops the remaining spend."

The strong version demonstrates the stress test happened. That demonstration, repeated across a few hard questions, changes the meeting's dynamic from prosecution to collaboration, because stakeholders stop hunting for the weakness you are hiding once it is clear you are not hiding any.

The question you cannot answer

There will be one. The protocol has three rules:

  1. Never improvise a number. A made-up figure in a board meeting is the single most expensive sentence you can say, because the room can now check it before the meeting ends.
  2. Bound what you do know. "I do not have the churn figure by market. What I can tell you is the blended rate and its trend, and that the market-level split has not moved the blended number more than two points historically."
  3. Commit to a date, not an intention. "You will have the full breakdown by Thursday" beats "we will look into that." Then deliver Thursday, because the follow-through is itself evidence for everything else you claimed.

The murder board

The final preparation step, two or three days out: a live simulated hostile Q&A. With a colleague if you can get one; with AI in persona if you cannot, and the AI version has one real advantage: it does not soften out of politeness.

Run it in character, out loud, answering in real time. Rules: no restarting answers, no checking notes for the first response, and every answer scored afterward against the architecture. Two rounds of twenty minutes produces more improvement than any amount of silent review, because the failure mode in the room is not knowledge, it is retrieval under pressure, and retrieval is trainable.

The decision one-pager

The last artifact, and for recurring stakeholder relationships the most valuable: a single page that travels with your recommendation.

Recommendation: one sentence, including the ask and its structure.
The case: the three to five claims, one line each.
Load-bearing assumption: named explicitly, with the evidence behind it and the tripwire that tests it.
Alternatives considered: each with the one-line reason it lost.
Risks and responses: top three, each with its tripwire.
What would change my mind: the evidence that would cause you to withdraw or restructure the recommendation.

That final line is the one leaders resist and the one that buys the most trust. It converts your recommendation from a position to be defended into a decision process stakeholders can see, and decision processes are what boards and investors actually evaluate over time. Anyone can be right once. The one-pager is how you become someone whose recommendations get approved faster with each cycle.

Check your understanding
A board member asks for a figure you do not have. What is the first rule?
Exercise 412 minutes
Build your bank and run one round

Schedule the murder board for two to three days before your actual presentation. Put it on the calendar now.

Wrap-up

The pre-presentation protocol

15 minutes

The repeatable approach, on a timeline

WhenWhatTime
T-minus 7 daysThe stress test: strip, rate, attack, compare, break. Rebuild with assumptions visible, alternatives shown, tripwires set.60 to 90 min
T-minus 4 daysThe question bank: four personas, merged and ranked. Architecture-shaped answers to the top ten. Start closing the can't-answer gaps.45 min
T-minus 2 daysThe murder board: live, in character, out loud, scored. Fix the answers that collapsed.40 min
T-minus 1 dayThe one-pager, written last because it distills everything the protocol surfaced. If any line is hard to write, that difficulty is your final warning.20 min
Day ofNothing new. No new numbers, slides, or arguments after the murder board. Untested claims are where prepared presenters get hurt.0 min

Total cost: roughly three and a half hours per high-stakes recommendation. The comparison price is one meeting where a stakeholder finds the flaw you did not.

What you built in this course

The next 30 days

  • Days 1 to 7: finish the full stress test on your real case: steps 3 through 5, then the rebuild.
  • Days 8 to 14: complete the preparation cycle: question bank, answers, murder board. If your presentation lands in this window, run the full T-minus protocol.
  • Days 15 to 21: run the seven-pattern checklist on incoming recommendations. One caution: use it to improve decisions, not to ambush colleagues. "Have we tested the do-nothing case?" lands better in prep than in the meeting.
  • Days 22 to 30: institutionalize. Add the claim stack and the one-pager to how your team prepares recommendations for you. The fastest way to raise the quality of what reaches your desk is to make the stress test the visible standard for getting there.

The one-month test

After your next high-stakes presentation, score it on one measure: how many questions in the room were already in your bank? Above 80 percent means the protocol is working. Below that, compare the missed questions against the four personas and find which lens you under-weighted. The bank improves every cycle: each defended recommendation makes the next one cheaper to defend.

Final principle: AI made scrutiny fast and cheap, so the leaders who thrive are the ones who point that scrutiny at their own thinking first. The advantage never belonged to whoever had better tools. It belongs to whoever is more willing to find out they are wrong while it is still cheap to be.
Appendix

AI toolkit

Keep this open while you work
Confidentiality first: follow your organization's policy on sharing strategy material with AI tools. Every prompt below works on an abstracted version of your case: keep the logic and structure, round the numbers, remove names and identifying details.

The stress test

Step 1 · Strip
Here is a strategic recommendation: [paste document or
abstracted version].
Reduce it to a claim stack:
1. The recommendation in one sentence, including the
   ask and its structure
2. The 3 to 5 key claims that must be true for the
   recommendation to hold
3. Under each claim, every assumption it rests on,
   including unstated ones the text takes for granted
Flag any assumption that appears nowhere in the
document but is required by its logic.
Step 2 · Rate
For each assumption in this claim stack, rate:
1. Load-bearing: if this is wrong, does the
   recommendation survive? (fatal / damaged / survives)
2. Uncertainty: how strong is the stated evidence?
   (strong / thin / none stated)
List the assumptions that are both fatal-if-wrong and
thin-or-no evidence. That is my attack list. Rank it
by how easily an outside stakeholder could challenge
each one with public information.
Step 3 · Attack
Here is the evidence supporting a load-bearing claim:
[paste evidence summary].
Interrogate it adversarially:
1. Source: who produced each item, and what did they
   want to be true?
2. Age: when is each from, and what could have changed?
3. Denominator: what population does it describe, and
   is that the population being projected onto?
4. Counterfactual: compared to what? What baseline
   choice is doing silent work?
5. Survivors: does it look only at winners?
Do not soften. For each vulnerability, write the exact
question a hostile stakeholder would ask.
Step 4 · Compare
The recommendation is: [one sentence].
Build the strongest honest case for:
1. Doing nothing, including everything the money,
   time, and attention could do instead
2. A staged or smaller reversible version
3. The strongest genuinely different approach
Argue each as its best advocate would, not as a straw
man. Then state what evidence would have to be true
for each alternative to beat my recommendation.
Step 5 · Break
It is [18 months] from now and this initiative failed.
Write the 3 most plausible failure stories, each with:
1. The specific chain of causes
2. The earliest observable evidence that this failure
   was beginning (the tripwire)
3. What the reasonable response would have been at
   that tripwire moment
Prioritize failure modes that stem from the
assumptions in my attack list.
The rebuild check
Here is my revised recommendation: [paste].
Verify it now contains: assumptions stated with the
load-bearing one flagged; alternatives shown with the
reason each lost; risks with tripwires and responses;
and a defined condition under which the initiative
stops. List anything still missing, and identify the
weakest remaining claim.

The question bank

The verifier
You are a board member who independently checks
management's numbers. You have this document, public
benchmark data, and the presenter's track record:
[past initiatives and outcomes, abstracted].
Generate your 8 hardest questions. Prioritize places
where the document's numbers can be checked against
external data or the presenter's own history.
The capital allocator
You are a director who treats every ask as competing
with every other use of capital. Generate your 8
hardest questions about this recommendation,
prioritizing opportunity cost, the alternatives not
shown, and whether the ask's structure (size, timing,
staging) is justified.
The operator
You are an executive who has run implementations this
size. Ignore the strategy; attack the execution.
Generate your 8 hardest questions about capacity,
sequencing, hiring, timelines, dependencies, and what
this initiative does to the performance of existing
operations while it absorbs attention.
The historian
You are the longest-tenured person in the room. You
remember every similar initiative: [list prior
comparable efforts and outcomes, abstracted].
Generate your 8 hardest questions connecting this
recommendation to that history, especially "what did
we learn last time and where does this plan apply it?"
Merge and rank
Here are four question lists: [paste all].
Deduplicate, then rank the top 20 by how much damage
an unprepared answer would do to the recommendation's
credibility. Mark the 3 where you predict I currently
have no strong answer.

Answers and rehearsal

Answer architect
Question: [paste one hard question].
My raw material: [facts, numbers, reasoning].
Draft an answer in exactly three parts: the direct
answer in one sentence, the strongest evidence in one
or two sentences, and the implication for the decision
in one sentence. No preamble, no "great question."
The murder board
Run a live hostile Q&A on this recommendation: [paste
one-pager or rebuilt document].
Rotate through four personas: a verifier who checks
numbers, a capital allocator focused on opportunity
cost, an operator attacking execution, and a historian
citing past initiatives: [abstracted history].
Ask one question at a time and wait for my answer.
After each answer, follow up once the way a skeptical
stakeholder would: probe the weakest part of what I
said. Every 5 questions, pause and score my answers
against this standard: direct answer first, evidence
second, implication third. Do not be polite. Begin.
Bounded non-answer
I will likely be asked: [the question I cannot fully
answer]. What I do know: [adjacent facts, bounds,
trends]. Draft a response that: states plainly what I
do not have, bounds the uncertainty with what I do
know, and commits to a specific delivery date. No
bluffing, no filler.

The one-pager

One-pager generator
From this stress-tested recommendation: [paste rebuilt
document], generate a one-page decision summary with
exactly these sections:
1. Recommendation (one sentence, including the ask
   and its structure)
2. The case (3 to 5 claims, one line each)
3. Load-bearing assumption (named, with its evidence
   and its tripwire)
4. Alternatives considered (each with the one-line
   reason it lost)
5. Risks and responses (top 3, each with a tripwire)
6. What would change my mind
Keep it under one page. Where my document lacks the
content for a section, write [NEEDED: description]
rather than inventing it.

Standing instructions worth saving

Custom instructions / preferences.md
When reviewing my strategic recommendations or
decision documents:
- Act as a red team by default; do not validate
- Surface unstated assumptions before commenting on
  stated ones
- Always ask "compared to what?" about every
  favorable number
- Flag any projection that bends away from historical
  trend without a named mechanism
- Never invent facts or figures; mark gaps as
  [NEEDED: description]
- When I ask for a defense of my position, first give
  me the strongest attack on it