Pressure-Test My Strategy in the AI Era
Find the weak points before your stakeholders do.
How this course works
5 minutesGood strategy is no longer enough. It has to hold up under pressure, and the pressure has gotten faster and sharper. This course teaches a repeatable method for stress-testing your own recommendations before anyone else gets the chance.
You will work on two tracks throughout: the safe case, a fictional recommendation in Module 2 with deliberate weaknesses planted in it, and your real case, a recommendation you will need to defend in the next 90 days.
Name your real case
The new scrutiny: how AI armed your stakeholders
18 minutesWhat changed in the room
Ten years ago, a well-built deck had a structural advantage: management held the information. Board members, investors, and executive stakeholders could probe, but real verification took weeks they did not have. A confident narrative, delivered well, often carried the day on delivery alone.
That advantage is gone. Governance research now documents the shift directly: directors are using AI to independently benchmark company disclosures against peers, test management's numbers against outside market data, and analyze the historical board materials management itself provided. PwC reports that 35 percent of directors say their boards have begun using AI in their oversight work, including preparing for meetings and pressure-testing strategy. Board platforms now generate director-specific probing questions from the pre-read materials automatically.
Translate that into what happens to your recommendation:
- Your benchmark gets checked in real time. The peer comparison you selected can be re-run, with peers you did not select, before you finish presenting.
- Your numbers get tested against your own history. "This projects 25 percent growth; the last three initiatives you presented projected similarly and delivered 8. What is different?" is now a question anyone can construct in minutes from your old decks.
- Your pre-read gets red-teamed. The two-day pre-read window is enough time for a stakeholder to have AI generate the twenty hardest questions your document raises.
- Alternatives you did not present get surfaced. "What options did you consider and reject?" now comes with the stakeholder having already generated the option list themselves.
The asymmetry, and how to flip it
Challenging a recommendation with AI takes minutes, while building one still takes weeks. The attacker's cost dropped faster than the builder's. That asymmetry is why polished-but-fragile narratives are failing at higher rates in rooms that used to accept them.
The flip is equally available: every tool your stakeholders can point at your recommendation, you can point at it first. The same AI that generates their twenty hardest questions can generate them for you, two weeks earlier, while you still have time to fix what the questions expose. Scrutiny is coming either way. The only variable you control is whether it happens before the meeting, run by you, or during the meeting, run by them.
What survives scrutiny
Recommendations that hold up under AI-enabled pressure share a structural property: their reasoning is visible and already stress-marked. They state their assumptions rather than burying them. They show the alternatives considered and the honest reason each lost. They name their risks with responses attached, and they define in advance what evidence would change the recommender's mind.
Notice what that list does not include: certainty. Defensibility does not come from having no weaknesses. It comes from knowing your weaknesses better than your challengers do, and having decided what to do about each one. A leader who says "this assumption is the load-bearing one, here is why I believe it, and here is what we will see within 90 days if I am wrong" is harder to rattle than one defending a claim of confidence.
Name your exposure
Quick answers on your real case, no deep analysis yet. Do not fix anything: this is a baseline of where you already suspect the soft spots are.
The seven weaknesses of executive narratives
22 minutesStrategic recommendations fail under scrutiny in patterned ways. The patterns repeat because they are not lying, they are motivated reasoning: each one is a shortcut a smart, honest person takes when they already believe their conclusion. Knowing the seven gives you a checklist for finding them in your own work, where they are hardest to see.
- The buried assumption. The recommendation depends on something never stated, so it is never examined. "The rollout completes in Q2" quietly assumes the two key hires close in Q1 and vendor timelines hold.
- The chosen baseline. Every comparison requires choosing what to compare against, and the choice is doing silent work. Growth "versus last year" when last year dipped. Benchmarks against the top quartile when the honest peer set is the median. The number is accurate. The frame was selected.
- The hockey stick. History grows at 6 percent; the projection grows at 25, with the bend arriving conveniently after the decision. The tell is a projection whose driver is the initiative's assumed success rather than any mechanism you can name.
- The friendly pilot. Evidence from a test run under the conditions most likely to make it succeed: the flagship location, the strongest team, volunteer users. Real result, wrong denominator. The pilot proves the idea can work; the narrative treats it as proof it will work everywhere.
- The missing alternative. The case argues for the recommendation without arguing against the alternatives, including the strongest one: doing nothing, or a smaller reversible version first. When no alternative appears, stakeholders correctly suspect the comparison was not run, or was run and lost.
- The convenient source. Key evidence from a party with a stake in the conclusion: the vendor's ROI study, the consultant hired by the sponsor. Also in this family: the three-year-old market study still cited because updating it might change the answer.
- The all-or-nothing ask. Full commitment requested up front, no staged decision points, no defined conditions under which the initiative stops. The absence of kill criteria signals that the recommender has not seriously imagined being wrong.
The safe case
Read this fictional recommendation. It contains at least five of the seven weaknesses, planted deliberately. Find them before opening the reveal in Module 3.
Bluewater Coffee Roasters: Retail Expansion Recommendation (excerpt from CEO memo to the board)
Bluewater operates 38 cafes across the Pacific Northwest with $46M in annual revenue. I recommend the board approve an $18M investment to open 12 new locations across two new metro markets over 18 months, alongside launch of the Bluewater mobile app.
The specialty coffee retail segment continues to grow, with leading chains posting 22 percent annual revenue growth. Our own app pilot at the Portland flagship generated a 31 percent increase in repeat visits, demonstrating strong customer appetite for digital engagement across our footprint. A market study by Cascade Growth Partners, who supported our original expansion in 2023, identifies both target metros as high-opportunity markets.
We project the new locations reach $14M combined annual revenue by year three, reflecting 25 percent year-over-year growth after opening. Company revenue has grown steadily at 5 to 7 percent annually over the past four years, and this initiative moves us decisively beyond that plateau. Speed matters: delaying entry risks ceding both markets to competitors. I recommend we commit to all 12 locations now to secure favorable lease terms, with construction beginning in Q1.
Two audits
The reveal is in Module 3, inside the claim stack teardown.
The stress test: strip, rate, attack, compare, break
32 minutesThe stress test is five steps run in order, each with AI as your red team. It takes 60 to 90 minutes for a significant recommendation, which is why it belongs a week before the meeting, not the night before. This module compresses it into a guided pass on the safe case; the toolkit appendix carries the full prompts for your real work.
Step 1: Strip it to the claim stack
Reduce the narrative to its logical skeleton: the recommendation on top, the three to five key claims holding it up, and under each claim, the assumptions it rests on. Prose hides logic. Stacks expose it.
The Bluewater memo, stripped (this is the flaw reveal — open after Exercise 2)
Recommendation: invest $18M in 12 locations plus the app, all at once.
Claim 1: the market opportunity is large and growing. Rests on: segment growth applies to us (benchmark is "leading chains": pattern 2, the chosen baseline); the Cascade study is current and independent (commissioned by the sponsor of the last expansion: pattern 6, the convenient source; and its currency is unstated).
Claim 2: customer demand for our model is proven. Rests on: the flagship pilot generalizes to 12 new locations in metros where the brand is unknown (pattern 4, the friendly pilot).
Claim 3: the financial projection is achievable. Rests on: 25 percent growth from a company that has grown 5 to 7 percent for four years, with no stated mechanism for the bend (pattern 3, the hockey stick), plus unstated assumptions about hiring, construction, and same-store performance (pattern 1).
Claim 4: committing fully now is the right structure. Rests on: an urgency claim ("ceding markets") that carries no evidence, and the absence of any staged alternative or stopping condition (patterns 5 and 7).
If your audit caught five of those, you found the plants. The stack format is what makes them findable in your own work, where motivated reasoning hides them from the prose reader you become.
Step 2: Rate the assumptions
Not every assumption deserves attack. Rate each on two axes: how load-bearing (does the recommendation survive if this is wrong?) and how uncertain (how strong is the actual evidence?). Your attack list is the quadrant that is both load-bearing and uncertain.
For Bluewater, the pilot-generalization assumption and the 25 percent growth rate are both squarely in that quadrant. The lease-terms urgency claim is uncertain but less load-bearing. Spend your scrutiny where wrongness is fatal.
Step 3: Attack the evidence
For every piece of evidence supporting a load-bearing claim, run five questions:
- Source: who produced this, and what did they want to be true?
- Age: when is this from, and what has changed since?
- Denominator: what population does this actually describe, and is it the population I am projecting onto?
- Counterfactual: compared to what? What would this look like under the do-nothing case, or a different baseline?
- Survivors: am I looking at the winners only? (The "leading chains at 22 percent" benchmark excludes every chain that expanded and shrank.)
This is where AI earns its place. Hand it your evidence summary and have it run the five questions adversarially. It will not know your industry's ground truth, but it is relentless at spotting which claims rest on which sources, and it does not share your motivation to go easy.
Step 4: Compare against the steelman alternatives
Build the strongest honest version of at least two alternatives: the do-nothing case and the strongest different approach. For Bluewater: invest in same-store growth and the app across all 38 existing cafes, or stage it: open 3 locations in one metro with defined success gates before committing to the rest.
The test your recommendation must pass: state, in numbers where possible, why it beats each steelman. If the staged alternative is nearly as good with a fraction of the downside, the honest recommendation may be the staged one, and discovering that before the board does is the difference between leading the meeting and losing it.
Step 5: Break it, then set the tripwires
Run a pre-mortem: it is 18 months from now and the initiative failed. Write the three most plausible failure stories, specifically, with causes. Then convert each into two artifacts:
- A tripwire: the earliest observable evidence that this failure story is beginning. For Bluewater's hockey stick: "if the first 3 locations are below 60 percent of projected revenue at month 6."
- A response: what happens when the tripwire trips. Pause openings, revisit the model, stop.
Tripwires plus responses are your kill criteria, and they repair pattern 7. Presenting them does not weaken your ask. It is usually the single strongest credibility move in the room, because it demonstrates you have imagined being wrong and priced it.
Rebuild: the stress-marked narrative
The output of the five steps is not a pile of doubts. It is a stronger document: claims stated with their assumptions visible and the load-bearing one flagged; evidence that survived the five-question attack; alternatives shown, with the numeric reason each lost; risks named with tripwires and responses attached; and an ask restructured where the stress test demanded it (for Bluewater, almost certainly staged).
Stress-test your real case
Preparing for sharper questions
28 minutesThe question bank: generate their prep before they do
Stakeholders can have AI generate their hardest questions from your pre-read. Your move is to generate that question bank first, from personas that match how real scrutiny arrives. Four personas cover most rooms:
- The verifier: independently checks your numbers against public benchmarks and your own history. "Your last two initiatives projected above 20 percent and delivered under 10. Walk me through why this projection is different."
- The capital allocator: treats your ask as competing with every other use of the money. "Why is this the best $18M we can spend, versus the alternatives you have not shown me?"
- The operator: attacks execution, not strategy. "Twelve locations in 18 months means an opening every six weeks. Show me the hiring and construction plan that supports that cadence."
- The historian: remembers everything the organization has tried. "How is this different from the 2021 expansion, and what specifically did we learn from it that changed this plan?"
Run your rebuilt document through all four (prompts in the toolkit). Merge the output into a single ranked bank: the twenty hardest, ordered by how much damage an unprepared answer would do.
The answer architecture
For each question, prepare the answer in a fixed shape: the direct answer in the first sentence, the evidence in the second, the implication for the decision in the third. The shape matters under pressure because pressure produces rambling, and rambling reads as evasion even when it is not.
Weak: "That is a great question, and there are several factors to consider around the growth rate..."
Strong: "The projection assumes 25 percent; our four-year history is 6. The difference is the app-driven repeat rate, which is the load-bearing assumption in this case, and it is exactly what the month-6 tripwire tests. If the first locations track below 60 percent of the projection, the staged structure stops the remaining spend."
The strong version demonstrates the stress test happened. That demonstration, repeated across a few hard questions, changes the meeting's dynamic from prosecution to collaboration, because stakeholders stop hunting for the weakness you are hiding once it is clear you are not hiding any.
The question you cannot answer
There will be one. The protocol has three rules:
- Never improvise a number. A made-up figure in a board meeting is the single most expensive sentence you can say, because the room can now check it before the meeting ends.
- Bound what you do know. "I do not have the churn figure by market. What I can tell you is the blended rate and its trend, and that the market-level split has not moved the blended number more than two points historically."
- Commit to a date, not an intention. "You will have the full breakdown by Thursday" beats "we will look into that." Then deliver Thursday, because the follow-through is itself evidence for everything else you claimed.
The murder board
The final preparation step, two or three days out: a live simulated hostile Q&A. With a colleague if you can get one; with AI in persona if you cannot, and the AI version has one real advantage: it does not soften out of politeness.
Run it in character, out loud, answering in real time. Rules: no restarting answers, no checking notes for the first response, and every answer scored afterward against the architecture. Two rounds of twenty minutes produces more improvement than any amount of silent review, because the failure mode in the room is not knowledge, it is retrieval under pressure, and retrieval is trainable.
The decision one-pager
The last artifact, and for recurring stakeholder relationships the most valuable: a single page that travels with your recommendation.
Recommendation: one sentence, including the ask and its structure.
The case: the three to five claims, one line each.
Load-bearing assumption: named explicitly, with the evidence behind it and the tripwire that tests it.
Alternatives considered: each with the one-line reason it lost.
Risks and responses: top three, each with its tripwire.
What would change my mind: the evidence that would cause you to withdraw or restructure the recommendation.
That final line is the one leaders resist and the one that buys the most trust. It converts your recommendation from a position to be defended into a decision process stakeholders can see, and decision processes are what boards and investors actually evaluate over time. Anyone can be right once. The one-pager is how you become someone whose recommendations get approved faster with each cycle.
Build your bank and run one round
Schedule the murder board for two to three days before your actual presentation. Put it on the calendar now.
The pre-presentation protocol
15 minutesThe repeatable approach, on a timeline
| When | What | Time |
|---|---|---|
| T-minus 7 days | The stress test: strip, rate, attack, compare, break. Rebuild with assumptions visible, alternatives shown, tripwires set. | 60 to 90 min |
| T-minus 4 days | The question bank: four personas, merged and ranked. Architecture-shaped answers to the top ten. Start closing the can't-answer gaps. | 45 min |
| T-minus 2 days | The murder board: live, in character, out loud, scored. Fix the answers that collapsed. | 40 min |
| T-minus 1 day | The one-pager, written last because it distills everything the protocol surfaced. If any line is hard to write, that difficulty is your final warning. | 20 min |
| Day of | Nothing new. No new numbers, slides, or arguments after the murder board. Untested claims are where prepared presenters get hurt. | 0 min |
Total cost: roughly three and a half hours per high-stakes recommendation. The comparison price is one meeting where a stakeholder finds the flaw you did not.
What you built in this course
The next 30 days
- Days 1 to 7: finish the full stress test on your real case: steps 3 through 5, then the rebuild.
- Days 8 to 14: complete the preparation cycle: question bank, answers, murder board. If your presentation lands in this window, run the full T-minus protocol.
- Days 15 to 21: run the seven-pattern checklist on incoming recommendations. One caution: use it to improve decisions, not to ambush colleagues. "Have we tested the do-nothing case?" lands better in prep than in the meeting.
- Days 22 to 30: institutionalize. Add the claim stack and the one-pager to how your team prepares recommendations for you. The fastest way to raise the quality of what reaches your desk is to make the stress test the visible standard for getting there.
The one-month test
After your next high-stakes presentation, score it on one measure: how many questions in the room were already in your bank? Above 80 percent means the protocol is working. Below that, compare the missed questions against the four personas and find which lens you under-weighted. The bank improves every cycle: each defended recommendation makes the next one cheaper to defend.
AI toolkit
Keep this open while you workThe stress test
Here is a strategic recommendation: [paste document or abstracted version]. Reduce it to a claim stack: 1. The recommendation in one sentence, including the ask and its structure 2. The 3 to 5 key claims that must be true for the recommendation to hold 3. Under each claim, every assumption it rests on, including unstated ones the text takes for granted Flag any assumption that appears nowhere in the document but is required by its logic.
For each assumption in this claim stack, rate: 1. Load-bearing: if this is wrong, does the recommendation survive? (fatal / damaged / survives) 2. Uncertainty: how strong is the stated evidence? (strong / thin / none stated) List the assumptions that are both fatal-if-wrong and thin-or-no evidence. That is my attack list. Rank it by how easily an outside stakeholder could challenge each one with public information.
Here is the evidence supporting a load-bearing claim: [paste evidence summary]. Interrogate it adversarially: 1. Source: who produced each item, and what did they want to be true? 2. Age: when is each from, and what could have changed? 3. Denominator: what population does it describe, and is that the population being projected onto? 4. Counterfactual: compared to what? What baseline choice is doing silent work? 5. Survivors: does it look only at winners? Do not soften. For each vulnerability, write the exact question a hostile stakeholder would ask.
The recommendation is: [one sentence]. Build the strongest honest case for: 1. Doing nothing, including everything the money, time, and attention could do instead 2. A staged or smaller reversible version 3. The strongest genuinely different approach Argue each as its best advocate would, not as a straw man. Then state what evidence would have to be true for each alternative to beat my recommendation.
It is [18 months] from now and this initiative failed. Write the 3 most plausible failure stories, each with: 1. The specific chain of causes 2. The earliest observable evidence that this failure was beginning (the tripwire) 3. What the reasonable response would have been at that tripwire moment Prioritize failure modes that stem from the assumptions in my attack list.
Here is my revised recommendation: [paste]. Verify it now contains: assumptions stated with the load-bearing one flagged; alternatives shown with the reason each lost; risks with tripwires and responses; and a defined condition under which the initiative stops. List anything still missing, and identify the weakest remaining claim.
The question bank
You are a board member who independently checks management's numbers. You have this document, public benchmark data, and the presenter's track record: [past initiatives and outcomes, abstracted]. Generate your 8 hardest questions. Prioritize places where the document's numbers can be checked against external data or the presenter's own history.
You are a director who treats every ask as competing with every other use of capital. Generate your 8 hardest questions about this recommendation, prioritizing opportunity cost, the alternatives not shown, and whether the ask's structure (size, timing, staging) is justified.
You are an executive who has run implementations this size. Ignore the strategy; attack the execution. Generate your 8 hardest questions about capacity, sequencing, hiring, timelines, dependencies, and what this initiative does to the performance of existing operations while it absorbs attention.
You are the longest-tenured person in the room. You remember every similar initiative: [list prior comparable efforts and outcomes, abstracted]. Generate your 8 hardest questions connecting this recommendation to that history, especially "what did we learn last time and where does this plan apply it?"
Here are four question lists: [paste all]. Deduplicate, then rank the top 20 by how much damage an unprepared answer would do to the recommendation's credibility. Mark the 3 where you predict I currently have no strong answer.
Answers and rehearsal
Question: [paste one hard question]. My raw material: [facts, numbers, reasoning]. Draft an answer in exactly three parts: the direct answer in one sentence, the strongest evidence in one or two sentences, and the implication for the decision in one sentence. No preamble, no "great question."
Run a live hostile Q&A on this recommendation: [paste one-pager or rebuilt document]. Rotate through four personas: a verifier who checks numbers, a capital allocator focused on opportunity cost, an operator attacking execution, and a historian citing past initiatives: [abstracted history]. Ask one question at a time and wait for my answer. After each answer, follow up once the way a skeptical stakeholder would: probe the weakest part of what I said. Every 5 questions, pause and score my answers against this standard: direct answer first, evidence second, implication third. Do not be polite. Begin.
I will likely be asked: [the question I cannot fully answer]. What I do know: [adjacent facts, bounds, trends]. Draft a response that: states plainly what I do not have, bounds the uncertainty with what I do know, and commits to a specific delivery date. No bluffing, no filler.
The one-pager
From this stress-tested recommendation: [paste rebuilt document], generate a one-page decision summary with exactly these sections: 1. Recommendation (one sentence, including the ask and its structure) 2. The case (3 to 5 claims, one line each) 3. Load-bearing assumption (named, with its evidence and its tripwire) 4. Alternatives considered (each with the one-line reason it lost) 5. Risks and responses (top 3, each with a tripwire) 6. What would change my mind Keep it under one page. Where my document lacks the content for a section, write [NEEDED: description] rather than inventing it.
Standing instructions worth saving
When reviewing my strategic recommendations or decision documents: - Act as a red team by default; do not validate - Surface unstated assumptions before commenting on stated ones - Always ask "compared to what?" about every favorable number - Flag any projection that bends away from historical trend without a named mechanism - Never invent facts or figures; mark gaps as [NEEDED: description] - When I ask for a defense of my position, first give me the strongest attack on it