IPI Global AI Accelerator · Workbook
From Product
to Launch
Shipping AI products that survive contact with reality. This workbook follows the session outline so you can keep working with your team: short recaps, copy-pastable templates, and matrices you can print, project, or fill in right here in the browser. Please note that nothing is saved here, so if you type anything directly onto the page, be sure to copy and save elsewhere regularly. It was a pleasure getting to know you and your projects. All the best!
Marie Kilg · AI Strategy · kilg.de
01 · Diagnosis
Why AI adoption fails
AI projects in media organisations rarely fail because the technology doesn’t work. They fail twice, from two directions at once.
- Top down — the pace problem. Strategy, decision-making and implementation can’t keep up with the speed of AI development. By the time a roadmap is approved, the landscape has shifted.
- Bottom up — the trust problem. Fear of replacement, shadow usage, blocked middle management, insecurity. People quietly use tools they don’t admit to, and quietly resist tools they’re told to use.
The answer to the pace problem is grounding decisions in reality: short-term decentralised pilots that show what works and what users need, feeding a strategy that in turn selects the next pilots. The answer to the trust problem is honesty about what AI can and cannot do — which brought us to a story about a horse.
02 · Foundations
Don’t sell your newsroom a clever horse
Clever Hans was a horse that appeared to do arithmetic. He couldn’t — he was reading the involuntary body language of the people around him. The performance was real; the explanation everyone believed was wrong.
Large language models invite the same mistake. Language models don’t reason, they memorize — they are pattern machines trained on text, and they perform best on patterns they have seen often. That makes them genuinely useful for knowledge management, creative language tasks and information retrieval, and genuinely unreliable wherever correctness depends on reasoning nobody can check.
Two biases make LLMs look smarter than they are in everyday use:
- Optional stopping — if the result isn’t good, you ask again, and you stop when it looks good.
- Sampling bias — successful results get shared; unsuccessful tries are invisible. Result quality depends heavily on the knowledge of the user.
The sweet spot: don’t trust generative AI tools — have them work together with skilled humans who can ground them in reality, find the applications where that collaboration is most efficient (time, money, energy), and pick the right tool for the right job without getting fooled by people selling you a clever horse.
Run this on any AI use case before you commit. If you can’t answer all five, you’re not done deciding.
CLEVER HANS CHECK — [use case / tool name] 1. WHAT does the tool actually do, mechanically? (Not the vendor's claim — the plain description.) → 2. WHO is the skilled human grounding the output in reality? (Name a role. "The team" grounds nothing.) If there is no human, WHAT PROCESS or measurement can do the grounding and ensure quality? → 3. HOW would we notice when it's confidently wrong? (What does failure look like, and who sees it first?) → 4. WHERE does the human + machine combination actually save time, money, or energy — measured against doing it without? → 5. WHAT happens to our credibility if a wrong output is published? Who reviews before anything goes public? →
03 · Product Management
The feature list: P0 to P4
A good product manager’s first tool is a ruthless, prioritised feature list. The priorities have meanings — they are commitments, not vibes.
| Priority | What it means | Example — “Alexa Music” launch (fictitious) |
|---|---|---|
| P0 | The promise breaks without it. No P0, no launch. | “Alexa, play David Bowie” — playback, pause, stop on one device; volume by voice; next/previous track |
| P1 | Day-one expectation. Ship it, or fix within days. | Your playlists |
| P2 | Fast follow. The first iterations after launch. | Multi-room audio, recommendations, sleep timer |
| P3 | Nice to have. Build only if the signals demand it. | Mood stations, lyrics on screen |
| P4 | Parking lot. Ideas live here, not on the roadmap. | Karaoke mode, DJ chatter between songs |
You launch with every P0 and most P1 — everything below that line is what iteration is for.
Fill in directly (cells are editable), then print — or copy the plain-text version into your own doc. Status: GREEN on track · YELLOW at risk · RED blocked · GREY parked.
| Priority | Feature | Status | ETA | Owner |
|---|---|---|---|---|
| P0 | ||||
| P0 | ||||
| P0 | ||||
| P1 | ||||
| P1 | ||||
| P2 | ||||
| P2 | ||||
| P3 | ||||
| P4 |
LAUNCH FEATURE LIST — [product name] Date: [____] Status key: GREEN on track / YELLOW at risk / RED blocked / GREY parked Rule: launch with every P0 and most P1. Everything below the line is what iteration is for. P0 (the promise breaks without it — no P0, no launch) [ ] Feature: __________________ Status: ____ ETA: ____ Owner: ____ [ ] Feature: __________________ Status: ____ ETA: ____ Owner: ____ [ ] Feature: __________________ Status: ____ ETA: ____ Owner: ____ P1 (day-one expectation — ship it, or fix within days) [ ] Feature: __________________ Status: ____ ETA: ____ Owner: ____ [ ] Feature: __________________ Status: ____ ETA: ____ Owner: ____ P2 (fast follow — the first iterations after launch) [ ] Feature: __________________ Status: ____ ETA: ____ Owner: ____ [ ] Feature: __________________ Status: ____ ETA: ____ Owner: ____ P3 (nice to have — build only if the signals demand it) [ ] Feature: __________________ Status: ____ ETA: ____ Owner: ____ P4 (parking lot — ideas live here, not on the roadmap) [ ] Feature: __________________ Notes: ______________________
04 · Working Backwards
Write the press release before you build
Amazon’s “working backwards” method: if you can’t write the launch headline, you’re not done deciding. The headline forces one promise to one audience — features that don’t serve it are scope creep.
The same exercise we did in the session — repeat it whenever scope starts to drift.
LAUNCH HEADLINE EXERCISE Imagine your product launched today and one outlet covers it. 1. Write ONE headline (max 15 words) naming the audience and the benefit: "________________________________________________" 2. Test: does every feature you are currently building serve this headline? Features that serve it (→ P0/P1): Features that don't (→ P2–P4): - ____________________ - ____________________ - ____________________ - ____________________ - ____________________ - ____________________
One page of press release, then FAQs. Write it in plain language your audience would actually read — if the press release is boring, the product probably is too. Keep the whole document under six pages; most of the value is in the arguments you have while writing it.
PR/FAQ — [working title] Owner: [name] · Date: [____] · Status: draft / reviewed / approved ================ PRESS RELEASE (max 1 page) ================ HEADLINE [Audience] + [benefit] in max 15 words. e.g. "ExampleNews launches WhatsApp briefing that explains the city budget in 3 minutes a day" SUBHEADLINE One sentence: who it's for and why they'll care today. [CITY], [launch date] — Opening paragraph (3–4 sentences): what launched, for whom, and the single most important benefit. A reader should be able to stop here and understand the product. THE PROBLEM (1 paragraph) Describe the customer's problem in their words, not yours. What do they do today instead, and why does it frustrate them? THE SOLUTION (1–2 paragraphs) How the product solves the problem. What the experience feels like the first time someone uses it. Mention the human review step if AI is involved — your audience will ask. QUOTE FROM YOUR LEADERSHIP "[Why we built this — one promise, no buzzwords.]" — [Name, role] HOW TO GET STARTED (1 paragraph) The first step a reader takes today: where to sign up, what it costs, what they'll see in the first five minutes. QUOTE FROM A (HYPOTHETICAL) CUSTOMER "[What changed for me — concrete, before vs. after.]" — [Name, who they are] CALL TO ACTION One link. One action. ======================== FAQ =============================== EXTERNAL (customer-facing) — answer honestly: Q1. What exactly does it do, and what does it NOT do? Q2. How much does it cost / what's the catch? Q3. Where does the AI come in, and who checks its output? Q4. What happens to my data? Q5. How is this different from [the obvious alternative]? INTERNAL (team-facing) — the hard ones: Q6. What are the P0s, and what did we cut to get here? Q7. What does it cost to run per month (incl. model/API costs)? Q8. Which decisions are one-way doors? (see Template 05) Q9. What's our North Star metric and the 2–3 input metrics? Q10. Who owns it the day after launch — by name? Q11. What would make us kill this in 6 months?
05 · Decisions
Decide and iterate: doors and rhythms
One-way vs. two-way doors
Before deciding, ask: is this reversible?
- Two-way doors (naming, UI, prompts, price tests): decide fast, launch, adjust.
- One-way doors (platform lock-in, data architecture, public promises to your audience): slow down, decide deliberately.
When have you done enough?
With LLMs, if the result isn’t good you ask again — and you stop when it looks good. With products, the same bias makes you polish forever: it always still “could be better”. Stop when new iterations stop changing user behaviour — not when it stops feeling improvable. Launch is a learning event, not a finish line: iterate on real usage, not on your own taste. And kill your darlings — it’s OK to park ideas on the “later” list.
Use the door check for any decision that feels heavy; use the rhythm block once, when you set up your team’s cadence.
DOOR CHECK — [decision] Date: [____]
Is this reversible? [ ] yes → TWO-WAY DOOR
[ ] no → ONE-WAY DOOR
If TWO-WAY: decide fast. Who decides alone? [name]
Decision: ______________________ Revisit on: [date]
If ONE-WAY: slow down.
- What exactly becomes irreversible? ______________________
- What would we pay later to undo it? ______________________
- Who must be in the room? ______________________
- Decision + rationale (3 sentences max):
________________________________________________________
ITERATION RHYTHM — set once, then protect it
- Review meeting: every [__] weeks, [day/time], owner: [name]
- Bug triage: [when], owner: [name]
- Every review: look at the signals, pick ONE change, ship it.
- "Later" list lives at: [link] (parked ≠ dead)
- Stop polishing when: new iterations stop changing user
behaviour — not when it stops feeling improvable.
06 · Map Your Tools
Two matrices for every AI idea
In the session you mapped your current tools in pairs. Keep doing it: every new tool or feature idea gets a place on both matrices before it gets a budget.
The Application Matrix
Horizontal: from proven to work to over-hyped. Vertical: from unproblematic to ethically, legally or reputationally problematic. From the session, roughly: transcription, translation, search & research tools, speech enhancement and verification tech sit in the proven/unproblematic corner. “Reasoning”, complex research, agentic “thinking and acting” need exploration. Avatars, voices cloned from real humans, photorealistic AI images and autonomous publication without review sit in the problematic zone — that’s where your ethics board, legal team and senior leadership belong, while average on-the-ground teams work the safe quadrant and the lab explores the rest.
Type directly into the quadrants, or print the page and use sticky notes. For each tool, ask: what evidence would move it from “explore” to “proven” — and who decides?
Adopt
Proven · unproblematic — roll out to on-the-ground teams.
Explore
Hyped · unproblematic — one for the lab, with an exit date.
Govern
Proven · problematic — ethics board / legal / leadership decide.
Avoid (for now)
Hyped · problematic — name it, park it, revisit deliberately.
APPLICATION MATRIX — [team / date] Axes: X: proven to work ←→ over-hyped Y: unproblematic ←→ ethically/legally/reputationally problematic ADOPT (proven + unproblematic) → on-the-ground teams - ____________________ EXPLORE (hyped + unproblematic) → the lab, with an exit date - ____________________ GOVERN (proven + problematic) → ethics board / legal / leadership - ____________________ AVOID FOR NOW (hyped + problematic) → parked, revisit on [date] - ____________________ For each tool in EXPLORE: What evidence moves it to PROVEN? ______________________ Who decides? ______________________
The Impact/Effort Matrix
Speech enhancement was the session’s example of a NOW. The super-agent that does everything for everyone is a tough “HOW?”.
Wow!
High impact · easy — do these first.
Now!
Low impact · easy — quick wins; if they’re truly free.
How?
High impact · hard — plan deliberately, break into steps.
Ciao
Low impact · hard — say goodbye.
IMPACT / EFFORT MATRIX — [team / date] Axes: X: high impact ←→ low impact Y: easy ←→ hard WOW! (high impact, easy) - ____________________ HOW? (high impact, hard) - ____________________ NOW! (low impact, easy) - ____________________ CIAO (low impact, hard) - ____________________
07 · Choosing the Model
You build on rented land
A model you depend on can be deprecated, repriced or changed overnight — without asking you.
- Separate your core (data, workflows, editorial value) from the replaceable (the model behind it).
- Design the connecting layer so you can swap models: thin wrappers, own prompts, own evaluation data.
- Ask of every part of your stack: where can I bet on one horse — and where do my solutions need to be model-agnostic?
Run this once per AI product, and again whenever your provider changes pricing or deprecates a model.
MODEL INDEPENDENCE AUDIT — [product] Date: [____]
OURS (the core — this is the business):
[ ] Our data (training/eval/feedback) is exported and stored
where we control it: [location]
[ ] Our prompts live in our repo, not in a vendor dashboard
[ ] Our evaluation set exists: [N] real examples with expected
outputs, owned by: [name]
[ ] Our workflows/editorial value are documented independently
of any tool
RENTED (the replaceable):
Current model/provider: ______________ Monthly cost: ______
[ ] The model is behind a thin wrapper / single interface
[ ] Swapping to an alternative takes days, not months
Tested alternative: ______________ Last tested: [date]
TRIPWIRES (review quarterly):
[ ] Price change > ___% → trigger: re-run eval on alts
[ ] Deprecation announced → trigger: swap plan within 2 wks
[ ] Quality drop on our eval → trigger: [name] investigates
ONE-HORSE BETS (deliberate lock-in, eyes open):
We accept lock-in for: ______________ because: ______________
08 · Metrics
Count what you can move
- Output metrics (subscribers, revenue, reach) tell you if you won — but you can’t move them directly.
- Input metrics (returning users, completion rate, corrections per AI output) are what you can act on.
- Pick one North Star and 2–3 input metrics that drive it — and build the measurement before launch.
And remember: at your scale, dashboards lie. With early-stage traffic, most analytics movement is noise — don’t redesign because Tuesday dipped. Five real user conversations beat any dashboard: watch someone use it, don’t ask if they like it. Treat qualitative signals as data — log every complaint and workaround in one place, review weekly.
METRICS WORKSHEET — [product] NORTH STAR (one output metric that means "we won"): Metric: ______________ Current: ______ Target (6 mo): ______ INPUT METRICS (2–3 things we can directly move): 1. ______________ measured how: ______________ owner: ______ 2. ______________ measured how: ______________ owner: ______ 3. ______________ measured how: ______________ owner: ______ BUILT BEFORE LAUNCH? [ ] Each input metric has a working measurement [ ] Someone looks at them on a schedule: [who], [when] THE MONDAY QUESTION The ONE number I check next Monday morning: ______________
One row per complaint, workaround or observation. Review weekly; pick ONE change; ship it.
| Date | Source | What we saw / heard (their words) | Pattern? Action? |
|---|---|---|---|
QUALITATIVE SIGNAL LOG — [product]
Review weekly. Pick ONE change. Ship it.
Date | Source (user call, support, observed) | What we saw/heard (their words) | Pattern? Action?
-----|---------------------------------------|----------------------------------|------------------
| | |
| | |
| | |
09 · After Launch
The launch is not the end: who keeps it alive?
Products die quietly when nobody is scheduled to look. Integration beats adoption: the question isn’t “do we use AI?” but how different AI developments affect different parts of the business, and what restructuring the interplay needs.
- A named owner who wakes up when it breaks — “the team” owns nothing.
- A runbook: what it does, what it costs, what to check weekly, who to call when the model misbehaves.
- A standing ritual: a monthly slot in an existing meeting.
Communication resolves insecurity: make the roles, guardrails and guidelines explicit, give teams a place to openly share learnings and failures, and connect experiments to strategic decisions in both directions.
RUNBOOK — [product] Last updated: [____] OWNER (by name, not team): ______________ Deputy when owner is away: ______________ WHAT IT DOES (3 sentences, plain language): ______________________________________________________ WHAT IT COSTS Model/API: ______ /mo Infra: ______ /mo Human review: ______ hrs/wk Total: ______ /mo WEEKLY CHECKS (15 min, every [day]) [ ] Input metrics within normal range (see metrics sheet) [ ] Spot-check [N] AI outputs for quality [ ] Signal log reviewed, new entries triaged [ ] Costs within budget WHEN THE MODEL MISBEHAVES First call: ______________ (owner) Escalation: ______________ (editorial) / ______________ (tech) Kill switch — how to pause AI output safely: ______________________________________________________ What users see while paused: ______________ STANDING RITUAL Monthly slot in: [existing meeting] on: [date/recurrence] Agenda: signals → ONE change → who ships it by when SUNSET CRITERIA (decide now, while you're calm) We retire this product if, for [3] consecutive months: ______________________________________________________
The three questions that closed the session. Answer them in writing, with your team, once a quarter.
REFLECTION ROUND — [date] 1. YOU: Who owns your product the day after launch — by name? → ______________ 2. YOUR METRICS: What is the ONE number you would check next Monday morning? → ______________ 3. YOUR SCOPE: What would you cut to launch two weeks earlier? → ______________
10 · Addendum
Staying sane with AI
FOMO is a sales tactic. Three tips for keeping your head while everyone around you is losing theirs:
- Find a manageable challenge. Pick one small, real problem and solve it end to end — confidence comes from finished things, not from keeping up with every release.
- Always think alone first. Form your own view before you ask the model — otherwise its first answer becomes your anchor.
- Tell the AI to prompt you. Flip the roles: have it ask you the questions. You’ll find out what you actually know — and what you’ve been outsourcing.