nSpire AI · Product Lead (0→1)
From "it demos well" to "it's defensible"
I joined a pre-revenue AI practice platform that everyone described confidently and nobody could describe precisely. Six weeks later: a corrected product definition, a falsifiable readiness bar, a repositioned competitive strategy, and analytics that admit what they can't yet prove.
The situation
A conversational AI practice platform for enterprise L&D — learners rehearse hard conversations against an AI persona, then get a scored report. Pre-revenue, pre-launch. Engineering considered it feature-complete: "Functionality and core offerings-wise we are ready to go." That was true and also not decision-useful. The first external design-partner test — a large enterprise data-infrastructure company — was 2–4 weeks out, and the honest state of the product existed only as scattered intuition across four or five people. My first job wasn't to fix anything. It was to find out what was true.
The operating system
A five-decision operating system, documented in 20+ living documents. Build the capability map before touching the roadmap. Convert "it's buggy" into a readiness argument with a written exit criterion and a 3-tester QA round. Re-tier the competitive landscape — 23 competitors across a 70-feature taxonomy — and discover the real competition wasn't who looked similar. Instrument the product before it had users, then audit the instrumentation instead of assuming. And name the unsolved problem (the module description *is* the rubric) instead of letting it hide in a normal backlog.
What it changed
A product that can survive an enterprise buyer's scrutiny, and a team that can say precisely what works. The readiness story went from an unbounded blocker list to two verifiable bugs with a written exit criterion. The pitch moved from "AI roleplay tool" (crowded, funded) to "practice-and-feedback engine for leadership and people skills" (an open corner — empty on our map and on the analysts' maps too). And the documentation method itself became the deliverable: correction trails kept visible, uncertainty made machine-readable, honest audits over confident dashboards.
The numbers
-
20+
Living documents produced
-
23 × 70
Competitors × features mapped
-
37 · ~149
Events · properties designed
-
32-item · 3
QA script · testers
-
21/37
Events actually firing (the audit)
-
8
Demo personas built
-
16
Redesign requirements (phased PRD)
All company and partner names are anonymized. Competitor names are retained — that research was conducted entirely from public marketing sites.
Decision 1 — Build the map before touching the roadmap
I inventoried every visible capability in staging against a six-level status legend (✅ functional → ⚪ unvalidated), separated by role. Sixteen manager capabilities, twelve learner capabilities, seven core objects. Explicitly scoped as “a capability inventory, not a prioritized bug list” — prioritization got its own document, because mixing the two is how status arguments become bug arguments.
Then I ran a full end-to-end walkthrough with the CTO, and the map was wrong in ways that mattered.
The top-level object didn’t exist. Every internal doc, including mine, described a “Project” as the container for a learning initiative. The product has no such concept. The real hierarchy is Workspace → Track → Module, with Cohorts as sub-groups and Assignments as the gate that makes content visible to a learner. A whole open-questions section evaporated. (The dead concept outlived the correction: a day later the QA round surfaced an upload failure still returning “Could not load projects. Please try again.”)
Most of what I’d marked “unvalidated” was working. Cohorts and assignments were rated ⚪ untested. The demo exercised both live. They were fine.
And the framing itself was wrong. I’d built a document that read as “many open blockers.” The accurate read was “polish a near-complete build” — the remaining P0s were mostly scoping and UX decisions, not missing functionality.
I kept the corrections visible rather than editing them away. In the master doc and the problem list, superseded beliefs are marked in place with a ⚠️ and a date rather than deleted. The correction trail turned out to be more useful to the team than the conclusions — it showed which claims had been tested and which were still inherited assumption.
Decision 2 — Convert “it’s buggy” into a readiness argument someone can act on
Gut feel doesn’t survive contact with a launch date. I built a P0/P1 problem list where every issue carried an ID and an observed behaviour, the significant ones a stated risk, and the most dangerous one an exit criterion.
The most useful line I wrote in the whole project was P0.1’s:
“all 5 internal testers can end AND cancel a session cleanly, twice, before the test date.”
That converts an argument into a test.
Then I ran an internal QA round against a 32-item script, with testers role-playing the actual prospect rather than testing abstractly. Three of five completed. Results: 6, 17 and 6 items failed. Every tester got the same closing question — would they show this to the partner’s managers today? Two yes, one no.
The finding that mattered wasn’t a new bug. It was that three independent testers had confirmed, 3-for-3, the exact two bugs the CTO had already named. That reframed the readiness conversation from a debate into a checklist.
It also caught what a single tester never would have. The generic “cohort scoping issue” engineering had described was actually much worse and much more specific: no one could create a cohort at all. No ”+ New cohort” button existed. That blocks the entire authoring-to-learner demo — you can publish content and then have nothing to prove a learner ever sees it. I escalated it from P1.8 to P0.7.
Net effect: the readiness story went from an unbounded blocker list to fix and verify two bugs, curate the demo content, bound the demo path.
The starting state
“It’s buggy” — scattered intuition, no launch bar
What that produced
Unbounded blocker lists and status arguments
The risk
A readiness debate nobody could win before the test date
The flip
A written exit criterion + a 3-tester QA round → two verifiable bugs
Decision 3 — We had been benchmarking against the wrong competitors
The team’s mental competitive set was the AI sales-roleplay category. I mapped 23 competitors across a 70-feature taxonomy, sourced strictly from vendor marketing sites — a constraint I stated twice in the document, because the matrix records what a vendor claims, not what a product does.
The result was counterintuitive:
- Tier 1 (direct AI roleplay — sales-skewed) dominated on raw feature breadth. Hyperbound 51/70, Yoodli 44, Quantified 43.
- Tier 2 (AI leadership coaching) clustered low. Valence 29, BetterUp 26–28, CoachHub 22–23.
A naive read flags Tier 1 as the threat. But our first wedge was leadership development — which meant we were competing for the same budget, the same buyer (Head of L&D), and the same use case as Tier 2. The tools that looked most similar to us weren’t our competition. The tools that looked least similar were.
Tier 1 · Direct AI roleplay — sales-skewed
Highest feature counts. Different buyer.
Tier 2 · AI leadership coaching
Lower counts. Same buyer, same budget, same use case as us.
What it changed: the pitch moved from “AI roleplay tool” (crowded, funded) to “practice-and-feedback engine for leadership and people skills” (open corner). It also produced a deliberate non-goals list — conversation intelligence on real calls, CRM integrations, VR, body-language scoring — that would look like negligence under the old framing and is obviously correct under the new one.
The sharpest supporting evidence came from analyst structure rather than from competitors. Forrester carved out “Leadership and Human Skills Development Platforms” as a category in 2024 — and published only a Landscape. No Wave. No scored leaders. Analysts had documented the gap, proven the mechanism works in sales, and named the category nobody yet leads.
“The corner isn’t just empty on our competitive map — it’s empty on theirs.”
Decision 4 — Instrument the product before it had users
I built the v1 analytics plan from a screen-by-screen walkthrough rather than a feature list: 37 events, ~149 properties, prioritized P0–P3, architected so that Group Analytics keyed on workspace would let every report slice per account.
Two details I’m still happy with.
A session_id generated at configuration and carried through Started → Completed → Feedback Viewed, which makes a practice attempt a joinable object instead of four unrelated events — so scores can be analyzed against the answer mode and session length the learner chose.
Learning outcomes ride on the feedback event. overall_score, per-skill skill_scores, is_passed, mastery_achieved, plus a pre-computed lowest_skill. This is the bit that matters commercially: “did the training change behaviour?” is the question L&D buyers can never answer, and this schema turns it into a query.
Session Configured
session_id generated here
Roleplay Started
answer mode, persona, length
Roleplay Completed
the middle of the spine
Feedback Viewed
scores, mastery, lowest_skill
Three flows I hadn’t observed got a [CONFIRM] marker rather than a confident guess — honest uncertainty made machine-readable and routable to engineering.
The audit is the part worth reporting. When instrumentation shipped, I checked the plan against reality instead of assuming. 21 of 37 events firing. Sixteen missing — and the one that mattered was Roleplay Completed, the middle of the spine. Its absence breaks the activation funnel, the retention curve, and every learning-outcome report simultaneously. I built the five dashboards anyway, documented exactly which numbers were untrustworthy and why — the per-account architecture was designed but never switched on, and demo workspaces were still counted as production — and escalated.
Shipping a plan is not the same as shipping instrumentation. Checking is cheap. Not checking means a quarter of confident, wrong reporting.
Decision 5 — Name the unsolved problem instead of designing around it
The hardest thing I found wasn’t a bug. It’s this: the module description is the rubric. Scoring is an LLM judging the transcript against whatever prose the author happened to write. There is no rubric engine.
It’s an elegant shortcut — one field does authoring, learner briefing and scoring at once, and it shipped fast. But scores are only as consistent as one manager’s paragraph, and a title-only module can be published, meaning an empty description produces both a blind learner and an unscoreable session. For a product whose enterprise value proposition is measuring skill, that’s the credibility gap.
The related one: the “Graded = counts toward readiness” label has no rollup behind it. No per-track percentage, no threshold, no certification. A UI promise the backend can’t keep — and the exact thing an L&D buyer justifies spend with.
I refused to let either hide inside a normal backlog. In the user stories I introduced a [label-only] tag for stories where the interface exists and the logic doesn’t, and a closing section titled “Known gaps to resolve (not stories — definition needed).” You cannot estimate your way out of an undefined concept, and letting one sit in a sprint board as a normal ticket is how it stays undefined for a year.
Both then became the spine of a redesign PRD — planning only, scoped as post-test work, none of it built: a proposed first-class rubric engine with structured, individually-scored criteria; configurable feedback dimensions; and a defined readiness rollup where retakes count at best score, so repeat practice carries no penalty (a default call, still to be confirmed with engineering).
Outcomes
- Product definition corrected and documented — object model, four module axes, role scoping, three answer modes. The reference the team now works from.
- Readiness reframed from an open blocker list to two verifiable bugs with a written exit criterion, backed by a 3-tester QA round.
- Strategic repositioning — competitive set moved a full tier, producing a differentiated position and an explicit non-goals list.
- Analytics in production — 5 dashboards live, with an honest audit of what the data could not yet support.
- GTM enablement — 8 demo personas covering all 12 content tracks, plus a prospect-branded content pack written the day after the discovery call. A document, not yet a live build — standing the trial up was still open work.
- A redesign PRD (planning stage) with 16 requirements, phased, each scope decision recorded with its rationale.
Not a win: the flagship design-partner test slipped its mid-July date and hadn’t run when I wrote this up. The docs say so plainly, including the open decision about whether it should be deprioritized behind a faster-moving prospect. A living document that only records wins isn’t a living document.
What I’d do differently
I’d open with the walkthrough, not close with it. I built the capability map from solo exploration and only then sat down with the CTO. It was wrong in structural ways — a whole object that didn’t exist — and one hour of someone else’s screen share caught all of it at once. Independent exploration finds questions efficiently and answers very slowly.
I’d have audited the instrumentation the day it shipped. I only caught the missing P0 event when I sat down to audit on 21 July, well after go-live — which means a stretch of dashboards that looked fine and weren’t.
And I’d design the drift problem in from the start. By late July the docs had gone three weeks without a roll-up, and two claims in our own recorded tutorials contradicted the source of truth — one of them, whether video is actually scored, made to prospects on camera. Documentation that isn’t maintained on a cadence doesn’t decay quietly. It decays into things your team says out loud to customers.