Skip to content

The research that told a roadmap 'not that' before six months were spent

Jobs-to-be-done interviews and demand signals redirected a planned vendor-risk module toward the problem customers actually had, expense reconciliation, before a single sprint was spent building the wrong thing.

The verdict
Two candidates for the next build, and all six jobs scored the same way underneath
Export briefOpen research briefRM
On the roadmap · planned nextStruck
Vendor-risk module
Where the ask came fromTwo enterprise prospects, in sales calls
Interviews3 of 12 · 1 unprompted, 2 when asked
Demand outsideNear-invisible
MarketCrowded · 4 of 5 tools already do it
Recommended insteadIn its place
Expense reconciliation
Where it came fromNever requested as a feature
Interviews10 of 12, unprompted
Demand outsideLarge and growing
The gap, namedMulti-currency card match · no tool does it

All six candidate jobs, scored the same way Unprompted Only when asked

Same frame for prospects and current buyers
Job · roadmap areaInterviewsSearch, 2 yrs (shape)Covered fully byVerdict
Reconcile card spend across currenciesExpense reconciliation100 of 5 toolsRecommend
Get spend approved before it's committedSpend approvals64 of 5 toolsCrowded
Explain a budget variance to leadershipVariance reporting52 of 5 toolsWatch
Vet a new vendor before signingVendor risk14 of 5 toolsDrop from roadmap
Forecast next quarter's cashCash forecasting21 of 5 toolsToo rare
Set card limits per teamCard controls25 of 5 toolsCrowded
Interviews that surfaced the real pain point10 of 12Unprompted · counted from coded transcripts
Engineering time redirected before being spent~2 quartersThe module's own estimate · time not spent
Research, start to brief5 weeksFixed price · ended in a brief, not a build

What the engagement involved

Industry
Professional Services
Duration
5 weeks
Cooperation model
Fixed price
Services
JTBD interviewsDemand signal analysisMarket gap analysis
Integrations
HubSpotDocuSignXeroGoogle Workspace
Technologies
Jobs-to-be-done interviewsSearch demand analysisCompetitive feature mappingInterview codingResearch synthesis
Team
1 Project lead2 Frontend engineers1 Backend engineer

Client name withheld under NDA. Engagement details are shown to the extent our agreement permits.

Introduction

The question we were asked

The roadmap had a vendor-risk module penciled in as the next big build, based on two enterprise prospects who'd asked for it in sales calls.

A B2B product company with two quarters of engineering time about to be committed to a vendor-risk module, on the strength of two enterprise prospects who had asked for it in sales calls. Leadership wanted evidence before committing, which is an unusually good instinct and an unusually cheap one to satisfy: five weeks of research against six months of building.

Product & Market Research

The decision, first

  1. 01

    The planned feature was a loud ask from a narrow segment; the real one was never requested directly.

    Two prospects asking loudly is a real signal about two prospects; the market signal was in what ten of twelve raised without being asked.

  2. 02

    Coding for the job, not the feature, is what surfaced a pain point nobody named.

    Buyers name features and describe jobs, and the two rarely match. Coding for the job is what let a pain point nobody requested become the recommendation.

  3. 03

    Five weeks of research replaced six months of building the wrong module.

    The saving is asymmetric: research that confirms the plan costs five weeks, and research that redirects it saves two quarters, so the expected value isn't close.

The problem

What the numbers couldn't answer

  • Before committing two quarters of engineering time, leadership wanted evidence that the broader market actually had this problem, beyond the two loudest accounts in the pipeline.

    The evidence behind the roadmap decision was two conversations, both with prospects at the top of the pipeline, both mediated by a salesperson with an interest in the answer. That's a real signal about two accounts and no signal at all about a market. Nobody had asked whether the ask generalized, and the roadmap had no mechanism that would have surfaced it if it didn't.

    We ran jobs-to-be-done interviews across a dozen prospective and current buyers, pulled demand signals from search and adjacent tool usage, and mapped the market gap. Vendor risk turned out to be a niche ask from a narrow segment, while expense reconciliation showed up as an unprompted pain point in nearly every conversation and had a clear, underserved gap in the market.

The solution

How we worked it through

Ran jobs-to-be-done interviews with 12 prospective and current buyers

Interviews were anchored on the last real occurrence of the problem, because recalled behavior is evidence and predicted behavior isn't.

Pulled demand signals from search volume and adjacent-tool category trends

Search volume and adjacent-tool trends were pulled independently of the interviews, so the two signals could confirm each other or fail to.

Mapped the market gap across existing competitors' feature coverage

Competitor coverage was mapped feature by feature, which separates an unmet need from one already answered four times over.

Delivered a research brief recommending expense reconciliation over vendor risk

The brief named a specific gap (the multi-currency case buyers kept describing) instead of a category, which is what made it buildable.

Process

Phase by phase

  1. Phase 1: Interview

    Twelve buyers, one question

    Ran jobs-to-be-done interviews with twelve prospective and current buyers, coding for the underlying job.

    • Interview guide
    • Coded transcripts
  2. Phase 2: Size

    Signals outside the room

    Pulled demand signals from search volume and adjacent-tool category trends to size each candidate independently.

    • Demand analysis
    • Category trends
  3. Phase 3: Map

    Where the gap actually is

    Mapped competitors' feature coverage to separate an underserved need from a well-served one.

    • Competitive map
    • Gap analysis
  4. Phase 4: Recommend

    Say the unwelcome thing

    Delivered a research brief recommending expense reconciliation over the planned vendor-risk module.

    • Research brief
    • Recommendation memo
Four recurring jobs
Every card is a sentence a buyer said, sorted into the job it belongs to
Unprompted onlyRM
Said unprompted Said only when askedCoded for the job, not the feature named
Expense reconciliation10/12Reconcile card spend across currencies
“Every close, someone rebuilds the exchange rate for each card line in a spreadsheet.”INT-01Unprompted
“The feed says pesos, the ledger says dollars, and nothing matches until I do it by hand.”INT-03Unprompted
“Partners travel, the fees land in three currencies, and my week goes.”INT-08Unprompted
“The grant is in euros and the card is in dollars. I match them line by line.”INT-09Unprompted
“Two entities, one card program, and the totals never agree on the first pass.”INT-12Unprompted
Spend approvals6/12Get spend approved before it's committed
“I find out about the spend when the statement arrives.”INT-05Unprompted
“A request sat in someone's inbox for nine days.”INT-07Unprompted
“I want the no to happen before the card is swiped.”INT-10Unprompted
Variance reporting5/12Explain a budget variance to leadership
“The board asks why we're over, and I need an afternoon to answer.”INT-06Unprompted
“I can see the overspend. I can't see which client caused it.”INT-02Unprompted
“Explaining the variance takes longer than fixing it.”INT-04Unprompted
Vendor risk1/12Vet a new vendor before signing
“We onboarded a supplier the week of close and nobody checked their insurance.”INT-04Unprompted
“If you had vendor risk, we'd look at it.”INT-07After roadmap question
“Sure, a risk score would be useful.”INT-11After roadmap question
One card unprompted. The module on the roadmap was built on this column.
On screen

Four recurring jobs: every card a sentence a buyer said, sorted into the job it belongs to, with vendor risk alone in the last column and the two mentions that came only after a roadmap question shown apart.

What it changed

10 of 12, unprompted

Interviews that surfaced the real pain point

~2 quarters

Engineering time redirected before being spent

Expense reconciliation, underserved

Market gap identified

Ten of twelve is a count of interviews in which expense reconciliation surfaced without being prompted. The redirected engineering time is the module's own estimate, not a measured saving: it's time not spent, which isn't the same as time recovered. The market gap is a judgment supported by the competitive map, and is reported as one.

Client name withheld under NDA. Figures are approximate, drawn from the engagement’s own reporting.

Interviews, not a survey

Twelve buyer conversations coded for the job being done, whatever feature was requested.

Twelve conversations, each coded for the job the buyer was hiring a tool to do, set apart from the feature they asked for. The two are routinely different, and the feature request is the less useful of them. Interviews were structured around the last time the problem occurred, never around hypotheticals, because recalled behavior is evidence and predicted behavior isn't.

What shipped
  • Twelve interviews coded for the job, not the feature request
  • Anchored on the last real occurrence, not on hypotheticals
  • Same coding frame applied to prospects and current buyers
Who we spoke to
Twelve buyers, each transcript coded for the job being done, not the feature requested
All buyersRM
Interviews12Recorded and coded
Current buyers6Same coding frame
Prospective buyers6Same coding frame
Anchored onLast real occurrenceRecalled behavior, not predictions

All twelve interviews

Jobs coded = raised unprompted
IdBuyerRelationshipAnchored onJobs coded
INT-01ControllerEngineering consultancy · 240 staffCurrentJanuary close, Toronto office cards2
INT-02Finance managerMarketing agency · 130 staffProspectiveClient rebill for a London shoot2
INT-03Head of financeStaffing firm · 410 staffCurrentContractor per diems in pesos2
INT-04VP financeHealthcare group · 1,900 staffProspectiveA new lab supplier, the week of close3
INT-05Accounting leadArchitecture practice · 85 staffCurrentSite-visit travel billed in euros2
INT-06CFOSoftware company · 320 staffProspectiveBoard pack, Q4 overspend2
INT-07Procurement directorManufacturing · 2,600 staffProspectiveA capex request held for sign-off2
INT-08ControllerLaw firm · 210 staffCurrentFiling fees on partner cards abroad2
INT-09Finance directorNonprofit · 160 staffProspectiveGrant spend in two currencies2
INT-10Senior accountantAccounting firm · 95 staffCurrentConference trip, three card feeds2
INT-11Head of procurementLogistics · 3,400 staffProspectiveFuel cards over limit in March2
INT-12ControllerConstruction · 560 staffCurrentSubcontractor spend across two entities3
Expense reconciliation Any other job, unprompted10 of 12 raised reconciliation

INT-03 · coded transcript

52 min · 02/04/2026
Head of finance · Staffing firmCurrent buyer
QTake me back to the last time you closed the month. Where did the time go?
AThe contractor per diems. They're paid on cards in pesos and booked in dollars.Last real occurrence · January close
ASomeone pulls the rate for each line and rebuilds the match in a spreadsheet. Two days, every month.Job · reconcile card spend across currencies
QWhat did you try before that?
AWe asked for a better export. What we need is for the lines to just agree.Feature named · CSV export
AAnd I'd like approvals before the card gets used, not after the statement.Job · approve spend before it's committed
Coded to 2 jobs · 1 feature request kept, not counted
On screen

Who we spoke to: all twelve buyers, six current and six prospective, with sector, relationship, the last real occurrence each interview was anchored on and how many jobs the transcript coded to, beside one transcript coded for the job, with the feature it named kept but not counted.

Demand outside the room
Search and adjacent-tool trends, pulled independently of anything a buyer said, then compared
US · EnglishRM
Pulled before the interview coding was shared with this workspaceAgreement between the two signals is the test. Shapes only: relative interest, no volumes stated.InterviewsComparedSearch + tools

Search interest, two years

Relative interest · same scale for every term
Q1 '24Q2Q3Q4Q1 '25Q2Q3Q4
expense reconciliationspend approval workflowcard limits by teamvendor risk management

Adjacent-tool category trends

Shape only
Card-feed matching add-onsGrowing
Exchange-rate lookup toolsGrowing
Approval workflow apps—
Supplier risk registersNear-invisible

Do the signals agree?

Each candidate, both ways
CandidateIn the roomOutside itResult
Expense reconciliation10 of 12unpromptedLargeand growingAgree
Vendor risk3 conversationsloud, narrowNear-invisibleeverywhere elseDisagree
On screen

Demand outside the room: two years of search interest and adjacent-tool category trends pulled independently of the interviews and drawn to shape, then checked against them. Expense reconciliation was large and growing in both; vendor risk was the smallest term on the page.

Demand signals

Search volume and adjacent-tool trends used to size the ask independently of what buyers said.

What buyers say and what they search for are separate signals, and agreeing is what makes either trustworthy. Search volume for the two problem areas and usage trends in adjacent tools were pulled independently of the interviews, then compared. Expense reconciliation was large and growing in both; vendor risk was loud in three conversations and near-invisible everywhere else.

What shipped
  • Demand sized independently of what interviewees said
  • Search volume and adjacent-tool usage compared against interview coding
  • Agreement between the two treated as the test
Wanted often, served badly
Competitor coverage mapped feature by feature, so an unmet need separates from one already met
Export mapRM

Six candidate jobs

How often raised × how well served
↑ Interviews raising it unprompted (of 12)Tools doing it fully (of 5) →

The gap, named

What the brief recommends building
Multi-currency card matchPartial answers exist for reconciliation. None of the 5 tools matches a card line booked in one currency against a ledger in another.
Buyers describing it
INT-03“The feed says pesos, the ledger says dollars, and nothing matches until I do it by hand.”INT-09“The grant is in euros and the card is in dollars. I match them line by line.”

Coverage, feature by feature

5 tools · anonymized
FeatureABCDE
Expense reconciliation
Card feed matching
Receipt capture and match
Multi-entity close
Multi-currency card matchThe gap
Vendor risk
Supplier onboarding questionnaire
Sanctions and watchlist screening
Document expiry tracking
Supplier risk scoring
Expense reconciliation · partial answers, one specific gapUnderserved
Vendor risk · already met 4 times overCrowded
Does it fully Partial Not at all

A mapped gap

On screen

Wanted often, served badly: six candidate jobs plotted by how often buyers raised them against how many tools already do them fully, vendor risk in the corner four tools own, and the feature-by-feature coverage map that names the gap: multi-currency card matching.

Competitor feature coverage mapped, so an underserved need could be distinguished from a crowded one.

Competitor coverage was mapped feature by feature so an unmet need could be told apart from a need already met four times over. Vendor risk was crowded; expense reconciliation had partial answers with a specific gap: nobody handled the multi-currency case the buyers kept describing. The recommendation named that gap instead of the category, which is what made it buildable.

What shipped
  • Competitor coverage mapped feature by feature
  • Crowded needs separated from genuinely underserved ones
  • Recommendation named a specific gap, not a category

How the engagement ran

01
  1. 01

    A cross-functional team of 3 worked on a fixed price basis over 5 weeks, covering JTBD interviews, Demand signal analysis, Market gap analysis. We ran a weekly demo and a shared board they could read at any time. Their team took over day-to-day operation before the engagement ended, with handover built into the last phase.

    Five weeks, fixed price, ending in a brief instead of a build. Interviews were anchored on the last time the problem actually occurred, not on hypotheticals, and were coded for the job the buyer was hiring a tool to do, not the feature they named. The demand analysis ran independently of the interviews so the two could agree or fail to.

Twelve interviews, four recurring jobs

The pain nobody requested was the one nearly everybody described

Pick a recurring job to see which of the twelve buyers raised it unprompted, which only when asked, and which never did. Below it, the three tests that put expense reconciliation ahead of vendor risk. Switch tabs, or use the arrow keys once one is focused.

Twelve interviews, one recurring job at a time

“Reconcile card spend across currencies”Ten of twelve buyers described this without being asked, current and prospective alike, and none of them asked for it as a feature. It came out of the last close they had actually worked through: card lines in one currency, a ledger in another, a spreadsheet in between.
Current buyers
  • INT-01ControllerCurrent · Engineering consultancy
  • INT-03Head of financeCurrent · Staffing firm
  • INT-05Accounting leadCurrent · Architecture practice
  • INT-08ControllerCurrent · Law firm
  • INT-10Senior accountantCurrent · Accounting firm
  • INT-12ControllerCurrent · Construction
Prospective buyers
  • INT-02Finance managerProspective · Marketing agency
  • INT-04VP financeProspective · Healthcare group
  • INT-06CFOProspective · Software company
  • INT-07Procurement directorProspective · Manufacturing
  • INT-09Finance directorProspective · Nonprofit
  • INT-11Head of procurementProspective · Logistics
Raised unprompted10of 12 Unprompted · 10 Only when asked · 0 Not raised · 2
“Every close, someone rebuilds the exchange rate for each card line in a spreadsheet.”INT-01 · unprompted
Only reconciliation’s 10 of 12 and vendor risk’s three conversations are the study’s figures; the rest of the coding is illustrative.

Why expense reconciliation won over vendor risk

Did buyers raise it without being asked?The roadmap was built on two enterprise prospects asking in sales calls: a real signal about two accounts. Across twelve buyers coded the same way, the pain nobody requested was the one nearly everybody described.
Expense reconciliationRecommended10 of 12Unprompted, never named as a feature
Vendor-risk modulePlanned3 of 12Loud in three conversations, one unprompted
How the evidence was built

From two loud sales calls to a gap the market actually has

Three independent lines of evidence (what buyers described, what the market searches for, and what competitors already cover), each able to disagree with the others, ending in a brief instead of a build.

  1. 01 · Source
    Twelve buyer interviewsProspective and current buyers, each anchored on the last real occurrence of the problem, because recalled behavior is evidence and predicted behavior is not.
  2. 02 · Code
    Job, not featureEvery transcript coded for the job the buyer was hiring a tool to do, with the same frame applied to prospects and current buyers.
  3. 03 · Size
    Search + adjacent-tool trendsPulled independently of the interviews, so the two signals could confirm each other or fail to. Agreement is the test.
  4. 04 · Map
    Competitor coverageMapped feature by feature, which separates an unmet need from one already answered four times over.
  5. 05 · Deliver
    Research briefNames a specific gap (the multi-currency case buyers kept describing) instead of a category, so the recommendation is buildable.

So the answer isn’t the one you walked in with

Evidence that could have said "build it," and didn't

Nobody led the witness

Interviews were anchored on the last time the problem actually occurred and coded for the job, not the feature. A pain point counted when the buyer raised it unprompted, and ten of twelve did for reconciliation.

Signals able to disagree

Demand was pulled from search volume and adjacent-tool trends independently of the interviews, then compared. Vendor risk was loud in three conversations and near-invisible everywhere else, and that disagreement was reported.

Figures labeled for what they are

Ten of twelve is a count. The ~2 quarters is the module's own estimate: time not spent, not a measured saving. The market gap is a judgment supported by the competitive map, and is reported as one.

Is your next build backed by the market, or by the two loudest accounts? Scope your build in 3 minutes.

Scope your build
Have a project?

Let's talk

Running a large platform, shaping a first MVP, or getting a product ready for a funding round? Tell us where you are. We'll shape the process around it, and stay with you after launch.