Skip to content

A scored roadmap that picked the one AI project worth funding first

A use-case map and reference architecture that turned eleven competing AI ideas into a single ranked roadmap, and named the two not worth building at all.

Funding sequence

Three funded builds staged across four quarters on one shared architecture · agreed by the partnership

Export board paper Fri, 10/16EH
Candidates scored11From 6 department interviews
Funded builds3On one reference architecture
Not building2Declined in writing, before spend
Time to funded first build3 wksFrom roadmap presented to budget
First valueQ2Drafting assistant live

Sequence

Shared layers first · each build lands on them
WorkstreamQ1Q2Q3Q4
Shared layersspecified once 4 layersdocument store · retrieval layer · evaluation harness · access model
Proposal & engagement-letter drafting#1 · 88 · Luis CarreraBuild Live
Billing narratives & WIP automation#2 · 83 · Tom KesslerBuild Live
Precedent research assistant#3 · 78 · Helen OstrowskiBuild Live
Next on the layersTax memo first drafts · #4then 5 more in backlog, re-scored each quarter
First value · Q2

Funded

Committed spend · approved by the partnership
RankBuildScoreStagedBudget
#1Proposal & engagement-letter draftingBusiness Dev. · Luis Carrera88Q1–Q2Approved
#2Billing narratives & WIP automationFinance & Billing · Tom Kessler83Q2–Q3Approved
#3Precedent research assistantKnowledge · Helen Ostrowski78Q3–Q4Approved

Not worth building

Reasoning on file
Always-on meeting summarizerUC-05 · scored 56Cost per query at realistic volume outweighs the hours it saves.
Matter profitability forecasterUC-07 · scored 55The time and matter data it depends on is captured inconsistently across offices.

Who, what, and how long

Industry
Professional Services
Duration
6 weeks
Cooperation model
Fixed price
Services
Use-case discovery & scoringReference architectureRisk & governance review
Integrations
HubSpotDocuSignXeroGoogle Workspace
Technologies
Scoring frameworkReference architectureRetrieval-augmented generationVector databaseGovernance framework
Team
1 Project lead1 ML engineer1 Backend engineer1 Data engineer

Client name withheld under NDA. Engagement details are shown to the extent our agreement permits.

Introduction

The question we were asked

The firm's partners had eleven AI ideas floating between departments and no shared way to compare them. We ran a scored use-case map, designed the reference architecture the winning ideas would share, and handed back a sequenced roadmap with an honest verdict on which ideas to shelve.

A mid-sized professional services firm where six departments had each arrived at an AI idea independently, and the partnership had been discussing them across six months of meetings without converging. Nothing had been funded and nothing had been refused, which is the worst of both: the enthusiasm was real, and the comparison was impossible because no two ideas had been described in the same terms.

AI Strategy & Architecture

The decision, first

  1. 01

    The most valuable output was the two ideas we argued against, not the one we ranked first.

    Two documented refusals stopped spend that would have run for months before failing, which is worth more than accelerating the one idea that was always going to be funded.

  2. 02

    Eleven ideas scored on one rubric ended a debate that six months of meetings hadn't.

    The rubric ended the debate because it made the ideas comparable, not because it made them measurable. The six months of meetings had never had a shared axis.

  3. 03

    A shared architecture is what makes a roadmap affordable. Without it, each project pays for its own foundation.

    Four teams each standing up a document store, a retrieval layer, an evaluation harness and an access model is four times the foundation. One shared roadmap avoids it.

What the numbers couldn't answer

01
  1. 01

    Every department had its own AI wishlist (a drafting assistant, a research tool, a billing automation), each pitched with enthusiasm but no shared way to compare cost, risk or payoff. Leadership couldn't tell which idea to fund first, or whether any were worth funding at all.

    Each idea was pitched in the vocabulary of its own department: one as hours saved, one as risk reduced, one as a competitive necessity. There was no shared measure, so the argument was resolved by whoever presented most recently. The billing automation and the research tool had both been described as the obvious first project, by different partners, in the same quarter.

    We interviewed each department, scored all eleven candidates on value and feasibility, and designed one reference architecture that the top-ranked ideas could share instead of each building its own stack. Two ideas were named unfit to build given data readiness and cost exposure.

How we worked it through

  1. 01

    Interviewed six department heads to surface every candidate use case

    Six department heads were interviewed about what they were trying to achieve, not what they wanted built, which is how eleven candidates surfaced instead of six.

  2. 02

    Scored all eleven ideas on business value and technical feasibility

    One rubric scored all eleven (value on hours and revenue, feasibility on data readiness, integration surface and regulatory load) in a joint session, out in the open.

  3. 03

    Designed one shared reference architecture for the top four ideas

    The top four turned out to need the same four layers, so those are specified once with named boundaries and each use case is an application on top.

  4. 04

    Ran a data and cost risk review and named two ideas not worth building

    The refusals were the hardest conversation of the engagement, and were delivered in writing so they could be revisited without being relitigated.

Process

Phase by phase

  1. Phase 1: Surface

    Every candidate, on the table

    Interviewed six department heads to collect every AI idea in circulation, including the ones that had never been formally proposed.

    • Use-case inventory
    • Interview notes
  2. Phase 2: Score

    Value against feasibility

    Scored all eleven candidates on business value and technical feasibility using one rubric, so enthusiasm stopped being the ranking mechanism.

    • Scoring rubric
    • Ranked candidate list
  3. Phase 3: Architect

    One design for the winners

    Designed a reference architecture the top four could share: retrieval, access control and evaluation built once instead of four times.

    • Reference architecture
    • Integration map
  4. Phase 4: Decide

    Sequence, and what to shelve

    Ran a data-readiness and cost review, sequenced the roadmap, and documented why two ideas shouldn't be built.

    • Sequenced roadmap
    • Risk register
    • Shelved-idea rationale
Use-case inventory

All eleven candidates as first surfaced, with the department, the sponsor and the interview each came from

Interview notes Tue, 09/22EH
1Surface · Wk 1–2Use-case inventory
2Score · Wk 2–3Rubric · ranked list
3Architect · Wk 4–5Reference architecture
4Decide · Wk 6Roadmap · risk register

Interviews

6 department heads
Dana WhitcombINT-1Tax partner · Tax · Wk 1 · Mon2 ideasUC-01 · UC-02
Marcus IlesanmiINT-2Head of audit · Audit & Assurance · Wk 1 · Tue2 ideasUC-03 · UC-04
Priya RamanINT-3Advisory partner · Advisory · Wk 1 · Wed1 ideaUC-05
Tom KesslerINT-4Finance director · Finance & Billing · Wk 1 · Thu2 ideasUC-06 · UC-07
Helen OstrowskiINT-5Head of knowledge · Knowledge · Wk 2 · Mon2 ideasUC-08 · UC-09
Luis CarreraINT-6BD director · Business Dev. · Wk 2 · Tue2 ideasUC-10 · UC-11

Candidates

In the order surfaced · not yet scored
IDCandidateDepartmentSponsorFrom
UC-01Tax memo first draftsPitched as: a slideTaxDana WhitcombINT-1
UC-02Timesheet capture nudgesPitched as: a hallway askTaxDana WhitcombINT-1
UC-03Audit sampling document reviewPitched as: a spreadsheetAudit & AssuranceMarcus IlesanmiINT-2
UC-04Client intake & conflict triagePitched as: an email paragraphAudit & AssuranceMarcus IlesanmiINT-2
UC-05Always-on meeting summarizerPitched as: a slideAdvisoryPriya RamanINT-3
UC-06Billing narratives & WIP automationPitched as: a spreadsheetFinance & BillingTom KesslerINT-4
UC-07Matter profitability forecasterPitched as: an email paragraphFinance & BillingTom KesslerINT-4
UC-08Precedent research assistantPitched as: a slideKnowledgeHelen OstrowskiINT-5
UC-09Internal policy Q&APitched as: a hallway askKnowledgeHelen OstrowskiINT-5
UC-10Proposal & engagement-letter draftingPitched as: a slideBusiness Dev.Luis CarreraINT-6
UC-11Pitch imagery generatorPitched as: a meeting mentionBusiness Dev.Luis CarreraINT-6
Asked what each department was trying to achieve, not what it wanted built · 11 candidates from 6 heads
On screen

The use-case inventory: the four phases of the six-week engagement, the six department-head interviews, and all eleven candidates as first surfaced with their department, sponsor and source interview.

What it changed

11

Candidate use cases scored

2

Ideas shelved before spend

3 wks

Time to funded first build

Eleven candidates and two shelved are counts from the engagement itself. Time to funded first build is measured from the roadmap being presented to the first project being approved with a budget. That's the outcome the firm was actually buying: the six preceding months had produced no such decision.

Client name withheld under NDA. Figures are approximate, drawn from the engagement’s own reporting.

A comparable score

One rubric across value and feasibility, so eleven ideas pitched in eleven ways became directly comparable.

Eleven ideas arrived in eleven shapes (a slide, a spreadsheet, a paragraph in an email), which made the argument about which to fund unresolvable. Each was restated against one rubric: value scored on hours saved and revenue exposed, feasibility on data readiness, integration surface and regulatory load. Scores were set in a joint session with the sponsoring departments present, so the ranking was agreed, not handed down.

What shipped
  • Eleven proposals restated against a single rubric
  • Value on hours and revenue; feasibility on data, integration, regulation
  • Scored jointly with sponsors in the room
Scoring sheet

Eleven candidates restated against one rubric · weighted across six criteria · sorted by total

Joint session · sponsors present Fri, 10/09EH
Agreed weights · out of 100 Value 40 Feasibility 60Hours saved25Revenue exposed15Data readiness20Integration surface15Regulatory10Cost at volume15

All candidates

Each score argued against the rubric definition · 1–5
#CandidateHrs 25Rev 15Data 20Intg 15Reg. 10Cost 15TotalVerdict
1Proposal & engagement-letter draftingBusiness Dev.88Fund
2Billing narratives & WIP automationFinance & Billing83Fund
3Precedent research assistantKnowledge78Fund
4Tax memo first draftsTax70Backlog
5Audit sampling document reviewAudit & Assurance64Backlog
6Internal policy Q&AKnowledge58Backlog
7Always-on meeting summarizerAdvisory56Not building
8Matter profitability forecasterFinance & Billing55Not building
9Client intake & conflict triageAudit & Assurance52Backlog
10Timesheet capture nudgesTax44Backlog
11Pitch imagery generatorBusiness Dev.24Backlog
Total = Σ weight × score ÷ 5Verdicts set after the risk review, not by the total alone11 of 11 scored · 0 disputed
Eleven shapes, one formA slide, a spreadsheet, an email paragraph: each restated in the rubric's terms before scoring.
Value against feasibilityValue on hours and revenue; feasibility on data readiness, integration surface, regulation, cost.
Agreed, not deliveredScores set in a joint session with the sponsoring departments in the room.
On screen

The scoring sheet: eleven candidates restated against one rubric, weighted across six value and feasibility criteria and sorted by total from 88 down to 24, each carrying a fund, backlog or not-building verdict.

Reference architecture

Four layers specified once, with named boundaries · the top-ranked ideas are applications on top

Integration map Tue, 11/17EH

One design for the winners

Build-versus-buy stated per layer
#1 · UC-10Proposal & engagement-letter draftingFunded
#2 · UC-06Billing narratives & WIP automationFunded
#3 · UC-08Precedent research assistantFunded
#4 · UC-01Tax memo first draftsNext
Shared layers · specified once · applications build on top, not beside
Access modelboundary: Who may retrieve whatMatter-level permissions over the firm's existing directoryused by all fourBuild
Evaluation harnessboundary: Is the answer good enoughFirm-specific test sets, run before any releaseused by all fourBuild
Retrieval layerboundary: Find the right passagesRetrieval-augmented generation over a bought vector databaseused by all fourBuild on bought
Document storeboundary: One copy of the sourceManaged store in the firm's own tenancyused by all fourBuy
Four teams, four stacksFour times the foundation
One roadmap, one foundationA new use case is an application on top

The funded builds on these layers

Planned, per the funding sequence
3funded builds, one foundationEach an application on the shared layers · none stands up its own stack
Proposal & engagement-letter drafting#1 · UC-10 · Q1–Q2Funded
Billing narratives & WIP automation#2 · UC-06 · Q2–Q3Funded
Precedent research assistant#3 · UC-08 · Q3–Q4Funded
Staged across four quarters · first value in Q2

Rules for every build

From the reference design
Each layer is specified once, with a named boundary
A new use case is an application on top, not a fresh stack
Build-versus-buy is stated for every layer
Firm-specific test sets run before any release
Retrieval respects matter-level permissions
One copy of the source, in the firm's own tenancy
Four teams never stand up four stacks
On screen

The reference architecture: document store, retrieval layer, evaluation harness and access model specified once with named boundaries and a build-versus-buy line, the top four ideas on top, and the rules every funded build follows on those layers.

One shared architecture

A single reference design the top-ranked ideas all draw on, so four teams never stand up four stacks.

The top-ranked ideas turned out to need the same four things: a document store, a retrieval layer, an evaluation harness and an access model. Instead of letting four teams each stand one up, the reference design specifies those once with named boundaries, so a new use case is an application on top, not a fresh stack. It names the build-versus-buy line for each layer explicitly.

What shipped
  • Four shared layers specified once, not per team
  • New use cases build on top, never beside
  • Build-versus-buy line stated for every layer
Risk register

All eleven ideas reviewed for data readiness, cost exposure and regulatory reach · a revisit condition on each

Export register Wed, 10/14EH

Register

11 entries · 2 declined in writing
IDIdeaDepartmentData readinessCost exposureRegulatoryVerdictRevisit when
UC-01Tax memo first draftsTaxLowLowElevatedBacklogStarts once the research assistant is live on the retrieval layer
UC-02Timesheet capture nudgesTaxElevatedModerateLowBacklogRevisit if time-entry lag becomes a billing issue
UC-03Audit sampling document reviewAudit & AssuranceLowLowModerateBacklogRevisit when the harness has an audit test set
UC-04Client intake & conflict triageAudit & AssuranceModerateLowModerateBacklogRevisit when intake moves to one system
UC-05Always-on meeting summarizerAdvisoryModerateHighModerateNot buildingRevisit when cost per query at realistic volume falls below the value of the hours saved
UC-06Billing narratives & WIP automationFinance & BillingLowLowLowFundFunded · review at first value
UC-07Matter profitability forecasterFinance & BillingHighModerateModerateNot buildingRevisit when matter and time codes are captured the same way in every office
UC-08Precedent research assistantKnowledgeLowModerateLowFundFunded · review at first value
UC-09Internal policy Q&AKnowledgeLowModerateLowBacklogRevisit as a by-product of the research assistant
UC-10Proposal & engagement-letter draftingBusiness Dev.LowLowLowFundFunded · review at first value
UC-11Pitch imagery generatorBusiness Dev.ElevatedHighHighBacklogRevisit only if brand guidelines allow generated imagery
lowmoderateelevatedhighRatings read from the feasibility scores on the sheet

Decline · Always-on meeting summarizer

UC-05 · scored 56 · #7 of 11
Why not nowCost per query at realistic volume outweighs the hours it saves.
What would have to changeRevisit when cost per query at realistic volume falls below the value of the hours saved.
Written, not verbal · sponsor Priya RamanCost per query

Decline · Matter profitability forecaster

UC-07 · scored 55 · #8 of 11
Why not nowThe time and matter data it depends on is captured inconsistently across offices.
What would have to changeRevisit when matter and time codes are captured the same way in every office.
Written, not verbal · sponsor Tom KesslerData readiness

A stated no

On screen

The risk register: all eleven ideas rated for data readiness, cost exposure and regulatory reach with a revisit condition on each, and the two written declines stating what would have to change.

Two ideas named unfit given data readiness and cost exposure, with the reasoning written down.

Two candidates were named unfit to build now, with the reasoning written down instead of delivered verbally. One depended on data captured inconsistently across offices; the other cost more per query at realistic volume than the hours it saved. Both entries state what would have to change for the answer to become yes, so the decision can be revisited without being relitigated.

What shipped
  • Two candidates declined, with the reasoning documented
  • Declines grounded in data readiness and cost per query
  • Each records what would have to change to reverse it
Ways of working

How the engagement ran

  • A cross-functional team of 4 worked on a fixed price basis over 6 weeks, covering Use-case discovery & scoring, Reference architecture, Risk & governance review. We ran two-week increments, each one shippable, reviewed with them before it merged. Decisions were recorded as they were made, so the reasoning survived the people who made it.

    Six weeks, fixed price, ending in a document, not a system. Scoring happened in a joint session with the sponsoring departments present, which is the difference between a ranking that's agreed and one that's disputed the following week. The two refusals were written down with their reasoning, not delivered verbally.

One set of scores, three weightings

The two ideas we declined were never the lowest scores

The eleven candidates ranked on the rubric the firm agreed, then on payoff alone and on feasibility alone. Watch where the two declines land, and why the verdict came from the risk review and not the total. Switch tabs, or use the arrow keys once one is focused.

Eleven candidates, ranked by weighted total

The weighting agreed with the sponsors in the room. The top three are funded at 88, 83 and 78. The two declines sit mid-table at #7 and #8: they were not refused for scoring low.

Hours saved25
Revenue exposed15
Data readiness20
Integration surface15
Regulatory load10
Cost at volume15
  1. #1 Proposal & engagement-letter draftingBusiness Dev.88 Fund
  2. #2 Billing narratives & WIP automationFinance & Billing83 Fund
  3. #3 Precedent research assistantKnowledge78 Fund
  4. #4 Tax memo first draftsTax70 Backlog
  5. #5 Audit sampling document reviewAudit & Assurance64 Backlog
  6. #6 Internal policy Q&AKnowledge58 Backlog
  7. #7 Always-on meeting summarizerAdvisory56 Not building
    Cost per query at realistic volume outweighs the hours it saves.
  8. #8 Matter profitability forecasterFinance & Billing55 Not building
    The time and matter data it depends on is captured inconsistently across offices.
  9. #9 Client intake & conflict triageAudit & Assurance52 Backlog
  10. #10 Timesheet capture nudgesTax44 Backlog
  11. #11 Pitch imagery generatorBusiness Dev.24 Backlog

Why two were declined anyway

The verdict came after a data and cost review, not from the total. Each decline hinged on one feasibility score of 1 out of 5:

  • Matter profitability forecaster: data readiness 1/5
  • Always-on meeting summarizer: cost at volume 1/5

Same scores, different weights

Every view uses the same 1–5 scores per criterion; only the weights change. Total = Σ weight × score ÷ 5.

The weighting the firm scored and decided on.

Architecture

From eleven pitches to one funded build on one foundation

Six weeks, ending in a document, not a system. Each stage hands the next something written down, so the ranking and the refusals can be revisited without being relitigated.

  1. 01 · Source
    Six department interviewsHeads asked what they were trying to achieve, not what they wanted built, so eleven candidates surfaced instead of six.
  2. 02 · Score
    One rubric, joint sessionEvery idea restated as value (hours, revenue) against feasibility (data, integration, regulation), scored with sponsors present.
  3. 03 · Design
    Four shared layersDocument store, retrieval, evaluation harness and access model specified once, with a build-versus-buy line per layer.
  4. 04 · Record
    Risk register & written declinesData readiness and cost per query reviewed before spend; each decline states what would have to change.
  5. 05 · Deliver
    Sequenced roadmapThe top-ranked ideas staged as applications on the shared layers; a first build funded three weeks after presentation.

Before any money moves

Checked for data, cost & a documented no

Data readiness checked before spend

Feasibility is scored on whether the data an idea depends on is actually captured consistently. One candidate was declined because its data is recorded differently across offices.

Cost per query at realistic volume

Cost is judged at the volume the firm would really run, not at pilot scale. One candidate was declined because that cost outweighed the hours it saved.

Declines in writing, with a way back

Both refusals were written down with their reasoning, not delivered verbally, and each states what would have to change for the answer to become yes.

Too many AI ideas and no way to tell which to fund first? Scope your build in 3 minutes.

Scope your build
Have a project?

Let's talk

Running a large platform, shaping a first MVP, or getting a product ready for a funding round? Tell us where you are. We'll shape the process around it, and stay with you after launch.