A scored roadmap that picked the one AI project worth funding first
A use-case map and reference architecture that turned eleven competing AI ideas into a single ranked roadmap, and named the two not worth building at all.
Who, what, and how long
- Industry
- Professional Services
- Duration
- 6 weeks
- Cooperation model
- Fixed price
Client name withheld under NDA. Engagement details are shown to the extent our agreement permits.
The question we were asked
The firm's partners had eleven AI ideas floating between departments and no shared way to compare them. We ran a scored use-case map, designed the reference architecture the winning ideas would share, and handed back a sequenced roadmap with an honest verdict on which ideas to shelve.
A mid-sized professional services firm where six departments had each arrived at an AI idea independently, and the partnership had been discussing them across six months of meetings without converging. Nothing had been funded and nothing had been refused, which is the worst of both: the enthusiasm was real, and the comparison was impossible because no two ideas had been described in the same terms.
AI Strategy & Architecture
The decision, first
- 01
The most valuable output was the two ideas we argued against, not the one we ranked first.
Two documented refusals stopped spend that would have run for months before failing, which is worth more than accelerating the one idea that was always going to be funded.
- 02
Eleven ideas scored on one rubric ended a debate that six months of meetings hadn't.
The rubric ended the debate because it made the ideas comparable, not because it made them measurable. The six months of meetings had never had a shared axis.
- 03
A shared architecture is what makes a roadmap affordable. Without it, each project pays for its own foundation.
Four teams each standing up a document store, a retrieval layer, an evaluation harness and an access model is four times the foundation. One shared roadmap avoids it.
What the numbers couldn't answer
- 01
Every department had its own AI wishlist (a drafting assistant, a research tool, a billing automation), each pitched with enthusiasm but no shared way to compare cost, risk or payoff. Leadership couldn't tell which idea to fund first, or whether any were worth funding at all.
Each idea was pitched in the vocabulary of its own department: one as hours saved, one as risk reduced, one as a competitive necessity. There was no shared measure, so the argument was resolved by whoever presented most recently. The billing automation and the research tool had both been described as the obvious first project, by different partners, in the same quarter.
We interviewed each department, scored all eleven candidates on value and feasibility, and designed one reference architecture that the top-ranked ideas could share instead of each building its own stack. Two ideas were named unfit to build given data readiness and cost exposure.
How we worked it through
- 01
Interviewed six department heads to surface every candidate use case
Six department heads were interviewed about what they were trying to achieve, not what they wanted built, which is how eleven candidates surfaced instead of six.
- 02
Scored all eleven ideas on business value and technical feasibility
One rubric scored all eleven (value on hours and revenue, feasibility on data readiness, integration surface and regulatory load) in a joint session, out in the open.
- 03
Designed one shared reference architecture for the top four ideas
The top four turned out to need the same four layers, so those are specified once with named boundaries and each use case is an application on top.
- 04
Ran a data and cost risk review and named two ideas not worth building
The refusals were the hardest conversation of the engagement, and were delivered in writing so they could be revisited without being relitigated.
Phase by phase
Phase 1: Surface
Every candidate, on the table
Interviewed six department heads to collect every AI idea in circulation, including the ones that had never been formally proposed.
- Use-case inventory
- Interview notes
Phase 2: Score
Value against feasibility
Scored all eleven candidates on business value and technical feasibility using one rubric, so enthusiasm stopped being the ranking mechanism.
- Scoring rubric
- Ranked candidate list
Phase 3: Architect
One design for the winners
Designed a reference architecture the top four could share: retrieval, access control and evaluation built once instead of four times.
- Reference architecture
- Integration map
Phase 4: Decide
Sequence, and what to shelve
Ran a data-readiness and cost review, sequenced the roadmap, and documented why two ideas shouldn't be built.
- Sequenced roadmap
- Risk register
- Shelved-idea rationale
The use-case inventory: the four phases of the six-week engagement, the six department-head interviews, and all eleven candidates as first surfaced with their department, sponsor and source interview.
What it changed
11
Candidate use cases scored
2
Ideas shelved before spend
3 wks
Time to funded first build
Eleven candidates and two shelved are counts from the engagement itself. Time to funded first build is measured from the roadmap being presented to the first project being approved with a budget. That's the outcome the firm was actually buying: the six preceding months had produced no such decision.
Client name withheld under NDA. Figures are approximate, drawn from the engagement’s own reporting.
A comparable score
One rubric across value and feasibility, so eleven ideas pitched in eleven ways became directly comparable.
Eleven ideas arrived in eleven shapes (a slide, a spreadsheet, a paragraph in an email), which made the argument about which to fund unresolvable. Each was restated against one rubric: value scored on hours saved and revenue exposed, feasibility on data readiness, integration surface and regulatory load. Scores were set in a joint session with the sponsoring departments present, so the ranking was agreed, not handed down.
- Eleven proposals restated against a single rubric
- Value on hours and revenue; feasibility on data, integration, regulation
- Scored jointly with sponsors in the room
The scoring sheet: eleven candidates restated against one rubric, weighted across six value and feasibility criteria and sorted by total from 88 down to 24, each carrying a fund, backlog or not-building verdict.
The reference architecture: document store, retrieval layer, evaluation harness and access model specified once with named boundaries and a build-versus-buy line, the top four ideas on top, and the rules every funded build follows on those layers.
One shared architecture
A single reference design the top-ranked ideas all draw on, so four teams never stand up four stacks.
The top-ranked ideas turned out to need the same four things: a document store, a retrieval layer, an evaluation harness and an access model. Instead of letting four teams each stand one up, the reference design specifies those once with named boundaries, so a new use case is an application on top, not a fresh stack. It names the build-versus-buy line for each layer explicitly.
- Four shared layers specified once, not per team
- New use cases build on top, never beside
- Build-versus-buy line stated for every layer
A stated no
The risk register: all eleven ideas rated for data readiness, cost exposure and regulatory reach with a revisit condition on each, and the two written declines stating what would have to change.
Two ideas named unfit given data readiness and cost exposure, with the reasoning written down.
Two candidates were named unfit to build now, with the reasoning written down instead of delivered verbally. One depended on data captured inconsistently across offices; the other cost more per query at realistic volume than the hours it saved. Both entries state what would have to change for the answer to become yes, so the decision can be revisited without being relitigated.
- Two candidates declined, with the reasoning documented
- Declines grounded in data readiness and cost per query
- Each records what would have to change to reverse it
How the engagement ran
A cross-functional team of 4 worked on a fixed price basis over 6 weeks, covering Use-case discovery & scoring, Reference architecture, Risk & governance review. We ran two-week increments, each one shippable, reviewed with them before it merged. Decisions were recorded as they were made, so the reasoning survived the people who made it.
Six weeks, fixed price, ending in a document, not a system. Scoring happened in a joint session with the sponsoring departments present, which is the difference between a ranking that's agreed and one that's disputed the following week. The two refusals were written down with their reasoning, not delivered verbally.
The two ideas we declined were never the lowest scores
The eleven candidates ranked on the rubric the firm agreed, then on payoff alone and on feasibility alone. Watch where the two declines land, and why the verdict came from the risk review and not the total. Switch tabs, or use the arrow keys once one is focused.
Eleven candidates, ranked by weighted total
The weighting agreed with the sponsors in the room. The top three are funded at 88, 83 and 78. The two declines sit mid-table at #7 and #8: they were not refused for scoring low.
- #1#1 Proposal & engagement-letter draftingBusiness Dev.88 Fund
- #2#2 Billing narratives & WIP automationFinance & Billing83 Fund
- #3#3 Precedent research assistantKnowledge78 Fund
- #4#4 Tax memo first draftsTax70 Backlog
- #5#5 Audit sampling document reviewAudit & Assurance64 Backlog
- #6#6 Internal policy Q&AKnowledge58 Backlog
- #7#7 Always-on meeting summarizerAdvisory56 Not buildingCost per query at realistic volume outweighs the hours it saves.
- #8#8 Matter profitability forecasterFinance & Billing55 Not buildingThe time and matter data it depends on is captured inconsistently across offices.
- #9#9 Client intake & conflict triageAudit & Assurance52 Backlog
- #10#10 Timesheet capture nudgesTax44 Backlog
- #11#11 Pitch imagery generatorBusiness Dev.24 Backlog
Why two were declined anyway
The verdict came after a data and cost review, not from the total. Each decline hinged on one feasibility score of 1 out of 5:
- Matter profitability forecaster: data readiness 1/5
- Always-on meeting summarizer: cost at volume 1/5
Same scores, different weights
Every view uses the same 1–5 scores per criterion; only the weights change. Total = Σ weight × score ÷ 5.
The weighting the firm scored and decided on.
From eleven pitches to one funded build on one foundation
Six weeks, ending in a document, not a system. Each stage hands the next something written down, so the ranking and the refusals can be revisited without being relitigated.
- 01 · SourceSix department interviewsHeads asked what they were trying to achieve, not what they wanted built, so eleven candidates surfaced instead of six.
- 02 · ScoreOne rubric, joint sessionEvery idea restated as value (hours, revenue) against feasibility (data, integration, regulation), scored with sponsors present.
- 03 · DesignFour shared layersDocument store, retrieval, evaluation harness and access model specified once, with a build-versus-buy line per layer.
- 04 · RecordRisk register & written declinesData readiness and cost per query reviewed before spend; each decline states what would have to change.
- 05 · DeliverSequenced roadmapThe top-ranked ideas staged as applications on the shared layers; a first build funded three weeks after presentation.
Before any money moves
Checked for data, cost & a documented no
Data readiness checked before spend
Feasibility is scored on whether the data an idea depends on is actually captured consistently. One candidate was declined because its data is recorded differently across offices.
Cost per query at realistic volume
Cost is judged at the volume the firm would really run, not at pilot scale. One candidate was declined because that cost outweighed the hours it saved.
Declines in writing, with a way back
Both refusals were written down with their reasoning, not delivered verbally, and each states what would have to change for the answer to become yes.
Too many AI ideas and no way to tell which to fund first? Scope your build in 3 minutes.
Scope your buildNearby engagements
AI & AutomationA private legal assistant grounded in verified precedents
A private knowledge assistant that searches internal case files and precedents, providing cited answers legal teams can verify in seconds.
Legal & Law Firms · 14 weeks
Product DesignAn onboarding flow that guides trial users to value
A redesigned SaaS trial onboarding experience with progressive checklists, sample data, and inline guidance that turns signups into active subscribers.
Professional Services · 10 weeks
Product DesignA design system that brought speed and consistency to 4 product teams
A token-based design system in Figma and React that eliminated component duplication across 4 product squads and cut the time from design handoff to merged frontend.
Professional Services · 14 weeks
Let's talk
Running a large platform, shaping a first MVP, or getting a product ready for a funding round? Tell us where you are. We'll shape the process around it, and stay with you after launch.














