← Blog Posts
September 24, 2026 14 min read Salesforce

The Use Case Lives in the Overlap

Every discovery conversation I’ve been part of starts the same way. Someone has a goal that’s exactly right, and still too broad to build against. “Reduce cost and improve the customer experience.” It’s the right direction. It just isn’t a destination yet.

And you can’t build toward a direction. If I take that sentence at face value and start designing, I end up with an agent that does a lot of things adequately and doesn’t clearly move any single number. Those projects don’t fail loudly. They get quietly shut down at the next budget review, when someone asks what it actually did and nobody has a good answer.

So before I design anything, I try to find out two things. What are you hoping to accomplish, and what are we working with?

The real use cases live where those two overlap.

Two tests

Take every request, every idea, every “wouldn’t it be great if,” and run it through two questions. Both have to come back yes.

Is there a measurable KPI? Not “would this be nice.” Is there a number the business already tracks that this work can move, and could you credit the movement to it?

Is the data available to the agent? Does the information needed to do the job exist, and can the agent actually get to it?

Venn diagram: one circle asks whether there is a measurable KPI, the other whether the data is available to the agent. The overlap is the use case. Product troubleshooting sits in the KPI-only side, nice-to-haves with no number in the data-only side, and disputes and exceptions fall outside both as a judgment call.

If the data isn’t available, you can’t make it work. If the outcome isn’t measurable, you can’t prove it worked. Anything that passes both is a use case. Anything that fails one isn’t, at least not yet.

That sounds almost too simple to write down. In practice it’s the most useful filter I have, because most AI projects that stall have failed one of those two tests and it isn’t obvious from the outside which one.

Start with the number, and let them pick it

I come into the first conversation with a pretty good idea of which metrics matter for the kind of problem in front of me. I don’t lead with that list.

I want the KPIs to come from the sponsor. They know what their team is on the hook for, and a number they chose is a number they’ll defend later, when someone questions the spend. A scorecard I hand them has nobody defending it when budgets tighten. My list is there so I can recognize what they’re telling me, steer if they’re vague, and catch anything they leave out.

I also ask two levels, not one. The person who runs the service center knows what they’re measured on. The frontline manager knows what’s actually eating the day. Those are often different, and both answers matter. The leader gives you the goal. The frontline gives you a first read on where the volume really is.

The questions are plain:

  • What KPIs are you and your team measured against?
  • Which one needs to move, and what does it need to be?
  • What does your leadership judge you on? (Sometimes that’s different from what you personally care about, and the project needs to serve the one that gets it renewed.)
  • What have you already tried, and why didn’t it move the number?
  • What are your most time-consuming requests? Your most routine ones?

Then I push, but not on which metric. On making it concrete. “Improve CX” isn’t something I can build toward. “Satisfaction on these contact types from 78% to 85% by the end of the first phase” is. We leave that conversation with a target and a date, in writing. And I want at least one number stated in dollars, because that’s the one that travels up to their leadership.

There’s pushback in the other direction too. If the sponsor names company-wide NPS, I’ll say that’s influenced by ten things the agent will never touch. Measure satisfaction on the contacts it actually handles and keep NPS as context. If they name raw deflection with no quality check, I’ll pair it with repeat-contact rate. Deflection on its own just rewards the agent for dead-ending people, and nobody wants to explain that metric in a quarterly review.

Sometimes a sponsor won’t commit to a number at all. That’s a signal. Either there’s no baseline yet, in which case the first phase establishes one and the target gets set a couple of weeks later, or the outcome doesn’t have a clear owner yet, which is worth raising before anything gets built. Flexible on the timing, not on the principle.

Then look at what they actually have

The sponsor gives you the goal. The data gives you the use case.

The best first move I know is to pull six to twelve months of whatever records the work leaves behind (call transcripts, case notes, tickets) and cluster them to see what people actually contact the business about. The real distribution, not the case-reason picklist. The picklist is always stale, and it’s picked in a hurry at the end of a call. That’s no knock on anyone. It just wasn’t built to be an analytics source.

Then rank those clusters: volume, average handle time, how self-contained the request is, and whether the data to answer it already exists.

“How good is their data” is where discovery gets vague if you let it, so I make it specific:

  • Is there a key that exists in both systems, so a join is clean? (I’ve written about this one before. It’s usually the whole ballgame.)
  • Who owns the source, and how often does it update?
  • Is consent captured, and where?
  • Is the content authoritative? Policy documents that aren’t version-controlled are a risk to ground on, not an asset.

Naming the actual questions is what keeps discovery from staying vague. If the honest answers come back rough, that’s fixable, and it’s worth starting now.

One important clarification: available means available or connectable. Data sitting in a warehouse the agent can’t reach yet still passes the test. How much work it takes to reach is exactly what the roadmap is for. If you define “available” as “already wired up,” the method collapses to one use case and you’ve talked yourself out of everything interesting.

The ones that don’t make it

The use cases that fail are just as useful as the ones that pass, because they fail in different ways.

Take a servicing agent. Product troubleshooting sounds like an obvious fit, until you notice there’s no diagnostic data the agent can reach. That fails the data test, and it’s fixable. Connect the right system later and it may come back around.

Disputes and exceptions fail differently. The data’s all there. The problem is that the work is a judgment call, and that isn’t a technology gap. No amount of integration makes that one a fit. Disputes are the reminder that passing both tests is necessary, not sufficient.

The distinction matters because the first one belongs on the roadmap and the second one doesn’t. Miss either test and it’s not a use case yet. “Yet” is what the roadmap is for.

Where it usually lands

For a lot of service work, the answer ends up being less exciting than the kickoff deck, and that’s a good sign.

The first use case that survives both tests is usually the high-volume, self-contained “how do I” questions, answerable from content the business already owns: product guides, policy documents, help articles. No external integration, so the build is short. Low-risk enough to run on its own, because it’s answering from curated documents rather than changing a record, with a clean handoff to a person when it isn’t sure. And it maps straight to a number: contacts handled, times average handle time, times loaded cost per minute. That’s a dollar figure someone can check.

Account questions and request status usually come next, once the customer data is connected. Anything that writes to a record comes after that, and the gate there isn’t engineering, it’s trust. Each phase only happens because the last one hit its number.

That’s the other thing the two tests buy you. The outcomes over features argument gets a lot easier to make when every phase was picked because it could move something measurable with data you could actually reach. You’re not defending a feature list. You’re pointing at a number.

Turning the tests into a design

Passing both tests gets you a use case. It doesn’t get you a design. The bridge between the two is simple to say: every KPI the sponsor agreed to should map to a specific part of the build. If you can’t point at the piece that moves a number, that piece is decoration.

For a servicing agent, it usually maps like this:

  • The cost number comes from the agent handling routine questions on its own, which takes that handle time off the human queue. The math is direct enough that finance can check it without me in the room.
  • The satisfaction number comes from answers grounded in current policy documents and cited, so they can be checked. When retrieval comes back thin or the question falls outside scope, the agent routes to a person instead of guessing. And the handoff carries the context forward, so the customer doesn’t have to explain themselves twice.
  • The guardrails watch the failure modes the headline number hides. Repeat-contact rate catches the agent closing conversations it didn’t actually resolve.

Then there’s scope. Once customer data is connected, retrieval has to be filtered to the customer in front of the agent, so it can only pull that person’s records. That sounds obvious. It gets harder than people expect once many customers, locations, or brands share one index, and it’s much cheaper to design in than to bolt on.

And measurement doesn’t stop at launch. A test suite tells you an answer looked good. It doesn’t tell you the answer came from the right place, and a confident answer from the wrong document reads as a pass. So I keep a set of questions where I know which source holds the answer and check that it came back. In production, every conversation gets logged with what it retrieved, and an LLM reviews the transcripts continuously. That catches drift in days instead of at the next quarterly review.

Where Salesforce fits

I’ve kept this vendor-neutral on purpose. The two tests don’t care what platform you’re on. But on Salesforce, the pieces line up neatly.

“Connectable” is what Data 360 zero-copy is for. Data 360 can query data in place in a warehouse like Snowflake, and the agent reaches it from there, with no pipeline to build and babysit and no copy to keep in sync. (If you’re weighing federation against ingestion, I walk through the zero copy options and what each costs.) Grounding on policy and product content goes into an Agentforce Data Library, at least for content that doesn’t change every release. (For what happens under the hood, from chunking to hybrid search, see vector search, visualized.) The Einstein Trust Layer wraps all of it with zero data retention on the model side and an audit trail. And Testing Center lets you run thousands of phrasings through the agent before a customer ever sees it.

The platform shortens the distance between connectable and connected. It doesn’t change which use cases are worth building.

What it looked like in practice

Not every use case is a service desk. One of my favorite examples of the two tests was a job-matching engine for a nonprofit workforce partner that places young people into jobs.

The process was entirely manual. Counselors met with each young person, wrote down their interests and how they could get to work, then compared that by hand against lists from employers. It was slow, and it made mistakes, like sending two kids to the same opening. The workload had pushed matching into a quarterly cycle, so a young person could wait months to hear about a job.

Run that through the tests. The KPI already existed and mattered to everyone in the building: how long it takes to place a young person, and how many placements the team can make. The data was the interesting part. Interests and transportation were easy to capture. Employer addresses were a mess, and a lot of them couldn’t be turned into a reliable distance.

That’s where the data test got interesting: available doesn’t mean clean. Clean addresses went through a deterministic calculation that surfaced the closest jobs, no model involved. Messy ones didn’t fail and didn’t guess. The agent showed what it had and asked: “I don’t have exact mileage for this one, but it’s on 5th Street. Is that a distance you can manage?” The data didn’t have to be perfect. It had to be honest about where it wasn’t.

The design followed the number. Forms captured the simple, structured details, which kept token costs down on a nonprofit budget, and the agent’s reasoning was saved for the actual matching. And the agent never finalized anything. It created a proposed match, a counselor reviewed the reasoning, and only an approval sent the confirmations out.

Matching went from quarterly to continuous. Placements ran at roughly four times the previous pace with the same staff, and counselors got their paperwork time back for the mentoring they were actually there to do.

The honest footnote is the useful part. My first design was far more ambitious: highly autonomous, with confidence scoring to minimize how often a counselor had to step in. Stepping back, that wasn’t what they needed. Going from a paper process to any AI-assisted matching, with one person approving each match, was already a massive leap for that organization. The bigger system would have solved a problem they didn’t have, with cost and risk they couldn’t absorb. The right-sized version wasn’t a compromise. It was the win. Nonprofits have taught me that lesson more than once, and I wrote about it here.

Deployed isn’t done

The last test isn’t on the diagram, but it decides whether any of this sticks. Deployment isn’t the finish line. Adoption is.

The people working alongside the agent will see its weak spots before any dashboard does, so they need a simple way to flag them from day one. I track adoption as its own number next to containment: are customers choosing the agent, and are the people receiving its handoffs trusting them? And I roll out in waves. A pilot group first, who become the internal advocates, then broader enablement on what the agent does and doesn’t handle.

On the matching engine, the counselors had never used AI, and they were being asked to trust it with kids they cared about. The approval step was built as governance, but it doubled as the adoption plan. Nothing reached a young person without a counselor’s sign-off, so the agent’s job was to earn trust one reviewed match at a time.

The measure I care about isn’t usage. It’s confidence: whether the team believes the thing works.

The short version

Ask what they’re trying to accomplish. Ask what they’re working with. Build where those two overlap, starting with the smallest thing that moves the number.

It’s not complicated. It fits on a napkin, if you draw small. But when an AI project stalls and you trace it back, it nearly always skipped one of those two questions.