Since generative AI became a global phenomenon, almost every new project we have worked on has required some form of AI application.

For quite a while, our discussions kept returning to which model to choose and what hardware to buy to meet each project's requirements.

Then, after we demonstrated one project's results to a senior executive at a client organization, his reaction made me seriously reconsider the entire process:
Did this project really need AI, or did we all simply want to use it?

The executive didn't dismiss what we delivered. He knew the project had been completed as planned and that the deliverables met the requirements.

During the demonstration, the system successfully interpreted users' questions, queried a database containing years of historical data, and produced analytical results and charts.

The system was built around an open-source large language model deployed on-premises. To make it perform these tasks, the team had put considerable effort into integrating language understanding, data queries, and chart generation.

So when the entire demonstration went smoothly, without a single error, we felt proud and excited.

But the executive wanted to know: beyond making reports easier to produce, what additional benefits could the system deliver?

His point was that if the system could only organize historical data after someone asked a question, could it go a step further and help people identify important changes earlier, so they could act in time?

At that moment, we were not prepared to answer from that perspective.

After the demonstration, our project contact reassured us that the executive often had ideas beyond the original project scope and told us not to take it to heart.

Yet I kept thinking about what he had said. His expectations were not unreasonable: we had prepared a feature demonstration, while he wanted to see business value.

Letting users get analytical results through natural language can certainly create value.
But that value didn't resonate with the executive. Perhaps we had not clearly explained how the system helped people do their work. Or perhaps the AI feature didn't address the problem that mattered most to him.

Setting aside how the project had originally been approved and its requirements defined, I was more interested in another question: What could we learn from this experience?

Looking back, both our client contact and our team focused on building the features. We didn't ask enough questions about what work those features were meant to improve, or how we would demonstrate that improvement to the client.

How could we do better next time? How could we help clients direct their budgets toward problems where AI was a suitable tool and could deliver meaningful benefits?

The issues that repeatedly slowed us down during development might deserve more attention than the charts we had successfully produced during the demonstration.

Misinterpreted questions, data-entry errors, and changes in classification definitions over the years had all affected the reliability of query results.

That made me wonder: Could data quality and the way people used the data be the issues we should address first?

But if we immediately set our next goal as “using AI to clean historical data,” would we be repeating the same habit: choosing AI first, then finding a job for it?

Eventually, we brought our thinking back to the project's starting point: clarifying requirements and evaluating technology options.

Perhaps we should have started by identifying the pain points, understanding the expectations of different stakeholders, and only then evaluating which technologies could address those problems and meet those expectations. Otherwise, we risk shooting an arrow first and drawing the target around it, simply to give AI a place in the project.

Six Questions to Identify Pain Points, Clarify Expectations, and Assess Whether AI Is a Good Fit

Organizations naturally consider AI when planning new projects. As with earlier efforts to digitize operations and go paperless, we expect new tools to improve how we work.

But a paperless initiative still succeeds only if documents are easier to process, information moves more efficiently, and necessary records are properly maintained. Adopting AI requires us to ask the same question: what does it actually improve?

Before discussing models and hardware, use the following six questions to identify pain points, clarify role expectations, and make an initial assessment of whether AI is a good fit.

The examples below use a hypothetical scenario involving internal data queries and reporting. The examples and figures are illustrative; they do not represent actual results from the project described earlier.
  1. What specific business problem is this AI use case intended to solve?
  2. Who is affected by the problem, and in what situations does it occur?
  3. How much improvement would count as success? Are the success criteria and KPIs measurable?
  4. If we introduce AI, which errors are acceptable under specific conditions, and which are not?
  5. How can we translate those success criteria into verifiable acceptance criteria for an AI system?
  6. Compared with other viable options, is AI the best fit for this problem?
These questions are adapted from the Business & Use-Case Fit domain of the AI Project Health Check Assessment Framework I am developing. I rephrased them for early-stage discussions to clarify the specific problem we want AI to solve, the benefits we expect, and what would justify the investment.

For each question, I will explain its purpose, provide an example, and suggest which roles to involve in the discussion.

1. What specific business problem is this AI use case intended to solve?

“We want to build an AI query system” describes a proposed solution. We still need to ask: who struggles with the current way of querying data, and what difficulties do they face? What impact do those difficulties have, and why are they worth addressing?

For example, users might complain that “the reports are difficult to use.” Further discussion may reveal that standard reports actually run quickly. The real problem is that whenever a manager requests analysis with new criteria, someone from IT has to handle it separately. Users must wait, while IT staff are repeatedly interrupted by ad hoc requests.

The problem can then be stated more precisely: “Ad hoc analysis depends on IT staff, creating long wait times and consuming resources needed for maintenance and development.” This gives us a useful basis for comparing solutions.

For this question, we would usually interview business managers or process owners to understand the organizational impact and priority, then check what actually happens with frontline users and IT support staff. Recent examples, request records, and waiting-time data help ground the discussion in evidence.

We also need to establish who owns the business activity, who can decide whether the problem deserves priority, and who will recognize the improvement as a success. The project contact may not have the authority to make those judgments. A submitted requirement does not necessarily mean that everyone agrees on the importance of the problem.

2. Who is affected by the problem, and in what situations does it occur?

The person who requests a feature, the person who operates the system, and the person who uses its output may be three different people. Their priorities may differ as well.

In reporting, a manager may want timely information for decisions. The colleague preparing the report may care about access to data and clear definitions. IT staff may want to reduce repetitive work on ad hoc queries. If we interview only one group, we might improve a particular step without addressing the main obstacle in the overall workflow.

We therefore need to describe the situation clearly: who asks the question, and when? How often does it happen? How is the task completed today? Where does it get stuck? Once the result is available, who uses it, and for what purpose? A fixed statistical report run once a month may call for a very different solution from daily ad hoc queries with changing conditions.

This question calls for conversations with the people who operate the system, the managers who receive the results, and those involved in the preceding and following steps of the workflow. Asking a user to walk through a recent report, from the initial request to the final delivery, often reveals more than simply asking, “What features would you like us to add?”

3. How much improvement would count as success? Are the success criteria and KPIs measurable?

Once we understand the pain point, we need to define the improvement we want. “Make work more efficient” gives us a direction, but it is not enough to determine whether the project succeeds.

Suppose an ad hoc query currently takes two working days from the initial request to a result that has been checked and confirmed as usable. The team could discuss with the business unit whether reducing that time to half a working day for selected common query types would be a reasonable pilot target. That target should reflect current records, business needs, and feasibility, not just an attractive number.

The measurement should include time spent waiting, operating the system, checking results, and making corrections. If answers arrive quickly but users spend more time verifying them, the overall workflow may not have improved.

Typically, the business or process owner confirms the improvement target, actual users describe current performance, and the project manager helps define the measurement method. A manager authorized to approve the outcome should agree to it. Each metric should specify what is measured, the baseline, the target, the observation period, and who collects and verifies the data.

4. If we introduce AI, which errors are acceptable under specific conditions, and which are not?

This question asks how the system's output will be used, how serious an error could be, and whether it can be detected and corrected before it causes harm.

For example, awkward wording in a draft report that a colleague will revise may be acceptable as part of the editing process. But an incorrect reporting period, omitted records, or figures based on incompatible definitions could affect later judgments if they go straight into a formal report. Even within the same text output, different errors can have very different consequences.
In some situations, asking the user to clarify an ambiguous question is more appropriate than guessing an answer. We should also discuss whether a delayed response or referral to a person is acceptable, and how much time human review would require. These conditions help determine which tasks AI can take on.

This discussion should involve the people who use the results, the business owner, and domain experts, with relevant risk or management functions joining when needed. Technical staff can explain the system's capabilities and limitations. But people who understand the business impact and are authorized to take responsibility must agree on the acceptability of the consequences.

5. How can we translate those success criteria into verifiable acceptance criteria for an AI system?

Question 3 asks how much improvement would count as success. This question goes a step further: what must the AI system be able to do to support that improvement?

If we want users to obtain trustworthy analysis more quickly, “the system generates a chart after a question is entered” is not enough as an acceptance criterion. The date range, the records included in the analysis, and the counting and calculation rules must also be correct. Rephrasing a question without changing its meaning should not produce unexplained differences in the results.

We also need to agree in advance on how the system should respond when a question lacks information or contains ambiguous conditions. It might ask for clarification, explain that it cannot currently answer, or refer the question for human review. These responses are part of what makes a system usable in practice.

Question 4 establishes which errors are unacceptable. Question 5 should incorporate those boundaries into the acceptance criteria. If a particular statistic must not combine data from different reporting periods, the criteria should explicitly require the system to identify the correct period. If it cannot, it should seek clarification rather than guess.

These criteria must be specific enough to verify through actual cases. Next, plan the full testing process.

This discussion should involve business owners, domain experts, actual users, and representatives from the technical and testing teams. The business side explains what makes a result usable. Technical and testing staff check whether the conditions are clear and verifiable. The project manager helps everyone agree on a common basis for acceptance.

6. Compared with other viable options, is AI the best fit for this problem?

After working through the first five questions, we have a stronger basis for judging whether the improvement AI could deliver justifies the additional cost of implementation, verification, and maintenance.

If a standard report simply lacks a few filters, extending the existing report may be enough. If missing fields or formatting errors are the problem, we should also consider field validation and changes to the data-entry process. If users need to ask questions in many different ways and conventional interfaces struggle to support them, there is a stronger reason to test whether AI can make the task easier.

A combination of approaches may be appropriate, with AI handling only part of the work. For example, AI could interpret worded questions differently, existing query programs could handle calculations, and users could clarify unclear definitions. Fixed statistics could continue to use existing reports. If this division of work achieves the agreed improvement, there is no need to change every step simply to make the whole project feel more “AI-powered.” Choose each tool for the task it needs to perform.

This discussion should include business or process owners, IT architects and engineers, data owners, operations staff, and the decision-makers responsible for the budget. The comparison should return to the success criteria everyone agreed on earlier.

These six questions do not require six disconnected meetings. The project manager can first use individual interviews to gather needs, evidence, and differences in expectations, then bring the key stakeholders together to agree on improvement targets, boundaries for use, and how the results will be verified.

If people still have different understandings of what needs to improve, that is a finding worth addressing first. Resolving that gap gives subsequent technology choices a shared basis.

Conclusion

Looking back on that demonstration, I remain proud of what our team accomplished. But the executive’s question also reminded me that, even after delivering the promised features on time, we still need to explain what work the system has improved and what benefits it has brought to the organization.

These questions deserve a shared discussion among business leaders, end users, and the technical team before we choose models or purchase equipment. Whether the problem is clearly defined, expectations are aligned, and outcomes can be measured will all influence which solution we ultimately choose.

The six questions above may not immediately tell us whether to use AI. But they can help us distinguish between ideas supported by evidence, assumptions we have yet to test, and areas that need further clarification. Even if a project is already underway—or has been delivered—we can still use these questions to reassess where to invest next.

The answer might be to adopt AI, improve existing reports, organize the data, or adjust workflows. Sometimes, a combination of approaches will better meet the actual need. Any solution that addresses an important problem at a reasonable cost and provides evidence of improvement deserves consideration.

For me, the most important lesson from this experience is to ask one more question the next time someone says, “We want to launch an AI project”:
“Who do you most want it to help, and what problem should it solve? How much improvement would make the investment worthwhile?”
Define the problem first, then choose the technology.