E-book
Before you commit: Evaluating AI in indirect tax
- Introduction
- 1. The evaluation gap
- 2. Why evaluation in indirect tax is different
- 3. Three triggers driving AI evaluation
- 4. How effective evaluations are structured
- 5. Six questions that anchor evaluation
- 6. The stakeholders that determine the outcome
- 7. Signals of readiness and risk
- 8. Conclusion: This isn’t a technology decision
The AI evaluation problem in indirect tax
Most AI vendors in indirect tax make similar promises of automation, accuracy, and scalability.
The claims, and the demos, are nearly indistinguishable, especially under controlled conditions.
In indirect tax, the gap between evaluation and real-world performance matters. These systems operate across jurisdictions, depend on fragmented data, and support compliance workflows. Traditional frameworks assume stable, predictable conditions. AI systems in indirect tax don’t.
The cost of getting evaluation wrong isn’t just poor fit; it’s compliance liability, audit exposure, and the difficulty of unwinding an embedded system.
Demos mask variability. Comparisons miss risk. And single-function decisions break down when broader stakeholders engage.
Addressing that requires a different approach: Effective AI evaluations start from operating constraints, testing under pressure, and aligning stakeholders before committing.
The result is a solution that holds up in production.
Chapter One
The evaluation gap
Every vendor claims AI leadership, and a consistent pattern emerges: Solutions are difficult to distinguish based on what vendors demonstrate.
In indirect tax, that creates a practical problem. Tax leaders are making multi-year platform decisions while ERP environments evolve and regulatory requirements expand.
The result is a widening gap between urgency and readiness. Organizations are moving forward without fully understanding how these systems will behave.
This gap exists because evaluation approaches test systems under stable, controlled conditions. Organizations rely on frameworks that compare features well but assess risk poorly under uncertainty.
Chapter Two
Why evaluation in indirect tax is different
Indirect tax isn’t an environment where “mostly correct” is acceptable.
It operates across jurisdictions with evolving rules. When content is incomplete or outdated, AI systems can produce confident but incorrect outputs, particularly across country, state, and local variations. Accuracy depends on consistently applied domain content.
In most enterprise systems, errors create inefficiencies. In tax, they create liability.
Why AI changes the risk profile
AI changes how risk is introduced. Traditional systems operate deterministically. AI systems generate outputs based on patterns rather than fixed rules. That introduces a different failure mode — outputs can be well-formed and persuasive, but incorrect.
Organizations often evaluate these systems as assistive tools, though they can perform work typically handled by junior team members.
That shift means tax leaders can’t evaluate AI the same way they evaluate traditional software. It’s no longer enough to assess capability alone. Evaluation must account for how work is performed, validated, and governed.
Where evaluation breaks down
Evaluation often begins within the tax function. While appropriate, this approach can underweight downstream requirements such as integration, data ownership, and workflow dependencies until later in the process.
Effective oversight requires transparency and human-in-the-loop controls. If outputs can’t be traced or validated, they introduce risk rather than remove it. In indirect tax, that lack of visibility constrains use.
“A black box solution is a liability,” says Mike Lucich, Indirect Tax Product Sales Specialist with Thomson Reuters. Systems must support intervention, not obscure it.
That risk doesn’t surface in demos. Instead, it appears in production when systems encounter incomplete data, conflicting inputs, and real-world variability.
Systems that perform well in selection processes can still encounter friction after implementation. At that point, correction costs increase significantly. Replacing a poorly aligned solution after integration, process change, and user adoption can be more disruptive than investing in more rigorous evaluation upfront.
The question is not whether AI can execute the work, but whether the system enables that work to be governed.
Chapter Three
Three triggers driving AI evaluation
Organizations are evaluating AI under increasing pressure as existing tax technology falls short. Recent Thomson Reuters research shows dissatisfaction rising sharply: 56% of professionals report being unsatisfied in 2025, up from 34% in 2024.
That pressure is driven by three primary triggers:
Catalyst 1: Regulatory pressure
What it looks like
Pressure to keep pace with expanding compliance requirements, jurisdictional complexity, and audit exposure.
How it skews evaluation
Compliance pressure often prioritizes coverage and speed over governance, validation, and integration readiness.
If this is your trigger
Focus on evaluating how the system handles exceptions, auditability, and regulatory change under real conditions.
Catalyst 2: Technology transformation
What it looks like
AI evaluation alongside ERP modernization, data migration, or architecture changes.
How it skews evaluation
Transformation-driven decisions often prioritize architectural fit and future-state alignment, while underweighting operational realities and workflow impact.
If this is your trigger
Test how the system performs within unstable or evolving environments, not just how it integrates on paper.
Catalyst 3: Efficiency pressure
What it looks like
Pressure to scale output without scaling headcount or operational cost.
How it skews evaluation
Efficiency-driven decisions often prioritize automation claims and productivity gains over transparency, control, and long-term validation.
If this is your trigger
Evaluate how work is governed, not just how much work is automated.
Across all three triggers, organizations often prioritize what vendors demonstrate most easily, not what matters most in production. Teams may overlook difficult-to-validate factors such as data handling, exception management, and long-term performance.
When tax teams lead evaluation without early input from finance and IT, organizations often revisit assumptions later in the process, slowing alignment and implementation planning.
Chapter Four
How effective evaluations are structured
Organizations that evaluate these solutions successfully start by defining constraints. Before engaging vendors, they clarify:
- Where human oversight must remain
- What data conditions must be tolerated
- What integration dependencies cannot change
These aren’t preferences. They’re operating realities.
Evaluation depends on how clearly vendors explain limitations. Solutions that acknowledge limitations and trade-offs reflect a clearer understanding of how they operate.
Evaluation shifts from comparing capabilities to testing how a system performs under real conditions.
Chapter Five
Six questions that anchor evaluation
Effective AI evaluations converge around questions that focus on risk.
- What domain expertise is embedded in the system, and how is it maintained?
Why it matters — Accuracy in indirect tax depends on current, complete domain content applied consistently across jurisdictions.
Weak answer — Vague references to “AI learning” or generic datasets, with little clarity on content provenance, update frequency, or jurisdictional coverage.
Strong answer — Clearly defined content sources, frequent updates, and transparent processes for maintaining regulatory accuracy across jurisdictions.
- How does the system handle real data conditions?
Why it matters — Indirect tax data is rarely clean. Systems must operate with incomplete and inconsistent inputs.
Weak answer — Relies on idealized scenarios or curated datasets that don’t reflect real operating conditions.
Strong answer — Demonstrates performance using real-world scenarios.
- Where does human review occur, and how is it structured?
Why it matters — Automation doesn’t remove responsibility. In regulated environments, accountability must remain visible and controlled.
Weak answer — Claims of fully autonomous operation or vague references to oversight without defined review points or workflows.
Strong answer — Explicit human-in-the-loop control points, with clear workflows for review, validation, and exception handling aligned to compliance requirements.
- How does the system integrate into existing workflows?
Why it matters — Indirect tax processes depend on data moving across multiple systems. Integration must reflect real workflows.
Weak answer — Focuses on theoretical integration or isolated system performance without validating end-to-end data movement.
Strong answer — Demonstrates how data flows across ERP, procurement, and finance systems in real conditions.
- How is performance validated and maintained over time?
Why it matters — AI systems evolve. Without ongoing validation, performance can drift over time.
Weak answer — Assumes performance at implementation persists indefinitely, with limited visibility into monitoring, validation, or correction mechanisms.
Strong answer — Clear processes for monitoring performance, detecting drift, validating outputs, and applying corrections, supported by auditability and traceability.
- How does the system protect data, and how is learning governed?
Why it matters — Indirect tax workflows contain sensitive operational and strategic information. Organizations need clarity on how AI systems use customer data and whether it informs outputs for other users.
Weak answer — Vague assurances about “enterprise security” without clear explanation of tenant isolation, model-training boundaries, or data usage.
Strong answer — Clear separation between tenant-level learning and model-level training, with explicit confirmation that customer data remains isolated.
These questions focus evaluation on what matters in production, from domain expertise to data protection and long-term validation.
Chapter Six
The stakeholders that determine the outcome
One of the most consistent patterns in indirect tax AI evaluations is this — deals rarely stall because of product fit alone. They stall because stakeholders are not aligned.
In most organizations, tax teams lead the initial evaluation, but finance and IT ultimately shape the decision, and each evaluates the solution through a different lens. When those perspectives emerge later, earlier assumptions are challenged, slowing progress.
Each stakeholder group is looking for something different:
- CFO or finance leadership
Focus — Cost, risk, and return on investment
What they need to see — A clear business case, including how the solution reduces exposure, improves efficiency, and operates as a predictable cost of compliance
- IT or CIO leadership
Focus — Integration, data governance, security, and architectural fit
What they need to see — How the solution handles data across systems, maintains control, and avoids introducing opaque or ungovernable processes
- Tax leadership
Focus — Accuracy, auditability, and operational burden
What they need to see — Embedded domain expertise, transparent logic, and the ability to retain control through human review
Alignment doesn’t require agreement upfront. But without it, even well-structured evaluations can fail, because decisions are made across the organization, not within a single function.
Indirect tax isn’t a point solution, it’s a process
Indirect tax operates less like a standalone function and more like a conveyor belt.
Data moves from upstream systems through determination, validation, and reporting. Each stage depends on the accuracy of what came before it.
On a conveyor belt, defects propagate. If one stage is misaligned, the error is carried through the entire process, affecting downstream outputs, increasing exceptions, and introducing risk.
This is why evaluation can’t be done in isolation.
A solution may perform well at one point in the process. But if it can’t operate consistently across the full workflow — integrating with upstream data and supporting downstream validation — it introduces friction rather than removing it.
Evaluation must reflect the full process, not just whether a solution works, but how it performs across the workflow. That means evaluating domain expertise, data handling, human review, integration, and ongoing validation.
Chapter Seven
Signals of readiness and risk
As organizations apply a structured evaluation framework, consistent patterns emerge. The distinction is not between “good” and “bad” solutions, but between those that are ready for real-world deployment and those that aren’t.
These signals help distinguish between them:
| Red Flags | Reassuring Signs |
|---|---|
| Outputs can’t be explained or traced | Clear, auditable decision-making at each step |
| Heavy reliance on generic AI models | Domain-trained systems grounded in tax expertise |
| Performance demonstrated only in controlled scenarios | Validation under imperfect, real-world conditions |
| Claims of fully autonomous operation | Explicit human-in-the-loop control points |
| Limited visibility into error detection and remediation | Transparent handling of exceptions and corrections |
| High-level claims not reflected in technical or operational detail | Consistent explanation across conceptual, technical, and operational levels |
Many solutions perform well under ideal conditions. What matters is how predictably systems perform when conditions are not.
Chapter Eight
Conclusion: This isn’t a technology decision
Adopting AI in indirect tax is not a race to implement advanced systems. It’s a question of readiness; whether the organization is prepared to operate these systems under real conditions.
That readiness isn’t just technical. It requires rigorous evaluation before committing.
The cost of that discipline is modest compared to the cost of correction. Once a system is embedded into workflows and tied to compliance outcomes, the effort required to unwind that decision increases significantly.
This is not a question of timing, but of structure.
By grounding evaluation in real operating conditions, stabilizing criteria early, and aligning stakeholders around what matters in production, organizations can move forward with confidence.
Because in indirect tax, the quality of the decision determines everything that follows.
About Thomson Reuters ONESOURCE
Thomson Reuters ONESOURCE delivers AI-powered indirect tax compliance solutions designed around the evaluation principles outlined in this framework. With decades of domain expertise, regulatory content spanning 190+ countries, and deep integration capabilities across leading ERP platforms, ONESOURCE solutions are built to operate under real-world data conditions while maintaining the transparency and control tax professionals require.
To explore how ONESOURCE addresses the evaluation criteria in this framework, from domain expertise and data handling to human review controls and ongoing validation, learn more about ONESOURCE or contact a Thomson Reuters representative.
Indirect compliance software
The manual era of indirect tax is over
Powered by CoCounsel, ONESOURCE helps global tax teams prepare faster with fewer errors and complete audit confidence across every jurisdiction