AI Transformation · D365 Copilot for Finance

The governance that made it deliverable

Inside a D365 Copilot for Finance transformation that slipped, recovered, and went live on the revised plan with all three use cases together.

★ The Key Insight
D365 Copilot for Finance · live since September 2026
"The programme slipped by 69 days. Two change controls were raised. An exception report went to the Steering Committee. All of it was expected, documented, and resolved."
£94,500
total programme investment, fixed fee plus two approved change requests
Approved programme budget
90%
modelled three-year return on investment, to be measured at the 30-day and 90-day reviews
Benefits model, estimate
19 months
modelled payback period, conservative at 60% realisation
Benefits model, estimate
69 days
go-live delay, managed through a formal exception process
Exception Report EXC-001
The Brief

A European manufacturer approached us with a specific ask: deploy Microsoft Copilot for Finance, and prove it had delivered value.

They had tried this before. They enabled the AI features in D365 Finance and Operations, watched the toggle flip to active, and waited. Nothing meaningful happened. The reconciliation suggestions were generic. The collections recommendations did not reflect how their business actually operated. After a few months, the features sat unused. No ROI. No adoption. Time written off.

What they had learned, the hard way, is that enabling Copilot in the Feature Management workspace takes approximately five minutes. Running a programme that makes those features deliver real financial return is an entirely different undertaking.

Flipping the switch turns on the AI engine. It does not configure the risk thresholds that tell the model when to act and when to flag. It does not define who owns a recommendation, who can override it, or what happens when the output is wrong. It does not clean the customer master data the model will act on. Without those foundations, Copilot analyses what is available and generates output against whatever it finds. Your finance team is then left with recommendations that are either useless, or plausible enough to action without being trustworthy enough to rely on.

The scope we agreed covered three use cases: automated Bank Reconciliation, Collections Intelligence for their outstanding receivables portfolio, and Period-End Close Acceleration. All three were operating on live data, inside a production environment, with 3,000+ customers on the AR ledger. The margin for a generic AI recommendation acting on uncleaned data was not small.

This is not a story about the technology. The technology worked, when it had what it needed to work on. This is a story about building that foundation.

Bank Reconciliation Collections Intelligence Period-End Close Acceleration

This is not a story about the technology. The technology worked. This is a story about what it takes to run a programme like this without it unravelling when the inevitable happens.

Starting from Truth

Before we agreed on scope, timeline, or investment, we ran a structured AI Readiness Assessment.

This is something I do at the start of every engagement now, and the reason is straightforward: the gap between what a client believes their environment can support and what it actually can support is almost always there. The question is whether you find it before or after the programme budget is committed.

The assessment covered four areas: data quality and completeness, process design, governance readiness, and organisational change capacity. What it surfaced was not catastrophic. The bank reconciliation data was cleaner than expected, matching at 94.6% in test conditions. Collections master data carried a 12.4% inconsistency rate across active customer records that needed resolving before the AI could act on it reliably. Period-end data was the strongest area of the three and needed no material remediation.

The governance framework was essentially absent. All eight control areas we assessed came back red: there were no policies for how AI output would be reviewed, no defined ownership of AI decisions, and no documented thresholds for when a human should override the model.

That last finding shaped everything that came next.

Snapshot · AI Readiness Assessment (AIRAM-AI-001)
94.6%
automated bank match rate in testing, 847 transactions sampled
12.4%
of 2,341 active customer records held material inconsistencies
97.3%
GL transaction completeness across the last six close cycles
0 of 8
AI governance control areas in place at assessment
Assessment dimension Current Target Key action required
Data quality and completeness Amber Green Remediate collections master data in Phase 2
Process design and maturity Amber Green Redesign the collections workflow and document period-end
Governance readiness Red Green Build the full AI governance framework in Phase 3, blocking
Organisational change capacity Amber Green Training programme and communications plan in Phase 4
Overall readiness rating Amber Proceed subject to completing the governance dimension and the data remediation tasks before Copilot activation
The blocking finding

Governance readiness was the one red rating, and it was the reason the programme was phased the way it was. Not a parallel workstream. A prerequisite.

The Framework

Phase 3 was dedicated entirely to governance. Not a governance appendix. Not a slide in a steering pack.

A working governance framework that would be in place before a single Copilot use case went near a user.

This produced three documents that became the backbone of the programme. An AI Output Review Policy (POL-AI-001), which defined how outputs from each of the three Copilot capabilities were to be reviewed, escalated, and overridden. An Ownership RACI (RACI-AI-001), which assigned explicit accountability for every AI decision pathway: who owned the data, who owned the model output, who had authority to override, and who was responsible for exception management. And a Confidence Thresholds Document (THR-AI-001), which set the quantitative boundaries below which the AI would flag rather than act.

Snapshot · AI Governance Ownership RACI (RACI-AI-001)
Activity Prog. Manager Technical Lead D365 Consultant Finance Systems Mgr AP Lead CFO IT Manager
Bank reconciliation
Configure bank matching rules in D365 I A R C I I C
Review daily Copilot bank rec output I I C C R I I
Investigate low-confidence matches I I C A R I I
Monthly reconciliation sign-off I I I R C A I
Collections intelligence
Configure customer risk scoring parameters I A R C C I I
Review and approve or reject draft communications I I I A R I I
Escalate risk score above 9.0 to the CFO I I I R C A I
Period-end close
Review AI-generated variance commentary I I I A/R C C I
Final management pack sign-off I I I C I A/R I
Governance and oversight
Manage incidents and AI errors A R R C C I C
Audit log review, monthly I I I R C A C
R responsible · A accountable · C consulted · I informed · A/R accountable and responsible

The confidence threshold work was the part the client found most uncomfortable, in a productive way. Setting a threshold is an act of honesty about what you trust the AI to do unsupervised. For bank reconciliation, we set the auto-approve threshold at 95%: matches below that confidence level go to individual review rather than batch approval, and anything under 60% cannot be processed at all. For collections, customer risk is scored from zero to ten, with anything above 7.5 reviewed the same day and anything above 9.0 escalated to the CFO within two hours. Every collections communication the AI drafts is reviewed by a person before it is sent, at any score. For period-end, the AI drafts and nothing it drafts reaches the management pack without Finance sign-off.

Bank reconciliation
95% and above auto-queues. 80 to 94% is reviewed line by line. 60 to 79% opens an investigation record. Below 60% is manual.
Collections intelligence
Above 7.5 is same-day review. Above 9.0 escalates to the CFO within two hours. All drafts reviewed by a human before sending.
Period-end close
Mandatory review of all commentary. Material variance set at €5,000 or 5% of the budget line. No auto-posting at any confidence level.

These were not arbitrary numbers. They came from stakeholder workshops, analysis of the client's existing exception rates, the Phase 1 data quality baseline, and a frank conversation about what happens when an AI gets it wrong inside a live finance system.

The governance pack was signed off at CFO level before Phase 4 began. That sign-off mattered more than the documents themselves. It meant that when things went wrong, and things did go wrong, we had an agreed framework to operate within.

When It Slipped

The first slip happened in Phase 2.

The Chart of Accounts review that underpinned the period-end commentary use case was estimated at 580 accounts. When we got into the actual structure, it was 634 active accounts, with a further 112 inactive accounts that needed an archive or reactivate decision against the client's own policy. Not large numbers in absolute terms, but enough to push the remediation effort outside the original scope. Change Request CR001 was raised on 3 April 2026, approved by the sponsor on 8 April, and added £8,000 to the programme.

The second, larger slip happened in June 2026.

The bank feed integration, connecting the client's banking partners to D365 for reconciliation, had a complexity that had not been fully surfaced during discovery. The banks transmit transaction references in a format the standard Copilot matching engine does not parse: a local-language payment description prepended to the reference number. Confidence scores were coming back systematically below the thresholds we had configured. The work was solvable, and three successive mapping iterations moved the scores in the right direction, but it was not a two-week fix. The realistic timeline pushed our planned 3 July 2026 go-live by roughly ten weeks.

On 8 July 2026, I raised Exception Report EXC-001 to the Steering Committee. The report documented the cause of the delay, quantified the impact, and presented three recovery options.

Snapshot · Exception Report EXC-001, recovery options
Option Description Assessment
A. Split go-live Go live on 3 July with Collections and Period-End only. Bank Reconciliation follows as a separate go-live once the mapping is fixed. Not recommended Two sets of user training, two hypercare periods, and governance complexity. The finance team would experience a disjointed rollout.
B. Single deferred go-live Continue configuration, resolve the bank feed mapping, and go live once all three use cases pass UAT together. Recommended Delivers the full intended scope in one clean go-live, with a single training programme, one hypercare period, and a cleaner benefits baseline.
C. Descope Bank Reconciliation Remove Bank Reconciliation from programme scope entirely and proceed with two use cases. Not recommended Bank Reconciliation is the highest-value use case and the primary driver of the efficiency saving. Descoping would materially reduce programme benefits.

The Steering Committee accepted the exception and confirmed Option B. Change Request CR002 was then raised to formalise the commercial impact: £11,500, covering the additional configuration iterations, the extended Phase 4 delivery period, UAT execution inside that extended window, and the shifted thirty-day hypercare. It was approved on 3 September 2026 alongside the revised go-live date of 10 September.

CR001 · approved 8 April 2026
£8,000. Extended Chart of Accounts and vendor remediation scope. Phase 2 extended by seven calendar days. No impact on the go-live date.
CR002 · approved 3 September 2026
£11,500. Extended Phase 4 delivery. Go-live moved from 3 July to 10 September 2026, a 69-day shift, with all phase durations unchanged.
The Recovery

What I want to note about that Steering Committee meeting is not the outcome. It is the quality of the conversation.

Because the exception process had been followed correctly, with a formal report, documented options, quantified impacts, and a clear recommendation, the committee was able to make a real decision rather than a reactive one. The question on the table was not "what do we do about this crisis." It was "which of these three options best serves the programme's objectives." That is a different conversation to be in.

The revised timeline came at a cost, and I want to be straight about which one. UAT had been planned as a three-week window from 15 June. To avoid pushing go-live any further, we compressed it into eight days, from 1 to 8 September. What we did not compress was the coverage: thirty-five test cases across all three use cases, executed by the client's own finance team, with defect triage and resolution running in parallel through the same week.

Snapshot · UAT execution summary (UAT-TP-001)
Use case Test cases Passed Failed Pass rate Defects raised
Bank Reconciliation 15 15 0 100% 4
Collections Intelligence 10 10 0 100% 3
Period-End Close 10 10 0 100% 3
Total 35 35 0 100% 10

Ten defects were raised during UAT, catalogued as D-001 through D-010 in the defect log: three high severity, four medium, and three low. None were critical path blockers. The three high-severity items were the ones worth having found: foreign currency transactions were not enforcing the mandatory individual review the governance framework required, collections risk scores were not refreshing after a payment was posted, and the period-end commentary was reversing the sign on variances, showing favourable as adverse. All ten were resolved and re-tested before the sign-off meeting on 8 September 2026.

High severity
3 raised. Governance rule not enforced, score refresh lag, and sign-reversed variance commentary. All fixed and re-tested.
Medium severity
4 raised. Approval notifications, a hardcoded signatory, a checklist that did not auto-complete, and an over-permissive user role.
Low severity
3 raised. Date format, a template typo, and dashboard filters that did not persist between sessions.
"Ten defects in UAT is not a warning sign. Ten defects in UAT that are all resolved before go-live is exactly what a working QA process looks like."
Where It Stands

The programme went live on 10 September 2026, and it went the way we had planned for.

All three use cases were switched on together, as agreed in the recovery plan. The go/no-go decision was taken against a completed readiness checklist rather than a general sense that we were ready. There were no critical or high-severity incidents at go-live, no rollback, and no change to the confidence thresholds we had set in Phase 3. The finance team started working with Copilot outputs on the first business day.

We are currently in hypercare, which runs for thirty days to 10 October 2026. This is the period where I want monitoring on high alert, because the questions that matter now are not whether the system works. They are whether the thresholds hold up against live volumes, and whether the team trusts what they are seeing.

One issue has been raised so far: a P3 medium, picked up by the AP Lead during the daily review rather than by a customer or an auditor, which is exactly the direction you want these to travel. It was triaged the same morning and resolved inside its three-day target, with no financial impact and no threshold adjustment required. Nothing has been escalated above that level.

The benefits case is built on a deliberately conservative model. We applied a 60% realisation rate across all projected savings, meaning the baseline assumes we capture six-tenths of the theoretical value, not the full figure. That conservatism is intentional: it builds in room for the learning curve, the edge cases the AI has not seen yet, and the natural friction of any new working pattern. The figures below are a model, not a result. They will be measured properly at the 30-day review and again at 90 days.

Snapshot · Return on investment model, estimated
Programme fixed fee £75,000
Change Request CR001, extended data remediation scope £8,000
Change Request CR002, extended Phase 4 delivery £11,500
Total programme investment £94,500
Modelled annualised benefit, 60% realisation rate applied £60,000 est.
Year 1 net position, savings less the investment (£34,500) est.
Year 2 net benefit £60,000 est.
Year 3 net benefit £60,000 est.
Three-year net benefit £85,500 est.
Three-year return on investment 90% est.
Payback period 19 months est.

On that model, a return of around 90% over three years on a £94,500 investment, with payback somewhere near nineteen months, the case holds. Whether it holds at those exact numbers is a question for the 30-day and 90-day reviews, not for this article.

The programme did not go to plan. It went to a better plan, because the governance framework gave us the mechanism to find a better plan when the original one needed to change.

"The governance framework did not prevent the programme from slipping. It prevented the slip from becoming a crisis."
What This Means

I get asked regularly what the single most important thing is to get right before deploying Copilot for Finance.

Or any enterprise AI capability on a D365 estate. The answer is always the same: governance architecture.

Not because it is the most technically interesting work, because it is not. Not because the AI tools require it, because they do not, they will operate without it. But because everything that can go wrong with an enterprise AI programme goes less wrong when the governance is clear before anyone starts. The change controls are cleaner. The exception process has a framework to operate within. The Steering Committee has a basis for decision-making that is not just instinct and cost. The users know what the AI is allowed to do, and what a human is expected to do instead.

In this programme, the governance framework we built in Phase 3 did not prevent the programme from slipping. It prevented the slip from becoming a crisis. When the bank feed complexity surfaced, we had a formal exception process to run it through. When the Steering Committee had to choose between three recovery options, they had documented criteria to evaluate them against. When UAT found ten defects, we had a defect log with agreed severity classifications and a resolution process that was already in place.

None of that is glamorous. All of it is the difference between a programme that delivers and one that does not.

Related reading

Why ERP programmes keep underdelivering, and what that means for your AI ambitions →

Running a D365 AI programme?
Worth comparing notes.

AI transformation in a D365 environment is a delivery question before it is a technology question.

If you are working through the governance, readiness, or delivery design for a Copilot programme and want to compare notes on what we found, I would genuinely welcome it.

I also run structured AI readiness assessments for D365 organisations before programmes begin. It is something I do, not something I am selling here.