AI Transformation · D365 · Technical

What I found when I built my own MCP server against D365

I built this to understand what it actually feels like to connect an AI agent to a live ERP and ask it to do something real. Here is what I found.

March 2026 · 8 min read
★ The APEX Build
June 2026 · D365 F&O + MCP Protocol
"Three data-layer failures. On Microsoft's own demo data, with every record clean, every field populated, every entity pre-configured."
TypeScript MCP Server 7 Tools Azure Container Apps 6/6 Tests Passing
7
tools built, exposing D365 F&O operations via OData
APEX build · 365Connect
6/6
Postman endpoint tests passing at project completion
APEX build · 365Connect
20%
response time improvement via $select query optimisation
APEX build · 365Connect
3
data-layer failures encountered on Microsoft's own pre-seeded clean data
APEX build · 365Connect - the important number
Why I Built It

In May 2026 I watched Patrick Mouwen, our Principal Architect at 365Connect, present a live demo of an AI agent taking real actions inside a D365 Commerce environment at DynamicsMinds in Slovenia.

Not a chatbot. Not a suggestion engine. An agent completing actual B2B commerce workflows - pricing, orders, master data - through a governed API layer our team had built.

The room was engaged. I was engaged. And I left with a question I could not answer from watching: what does it actually feel like to build one of these, and what do you find when you do?

I am a Programme Manager. I am not a developer. I have never written production code in a professional capacity. But I have spent eighteen years inside technology programmes, and I have learned that there is a category of understanding that only comes from doing - from sitting in front of a blank terminal and making something work, rather than reading about it or watching someone else.

So when I got home from Slovenia, I built one myself.

Not to publish a tutorial. Not to claim technical credit alongside the people who build these things professionally. To understand, from the inside, what the barriers actually are - because my job is to help organisations navigate those barriers, and I wanted to know what I was talking about.

What I Built

The project is called APEX. It is a TypeScript MCP server with seven tools, exposing Dynamics 365 Finance and Operations data through OData APIs.

It runs on Azure Container Apps in West Europe, authenticated via Azure Entra ID OAuth2 with two separate service principals - one for reading, one for writing - and it connects to Claude Desktop.

The seven tools cover the operations a finance or supply chain team would actually use: searching vendors, querying open purchase orders, checking inventory levels by site, pulling a financial summary from general ledger journals, surfacing overdue vendor invoices, retrieving customer balances, and creating draft purchase orders.

The constraint on that last one is deliberate: the server can create a draft purchase order, but it cannot submit one for approval. That was a governance decision, not a technical limitation. I will come back to why.

I built on Microsoft's Contoso USMF sandbox - the standard pre-seeded demo environment that comes with D365. Every vendor record is clean. Every price line exists. Every field is correctly populated. I chose it precisely because I wanted to start with the best possible data before introducing any real-world complexity.

It took a week.

TypeScript / Node.js v24 MCP SDK v1.29.0 D365 F&O v10.0.46 Azure Entra ID OAuth2 Claude Desktop
What I Actually Encountered

I want to be specific here, because the generalities about AI being "hard to implement" are not very useful.

What is useful is knowing exactly where it gets hard and why.

01

The field names were not what I expected

D365 Finance and Operations uses specific OData entity names and field names that do not always match what you would assume from the D365 user interface or documentation. The entity for inventory on-hand is not called InventoryOnHand. It is called InventoryOnHandForAI in the form I needed, and the field names within it do not match the labels visible in the D365 interface.

Every time I called an OData entity with the wrong name, the API returned a 400 error. The AI would attempt the query, get back an error, and have nothing useful to return. The AI was not wrong. My tool definitions were pointing at fields that did not exist.

I had to build a separate discovery utility - a script that queries the D365 OData metadata endpoint directly and returns the real entity and field names - before I could correctly define what the AI had access to. Without it I would have been guessing indefinitely.

This was on clean, fully populated, Microsoft-provided demo data. Every record correct. The problem was not data quality. It was the gap between what the documentation implied the fields were called and what they were actually called in the environment.

02

The data types were not what I assumed

The purchase order status field in D365 is not a string. It is an enum type. When I wrote a tool that filtered purchase orders by status using string logic, the API rejected the filter. The AI was applying logic that would be correct in most contexts. The D365 data model had a specific requirement that was not obvious from the field name alone.

The result was that queries which should have returned a filtered set of purchase orders returned nothing, or returned an error. Again: not an AI failure. A data model assumption that was wrong.

03

The data existed but was invisible in the wrong scope

When I built the tool to create a draft purchase order, the vendor could not be found. The vendor account I was querying existed in D365. But the cross-company OData query was returning results from a different legal entity - a different dataAreaId - and in that context, the vendor did not exist.

Adding the company scope explicitly to the query - specifying dataAreaId: 'usmf' in the POST body - resolved it immediately. The data was there the whole time. The context was wrong.

"Three data-layer failures. On Microsoft's own demo data, with every record clean, every field populated, every entity pre-configured."
What This Told Me About Readiness

Three data-layer failures. The failures were not about data quality in the conventional sense - dirty records, missing values, inconsistent entry.

They were about field naming conventions, data type specifics, and company scope logic that only became visible when something outside the system tried to act on the data.

A human user of D365 navigates these things invisibly. The interface handles the translation. The filter is applied through a dropdown, not a programmatic enum call. The company context is set by the session, not specified in a query body.

When the AI tries to do the same things, it has to specify everything explicitly, based on descriptions that someone has written. If those descriptions are incomplete, ambiguous, or wrong, the AI does not guess. It acts on what it has been told, and it returns whatever that produces.

Only around 7 percent of enterprises say their data is fully ready for AI. I read that figure before I built APEX. After building it, I understand what it means more concretely. The readiness question is not just whether your records are complete. It is whether the structure, the naming, the types, and the context of your data have been described accurately enough for a machine to act on them correctly - and in a live D365 estate with years of customisation, configuration changes, and process evolution behind it, that is a genuinely significant amount of work.

The Governance Decision I Had to Make

I want to return to the draft purchase order constraint, because it is the most interesting thing I built.

The tool creates a purchase order in draft status. It will not submit it for approval. That is not because the MCP protocol cannot do it, or because D365 does not have an API for workflow submission. It is because I decided it should not.

A draft purchase order is visible in D365. A human can review it, change it, approve it, or delete it. Nothing commits until a person looks at it. The AI's action is reversible.

A submitted purchase order has entered the approval workflow. It may trigger notifications, update demand planning, create downstream records. Unwinding a submitted order requires deliberate human action. The cost of a wrong submission is higher than the cost of a wrong draft.

When I was building the write governance for APEX, I had to think through what a non-human actor should be permitted to do in a live ERP system - not as a compliance exercise, but as a real design question. Where is the line between what the AI can do autonomously and what requires a human in the loop? What are the consequences if it does the wrong thing on either side of that line?

These are not questions that have obvious answers. They depend on the business process, the data quality, the audit requirements, and the organisation's tolerance for automated action at different stages of a workflow. The interesting thing is that most AI readiness conversations I observe do not get to these questions at all. They are focused on which model to use, which vendor to select, which use case to pilot.

"The governance layer - who controls what the AI can and cannot do - is treated as something to design later. It is not something you can design later. It has to exist before the AI acts, not after."
What the Build Produced

By the end of the week, six of six Postman endpoint tests were passing.

The AI was querying live D365 data and returning accurate results. The $select optimisation on the purchase order query reduced response time by 20 percent. The draft purchase order tool was creating correctly scoped records in the USMF entity.

The carousel below shows the actual output - real queries against a live D365 F&O environment, the responses as Claude returned them. The vendor list, the open purchase orders, the inventory levels, the financial summary, the overdue invoices, the customer balances, the draft PO creation. And the final slide: what happens when you ask the AI to submit the PO for approval.

The data in these outputs is Microsoft's demo data, not a real client estate. But the mechanism is real. The queries are real. The errors I encountered before getting to these outputs were real.

What I Took From It

I went into this build to understand the gap between watching an AI agent work and knowing what it actually requires.

I came out of it with something more specific than I expected.

The gap is not technical, in the sense that the tools are not difficult to use or the code is not achievable. The gap is descriptive. It is the work of accurately describing your systems - the entities, the fields, the types, the contexts, the rules, the constraints - in enough detail that a machine can act on them correctly and safely.

That work does not currently exist in most D365 estates. Not because anyone made a bad decision, but because it was never needed before. Humans can navigate D365 without it. Machines cannot.

The organisations that will get the most from agentic AI on their ERP estates are the ones that treat this descriptive work as a genuine investment - not a configuration task that happens alongside the AI deployment, but a prerequisite that happens before it. Field naming conventions documented. Data types mapped. Company scope logic explicit. Governance boundaries defined.

That is a different kind of readiness from what most AI readiness conversations address. And it is, in my experience of building APEX and of running delivery programmes for the past eighteen years, the part that gets skipped.

"The gap is not technical. It is descriptive. It is the work of accurately describing your systems in enough detail that a machine can act on them correctly and safely."

APEX Build Specifications
Component Detail
Server TypeScript / Node.js v24
MCP SDK @modelcontextprotocol/sdk v1.29.0
D365 environment Microsoft Dynamics 365 F&O v10.0.46 (PU70)
Legal entity USMF - Contoso Entertainment System USA
Tools 7 (read: 6, write: 1)
Authentication Azure Entra ID OAuth2 - two service principals (read/write separation)
Hosting Azure Container Apps - West Europe
AI interface Claude Desktop v1.5354.0 - custom connector
Write governance Draft-only - createDraftPurchaseOrder enforces Draft status, cannot submit or approve
Full case study Available on request
Related reading: What organisations get wrong about AI readiness, and what it actually takes →

Open to comparing notes on what
this actually involves

If you are thinking about what MCP or agentic AI might mean for your D365 estate, I am genuinely interested in the conversation.

These are early days and the most useful thing I have found is comparing notes with people who are working through the same questions.

I also run structured AI readiness assessments for D365 organisations. It is something I do, not something I am selling here.