Before its AI could write the monthly report, Microsoft built an agent to fix the project data
Microsoft's infrastructure engineering team pairs AI reporting over Azure DevOps with a hygiene agent that flags gaps in planning records. The second one is why the first works.
Listen to this article · 5 min
AI-generated narration of the full article.

July’s leadership report arrived at the end of August. That was the old rhythm for Microsoft’s Infrastructure Engineering Services team, part of the company’s IT organization: weeks of manual work to turn a month of engineering scenarios into something executives could read. Now, says principal PM manager Martin O’Flaherty, “We can now send that report on the first day of August.”
The case study, published on October 8 by Microsoft Digital, is easy to read as a story about AI writing reports. The more useful story is in the second capability the team built, and in why the first one did not work without it.
The first attempt was inaccurate
The reporting capability reviews all scenarios completed in the previous month in Azure DevOps, groups them under strategic priorities and summarizes them into leadership reports and visualizations of engineering investment and outcomes.
Early output, Microsoft says, was sometimes inaccurate or incomplete. The team fixed three things: the reporting logic, the underlying planning data, and how the system weighs documented evidence of delivery. O’Flaherty’s summary is the line every reporting project should start with: “At the end of the day, the AI can only be as accurate as the data it has in front of it.”
The second agent fixes the source
So the team built a data hygiene agent. It continuously checks in-scope scenarios against a defined rule set, identifies gaps and prompts the responsible team members with specific guidance on how to fill them in Azure DevOps.
Three design choices make it work.
The agent asks; people fix. The hygiene agent does not write planning data itself. It tells the owner what is missing, and the owner corrects the record in the system of record. The source of truth stays owned by the people accountable for it, and a wrong guess by the agent cannot quietly become a fact.
Completeness is a rule set, not a judgment. “In scope” and “complete” are defined in advance. That is a data contract, checked continuously, and it is deterministic where it can be.
Evidence of delivery carries weight. Changing how the system weighs documented evidence is the same move we saw at Chatham Financial, which redesigned trade validation around evidence rather than around where an agent could sit. A report built on recorded evidence can be checked; one built on inferred progress cannot.
The two capabilities form a loop. Better data makes the reports better; the reporting effort keeps the pressure on the data.
Where to be careful
All the figures are anecdotal. Microsoft’s numbers are individual estimates from the team: a report on day one instead of weeks later, partner reporting that “could take hours daily” now generated on a schedule, prototypes in under an hour instead of a day or two. There are no baselines, sample sizes or accuracy rates. The models and the agent platform behind the two capabilities are not named; the tools mentioned include Azure DevOps, GitHub Copilot, MCP servers and Copilot Studio.
Less review needs a replacement. O’Flaherty says the reports now “don’t really require much human review anymore”, and that the latest one drew almost no feedback. Silence is not accuracy. When human review shrinks, something else has to detect drift: periodic sampling against the raw records, or a check that the totals in the report reconcile with the system. The question we raised in IBM’s study of human oversight applies: review is a control only if someone is still able and expected to challenge the output.
It transfers where a system of record exists. The pattern depends on planning data that lives in one structured system, with fields that can be checked by rule. Teams whose status lives in slides and chat threads would need to fix that first, and that is a project of its own.
The reporting agent is the visible result. The hygiene agent is the more transferable one: it makes every later use of the same data better, whichever tool or model reads it next.