Back home

About

I lead Copilot Mobile product work and build AI development systems hands-on.

I am Daniel Chu, a Partner Group Product Manager at Microsoft leading M365 Copilot Mobile product work. I also build multi-agent development, eval, review, and recovery systems, and use Atlas as a personal lab for studying agent reliability.

A warm, messy computer desktop with AI workflow sketches and product notes. Daniel Chu
Current focus Agent development, eval, engineering review, side effects, and recovery.

Product approach

How I turn model capability into dependable products

I stay close to both product direction and the system itself, especially across agent workflows, evals, review, reliability, and recovery.

  • Agent products Ground model capability in real workflows with tools, state, permissions, review, and recovery.
  • Evals and reliability Use workflow-level evals, failure replay, and quality gates as product infrastructure.
  • Research to product Translate capabilities and limits across research, engineering, users, and launch decisions.
An open notebook with AI product diagrams, clipped notes, color swatches, and a pencil.
How the notes are made The notes draw from public-safe project artifacts, workflow traces, and lessons from failures.

What I am usually trying to make visible

The parts of AI products that interest me most are rarely the demo moment. They are the follow-through: which files changed, which tool was trusted, what review needs to happen, what happens if the system is wrong, and how a person can recover without starting over.

That is why a lot of my writing circles around agents, evals, workflow design, and product judgment. A useful AI system should not only sound smart in the chat window. It should make the work easier to inspect, improve, and trust.

Career arc

I came to this work through product leadership, startups, and international teams. My background includes Booth, a GNVC Asia win, and work across North America and China. That is context for my career path; the systems and public artifacts above are the evidence for my current AI work.

How I work

I like building close to the actual workflow. If the product changes files, the eval should care about files. If the system can write data, the interface should include undo, audit, or repair. If more people can create changes, the review path has to become part of the product.

My current focus is practical: constrain agents where repeatability matters, keep human review for ambiguous or high-consequence work, and make failures inspectable so a person can recover.

Elsewhere