01 / 01
My work / Custom Agents
Client: M-Files
PROCESS AGENTIC AI TRUST BY DESIGN

Custom Agents: the right step to hand to AI.

The first M-Files feature where AI does the work rather than suggests it — an agent that finishes a slow, judgement-heavy step on its own, without asking the person anything. Public beta on M-Files 26.6. I owned the design end to end: which steps were safe to hand over completely, and how to make work nobody watched happen still worth trusting.

GoalLet AI finish a slow, judgement-heavy step on its own, without losing the trust people already have in Workflow.

Google PAIR framework Double Diamond Spec-driven design
Scope
1st autonomous
capability shipped inside M-Files Workflow — public beta, Cloud-hosted vaults only.
What I owned
  • Judging whether AI was even the right fit for the task
  • Choosing which type of AI to design for — and building the trust it needs
  • Admin configuration UI — legacy platform, out of my scope
The real question was never whether AI could do it — it was whether it could be trusted to.
Overview
A workflow in M-Files is the path a document travels — a contract moving from Draft to Legal review to Approved, with rules deciding what happens at each step. Custom Agents adds one thing to that: a step the AI completes on its own.
The product's first autonomous AI capability — public beta, 26.6. An object reaches a configured state, the agent reads only what the admin gave it, and writes structured output back. Nobody is asked anything.
Every value it sets carries its reasoning, one click away, kept for as long as the object exists. That one detail does most of the work of making it trustworthy.
Role
As a senior product designer in an AI era, there was no playbook for this. Designing an autonomous feature meant learning the craft while doing it — reading how the field was answering these questions, borrowing what held up, and testing the rest against a real product.
Designing for AI turned out to be a different job than designing a screen. The material is non-deterministic — the same input can produce a different answer twice — so you can't specify an outcome, only the conditions around it: what the agent may touch, what it must reveal, and where a person still decides. Most of my work here was on those conditions, not on the interface.

Business goal

1

Put AI into Workflow — and be trusted for it

The business wanted AI inside Workflow, the most used feature in the platform. With one condition: it had to earn people’s trust. The market is full of AI bolted onto products that nobody relies on twice, and touching the feature customers depend on every day meant we had one chance to be believed.

2

Ship fast enough to learn from real use

A better setup screen a year from now wouldn’t have taught us what customers actually need. Real use — even through a rough legacy setup — was worth more than waiting on a platform migration that’s already slipped five years.

WHAT IT CHANGES, IN ONE PICTURE

A contract approval, with one step handed to an agent.

Initiatecontract Legalapproval Human still in the loop Final signatory Contract reviewagent PLUGGED INTO THE STEP
BeforeSomeone read the contract end to end, checked it against the compliance checklist, and wrote up what they found — then the approvals began.
AfterThe step still exists and still belongs to Legal. An agent is plugged into it, and the moment a contract arrives it reads, checks, and writes pass or fail with a summary onto the contract object in M-Files.
UnchangedNo approver was removed. Legal still approves — now starting from findings instead of from page one.
01
Discover stage
DISCOVER DEFINE DEVELOP DELIVER
Double Diamond
Design Framework
01Discover stageexplore the problem widely

Which steps in a workflow eat the most time today — and could AI do them instead?

Two questions, answered in that order. I used Google PAIR’s Evidence of User Need activity to keep them apart and to write down the reasoning behind each.

Lightning talks

Short share-outs across the trio, each bringing the part we owned. The architect walked where Workflow stops: a rule can check whether a value passes a threshold, but not whether a contract says what it should. Anything needing a document read and interpreted falls to a person. AI engineers were in the room from day one — feasibility was the open question, not a later check.

Evidence of User Need

A cross-functional session on the PAIR worksheets: aggregate the research, write one need statement, run it through the AI suitability checklist — both columns, including the reasons AI is the wrong tool — and record the verdict.

"How might we" + affinity mapping

Grouped the evidence until the same few jobs kept showing up, and wrote down what we would not solve. Those exclusions became limits we built in on purpose — not things we ran out of time for.

WORKSHEET 01 — USER RESEARCH SUMMARY

What we already knew, gathered in one place before anyone argued.

SOURCE
SUMMARY OF FINDINGS
OngoingAggregated customer feedbackvia product management
Customers kept asking for smarter automation, without being asked. The example was always the same kind of job: reading a 40-page contract to check a clause is there. People do the first one carefully and the fortieth badly.
ExistingCustom code already in customer vaultsVAF — the Vault Application Framework, M-Files’ way of extending the product with custom code
The strongest evidence we had: customers were already paying developers to automate these steps. Slow to iterate, expensive to maintain, dependent on a technical team the process owner doesn’t control. People clearly wanted this. How they had to get it was the problem.
ExistingLive workflow configurations
Where documents sit and wait. Not because anyone is confused about the process — because the next decision needs someone to read something and think about it.
DiscoveryProcess walkthroughsquality, contract, project managers
Described first-hand: checking inspection results against a spec, reviewing contract compliance before approval, pulling action items out of meeting notes. Each is clearly defined and repeatable — and each was described as tedious, not difficult.
Apr 2026Market scan
Only 14% of organizations report high confidence their content is AI-ready (Gartner). The market was moving quickly toward agentic AI — which made the need urgent, though it was never the reason for it.

Aggregated by product management ahead of the session, per the activity’s instruction — existing evidence, not new research commissioned to justify a decision already made.

Across every source, the same job kept coming back in five forms.

01

Review & validate. Check a contract against compliance requirements; record pass/fail with a review summary.

02

Compare against a specification. Check inspection measurements against tolerances; calculate an accuracy score.

03

Pre-screen a request. Evaluate an access request against business rules and the requester’s role; return it for clarification if the reasoning is thin.

04

Create objects from content. Extract action items from meeting notes; create assignments with owners and deadlines.

05

Initialize structured data. Read a project plan; create the milestone, team, and task objects that populate the workspace.

WHAT THE EVIDENCE SAID

There was no middle ground between “assign it to a person” and “build custom code.”

Five sources, one gap. Which meant the problem wasn’t that these steps were too hard to automate — it was that automating them cost more than doing them by hand. That is what this feature exists to change — a problem customers were stuck with long before AI could do anything about it.

Three needs, distilled from the evidence.

1

Automate the heavy thinking, not just the checking

The part that stays manual isn’t the comparison — it’s holding forty pages in your head long enough to work out what they mean against a set of criteria. Automate only the checking and that load stays exactly where it was.

2

Set up by the person who knows the process

The people who understand these rules are administrators, not engineers. They should be able to change an instruction the same week the policy changes — not file a ticket and wait a quarter.

3

An answer you can check

Two reviewers, two verdicts — and no record of how either of them got there. What a compliance review actually needs isn’t an identical answer every time. It’s an answer that arrives with its evidence attached, so someone can see what it was based on and disagree with it.

WORKSHEET 02 — STATEMENT OF USER NEED

How might we take the reading off people — without taking away their ability to defend the outcome?

How pervasive is the problem?

Five kinds of work, in every customer we looked at. Inside one company it runs hundreds of times a month. No single review is expensive — doing the same one forever is.

Whom does it affect?

Two groups. Business users do the reading and the deciding. Administrators have to choose between giving that to a person or paying for custom code. We designed for both — fixing it for one and not the other would have been useless.

How does the problem affect different individuals?

Not equally. The reviewer loses hours and carries the blame if something is missed. The administrator can see the fix and can’t deliver it. And whoever is waiting — the person who needs the access, the subcontractor waiting to be paid — just waits, with no idea why.

How does it vary across identity or marginalised groups?

Not the sharpest lens for an internal document tool, but the question surfaced a real one: the people affected by a wrong result are never the people who set the agent up. An admin writes the instruction; a subcontractor gets their invoice flagged. That imbalance is why the reasoning has to be visible to everyone downstream, not just to whoever configured it.

Do the problem’s effects change over time?

They get worse. Policies change and volume grows. Changing code means a developer and a release. Changing a written instruction means editing a sentence.

WORKSHEET 03 — AI SUITABILITY

Three good reasons to use AI here — and three warnings we chose to design around rather than ignore.

AI is probably better for

Personalization: The core experience requires personalizing content to individual users, or to a domain or input The same job in every company — never the same rules.
Predictions: The core experience requires prediction of future events
Scale: Need to recognize a general class of things that is too large to articulate every case Nobody can list every way an invoice fails to match a contract.
Volume: Need to detect low occurrence events that are constantly evolving
Language: User experience requires natural language interactions, or agent or bot experience for a particular domain The admin describes the task in a sentence. They don’t code it.
Diversification: The user experience is significantly enhanced when outcomes are variable, rather than predictable

AI is probably not better for

Consistency: The most valuable part of the core experience is its predictability, regardless of context or input Hardest one to accept: the same prompt can give a slightly different answer twice. We limited what it can write at all, and made its reasoning checkable when it varies.
Cost: Running an AI solution at the needed frequency is very high
Stakes: The consequences of errors is very high and outweighs the benefits of a small increase in success rate True of approving an invoice. Not true of narrowing down what a person then reads. We gave the agent the second kind of step only.
Explainability: Users, customers, or developers need to understand exactly everything that happens in the code and application “The system decided” is not an answer you give an auditor. So the reasoning had to come with every value the agent sets.
Momentum: Responsibly getting to market first is more important than other factors or the value using AI would provide
Directive: People explicitly tell you they don’t want a task automated or augmented

Suitability statements from Google PAIR, People + AI Guidebook. Both columns read, three warnings ticked, none waved away.

Worksheet 04 — summary of evidence

We think AI can help with the step where a person reads a document and judges it against a set of rules.

Because those rules are different in every company and keep changing, so nobody can write them as code — and because they only exist inside documents, which is the one thing a language model reads well.
contracts, invoices, specs it reads every line you get the findings
i.

What we handed to AI. Reading the documents, applying the criteria the process owner wrote in their own words, and recording a result — with its reasoning attached.

ii.

What we kept with people. The decision itself. Three warnings were ticked, not waved away, and each one turned into a limit we designed in on purpose.

02
Define stage
DISCOVER DEFINE DEVELOP DELIVER
From Discover: AI can review a contract, trace a renewal date, or compare clauses against a spec.
Double Diamond
Design Framework
CARRIED FORWARD FROM DISCOVER

The work already happens at a fixed point in a workflow. That is where the agent goes.

AI ENGINEERS’ APPROACH

Discover told us what tasks AI can do. It didn’t tell us where AI should sit to do them. But this work already happens somewhere — a contract sitting in Legal review, waiting for someone to get to it. So that’s where the agent naturally goes: attached to that step, firing the moment a document lands there.

Nothing changes for the reviewer. Same workflow, same place to look — except the reading is already done when they arrive, and saying yes is still their call.

it does the reading first same steps, same place you still say yes or no
01

A document enters the step

The workflow moves as it always did. Nobody waits on the agent.

02

The agent is plugged in here

This is the only step we hand over.

03

Then it has to earn trust

A person still approves. What they need to see — not designed yet.

The answer lands where the reviewer already looks — nothing new to open.

02Define stagedesign the thing right

Think of an AI agent as an employee: within its capabilities, how can we design it to succeed at the task?

Discover gave us the what — the heavy reading: checking a document against the rules it has to meet. And the where — the workflow step where that reading already happens. Neither told us how to build it, or why anyone would trust it. I used Google PAIR’s Specify and Align the AI Problem to work that out.

trust Input Output Context Guardrails
Specify & align the AI problem

A session with the same cross-functional group: name the primary goal, the sub-goals underneath it, the parts of the problem people leave unsaid, and the ways the agent could pursue the goal and still be wrong.

Capability mapping

Checked each requirement against what the platform could already do. This is what kept the solution small: most of what we needed already existed in Workflow and version history. Only the judgement was new.

THE REAL DIFFICULTY

No pattern existed for this.

Everything in M-Files assumed the user drives and the system responds. An agent that finishes a step unsupervised inverts that, and nothing told us how much of its reasoning to expose, when it should stop and ask, or what someone needs to see to trust a result they never watched happen. What people said, in one form or another, was always the same thing: “I don’t want a black box making changes I can’t explain.”

The challenge

How might we let an agent work unsupervised — and still be trusted with the work?

What people needTo stop doing the reading — and still be able to stand behind a result they never saw produced.
Solution space
What AI can doRead the documents and apply written criteria — but not reliably enough to be believed on its word alone.
The design lives in the overlap: everything the agent does has to arrive with a reason attached.
WORKSHEET 05 — PRIMARY GOAL

Get the document through its review step without doing the reading — and still be able to stand behind the outcome.

Our user’s primary goal is…

To move a document through a review step without personally reading everything it has to be checked against — and still be able to trust and defend the result if someone asks.

To meet it, our AI system must…

Read the documents it was handed. Apply the criteria the administrator wrote in their own words. Write its result only into the fields it was told to fill. Record why, for every value. And either move the document on, or leave it exactly where it is for a person to decide.

Reading and applying the criteria are capability. Naming what it may change, recording why, and knowing when to stop are the design.

WORKSHEET 06 — SUB-GOALS

What has to be true before and around the main goal.

Before

Pick a step that is safe to hand over at all. Write down criteria that currently live in somebody’s head. Make sure the documents the decision depends on are actually linked to the thing being reviewed.

While

Tell which values came from the agent and which came from a person. Judge whether the result is right. Decide whether it moves forward, goes back, or needs a question asked.

Consistent with the primary goal?

Mostly. Every sub-goal is a way to check the agent’s work. One isn’t: writing the criteria down is new work for the admin. We accepted it — an unwritten rule can’t be checked by anyone, and admins had already done this work when they built custom-coded versions.

WORKSHEET 07 — UNDERSPECIFICATION

The things everyone assumes, and nobody writes down.

What counts as a problem

A reviewer knows which mismatches matter and which are noise. Left unsaid, the agent flags everything or nothing. Specified by making the admin state the check as a numbered list, with worked examples shipped alongside the feature.

Which documents are relevant

People find the right contract or change order by knowing the project. Left unsaid, the agent reasons over whatever it happens to have. Specified by placeholders: the admin names the exact relationship to pull, so the evidence is chosen, not guessed.

How sure is sure enough

Nobody says out loud how confident you have to be before acting. Left unsaid, the agent moves a document forward on a weak read. Specified by letting the admin choose, step by step, whether the agent may advance the workflow at all.

What a person needs to see to believe it

The hardest one, because nobody can describe it until they don’t have it. Specified by attaching the reasoning to the value itself, one click away — not in a log somebody would have to go looking for.

DeepMind’s safety research calls this the gap between the ideal specification — what you meant — and the design specification you actually wrote.

Worksheet 08 — the agent can hit the goal and still be wrong.

PAIR calls this reward hacking; DeepMind calls it specification gaming. Either way, the system pursues exactly what you asked for in a way you never intended. Naming the bad versions early is how the guardrails got decided.
guardrails it only goes where you let it
i.

Meeting our goal can…

Cut review time most on the documents there are most of.

Leave a written trail on reviews that used to leave none.

Let the reviewer spend their attention on the exceptions instead of the routine.

ii.

But what if the agent…

Fills in every field it’s offered, even when the document doesn’t support an answer — and produces something plausible instead of nothing.

Treats “keep the process moving” as the goal, and advances a document it should have paused on.

Reads text somebody typed into an ordinary field as an instruction meant for it.

The last one is a different failure: not a bad goal but an adversarial input — data read as instruction.

Design strategy: three principles, applied everywhere.

1
INTERPRETABILITY

Show the work

Every value the agent sets carries the reasoning that produced it, one click away, kept for as long as the object exists. Not a separate log nobody opens — attached to the thing it changed.

2
SPECIFICATION

Name what it may touch

The admin lists exactly which fields the agent may fill in. It writes those and nothing else — and it can’t go looking around the vault for anything it wasn’t handed.

3
INTERRUPTIBILITY

Let the admin choose when it stops

Not one setting for the whole product, because not every step carries the same risk. It’s chosen step by step, by the person who knows what that step costs if it goes wrong.

Interpretability and interruptibility are DeepMind’s assurance category; naming what it may touch is specification.

DEFINE, SO FAR

Three moves. One decision left.

01

Named the task

The heavy reading — at the step where it already happens.

02

Named the failures

What everyone leaves unsaid, and the ways it could hit the goal and still be wrong.

03

Named the limits

Show the work. Name what it may touch. Let the admin choose when it stops.

All three say how the AI agent behaves and respects unique boundaries — which suggests exactly what kind of AI this agent can be.

WHAT KIND OF AGENT

Not an assistant. Not a copilot. An automation that runs when nobody is watching.

Assistant

Person asks, agent answers

A conversation. The person is present, judges each answer, and acts on it themselves.

Copilot

Agent suggests, person accepts

Inline suggestions. Nothing changes until somebody approves it, so a wrong answer costs a click.

Automation

A step fires, the agent finishes it

No conversation, no approval queue, nobody watching. The person meets only the result — which is why the result has to explain itself.

Input

What it may read

Only what the prompt names. Values, related documents and files are resolved before the agent sees anything. It cannot go looking.

Output

What it may touch

Only the fields listed as outputs. It fills those, or creates the objects it was told to create. Nothing else.

Trace

What it must reveal

Every value arrives with the reasoning that produced it, attached to the value and kept for as long as the document exists.

Authority

Where it must stop

Whether it may move the document forward at all is decided per step, by the administrator. The judgement stays with a person.

Four boundaries — enough guardrail for an automation to run unwatched.

Each principle landed on something the platform already did.

What it may read

Placeholders already point at data without code. Values and related documents are resolved before the agent sees them — it never goes looking.

What it may touch

The same placeholders name the outputs. That list became exactly what the agent is allowed to write, and nothing else.

What it must reveal

A design for AI-generated values already existed. The agent reuses it — the same marking on the value, so nothing new for a user to learn.

Where it must stop

Workflow states already model “a step.” Building the agent as a workflow action meant the choice could be made per step, by the admin.

The gap

Nothing existed for judgement. The only piece the agent adds, and the only piece that needed a new pattern.

Solution

Guardrails let the AI agent work unsupervised. Traceability lets people trust what it did.

Initiatecontract Legalapproval Human still in the loop Final signatory Contract reviewagent AUTHORITY advance the step, or stop and wait INPUT only what the prompt names OUTPUT only the fields it was given TRACE every value carries its reason
TWO SIDES OF THE SAME VALUE

What the admin writes. What the reviewer sees.

Admin — agent configuration
## Main task You are reviewing a subcontractor invoice against the terms of the subcontract agreement. Check this: Are all billed labor roles and rates consistent with the contract rate schedule? ## Actions to perform - {Property Reference(Contract Validation)}: List unconfirmed items, one bullet per item. State what was billed, what the contract allows, and the difference. ## Files Invoice: {Files()} Contract: {Subcontractor.Agreement.Files()}
Invoice check agent Received Verify against contract Waiting for approval
Reviewer — the invoice in M-Files
Metadata
Subcontractor invoice 4417
SubcontractorContinent Construction
Contract ValidationScaffolding hire billed at 480/day; the contract rate schedule allows 410/day. Difference of 70/day across 12 days.
Click the value to edit it, or click X to remove it.
Source of this value

Line 4 of the invoice bills scaffolding hire at 480 per day. The rate schedule in the linked subcontract agreement sets 410 per day for the same equipment.

The suggested value was AI-generated.

Verify against contract

Humans always in the loop.

The agent removes the reading and the cross-referencing. It doesn’t remove the decision.
03
Develop stage
DISCOVER DEFINE DEVELOP DELIVER
Double Diamond
Design Framework

Two trade-offs, made on purpose. Each one cost something real.

1

Built on Workflow, not a new surface

Bought: users already had steps, states, and something moving along a defined path, so autonomy was the only unfamiliar idea. Cost: one agent per state, by design, to avoid conflicting concurrent writes — multi-step reasoning has to be split across consecutive states.

2

Shipped through the legacy admin tool

Cost: every configuration step happens in M-Files Admin, the twenty-year-old Windows-style app the company has been migrating off for about five years. Moving a setup between vaults means replicating the workflow and exporting the agent definition as JSON. Bought: real use, this release.

3

A pattern, not a feature

How reasoning is exposed and how confirmation is handled were decided here for the whole platform — so every AI capability that ships after this one inherits the answer instead of reinventing it.

“A better setup screen a year later wouldn’t have taught us anything about what customers need.”

This is why we accepted the old admin tool as a trade-off. Shipping the right thing first, without waiting for the perfection.
MODEL SELECTION

The cheaper model won on speed too.

An agent isn’t one inference — it’s one every time the state fires, on every object, for every customer. Before picking a default, we benchmarked gpt‑5‑mini against gpt‑5.1 on the same review task. gpt-5-mini shipped as the default — faster at every effort level, and cheap enough to run at the volume this feature is built for.

Bar charts comparing gpt-5-mini and gpt-5.1 on prompt duration and output tokens across high, medium and low effort levels for the complex task TC4. gpt-5-mini is faster at every effort level: 51s vs 78s at high effort, 34s vs 44s at medium, 21s vs 28s at low.

Benchmarked on task TC4 (complex), Custom Agents model-selection testing.

04
Deliver stage
DISCOVER DEFINE DEVELOP DELIVER
Double Diamond
Design Framework
04Deliver stageship the feature, set up the follow-up

Shipping it was the easy part. Knowing who used it wasn’t.

Public beta on 26.6, Cloud vaults only. Two systems went out alongside the feature itself — one to know who signed up, one to know what they actually did with it.

Enrollment tracking

An automated workflow flags every beta signup, every activation, and every subscription’s first real agent run — the moment someone starts trusting it with real work, not just switching it on.

Usage dashboard

A live view of active vaults, event volume, and token cost per subscription across every partner and customer running it. Not for the numbers themselves — for spotting who’s actually using it, so we know who to call.

Why it matters

Adoption alone doesn’t say whether the feature works. Interviewing the customers this surfaces — the ones running it for real — is what the next release gets built around.

What shipped, in the terms the business asked for.

1

The tedious, heavy-cognition steps — automated. Reading a document, interpreting it, and acting on what it finds — consistently and at scale, not one object at a time by hand.

2

Declarative, not developer-dependent. Admins describe the task in prose and constrain it with output placeholders. No custom code, no engineering queue, fast to change.

3

Traceable outcomes, built on trust. Every value the agent sets carries the reasoning behind it — the same record every time, which is what makes a quality or compliance review defensible at volume.

HONESTLY

Still public beta — and the honest gaps are named, not hidden.

No dry-run or preview yet — testing happens against real or test objects, with debug logging as the diagnostic of record. Authentication still takes one manual step from a Subscription Admin, and it’s already on the roadmap to go away. No adoption numbers yet either — beta only opened this release.

05
Reflect stage
What held up,
and what I’d change

2026’s UX consensus calls this “the first new UI paradigm in 60 years” — and names trust, not capability, as the bottleneck.

Nielsen Norman Group, on the shift from command-based to intent-based interfaces.“As organizations move from AI-assisted work to agentic workflows, trust becomes essential.” — Tony Grout, Chief Product Officer, M-Files, at the June 2026 launch.

Where the design landed against that consensus.

Match

“Show the work.” The AI-reasoning icon is exactly this pattern — reasoning attached to every value, not a separate audit screen nobody opens.

Match

Scoped consent. Output placeholders are the same idea as current guidance on agent permissions: name what the agent may touch, refuse everything else by default.

Match

Trust over capability. The bet from the start was that guardrails and traceability — not model accuracy — are what let an agent work unsupervised. A year on, that’s exactly the axis the field now says matters most.

What I'd push for next.

Two gaps I'd close first — both for the person configuring the agent, not the person reviewing its work.
i.

An admin experience built for this, not borrowed. Every configuration step still happens in M-Files Admin, the twenty-year-old tool the company is migrating off. Once the new admin tool is ready, Custom Agents deserves setup screens designed for it — not a form bolted onto software from a different era.

ii.

A pretest for the prompt. There's no dry-run today — the only way to know if a prompt works is to run it for real. Letting an admin test a prompt against a real or sample object and see the output before it goes live would turn writing a prompt from guesswork into iteration.

Takeaway 1

Defining the AI type is essential to creating trust — we chose automation.

Takeaway 2

Trust is a UI decision, not a marketing line.

Thank you.

Waiting for approval… just kidding.
Find me on LinkedIn. :)