Yu Mao← Back to portfolio
Agent Reliability · Systems case study

From business documents to auditable code.

Agents naturally optimize for textual closure; people care whether the task’s real-world actions actually close the loop. I built a document-to-code workflow that keeps business authority, implementation, tests, and change impact connected—so a green test suite cannot silently legitimize the wrong rule.

Read the Task World Model outline
30 / 30sample rows agreed with the wrong inferred rule
14 / 15expected values changed after the authoritative rule arrived
5 evidence stateskeep inference separate from authority
01 / The case

A rule can fit every example and still be wrong.

The useful artifact is not another generated pipeline. It is a system that remembers what the source actually said, what was inferred, and what must be invalidated when authority changes.

AUTHORITATIVE BUSINESS DOCUMENT · SANITIZED EXCERPT
“Target date = Max(Dependency A ready, Dependency B ready) + Service window”

The milestone clock starts only after both prerequisites are ready, then applies the governing service window.

Before the document stated the formula

Two example tables suggested a different rule:

max(Dependency B ready + 14d,
Downstream start date)

15/15 + 15/15 rows passed. It looked independently validated.

After the authoritative revision

The source named the real rule:

max(Dependency A ready,
Dependency B ready) + service window

On the retained 15-row case, 14 expected dates changed. The old screenshots encoded a parallel-work assumption, not the governing policy.

The executable regression case

Fourteen rows exercise the Dependency A branch; one row exercises the Dependency B branch. Reverting to the old inferred formula makes the case fail.

14 rows → Dependency A branch
1 row  → Dependency B branch
Green ≠ correct

Tests prove that code matches their expected values. They do not prove those expected values express the business rule—especially when code, tests, and expectations were generated from the same interpretation.

Public case study: internal identifiers, source links, and business data are omitted. Counts and rule evolution are preserved from the project evidence.

02 / State machine

A state machine exposes the assumptions hidden inside a formula.

It does not discover the rule. It makes missing evidence visible.

Rewriting the formula as behavior forces three questions:
  • What state are we in?
  • What event moves us forward?
  • What evidence supports the guard?
Before · wrong

Dependency B ready → advance

The generated logic treated one prerequisite as sufficient and advanced too early.

Dependency B ready → In progress
After · correct

Wait until both prerequisites are ready

The workflow advances only after Dependency A and Dependency B are both ready.

Waiting → Eligible → In progress
guard: dependency_a_ready && dependency_b_ready

Evidence status is a separate layer.

Once the logic is explicit, each interpretation still needs a recorded source of authority.

State 01

Unresolved

The source is incomplete or conflicting. Affected output is blocked or marked provisional.

State 02

AI inferred

A reading is possible, but no independent evidence or owner has authorized it.

State 03

Evidence derived

Examples support the rule. The exact rows and counterfactuals remain attached.

State 04

User answered

A workflow owner accepts the interpretation, with provenance and scope recorded.

State 05

Author answered

The rule enters the source document and leaves the exception ledger.

Unresolved → AI inferred
A reading is selected.
Unresolved → Evidence derived
Samples distinguish a rule.
Any open state → User answered
A workflow owner rules.
Any open state → Author answered
The decision enters the source.
On any source-block hash change: dependent judgments become stale, impacted cases rerun, and only affected rules return for review.
03 / Correctness

Correctness is a proof chain, not one model call.

The document compiler keeps the source, semantic decision, implementation, and runtime evidence independently challengeable.

Freeze the source

Mirror the original document with revision IDs, normalized block hashes, and asset hashes. A missing or changed source fails closed.

Compile semantics

Turn prose into explicit fields, keys, formulas, branches, and unresolved decisions. No silent default becomes a business rule.

Break shared failure

Use owner-reviewed hand cases and a separately built reference model. Code and expected outputs cannot share one unchallenged interpretation.

Bind evidence

Map each rule to DAG nodes, code symbols, cases, invariants, mutations, replay diffs, and row-level lineage.

ReconciliationInput and output totals reconcile within each rule scope.
CounterfactualsInject plausible wrong implementations and require at least one case or invariant to fail.
Controlled changeSame inputs reproduce the same output; one rule edit explains exactly which outputs move.
Boundary: this workflow can prove faithful implementation of an authorized rule. It cannot manufacture authority when the business document itself leaves a consequential choice undefined.
04 / Best Paper

Intent should define what an agent is allowed to do.

The companion research asks who may determine each field of an agent action when user requests, documents, tools, and runtime evidence carry different authority.

AgenticOS 2026
Best Paper Award

LLM Agent Capabilities Should Follow Task Intent and Context Source

Yusheng Zheng · Wenhui Zhang · Yu Mao

IntentCap composes task-scoped capabilities from multiple context sources using field-level ownership and monotonic narrowing, then validates short-lived leases with a deterministic checker before side effects commit.

05 / Technical deep dive

What an interviewer can challenge.

Why not generate code directly from the document?

Direct generation collapses extraction, interpretation, implementation, and validation into one correlated failure path. The compiler makes those boundaries explicit and retains unresolved semantics instead of guessing through them.

What exactly triggers invalidation?

Normalized source-block content, document revision identity, and referenced asset bytes. When a source hash moves, the system marks dependent judgments stale through the existing block → rule → operation → case graph.

How is the state machine different from evidence status?

The state machine organizes business behavior: states, events, guards, transitions, and actions. Evidence status answers a different question—who or what supports each extracted rule. “Code exists” and “tests pass” do not advance that authority.

What is the review unit?

A rule card: original text, executable interpretation, a small hand-checkable case, business impact if wrong, implementation anchors, evidence strength, freshness, and the exact decision being signed.