IOU Financial · External Design System

Making our design system work beyond the design team.

I turned our design system into a skill so PMs could explore ideas with AI, then tested and refined how it builds, flags changes, and handles decisions we haven’t made yet.

Role
Sole product designer
Who it’s for
Product managers
Scope
Design system skill and AI evaluation
Status
Piloting with PMs; system updates ongoing

Overview

How could PMs use AI agents to build with our design system while keeping mocks clearly separate from approved designs?

On our small team, I’m the only designer, so system-aligned screens tend to signal ready for development. I wanted PMs to use the system to explore ideas and check them against product requirements while keeping proposals clearly marked. I turned it into a skill that labels every screen as a mock with its creator and date, and keeps unapproved amounts, legal copy, and policy details as placeholders.

The Goal

Make our design system usable by PMs without needing me beside every mock.

01
Start with our system
Give PMs a skill that builds from our components, patterns, and behavior.
02
Make room for new ideas
Let them explore changes and new flows beyond the screens already defined.
03
Show what needs review
Label proposals and keep unapproved details as placeholders.

The Setup

I turned our design rules into a skill an agent could use, then tested where the guidance fell short.

For PMs to use the system independently, the decisions I knew needed to be explicit. I organized the guide into visual values, component behavior, and screen recipes, with rules I could check against the output. I tested eight tasks in fresh sessions, scoring screens and code against predefined checks. Seven covered defined screens and states. The eighth requested income verification, a step the system didn’t define. I expected the agent to stop and ask. Could it recognize a missing product decision when it had the components to build a convincing screen?

The first version combined visual styles, component behavior, and screen patterns to guide what the agent built.

Failure and Fix

The agent built a step the guide never defined. I made the supported steps explicit so it would stop and ask.

Attempt 1: it built the undefined step. The agent ignored the instruction to ask and created income verification with fields, dollar amounts, and legal copy the guide never defined. Attempt 2: a firmer rule still left a gap. I told it to stop when no recipe matched. It called the step a generic form and built it again. I needed to distinguish a reusable layout from a defined product step. Attempt 3: an explicit list made that distinction. I listed the supported steps and ruled out using generic recipes for anything outside them. Across five reworded requests in fresh sessions, the agent stopped and asked.

The list worked in all five follow-up tests. But PMs still needed to explore beyond the loan application.

The Result

One skill now supports existing screens, proposed changes, and new experiences.

Defined screens follow their recipes. Changes to existing screens are marked, with any rule conflicts explained in plain language. New experiences use system components and are identified as new to the system. I tested the broader behavior on a checkout flow using our team’s brainstorm transcript, running it on both Opus and Sonnet. The skill used the discussion to shape the flow while keeping unapproved requirements and timelines as placeholders. The skill is packaged and ready for a PM pilot. The next test is how it holds up during everyday exploration.

The Goal

Make our design system usable by PMs without needing me beside every mock.

01
Start with our system
Give PMs a skill that builds from our components, patterns, and behavior.
02
Make room for new ideas
Let them explore changes and new flows beyond the screens already defined.
03
Show what needs review
Label proposals and keep unapproved details as placeholders.

The Setup

I turned our design rules into a skill an agent could use, then tested where the guidance fell short.

For PMs to use the system independently, the decisions I knew needed to be explicit. I organized the guide into visual values, component behavior, and screen recipes, with rules I could check against the output. I tested eight tasks in fresh sessions, scoring screens and code against predefined checks. Seven covered defined screens and states. The eighth requested income verification, a step the system didn’t define. I expected the agent to stop and ask. Could it recognize a missing product decision when it had the components to build a convincing screen?

The first version combined visual styles, component behavior, and screen patterns to guide what the agent built.

Failure and Fix

The agent built a step the guide never defined. I made the supported steps explicit so it would stop and ask.

Attempt 1: it built the undefined step. The agent ignored the instruction to ask and created income verification with fields, dollar amounts, and legal copy the guide never defined. Attempt 2: a firmer rule still left a gap. I told it to stop when no recipe matched. It called the step a generic form and built it again. I needed to distinguish a reusable layout from a defined product step. Attempt 3: an explicit list made that distinction. I listed the supported steps and ruled out using generic recipes for anything outside them. Across five reworded requests in fresh sessions, the agent stopped and asked.

The list worked in all five follow-up tests. But PMs still needed to explore beyond the loan application.

The Result

One skill now supports existing screens, proposed changes, and new experiences.

Defined screens follow their recipes. Changes to existing screens are marked, with any rule conflicts explained in plain language. New experiences use system components and are identified as new to the system. I tested the broader behavior on a checkout flow using our team’s brainstorm transcript, running it on both Opus and Sonnet. The skill used the discussion to shape the flow while keeping unapproved requirements and timelines as placeholders. The skill is packaged and ready for a PM pilot. The next test is how it holds up during everyday exploration.

Reflection

What I learned

A polished screen can still be wrong.
Using our components didn’t mean the agent had the product decisions to complete a screen. The skill needed to make unapproved details visible rather than present them as settled.
Testing the agent also tested my system.
The guide said radio pills “fill the row”; the code specified 148px. Testing exposed inconsistencies like this, and I used them to improve the source guidance.
Useful exploration shouldn’t depend on the right prompt.
PMs shouldn’t need special instructions to explore an idea. I made labeled mocks the default, with unapproved details left as placeholders. The PM pilot will test how this holds up in everyday use before I extend the approach to our internal tools system.