DESIGN EVIDENCE · PART 2 OF 2
This is Part 2 of the Design Evidence series.
Start with Part 1: Design Without Evidence Is Mostly Confidence. →
Governance, capability, and the difference between knowing and acting
Most enterprise product failures are not caused by people who know nothing.
The researcher knows the evidence is weak.
The designer knows the requested feature does not address the problem.
The engineer knows the system cannot support the promised behavior.
The accessibility specialist knows the component should not ship.
The product manager knows the target was chosen without a credible baseline.
Everyone knows a piece of the truth.
The work proceeds anyway.
That is the capability problem most organizations misdiagnose.
They respond with another workshop, maturity assessment, certification program, playbook, or dashboard. People complete the training. Their scores improve. They return to the same deadlines, incentives, approval structures, and political constraints that prevented responsible action in the first place.
Nothing changes in the work.
Governance without capability becomes bureaucracy.
Capability without governance becomes optional training.
An organization needs both. Teams must know how to define evidence, challenge weak assumptions, set guardrails, and evaluate results. The delivery system must also require those behaviors, fund them, and give people enough authority to act on what they find.
Otherwise, the organization produces a dangerous kind of competence: people who can recognize bad decisions but cannot stop them.
The system made fragmentation easier than coherence
I learned this while working inside a global healthcare environment with more than 800 digital properties spread across brands, markets, languages, teams, vendors, and technology platforms.
The visible complaint was inconsistency.
Similar interactions appeared in different forms. Brand requirements were handled locally. Product teams built what they needed to meet the next deadline. Design intent weakened as work moved from Figma to Storybook to Adobe Experience Manager.
Accessibility defects appeared late, when they were expensive to correct.
The organization did not have a shortage of talented designers.
It had a delivery system that made fragmentation easier than coherence.
Teams recreated patterns that already existed. Vendors solved the same problems repeatedly. Local exceptions became permanent implementations. Documentation described the intended system but did not always control what reached production.
A prettier component library would not have fixed that.
We built a broader delivery system around shared Figma libraries, reusable design patterns, token-driven brand variation, documented behavior, accessibility requirements, design and engineering review, and clearer handoff into production.
The objective was not to make hundreds of digital properties identical.
It was to reduce unnecessary reinvention while preserving the differences that were genuinely required.
A brand should be able to look like itself without inventing a new form field.
A market should be able to meet a regional requirement without detaching every component from the system.
A product team should be able to move quickly without recreating accessibility, responsive behavior, interaction states, and content structure from memory.
Across the program period, compared with the earlier bespoke and heavily vendor-dependent delivery model, internal reporting associated the shift with approximately a 40 percent reduction in production time, a 50 percent reduction in vendor-dependency costs, and an 18 percent reduction in critical accessibility defects.
Those figures measured different parts of the system.
Production time reflected delivery efficiency. Vendor dependency reflected reuse and internal capability. Accessibility defects reflected whether serious failures were recurring less often.
No single number proved that the user experience was good.
Together, they showed that the organization was producing and maintaining it differently.
The results were also shared outcomes. Design, engineering, content, accessibility, operations, vendors, and leadership all contributed. Presenting them as the achievement of one person or discipline would be dishonest.
Teams still needed to adopt the system. Documentation still needed maintenance. Local requirements still created legitimate exceptions. Accessibility still required review, testing, and human judgment. Governance could reduce drift, but it could not eliminate organizational entropy.
Most case studies end when the dashboard turns green.
Real systems keep aging after the presentation is over.
One exception creates permission for the next one
Delivery pressure has a predictable effect on governance.
A deadline appears.
The shared system does not cover an exact need.
The review process will take too long.
A vendor promises to solve the problem immediately.
Someone detaches a component, hardcodes a value, creates a local pattern, or ships an exception without documenting it.
One exception rarely destroys a system.
It creates permission for the next one.
Soon, several versions of the same interaction exist. The Figma component no longer matches production. Storybook no longer represents every product. Accessibility corrections made in the shared component do not reach the detached versions. Vendors continue maintaining local code because replacing it now appears more expensive than patching it again.
The immediate project meets its deadline.
The organization pays for the decision repeatedly.
Teams often describe design-system governance as something that slows delivery. Months later, they spend far more time reconciling components, correcting inaccessible behavior, rewriting documentation, and explaining why the same control behaves differently across products.
Governance is abandoned because its cost is immediate and visible.
The cost of bypassing it is delayed, distributed, and assigned to somebody else.
That is why governance cannot depend on everybody remembering to do the right thing under pressure.
The evidence contract has to enter delivery
For a single project, the problem, baseline, hypothesis, measures, guardrails, and ownership can fit on one page.
Across a portfolio, those elements have to enter the places where work is requested, funded, assigned, built, reviewed, and maintained.
The problem and intended outcome belong in project intake.
The baseline and known limitations belong in the initiative or epic.
Instrumentation requirements belong in the analytics specification.
Critical user outcomes and accessibility expectations belong in acceptance criteria.
Reusable behavior belongs in the design system.
Exceptions belong in a decision record with an owner and expiration condition.
Post-launch evidence belongs beside the original prediction.
When those elements remain trapped inside a kickoff deck, delivery gradually strips them away. The team remembers the feature but forgets the condition it was meant to change. The target survives while the guardrail disappears. The deadline remains visible while the review date quietly dies.
The organization ships what it remembers.
Accessibility shows whether governance is real
Accessibility is a useful test because it exposes the difference between a stated value and an operating system.
A company may say accessibility matters.
The practical question is where that commitment appears.
Does it affect the tokens and component patterns designers can use?
Are focus states, labels, keyboard behavior, interaction states, and error treatment defined?
Do acceptance criteria include accessibility requirements?
Are automated checks part of delivery?
Does somebody still test complete workflows with assistive technology?
Can a release be delayed when the risk is serious?
WCAG 2.2 includes requirements such as a Level AA minimum target-size criterion of 24 by 24 CSS pixels, subject to defined exceptions. A design system can support that requirement through component dimensions and spacing rules. It can also constrain color combinations, define focus presentation, and document expected semantics.
But a compliant token does not guarantee an accessible experience.
Automation cannot determine whether the reading order is useful, whether a control name makes sense, whether an error message helps someone recover, or whether a complete journey works with assistive technology.
The system can prevent known failures.
It cannot replace human evaluation.
W3C’s organizational guidance treats accessibility as work that should be integrated throughout production and repeated over time, with responsibilities, skills, implementation practices, and monitoring established at both project and organizational levels.
An audit at the end can identify damage.
It does not constitute an accessibility program.
Governance without capability becomes bureaucracy
An organization can require every initiative to include a baseline, metric, guardrail, and review date.
That does not mean the assigned team knows how to create any of them.
A product manager may own the KPI but not understand what user behavior it represents.
A designer may understand task success but lack access to analytics.
An engineer may know how to instrument events but not which events matter.
A business leader may set a target without knowing how the baseline was calculated.
A researcher may identify the real problem but have no authority to affect the roadmap.
Without capability, the evidence contract becomes another template.
People learn what language gets a project approved. The metric is copied from a previous initiative. The guardrail says “accessibility and security.” The review date arrives after everyone has moved on.
Every field is complete.
Nothing is alive.
Capability without governance becomes optional training
The reverse failure is just as common.
An organization invests in research education, analytics training, accessibility champions, design-system onboarding, mentorship, and workshops on outcome thinking.
People learn useful things.
Then they return to a system that rewards speed, output, and obedience to the roadmap.
A capable designer recognizes that the requested feature does not address the problem. The project proceeds.
A researcher identifies that the target rests on a weak assumption. Leadership has already announced it.
An accessibility specialist finds a recurring defect. The component team has no capacity to fix the source.
An analyst explains that the available data cannot support the promised conclusion. The presentation has already been written.
Capability without authority becomes frustration.
The point of capability development is not to make employees more articulate about problems the organization continues to tolerate.
It is to change what happens in live work.
Measurement is a team capability
Evidence discipline should not be treated as a specialist skill owned only by analysts or researchers.
A mature product team needs collective coverage across several behaviors:
- translating business goals into observable user outcomes;
- distinguishing outputs from outcomes;
- establishing or reconstructing a baseline;
- choosing quantitative and qualitative evidence appropriately;
- defining measures and guardrails;
- recognizing weak attribution;
- planning instrumentation;
- documenting limitations;
- reviewing results after launch;
- changing direction when evidence contradicts the original belief.
That does not mean everyone needs to perform every task.
The team needs enough coverage to make and evaluate the decision responsibly.
One person may understand experimentation. Another knows customer operations. Another owns accessibility. Another can implement telemetry. Another understands the commercial model.
The useful question is not:
Is everyone data-driven?
It is:
Can this team collectively define, collect, interpret, challenge, and act on the evidence this project requires?
The answer may reveal a skill gap.
It may also reveal a staffing, authority, access, or leadership problem.
Assess behavior, not identity
Capability assessments become weak when they ask people whether they are strategic, collaborative, data-driven, or user-centered.
Most people know which answer sounds correct.
Useful assessment language describes behavior.
For example:
I can explain which user behavior a product metric represents and where that interpretation may fail.
I define what must not become worse when selecting a success measure.
I can identify when available data is too weak to support the conclusion being requested.
I document the baseline, hypothesis, limitations, and review point before implementation.
I can describe a recent decision that changed because of evidence.
These statements are harder to answer casually.
They also point toward action.
Someone who cannot yet define guardrails does not need another lecture on the importance of metrics. They need to practice defining them on real work with someone who understands the consequences.
Development has to touch a live project
A resistant product team rarely needs another abstract data-literacy course.
It needs to apply the discipline to the work already creating risk.
A focused development sprint could take one active initiative and require the team to reconstruct its baseline, state the intended outcome in behavioral terms, define measures and guardrails, identify what cannot currently be measured, and assign ownership for the post-launch review.
The work should then be reviewed by someone with enough expertise to challenge weak logic.
What moved?
What did not?
Which users benefited?
Did the guardrails hold?
What changed because of the evidence?
The output is not a training certificate.
It is a real project with a defensible evidence chain.
That is how capability becomes behavior.
Resistance is not always a skill problem
A team may know how to work with evidence and still avoid it.
Leadership has already selected the answer.
Bad results are punished.
The deadline rewards shipping, not learning.
Analytics ownership is unclear.
Nobody wants to challenge an executive target.
Research was ignored on the previous project.
The team can identify the problem but lacks authority to change the roadmap.
A healthy team should be able to say:
- the baseline is unreliable;
- the target was invented;
- the requested feature does not address the problem;
- the evidence contradicts the preferred narrative;
- the team cannot measure the promised outcome;
- the current staffing cannot support the quality being promised.
If people cannot say those things safely, capability remains theoretical.
A green capability dashboard beside a low-trust team is a lie.
People may know exactly what responsible practice looks like and still be unable to perform it.
A tool cannot be the enforcer
This is the point where many capability products overreach.
They promise to assess the team, expose gaps, recommend training, and transform delivery.
The missing question is authority.
Who can change the staffing?
Who can fund discovery?
Who can require a development sprint?
Who can stop an initiative from moving forward without a credible baseline?
Who can protect someone who challenges an executive target?
A tool cannot do those things.
Depending on the organization, authority may sit with Product Operations, DesignOps, an enablement function, portfolio governance, an executive sponsor, or a cross-functional review council. The title matters less than the decision rights.
Somebody must be able to act on the signal.
Otherwise, the assessment becomes a more sophisticated description of a problem everyone already knows exists.
The hypothesis behind the Team Capability Engine
I began developing the Team Capability Engine because capability, mentoring, staffing, trust, and live project needs are usually managed as separate systems.
A manager may identify a skill gap without having a mechanism to address it.
A team may complete training without applying it.
A project may require expertise the assigned pod does not possess.
Leadership may receive a green performance dashboard while the people closest to the work are afraid to challenge the plan.
The current TCE model combines three types of signal:
- structured capability assessment;
- focused mentoring and development sprints;
- team-level trust conditions that affect whether people can raise risk and ask for help.
The hypothesis is not that another dashboard will fix product delivery.
The hypothesis is that capability development becomes more useful when it is tied to live project needs, mentoring, staffing decisions, repeated evidence of behavior, and an operating authority capable of responding.
That is still a hypothesis.
TCE should be judged by the same evidence standard it asks organizations to adopt.
Apply the evidence contract to the product
Problem
Organizations often assess skills, deliver training, staff projects, monitor team health, and govern delivery through disconnected processes.
The result is visibility without coordinated action.
Baseline
The first baseline cannot be a generic claim that teams lack skills.
It has to show what happens today when a pod lacks a capability required by a project.
Does the gap emerge during intake, implementation, or after failure?
How long does it take to find support?
Do teams request help or conceal the weakness?
Does training appear later in delivery behavior?
TCE does not yet have enough mature pilot evidence to claim that it has transformed those conditions across organizations.
That limitation should remain visible.
Hypothesis
Connecting team-level capability information to live-project development, mentoring, staffing, and trust conditions will help organizations identify delivery risk earlier and close important gaps before those gaps become defects, delays, or repeated work.
The useful measures are not the number of completed assessments or polished reports.
They include whether:
- capability risks are identified earlier;
- projects receive missing expertise sooner;
- development sprints change real delivery artifacts;
- more initiatives begin with credible evidence;
- teams escalate uncertainty before release;
- known capability gaps recur less often;
- post-launch reviews influence the next decision.
A capability system can damage the organization it is supposed to help.
Scores can become labels.
Managers can use self-assessment as performance evidence.
Employees can learn to game the questionnaire.
Cultural and disciplinary differences can be mistaken for weakness.
Trust data can become a tool for identifying dissenters.
The system therefore needs data minimization, clear anonymity thresholds, human interpretation, transparent scoring logic, bias review, and an explicit prohibition against using a single capability score as an employment judgment.
A team-development system should not become employee surveillance with kinder language.
Ownership
The product can surface a capability risk.
It cannot decide whether the organization accepts it.
An accountable operating owner must review the signal, fund the response, change staffing when necessary, and connect development to live work. Executive sponsorship is required when the intervention challenges delivery dates, staffing commitments, or leadership assumptions.
Without ownership, insight becomes notification.
Self-assessment is not proof
TCE currently begins largely with structured self-evaluation.
That is useful for reflection and shared language.
It is not sufficient evidence of delivery capability.
“I understand outcome metrics” is a claim.
“I defined the outcome, baseline, guardrails, and review point for the last three initiatives” is demonstrated behavior.
Over time, capability assessment should incorporate real artifacts: evidence contracts, research plans, metric definitions, analytics specifications, accessibility criteria, decision records, and post-launch reviews.
This must be handled carefully.
The objective is not to create a permanent score attached to a person.
The objective is to understand whether the team has enough demonstrated capability to carry the work responsibly—and where the organization must provide support.
The standard is changed work
It should ask whether the operating behavior changed.
Are more initiatives entering delivery with a credible baseline?
Are teams defining guardrails?
Are accessibility and research gaps being addressed earlier?
Are pods requesting expertise before a release is already in danger?
Are people able to challenge unreliable targets?
Does evidence ever stop a project, narrow its scope, or change its direction?
If the answer is no, the organization may have created a useful learning platform.
It has not yet changed delivery.
That distinction matters for TCE and for every capability program sold with the language of transformation.
Evidence becomes real when it is inconvenient
The real test of an evidence system is not whether it works during kickoff.
It is whether it survives the deadline.
The executive escalation.
The vendor promise.
The reorganization.
The person who says, “We do not have time for this right now.”
Every organization supports evidence when the evidence agrees with what it already wants to do.
The discipline becomes visible when the result is inconvenient.
When research says the feature solves the wrong problem.
When the target has no credible baseline.
When accessibility remediation threatens the release date.
When the primary measure improves but the guardrail fails.
When the expensive intervention does not move the intended outcome.
That is the moment the organization decides whether evidence is a management tool or presentation material.
A capable team without authority is a warning system nobody listens to.
A governance process without capable people is paperwork nobody believes.
The work changes only when knowledge, permission, and responsibility meet inside the same decision.
Return to Part 1: Questions before the screens
Design Without Evidence Is Mostly Confidence →
