"Is our design system under control?" Most teams answer that question with a feeling. Designers feel the product has drifted. Engineers feel it's mostly fine. Leadership hears both and has nothing to compare them with.
Design system health is how closely the product in production follows the approved design system, and how quickly it returns to it after a change. This article shows how to measure it with a few signals that design, engineering and leadership can all read the same way.
Key takeaways
- Measure design system health with four signals: divergences, reach, bypasses and time to fix.
- Only 41% of teams measure adoption (zeroheight 2026), and only 16% had metrics tracking established in the 2022 Sparkbox survey.
- Code-side measurement lags: only 9% of respondents named a dedicated code adoption tracker among their measurement tools (zeroheight Design Systems Report 2026).
- Roll divergences and bypasses into one score weighted by reach, track time to fix next to it, and keep the full list of findings one click away.
Why a feeling doesn't work
Feelings come from whatever people happened to look at. A designer who spent the week on the settings screen sees its problems everywhere. An engineer who just cleaned up the token file sees a tidy system. Both are right about their corner and wrong about the whole.
Without a shared measurement, design system work loses to feature work, cleanups stall because nobody can show a result, and the same debate returns every quarter. Winning that debate is getting harder. The zeroheight Design Systems Report 2026 surveyed 147 practitioners, most at companies with 1,000+ employees. 40% of them were dissatisfied with their ability to get buy-in for their design system, up from 23% the year before.
What teams measure today
Most teams measure little. In the zeroheight 2026 report, 41% of teams measure adoption, 26% measure product consistency and just 5% measure ROI. Asked which tools they use to measure, 37% of respondents named code repository analytics and only 9% a dedicated code adoption tracker. The older Sparkbox Design Systems Survey 2022 (219 responses) found that only 16% had metrics tracking or reporting established for their design system.
Mature teams show what good looks like. Pinterest measures design-side adoption as the share of design system component layers among all layers on handoff pages, counting only files edited in the last two weeks (Figma Blog, 2023). Twilio's Paste team found that npm download counts said nothing about who used which parts, so they built a tool that reads import statements file by file (Twilio, 2021).
Atlassian goes further and blocks bypasses at the source. ESLint rules such as ensure-design-token-usage flag values that should be tokens, and codemods clean up the rest (Atlassian Design System).
Four signals of design system health
A useful measurement is one you can recompute automatically and compare over time.
Reach turns a list into priorities. A colour that drifted in a component used on every screen outweighs a dozen one-off values on a forgotten admin page.
Bypasses show whether the system is really used. A spacing value written by hand isn't wrong yet. It's simply outside the system, and it won't follow the next change.
Time to fix shows whether the process works. If divergences are found quickly but live for months, the problem is ownership, not visibility.
How to collect each signal
Each signal can be collected automatically, without asking anyone to fill in a spreadsheet:
- Divergences: diff token names, values and modes between Figma and code, matching by stable ID rather than by name. The token workflow guide explains why.
- Reach: count the files that import or use each token and component.
- Bypasses: use a lint rule or a code search for hex and pixel literals where a token exists. Exclude generated folders.
- Time to fix: record when each finding was first seen and when it was resolved.
Exclude generated code from bypass counts. FISYCO is the product we build. When we first measured a demo repository with it, the icons our own sync had exported counted as about 560 inline SVGs and 590 hard-coded colours. The code was fine; the metric wasn't.
One score, with the details behind it
Leadership needs one number. Teams need the list behind it. A good health score gives both: a single value that shows the trend, built from findings anyone can open, sort and fix.
A simple version weights findings by reach. For example (illustrative): if 2,400 usages of tokens and components are tracked and 120 of them are affected by open divergences or bypasses, health is 100 × (1 − 120 ÷ 2,400) = 95.
Time to fix stays outside the score. Track it next to the score as a trend, for example the median number of days a finding stays open.
Two rules keep the number honest:
- Weight by reach, not by count. Ten small findings should not outweigh one divergence in a component used everywhere.
- Move the baseline only on purpose. Measure against the approved design at a known point, and update that point only when an approved change lands in code.
How to present the score to each audience
The same data needs three views:
- Leadership: the score, its trend and the few divergences with the largest reach.
- Design: approved decisions that never reached production.
- Engineering: where code bypasses the system, with a proposed fix for each finding.
When everyone looks at the same findings, the conversation changes from "is it bad?" to "which of these do we fix first?"
Review the score on a fixed rhythm, for example once per sprint. A trend over several weeks says more than any single value.
If you haven't looked at where the gap comes from, start with why production quietly drifts from your design system.
Frequently asked questions
Should we measure adoption in Figma or in code?
Both, but they answer different questions. Figma adoption shows whether designers use the system. Code adoption shows whether customers get it. If you can measure only one, measure code, because that's what ships.
Is a single score too simplistic?
Only if it hides the details. A score is a summary for decisions; the findings behind it are what teams act on. Keep both.
How do we measure without adding work for the team?
Automate it. Recompute on every push or design publish. A metric that needs manual input will be stale by the second month.
How FISYCO measures health
We built FISYCO to make this measurement continuous. It indexes your code on every push and compares it with the design from your Figma libraries. The health report shows divergences, their reach and bypasses of tokens and components. Fixes are verified by rendering before they arrive as pull requests. See what FISYCO does today.

