Impeccable explained: design guidance inside an AI coding workflow
Measure Impeccable without mistaking rule counts for design quality
Build a comparison that keeps deterministic findings and human judgment apart.
What you will learn
- Choose a fixed screen set
- Record action on each finding
- Include time and operational cost
Before you start
- One screen and its user task
- Authority to inspect an agent skill and project hook
Produce an evidence-backed design change that another reviewer can accept or reject.
Key takeaways
- Rule count does not measure usability.
- False positives belong in the evaluation.
- Model critique and deterministic scans have different costs.
Choose a fixed screen set
Select a few representative pages: a marketing page, a form-heavy surface and a mobile view. Save screenshots, dimensions, task descriptions and the current design-system constraints before running any command.
The upstream count of detector rules is not a measure of recall or usefulness. A rule may be correct but irrelevant to the product, and a serious interaction issue may sit outside the rules entirely.
Record action on each finding
For every finding, write its location, evidence, proposed change and reviewer decision. Mark false positives and unresolved cases. If an agent edits the screen, compare the rendered before and after state at the same viewport.
Run the application’s build, accessibility checks and task-based review separately. A clean detector run can be evidence for one class of problems, but it cannot demonstrate that users finish a task.
Include time and operational cost
Track detector runtime, number of actionable findings and review time. If a model-powered critique is used, record its provider and spend separately from deterministic scanning.
This series did not run a benchmark or user test. The scorecard is a proposed evaluation design; fill its cells only with measurements from your own screens.
Decision guide
| Criterion | Option A | Option B |
|---|---|---|
| Best when | You need predictable behavior and easy auditing | You need adaptive optimization and have reliable telemetry |
| Main risk | May leave performance on the table | Can become difficult to explain or debug |
Implementation steps
- 1
Freeze a small set of screens and viewports.
- 2
Classify each finding with a reviewer decision.
- 3
Compare task outcomes and time apart from rule counts.
Copy-ready example
screen,viewport,finding,reviewer_decision,change,task_result,scan_ms
checkout,390x844,,,,,Frequently asked questions
Does a zero-finding scan mean the UI is ready?
No. It only means the scanner reported no primary findings for the inspected target.
What should be measured first?
Whether findings lead to justified changes on the screens your users actually use.
Sources
- Impeccable / README.mdSource checked 2026-09-29
- Impeccable / docs/CLI-CONTRACT.mdSource checked 2026-09-29