A Consumer Duty policy proves that a firm has designed a control. It does not prove that a customer received the intended outcome. The Financial Conduct Authority expects firms to monitor, assess, test, understand and evidence the outcomes customers are receiving – including customers with characteristics of vulnerability. That requires evidence from real journeys, not only documentation and training records.
The supervisory context is becoming sharper. In Enforcement Watch 2, published on 7 July 2026, the FCA reported 11 open operations involving potential Consumer Duty breaches, up from six in its first Enforcement Watch. The matters span insurance, pensions, wealth management, consumer investments, peer-to-peer lending and claims management.
Not every weakness becomes an enforcement case, and the FCA considers seriousness, harm, proportionality and deterrence. The direction is nevertheless clear: firms need to show what customers experienced, how poor outcomes were identified and what changed as a result.
The commercial reality: A policy explains what should happen. Consumer Duty evidence shows what did happen, for which customers, and whether the firm acted when the outcome was poor.
What evidence does Consumer Duty require?
The Duty requires firms to act to deliver good outcomes for retail customers. Outcome monitoring should help a firm understand performance across:
- products and services;
- price and value;
- consumer understanding; and
- consumer support.
The FCA’s insurance review states plainly that firms must regularly assess, test, understand and evidence customer outcomes. Without effective monitoring, a firm cannot know whether it is meeting the Duty.
Evidence should therefore do more than count completed training, approved scripts or complaints. It should help management answer:
- Did the intended customer group receive the product and support designed for it?
- Did communications enable customers to make informed decisions?
- Were customers able to obtain support without unreasonable friction?
- Were foreseeable harms identified and addressed?
- Did vulnerable customers receive outcomes at least as good as other customers?
- Where poor outcomes were found, what remediation followed?
Why policies and file reviews leave a blind spot
Most regulated firms can produce a Consumer Duty policy, vulnerable-customer framework, training material and monitoring reports. These are necessary controls. They mainly show design and declared process.
They may not reveal whether a frontline colleague:
- noticed a subtle vulnerability indicator;
- asked an appropriate follow-up question;
- adapted pace, channel, explanation or support;
- avoided an unsuitable or harmful recommendation;
- recorded the relevant information accurately;
- made a promised adjustment available to the next colleague; or
- resolved the customer’s need without creating avoidable friction.
A transcript review can show what was said, but not always the full context or the options that were not offered. A complaints analysis focuses on customers who complained. A customer survey may exclude those who abandoned the journey or did not realise that the outcome was poor.
Scenario-based customer experience testing can complement these sources by creating controlled journeys with known facts and expected behaviours.
The three evidence layers to test
Layer 1: compliance and process
Did the interaction follow the documented process? This can include identification, mandatory disclosures, consent, prescribed questions, vulnerability recording, complaints recognition, signposting and escalation.
Layer 1 shows whether the control was applied. It is necessary but not sufficient.
Layer 2: service quality and customer outcome
Was the customer’s need understood and resolved competently? Did the colleague provide relevant information, explain options, avoid unnecessary repetition and deliver the support promised by the firm?
A journey can follow the script and still produce a poor outcome.
Layer 3: emotional and situational fit
Did the support land appropriately for the customer in front of the firm? Vulnerability, distress, bereavement, low financial resilience, limited digital confidence, language and cultural expectations can affect how a technically correct interaction is understood.
The third layer does not replace objective criteria. It tests whether the firm’s process works in the human circumstances for which it was designed.
Copernicus view: Consumer Duty testing should separate process compliance from outcome quality. One overall percentage can conceal a firm that follows the script but fails the customer.
How mystery shopping can support Consumer Duty monitoring
Mystery shopping in a regulated environment should be a controlled evidence programme, not a general retail scorecard.
Seed a known customer scenario
Create a scenario with defined needs, objectives, knowledge, financial position and vulnerability indicators. The expected actions and unacceptable outcomes should be agreed before fieldwork.
Test the full journey
Include relevant digital, telephone, branch, adviser and follow-up stages. Consumer support often fails at handovers or channel boundaries rather than in a single conversation.
Use trained, appropriate evaluators
Complex or high-value journeys may require experienced evaluators who understand the product, regulatory context and difference between observing behaviour and soliciting regulated advice beyond the approved scenario.
Capture objective evidence
Record date, time, channel, journey steps, documents, disclosures, questions, answers, promises and outcome. Audio, screenshots or other records should be collected only where lawful and authorised.
Protect customers and staff
The programme should be designed with legal, compliance, data-protection and employment input. Scenarios should avoid causing real customer harm, consuming scarce emergency support or inducing conduct the firm would not reasonably encounter.
Feed findings into governance
Results should identify root causes, customer impact, affected groups, remediation, accountable owners and retest dates. Evidence is useful only when it changes decisions.
What to test before the next supervisory review
Vulnerability identification
- Are explicit and implicit indicators recognised?
- Do staff ask proportionate follow-up questions?
- Is information recorded accurately and respectfully?
- Does it remain visible to colleagues who need it?
Adapted support
- Is communication pace, format or channel adjusted?
- Are accessible options and reasonable support offered?
- Does the firm avoid asking the customer to repeat sensitive information?
- Are third-party or representative arrangements handled correctly?
Consumer understanding
- Are key risks, costs, exclusions and consequences explained clearly?
- Does the colleague check understanding rather than merely deliver wording?
- Are communications appropriate for the customer’s knowledge and circumstances?
Consumer support and friction
- Can customers reach the right team without unreasonable barriers?
- Do cancellation, complaints, claims and switching journeys work as readily as sales?
- Are vulnerable customers disadvantaged by digital-only or time-limited processes?
Outcome and follow-through
- Was the issue resolved or an agreed next action completed?
- Does the system record match what the customer was told?
- Are repeat contact, abandonment and failure demand visible in management information?
A real anonymised Copernicus engagement: 22% missed, then 9%
A tier-1 European retail bank commissioned a vulnerable-customer audit across six countries and 312 branches. Its documented compliance controls were strong, but management wanted evidence that they translated into consistent frontline advice.
Procedural compliance scored highly. The behavioural test produced a different picture. In a seeded vulnerable-customer scenario, more than one fifth of branch interactions – 22% – failed to identify the vulnerability indicator. A further 14% identified the indicator but did not adjust the advice that followed.
Training was rebuilt around the indicators fieldwork showed colleagues were actually missing, rather than those the policy assumed would be recognised. In the second audit wave, the missed-indicator rate fell to 9%. Findings were then incorporated into the bank’s quarterly conduct and compliance reporting.
This is an anonymised, completed Copernicus engagement. Details are presented in accordance with the client’s confidentiality requirements.
Turn individual findings into outcome evidence
A useful Consumer Duty audit report should show more than pass rates. It should connect:
- the scenario and customer need;
- the expected control and outcome;
- the observed behaviour;
- evidence supporting the finding;
- actual or potential customer harm;
- whether the issue is isolated or systemic;
- affected products, channels, markets or customer groups;
- root-cause hypothesis;
- corrective action and owner; and
- retest result.
Segment results by relevant customer groups, channels, products and locations. A satisfactory average can conceal poor outcomes for a smaller vulnerable cohort.
Build a repeatable monitoring cycle
Diagnose
Use an initial wave to identify where process, capability, technology or culture breaks down.
Remediate
Target the specific failure. The answer may be training, but it may instead be product design, workflow, system visibility, authority, staffing or incentive.
Retest
Repeat matched scenarios so the firm can show whether outcomes improved. A redesigned policy without a second measurement is still an untested control.
Govern
Report trend, customer impact, exceptions and overdue actions to the appropriate committee and board. Link evidence to the annual Consumer Duty assessment and other outcome-monitoring information.
Refresh
Update scenarios when products, channels, rules, customer risks or known failure modes change.
Where Consumer Duty evidence programmes usually go wrong
- Measuring policy existence rather than customer outcome. Documentation proves design, not delivery.
- Using only layer-one compliance scoring. A correct script can coexist with a poor result.
- Testing obvious vulnerability only. Frontline performance often breaks on subtle or emerging indicators.
- Relying on general shopper panels for specialist journeys. Evaluators need the competence to observe complex interactions accurately.
- Reporting one percentage without evidence. Management needs the behaviour, impact and root cause behind the number.
- Failing to segment results. Overall performance can hide weaker outcomes for vulnerable groups or particular channels.
- Running one diagnostic wave. Improvement must be demonstrated through retesting.