CBA-CRP · Management
Specimen paper
Twelve examination items for the CBA Certified Compliance & Risk Professional, with the answer and a rationale for every option.
Examiner’s note
This specimen is drawn from the live CBA-CRP bank and is representative of it, not a selection of its gentlest items. You will find four foundational items, six standard and two demanding, spread across the five domains in roughly their blueprint proportions. The scenarios are short, situated in real firms in real markets, and every option is written to be defensible by somebody. That is where marks are lost on this certification. The commonest failure is not ignorance but settling for the near miss: an option that is true, or that fixes a symptom, or that names the proximate trigger, when the question asks for the mechanism or the decisive test. Read what the stem has already told you, and ask which options the stated facts rule out.
Items
12
Domains
5
Questions in the examination
75
- 01Compliance Risk AssessmentFoundational
At Aravalli Credit, an Indian non-bank lender, the head of collections argues that the risk of lending against unverified income should be scored low, because a second reviewer checks every file before drawdown. What is the correct treatment?
- A
Score inherent risk low, because a check applied to every file removes most of the exposure.
You arrive here by scoring the world as you find it rather than as it would be without controls. Universal coverage may well reduce residual risk, but by definition it cannot touch inherent risk. This would be right only if inherent risk meant net exposure, which would leave the register with no way to price what the second review is worth or to notice it degrading.
- B
Score inherent risk low for checked files and high for the rest, so both populations appear.
This comes from splitting the population by control coverage, a residual-stage refinement borrowed into the inherent-stage column, and it appeals because two rows look more granular than one. It would be right only if inherent exposure genuinely differed between the two groups, for example because they were different products or channels, rather than differing only in whether a control happened to touch them.
- C
Score inherent risk before the check, then set residual from tested evidence that it operates.
Correct: it scores the exposure before any control and then derives residual risk from tested evidence that the control actually operates.
- D
Score inherent risk high, then set residual low because the check covers every file in the book.
This gets the inherent half right and then rates the control on its existence and coverage rather than on evidence that it works, which is the error of treating design as operation. It would be right only if the review had been tested and found effective; coverage tells you how often something happens, not how well it is done.
Why that is the answer
Inherent risk is the exposure before any control is applied; residual risk is what remains once you have tested that the control operates. Keeping the two apart is what allows the register to show what the second review is worth, and to warn you if that review later weakens. Collapse them into one score and you can never take the judgement apart again to see whether the low rating came from a small exposure or a strong control. The keyed option preserves both judgements and ties the second one to evidence rather than to the control's mere existence.
- 02Programme Governance, Policies and ControlsFoundational
At Nasma Capital, a Bahraini investment firm, the onboarding policy states the verification requirement, a named analyst performs the verification each day, and the team confirms that it happens. Asked to produce evidence for six named clients, the team can produce none. Which joint in the programme has failed?
- A
Obligation to policy: the requirement was never written in the firm's own words.
You get here by reading "we cannot produce evidence" as "it was never properly written down", which is the usual first guess when a programme fails. The stem forecloses it: the onboarding policy states the verification requirement. It would be right if the firm's documents were silent on verification, or merely reproduced the regulator's words without translating them into the firm's own obligation.
- B
Policy to control: the requirement is not assigned to any named person or system.
This is the near miss that stops one joint too early, and it tempts because an unassigned requirement is the more familiar failure and "nobody can produce evidence" superficially resembles "nobody owns it". The stem closes it: a named analyst performs the verification each day. It would be right if the policy set the requirement and no person or system had been made responsible for it.
- C
Control to evidence: performing the control leaves nothing that can be tested.
Correct: the control operates but produces no artefact, so nobody outside the team can test that it happened.
- D
Monitoring to remediation: the finding has no owner and no date set for closure.
You reach this by jumping to the end of the chain, where findings, owners and closure dates live. The failure here is upstream of monitoring, because there is nothing for monitoring to inspect in the first place. It would be right if the gap had already been identified as a finding and then allowed to drift without an owner or a closure date.
Why that is the answer
A compliance programme is a chain of joints: obligation to policy, policy to control, control to evidence, monitoring to remediation. When something breaks, your task is to identify which joint gave way rather than to restate that something is wrong. Here the policy states the requirement and a named person performs it daily, so the first two joints hold; what the control fails to do is leave a by-product that anybody outside the team can inspect. A control that leaves no trace cannot be tested, and therefore cannot be relied on by anyone who was not standing there watching.
- 03Financial Crime and Data Protection FundamentalsFoundational
Ransomware encrypts the customer document store of an Indian non-bank lender. Forensic work finds no sign that any data left the network. The head of technology records the event as an availability incident rather than a personal data breach. What is the correct classification?
- A
It is not a personal data breach: the definition turns on unauthorised disclosure of or access to personal data.
This comes from remembering only the confidentiality half of the definition, which is the half that dominates the headlines, and it is the commonest error on the topic. It would be right only under a regime whose definition was confined to unauthorised disclosure or access, whereas the mainstream definitions also cover destruction, loss and alteration.
- B
It is not a personal data breach yet: the classification would change only if forensics later found exfiltration.
This treats classification as provisional pending forensic work, which sounds prudent and is how many incident logs are in fact kept. The error is deferring a decision the known facts already settle, because availability has been lost now. It would be right if the only limb of the definition in play were unauthorised access, so that the facts genuinely were incomplete.
- C
It is a personal data breach only if the backups fail: a successful restore means no data was in fact lost.
This confuses recoverability with whether loss occurred, and it is the reasoning technology teams reach for because they measure success by service restoration. A successful restore shortens the period of unavailability and lowers the risk; it does not undo the destruction that took place. It would be right if the definition contained a duration or materiality threshold, which it does not.
- D
It is a personal data breach: loss of availability qualifies, and the absence of exfiltration goes to the risk.
Correct: loss of availability falls inside the definition, and the absence of exfiltration bears on the risk assessment rather than on the classification.
Why that is the answer
The definition of a personal data breach is deliberately broad: accidental or unlawful destruction, loss, alteration, unauthorised disclosure of, or unauthorised access to personal data. Encryption by an attacker destroys availability, which lands squarely inside it even though nothing left the network. The discipline this item is testing is the separation of two questions that firms routinely merge: whether it is a breach is settled by the definition, while whether data was taken feeds the risk assessment and therefore the notification decision. Recording it as an availability incident does not avoid the classification question, it answers it wrongly.
- 04Reporting, Culture and ImprovementFoundational
Batinah Exchange, an Omani exchange house, recorded 28 breaches last year and 84 this year. The increase followed the introduction of a simple online breach form that any member of staff can submit. The chief executive wants to know whether control has deteriorated. What should compliance say first?
- A
The number is ambiguous: split it by discovery source before reading it
Correct: the total conflates how often failures occur with how well they are found, and splitting by who discovered each breach is what separates the two.
- B
Control has deteriorated: a threefold rise needs immediate escalation
This is the reflex reading in which any rise in a bad number is bad news, and it is the version most board papers adopt. It ignores the one change in the stem that is the obvious rival explanation. It would be right if the reporting route had been stable across both years, so that the count was measuring the same thing twice.
- C
Detection has improved: the new form explains the rise in recorded breaches
This is the mirror error and the more sophisticated-looking one, and it is what a compliance function reaches for when it wants the number to be good news. It credits the form with the whole increase without testing whether the underlying failure rate also moved. It would become supportable only after the split has been done and shows the growth concentrated in self-reported breaches.
- D
The rise sits within normal variation for a business of this transaction size
This asserts a range of normal that the firm has never established. Expected ranges and control charts are real tools, but you cannot appeal to one you do not have. It would be right if the firm held a history of breach counts whose measured variation was wide enough to contain a movement from 28 to 84, which a threefold jump makes unlikely.
Why that is the answer
A breach count is the product of two things, how often failures occur and how well they are found, and a single total cannot separate them. When a firm changes its reporting channel and the recorded number triples, the honest first move is to break the count down by who discovered each breach: the first line, compliance testing, internal audit, a customer or the supervisor. That split tells you whether the extra 56 came from staff now reporting what they previously absorbed, or from genuine deterioration. Answering the chief executive before you have done that split is guessing, in whichever direction you guess.
- 05Programme Governance, Policies and ControlsStandard
At Selangor Cover, a Malaysian insurance broker, the sanctions check is recorded as a first line control. It is written in the team's working guide, it sits in the team's objectives, and the system captures the evidence automatically. Over the past year every failure of the check was found by second line monitoring and none by the team itself. What does this indicate?
- A
The control is not yet first line owned: the line finds none of its own failures.
Correct: three formal indicators of ownership are present, but the decisive one, self-detection, is entirely absent over a full year.
- B
The control is genuinely first line owned, because three of the four ownership tests pass.
This scores ownership as a checklist and takes the majority, which appeals because the first three conditions are the visible, auditable ones. The error is weighting four indicators equally when one of them, whether the line finds its own failures, is the outcome the other three are only inputs to. It would be right if the four were genuinely interchangeable evidence of the same thing.
- C
The monitoring plan is testing this check too often and should reduce its sampling rate.
This reads "second line found every failure" as an artefact of testing intensity rather than as a fact about the first line, and it is where the reasoning goes when you are defending the business unit. It would be right if the concern were duplicated effort or testing cost, and if the first line were also catching failures of its own between tests.
- D
The second line has taken the control over and should now hand it back to the business.
This confuses independent testing with performing the control. The second line is detecting failures, which is its proper job, not executing the sanctions check itself. It would be right if second-line staff were carrying out the screening, which is a real and separate defect worth watching for, but nothing in the stem says they are.
Why that is the answer
Ownership of a control is not established by documentation, objectives or automatically captured evidence; those are the conditions that make ownership possible. The test that actually discriminates is who finds the failures. A first line that detects none of its own breaks across a full year is running the check as a task it has been handed rather than as a risk it owns, and a self-detection rate of zero is not a minor shortfall but the whole of the evidence. Note also what follows: the remedy is to build first-line detection, not to move the control anywhere.
- 06Programme Governance, Policies and ControlsStandard
An Emirati free-zone trading company records its daily exception report as an automated control, because the system produces the report without human intervention. A monitoring test finds the report generated every day for six months and no evidence that anybody acted on it. How should the control be classified and tested?
- A
As a preventive control, tested on coverage rules.
This comes from reaching for the most familiar label in the taxonomy without asking when the control operates. A preventive control stops the exception arising; an exception report by definition runs after the event, which is what makes it detective. It would be right if the system blocked the transaction at the point of entry rather than listing it the following day.
- B
As a corrective control, tested on the fixes made.
This is the closest wrong answer, because the end state the firm wants is that exceptions get fixed. But the control's job is to surface them; correction is what should follow from acting on it. It would be right if the system automatically repaired or reversed the exceptions it identified, rather than reporting them for a person to work.
- C
As an automated control, tested on configuration.
This is the firm's own error restated, and it is the most seductive option because the stem's reasoning looks sound: the report really is produced without human intervention. The mistake is classifying by the input rather than by the step that can fail. It would be right if the system also disposed of the exceptions, so that no human action stood between report and outcome.
- D
As a manual control, tested on the action taken.
Correct: the human step is where this control can fail, so it is manual with an automated input and the test must examine what was done with the exceptions.
Why that is the answer
Classify a control by its failure mode, not by the technology that feeds it. The system here reliably produces a report every day, so the automated part is not where the control can break; the breakable step is the person who is meant to read the report and act on what it shows. Call it automated and you will test the configuration, find it correct, and rate the control effective for six months during which nothing at all was done with its output. Naming it a manual control with an automated input forces the test that matters.
- 07Financial Crime and Data Protection FundamentalsStandard
At a Saudi leasing company, staff who form a suspicion must raise it with their line manager, who then decides whether it is passed on to the nominated officer. What is the main defect in that route?
- A
It slows the route, so reports reach the nominated officer later than they should.
Delay is real and it is what you would observe in practice, but it is the symptom rather than the defect. Fixing timeliness alone, by imposing a 24-hour deadline on managers for instance, leaves the veto entirely intact. It would be the best answer if the manager were obliged to pass every report on and the only fault were how long that took.
- B
It gives line managers a task they have had no training to perform consistently.
This diagnoses a capability gap where the problem is a conflict of interest, and it carries a dangerous implication: that training could make the veto acceptable. It cannot, because a well-trained manager with revenue targets is still conflicted. It would be right if the manager's role were to add context to a report that had to be forwarded regardless of what he thought of it.
- C
It places a commercial decision maker between the individual and the officer.
Correct: it puts a person carrying revenue and relationship responsibilities at the one point in the route where those interests must not weigh.
- D
It creates two records of one suspicion, which complicates later retrieval.
This treats a records management inconvenience as the principal defect, and it inverts the position, because a second contemporaneous record is usually a benefit here: it evidences when the individual first raised the concern. It would be the answer only to a question about audit trail quality, and even there the duplicate helps rather than hinders.
Why that is the answer
The internal reporting route exists to give an individual a path to the nominated officer that nobody can close. Placing a line manager in that path as a decision maker inserts somebody with revenue and relationship responsibilities at precisely the point where such interests must carry no weight. The defect is structural rather than operational: it is not that the route runs slowly or is poorly administered, it is that the route can be stopped by exactly the person whose judgement the design exists to bypass. A manager may properly be informed; a manager may never be the gate.
- 08Monitoring, Testing and InvestigationsStandard
At a Qatari asset manager, monitoring has reported about 1% exceptions for three years, testing each file against the team's desk instruction. A review that tested the same files against the policy found 11%. What does the gap most likely indicate?
- A
The monitoring sample is too small to detect a ten percentage point gap
This is the reflex explanation for any discrepancy between two results, and here it is arithmetically wrong: a ten point difference is visible even in modest samples, and sampling error would not push the result in the same direction consistently for three years. It would be right if the two figures were close and the difference sat within ordinary sampling variation.
- B
The monitoring reviewers apply the pass criterion less strictly than the review team
Inconsistent human application is a plausible and common failure, but it produces scatter: results that move about between periods and between testers. A stable 1% across three years is the signature of a fixed rule applied correctly, not a loose rule applied variably. It would be right if the monitoring figures were erratic and diverged unpredictably from independent review findings.
- C
The review included older files that the monitoring population never covered
This changes the population rather than the standard applied, and it is the answer of somebody reaching for a sampling frame explanation. The stem closes it off by stating that the review tested the same files. It would be right if the review had drawn on a back book, or on a period or product line that the monitoring plan had never included.
- D
The desk instruction and the policy set different standards for the same file
Correct: the two documents demand different things, so testing against the instruction returns clean results while the firm sits in breach of its own policy.
Why that is the answer
An operating effectiveness test measures compliance with whatever document the tester is holding. If the desk instruction has drifted from the policy, testing against the instruction will keep returning clean results while the firm is in breach, and that is exactly the stable, long-running pattern described here. The shape of the gap is the tell: three years of consistent 1% against 11% on the very same files is systematic, not random. Testing should run against the governing document, with the instruction itself periodically checked for alignment to the policy above it.
- 09Monitoring, Testing and InvestigationsStandard
In a test of 30 files at Muscat Leasing, an Omani leasing company, the tester marks 10 files as not applicable while working through them, and reports no exceptions in 30 files. What is the main problem with that result?
- A
The excluded files should have been counted as exceptions, not removed
This treats every exclusion as a failure, which is too crude. Genuine non-applicability exists, for instance where an attribute does not apply to a particular product variant, and forcing those files to count as defects would produce a rate that means nothing. It would be right only if the test attribute were mandatory for every file in the defined population.
- B
The effective sample is 20, and the exclusions are a second selection
Correct: only 20 files were actually tested, and the ten exclusions amount to an undocumented second selection made with the files open.
- C
The replacement files needed to restore the sample were never drawn
This spots the arithmetic consequence and proposes the obvious repair, but drawing ten replacements would restore the count while leaving the real problem untouched, because nobody defined in advance what makes a file not applicable. It would be the main issue if the exclusion criteria had been written and agreed beforehand and the only fault were the shortfall in numbers.
- D
The pass criterion defined applicability too narrowly for the population
This presupposes that a documented applicability rule existed and was miscalibrated. The stem says the tester made the calls while working through the files, so there was nothing to calibrate. It would be right if the test plan had contained a written applicability rule that was excluding files it should have retained.
Why that is the answer
Two faults compound here. The report says 30 files when 20 were tested, which overstates the assurance the work provides. More seriously, the decision to exclude ten files was taken by the tester while the review was under way, against criteria never defined or recorded beforehand, and that is a second selection sitting on top of the first. Because those calls were made with the files open, the exclusions cannot be shown to be independent of what the files contained. Applicability criteria belong in the test plan, settled before anybody opens a file.
- 10Reporting, Culture and ImprovementStandard
Luzon Remit, a Philippine remittance operator, has found a control failure affecting several thousand customer records, with an estimated remediation cost close to a year's profit. The board asks whether the supervisor should be told now. Which reasoning should drive that decision?
- A
Whether the supervisor would expect to hear this from the firm first
Correct: it applies the general duty of openness to the facts as currently known, rather than waiting for a rule or a finished plan.
- B
Whether a published rule names this category of failure for notice
This is the legalistic reading, and it appeals because it promises certainty. Rules are drafted in categories while events arrive in particulars, so the search usually finds nothing and the firm concludes it owes no duty. It would be sufficient if the notification regime were exhaustive and closed, whereas most sit alongside a general obligation to deal with the supervisor openly.
- C
Whether the remediation plan is complete, costed and formally approved
This is the most respectable-sounding wrong answer, because arriving with a plan is genuinely better than arriving without one. The error is treating plan completeness as a precondition of disclosure rather than as something you bring to a conversation already opened. It would be right if the question were what to say at the meeting, not when to ask for it.
- D
Whether the failure would be found at the next routine supervisory visit
This converts a disclosure duty into a calculation about detection risk, which is a different question and one that does not survive being said aloud in a boardroom. It is where the reasoning goes when a firm is looking for permission to stay quiet. It bears only on the consequences of non-disclosure, never on whether the duty exists.
Why that is the answer
Notification regimes differ across jurisdictions, but beneath them sits one stable test that travels: would a reasonable supervisor expect to hear this from the firm, and to hear it now? You apply that test to the facts as you currently know them, and you record the decision and its reasons whichever way it falls. Its value is that it does not depend on finding a rule that happened to anticipate your situation, and it does not let the firm postpone the conversation while it tidies up. Firms are rarely damaged by the failure itself; they are damaged by the interval between knowing and speaking.
- 11Compliance Risk AssessmentDemanding
A Saudi fintech's alert triage standard allows 12 minutes an alert. The reported median is 6 minutes and the queue clears every month. The operations manager presents this as performance ahead of standard. What does it more likely indicate?
- A
The control is running at half its designed time budget, not ahead of it.
Correct: a designed time budget is part of the control's specification, so operating at half of it is a departure from design rather than an outperformance.
- B
The standard was set too high, so it should be reset to the observed handling time.
This recalibrates the ruler to the plank, taking current practice as the definition of the right answer and leaving a measure that can never reveal a shortfall. It would be right only after a design review had established that competent triage genuinely takes six minutes, in which case the standard changes because the design changed, not because the observed number did.
- C
The alert volume has fallen, so the team now has spare capacity to redeploy.
This confuses the queue clearing with less work arriving, and it does not explain the observation in any event: lower volume gives you more time per alert, or the same time with idle capacity, but it would not halve the minutes spent on each one. It would be relevant if the reported figure were total hours worked rather than time per alert.
- D
The rules are well tuned, so the false positive rate must have fallen materially.
This leaps from a handling time to a claim about alert quality, two things the data do not connect. Better-tuned rules change the mix of alerts, and a higher proportion of true matches would take longer to work rather than less. It would be supportable only with data on alert outcomes and dispositions, which the stem does not provide.
Why that is the answer
A designed time budget forms part of a control's specification, in the same way that a sample size or an approval level does. If the design says an alert requires 12 minutes of analyst attention and the median is 6, the control is not being performed as designed, and a clearing queue tells you about throughput rather than about quality. Faster is better only if the original 12 minutes was known to be unnecessary, which is a design question the firm has not asked. Sound appetite measures are two-sided for exactly this reason: they should be capable of breach from above and from below.
- 12Monitoring, Testing and InvestigationsDemanding
Selangor Cover tests 200 policies across two channels, taking 100 from each. Direct channel population: 4,000 Direct channel defect rate: 4% Agency channel population: 1,000 Agency channel defect rate: 14% A draft board paper quotes the raw sample rate of 9% as the firm's defect rate. What is the correct firm level rate?
- A
9%, because each policy in the sample counts once towards the rate
This is the unweighted sample average, obtained by treating every tested file as one unit of the firm, and it is what the draft board paper did. The design deliberately over-sampled the smaller and worse channel, so this figure carries that distortion straight into the reported result. It would be correct only if the sample had been drawn proportionally, at 160 direct files and 40 agency files.
- B
6%, because each channel rate is weighted by its population share
Correct: weighting each channel's rate by its share of the 5,000 policies gives 6%.
- C
4%, because the direct channel holds four fifths of the population
This grasps the weighting insight and then over-applies it, discarding the smaller channel instead of giving it the 20% share it holds. Dominance reduces the agency channel's influence on the firm figure; it does not eliminate it. It would be right if the agency rate were also 4%, or if agency policies sat outside the population the board figure describes.
- D
14%, because the firm reports the higher of the two channel rates
This substitutes prudence for arithmetic: quote the worst case and nobody can accuse you of flattering the firm. But a firm level defect rate is a description of the whole book, and quoting the worst stratum misdescribes it as surely as quoting the best would. It would be the right figure only if the question asked for the rate in the higher-risk channel, which the board should also see separately.
Why that is the answer
When you sample equally from strata of unequal size, the sample stops mirroring the population and the raw combined rate is no longer the firm's rate. Here 100 files came from a 4,000-policy channel and 100 from a 1,000-policy channel, over-representing agency business fourfold, and because agency is also the worse channel the raw figure is pulled upwards. Weighting each stratum by population share gives (4,000 x 4% + 1,000 x 14%) / 5,000 = 6%. The wider lesson is that a design decision taken for sound testing reasons, over-sampling the risky channel, must be undone in the arithmetic before the number reaches a board.
About these items
These twelve items are written to the specification of the live CBA-CRP paper, and none of them will appear on one. Every item in the bank is reviewed by a named subject-matter expert and audited for answer cueing domain by domain.
