CBA-AIG · Technology
Specimen paper
Twelve examination items for the CBA Certified AI Governance Professional, with the answer and a rationale for every option.
Examiner’s note
This specimen is twelve items drawn from the live bank in the proportions the blueprint sets, with the same mix of difficulty you will meet on the day: four foundational, six standard, two demanding. Every item is a short scenario ending in a question about what a governance professional should conclude or do next, and nothing turns on recalling an article number. The commonest way to lose marks on this certification is not ignorance of the frameworks but answering a question adjacent to the one asked: reaching for the control that is usually right, treating a policy or a certificate as evidence that something was done, or classifying by technique when the question is about use. Read what the output does to the person affected, then choose.
Items
12
Domains
6
Questions in the examination
80
- 01Foundations of AI GovernanceFoundational
A customer of Banco Vinha do Norte, a Portuguese retail bank, complains that the bank's assistant quoted her the wrong overdraft fee. Data governance confirms the fee table is accurate, IT reports the platform unchanged, vendor management confirms the supplier is in contract, and compliance points to the board-approved AI policy. What does this pattern of responses indicate?
- A
The complaint is unfounded, since every relevant function has reported no exception.
This reads silence as an all-clear. None of the four controls was built to detect a wrong answer composed at inference, so their green reports are evidence about coverage, not about what happened to the customer. It would be right only if one of those functions held a control that actually tested the assistant's answers against the fee table and had run it.
- B
The supplier caused it, so third-party risk management should lead the remediation.
This is the bought-in module reflex: accountability is assumed to travel with the build. The wrong answer went to the bank's own customer, and the deployer owns that outcome. Third-party risk would lead only if the evidence pointed at a specific supplier obligation being breached, and even then the customer harm would stay with the bank.
- C
The gap sits between the disciplines, so the decision made has no owning function.
Correct: each discipline covers a slice, none covers the answer the assistant produced, so the complaint has no home.
- D
The answer was inaccurate, so data governance should reopen its lineage and quality review.
This assumes data governance reaches past the model boundary. It governs the inputs, and the fee table has already been confirmed accurate, so reopening lineage searches the one place the fault is known not to be. It would be right only if the inputs themselves had been shown to be wrong.
Why that is the answer
A governance gap rarely announces itself as a failure. It appears as complete-looking coverage that still misses the thing complained of: data governance holds the inputs, IT holds the platform, vendor management holds the contract, compliance holds the policy, and nobody holds the decision the assistant actually made. Because no control was designed to fire, every function reports green and the complaint has no owner. When the harm is real and every report is clean, look at the seams between the disciplines rather than deeper inside any one of them.
- 02Foundations of AI GovernanceStandard
At Nordvik Insurance, a Norwegian general insurer, the counter fraud lead owns the claims fraud model and also signed the approval accepting its residual risk. Flagged claims settle substantially later than unflagged ones. What is the substantive defect?
- A
Nobody outside the function weighed the delay imposed on honest claimants.
Correct: the person who gains from the flags also accepted the cost they impose, so that cost was never weighed by anyone able to refuse.
- B
Nobody recorded a challenge, which only a committee minute can evidence.
This names a forum where the requirement is a person. A committee can be the approving body, but a minute of a committee that never says no evidences challenge no better than a signature. It would be right only if the defect were the record of the decision rather than who was entitled to take it.
- C
Nobody with build knowledge approved it, as the model's authors should have.
This confuses technical competence with authority to accept risk, and hands approval to the people least able to conclude that their own work should not run. Build knowledge belongs in the evidence put before the approver. It would be right only if approval were a technical verification rather than a decision about who bears the residual harm.
- D
Nobody set a review date, which a separate approver would have imposed.
This substitutes a procedural omission for the substantive one. A review date is good discipline, but scheduling a repeat of an unweighed judgement changes nothing about the judgement. It would be right if the flagged-claim delay had been assessed once and then left unmonitored, which is not what happened here.
Why that is the answer
Accepting residual risk is a decision about who bears a cost, and it carries no weight when taken by the person who receives the benefit. Every flag helps the counter fraud lead and the delay lands on claimants who have done nothing wrong, so the trade was never made by anyone with an interest in the other side of it. Independence in approval means the approver could realistically decline, and here declining would mean refusing his own model. Note too that the harm is one nothing measures: settlement delay to honest claimants appears in no fraud metric.
- 03Risk Classification and the Regulatory LandscapeFoundational
A Chilean mining group's proposal review concludes that a proposed workforce monitoring use falls in the prohibited tier. The project team offers human sign-off, an appeal route and quarterly bias testing. How should the governance function respond?
- A
The added controls move the use into the high-risk tier, where those measures are the required set.
This treats the tiers as one sliding scale that mitigation walks you down. The three measures offered are indeed the high-risk control set, which is what makes the option attractive, but they are controls for something a firm may lawfully do. It would be right only if the original classification were wrong and the use was never prohibited.
- B
A prohibited practice is unacceptable however it is controlled, so the answer is that it stops.
Correct: the tier is defined by the practice being unacceptable whatever safeguards accompany it, so the only available control is to stop.
- C
The controls should be assessed for effectiveness before the classification is finally settled.
This applies control-effectiveness reasoning to a test that does not turn on effectiveness, implying the answer could depend on how well the safeguards work. That reasoning is exactly right in the permitted tiers, where you weigh residual risk after controls, and exactly wrong here. It would be right only if the use sat in a tier where mitigation can change the outcome.
- D
The tier should be reopened, since a use carrying an appeal route cannot meet a prohibition test.
This imports contestability, a factor that operates within the permitted tiers, into a test that does not use it. It arrives at a reasonable-sounding instinct, reopening the tier, for a reason the prohibition analysis does not recognise. It would be right only if the prohibition were defined by the absence of redress, which is not how prohibitions are drawn.
Why that is the answer
The prohibited tier is not the top of a ladder you can descend by bolting on safeguards; it is a statement that the practice itself is out of bounds. Because the classification does not turn on how good the controls are, the only control available is that the project stops, which is why prohibition analysis belongs at proposal stage before money and reputations are committed. The offer of sign-off, appeal and bias testing is a sincere but category-confused response: those are the measures that make a permitted high-risk use acceptable. You have to establish that the firm may do the thing at all before you start designing how to do it safely.
- 04Risk Classification and the Regulatory LandscapeStandard
An Italian energy retailer buys a customer-contact and forecasting system, fine-tunes it on its own customer records, and puts the result into service under its own brand. What does that most likely change?
- A
Its tier, which rises because a fine-tuned system carries more uncertainty than a bought one.
This moves the classification when the use has not moved. Tier follows what the output does to a person, and added technical uncertainty is a matter for testing and monitoring, not for reclassification. It would be right only if the fine-tuning had also changed what the system decides, or about whom.
- B
Its vendor's role, which may narrow to component supply once the retailer retrains it.
This assumes roles trade off, so that duties the retailer picks up must have been dropped by somebody. The vendor keeps its own obligations for what it placed on the market; the retailer's obligations are added, not transferred. It would be right only if downstream modification extinguished the original provider's position, which is not how role allocation works.
- C
Its role, which can move from deployer to provider and bring the heavier duties with it.
Correct: rebranding, substantial modification and fine-tuning on your own data are the recognised routes by which a deployer takes on provider duties.
- D
Nothing in its duties, since the use and the population the system meets are both unchanged.
This notices something true, that the use has not changed, and then assumes duties follow the tier alone. Obligations are allocated by role as well as by classification, and the retailer has just done three things that alter its role. It would be right only if role were fixed at the moment of purchase.
Why that is the answer
Two independent questions decide what a firm owes: what the system does to people, which sets the tier, and what the firm does with the system, which sets the role. Here the tier is untouched because the use is unchanged, but the retailer has fine-tuned the model on its own data and put it out under its own brand, which are precisely the acts that convert a deployer into a provider. Firms cross that line by ordinary commercial instinct, branding a product and adapting it to fit, usually without anyone recording that the obligations have changed. The governance question at procurement is therefore not only what tier is this, but what will we do to it after we buy it.
- 05Risk Classification and the Regulatory LandscapeDemanding
A Portuguese bank's counsel notes that an EU high-risk list names creditworthiness assessment but not arrears-collection prioritisation, and concludes that the collections model is therefore not high risk. Why does that conclusion not follow?
- A
The list is indicative, so a use closely resembling a listed one is treated as listed as well.
This converts a defined, closed category into an open one by resemblance, which would make the boundary of the list unknowable in advance. It would be right only if the instrument itself declared the list non-exhaustive, and the remedy for a use that is genuinely dangerous but unlisted is the effect-based assessment, not an elastic reading of the list.
- B
Absence from one instrument's list is not a finding on the factors or on the rules that do bind.
Correct: the listing conclusion answers one narrow question and leaves the effect-based factors and the binding conduct, data protection and equality rules entirely untouched.
- C
Collections prioritisation is a form of creditworthiness assessment and so sits on the list.
This stretches a listed use by analogy. Creditworthiness assessment decides whether credit is extended; prioritising which existing borrowers in arrears get contacted and how is a different operation on a different population. It is counsel's own undefended scope argument running in the opposite direction, and would be right only if the model were in fact scoring creditworthiness.
- D
The list reaches providers only, so a bank deploying such a model falls outside it in any event.
This mixes up role with scope, and would prove far too much: on this reading no deployer of any listed system would ever be caught. Deployers carry their own set of duties for listed high-risk systems. It would be right only if the instrument placed obligations on providers alone.
Why that is the answer
Not on the list and not high risk are different propositions, and a firm that treats them as identical governs whatever a statute happens to name and nothing else. The list settles one question under one instrument; it does not measure what the model does to a customer in arrears, and it does not repeal the conduct rules, data protection duties and equality obligations that already bind the bank whatever the list says. The defensible record states the listing conclusion, then separately records the effect-based assessment and the other regimes considered. Counsel's answer is not wrong so much as incomplete in a way that leaves the internal tier undecided.
- 06Governance Frameworks and Management SystemsFoundational
A Canadian credit union writes its monitoring thresholds and its list of required validation tests into the board-approved AI policy. What should it expect?
- A
Thresholds will not change, because every change needs a board decision.
Correct: volatile operational content placed in a board document inherits the board's change cadence and therefore stops moving.
- B
Thresholds will be unenforceable, because the board cannot audit them.
This confuses approval level with binding force. Policy content binds, and nothing requires the approving body to be the one that tests compliance; the second line audits against it. It would be right only if a requirement could be enforced solely by whoever approved it.
- C
Thresholds will be read as guidance, because policies state intent only.
This takes a sound description of what policies should contain and turns it into a rule about what they can contain. The difficulty here is the opposite: the policy carries operational detail and that detail binds. It would be right only if the hierarchy stripped binding force from misplaced content, which it does not.
- D
Thresholds will be duplicated, because standards will restate them anyway.
This identifies a tidiness cost rather than the control failure. Duplication does happen and creates conflicting versions, but it is a drafting nuisance, not the paralysis the hierarchy exists to prevent. It would be right if the question asked what the main documentation untidiness would be.
Why that is the answer
A document hierarchy sorts content by two things: how often it needs to change, and who has to approve a change. A policy sits at the top because it states intent and is approved rarely at board level, so anything written into it acquires the board's cadence. Monitoring thresholds and test lists are exactly the content that should move as the firm learns, and once frozen they drift out of line with practice. The predictable end state is a policy mandating a control the credit union no longer operates, which is worse than having no threshold written down at all. Put the intent in the policy and the numbers in the standard beneath it.
- 07Governance Frameworks and Management SystemsStandard
A Polish logistics firm receives a supplier's ISO/IEC 42001 certificate whose scope reads "development and operation of AI systems supporting warehouse robotics at the Gdansk site". The product being bought is the supplier's route pricing model. What does the certificate support?
- A
Nothing about the route pricing model, which sits outside the declared scope.
Correct: the assessment was performed within a scope the supplier itself declared, and that scope excludes the product being bought.
- B
Reasonable confidence in the model, since one management system covers the firm.
This assumes conformity travels across the whole organisation, which is precisely the assumption a narrowly drawn scope statement exploits. Certifying one site's robotics work is a much smaller undertaking than certifying everything the supplier builds. It would be right only if the declared scope named AI development and operation across the supplier as a whole.
- C
Confirmation that the supplier's high-risk products have passed conformity assessment.
This conflates certification of an organisation against a voluntary management standard with conformity assessment of a product against legal requirements. They are different instruments with different assessors and different questions. It would be right only if the supplier had produced conformity assessment documentation, which is what you should ask for if the product is in fact high risk.
- D
Evidence that the route pricing model was assessed by an accredited certification body.
This reads a management system certificate as a product test. Even for a system squarely inside the declared scope, the accredited body examines how the organisation governs its AI work, not whether a particular model is accurate or fair. It would be right only if certification bodies validated individual models, which they do not.
Why that is the answer
A certificate is a statement that an accredited body found an organisation's AI management system conforming to the standard within a scope the organisation wrote for itself. That makes the scope statement, and the statement of applicability behind it, far more informative than the certificate front page, and both are where narrowing happens quietly. Here the scope names warehouse robotics at one site, so it says nothing whatever about a route pricing model sold from elsewhere in the business. The practical move is to stop treating the certificate as an answer and ask for what you originally wanted: validation evidence, test results on populations like yours, and notice of material model change.
- 08Documentation, Transparency and AccountabilityStandard
A Gulf telecoms operator's retention model restricts a business customer's renewal terms. The customer asks why. The analytics team sends the model's global feature importance ranking, in which the highest ranked feature contributed nothing to this customer's score. What is the principal objection to sending it?
- A
It exposes weightings that the supplier regards as proprietary.
This treats a disclosure limit as the defect. Confidentiality is a real constraint on how much you can say, but the artefact would still describe the wrong subject if the operator were entirely free to publish it. It would be right if the objection under discussion were what may be shared, rather than whether what was shared answers the customer's question.
- B
It is accurate about the model and false about this customer's case.
Correct: a global artefact was sent in answer to a local question, so the response misstates the very case it purports to explain.
- C
It uses technical vocabulary the customer is likely to misread.
This assumes the fault is readability, which is the most common reflex when an explanation lands badly. A plainer, friendlier rendering of the same global ranking is still about the model in general and still wrong about her score. It would be right if the content answered her question and only the presentation defeated her.
- D
It omits the fairness testing that supports the model's use.
This substitutes population-level evidence for an individual explanation. Fairness testing answers whether groups are treated differently, which is a legitimate question but not the one asked. It would be right if the customer had alleged discrimination and sought assurance about outcomes across comparable customers.
Why that is the answer
Explanations are answers to questions, and the question determines which artefact is responsive. This customer asked a local question, why my terms, and received a global one, what generally drives this model. Because the top-ranked feature contributed nothing to her score, the document sent is simultaneously true about the model and false about her, which is worse than sending nothing: it has the form of a formal answer and so closes the exchange. What she is owed is a local attribution for her own case, expressed in terms she can act on, together with the route to contest it.
- 09Documentation, Transparency and AccountabilityDemanding
A Portuguese bank maintains that its account closure decisions are not solely automated, because a reviewer confirms each recommendation. The reviewer handles several hundred cases a shift and cannot reverse a recommendation without a manager's agreement. What does the bank need in order to sustain that position?
- A
Evidence of time taken, information available and authority to overturn.
Correct: those are the observable properties that distinguish a decision from a countersignature, and they can be tested case by case.
- B
A policy requiring that every recommendation is reviewed by a person.
This offers the existence of a rule as evidence that the rule operates, which is the single most common failure in oversight evidence. Nobody disputes that the bank requires review; the dispute is about what the review consists of. It would be right only if the challenge were that no such requirement existed.
- C
An attestation from reviewers that they applied their own judgement.
This substitutes assertion for operating evidence, and asks the people whose conduct is in question to certify it. The arithmetic of the shift already contradicts the assertion. It would carry weight only as corroboration of an operating record showing time taken and overturns actually made.
- D
A record of the qualifications and training each reviewer completed.
This evidences competence, where the missing ingredients are opportunity and authority. A well-trained reviewer with a minute a case and no power to reverse still does not decide the outcome. It would be right if the objection were that reviewers lacked the knowledge to understand what they were confirming.
Why that is the answer
If your defence for not giving an explanation is that a human took the decision, then the human deciding is the fact you must evidence, and you must evidence it for the individual case rather than for the process in the abstract. Time taken, the information actually in front of the reviewer, the overturn rate and the authority to overturn are the properties that make involvement meaningful and, importantly, that make the claim falsifiable. Several hundred cases a shift with no power to reverse alone points hard the other way. The honest options are to resize the review so the evidence can exist, or to accept the decision is automated and provide the explanation and review rights that follow.
- 10Bias, Robustness and Testing ConceptsFoundational
An Australian utility's restoration-cost model receives the same mix of fault descriptions, in the same volumes, as it did at build. Contractor rates and parts prices have risen sharply since then, and the model's cost estimates now run consistently low. What does this describe?
- A
Distribution shift, because the input population has moved since build
This assumes any post-build decay must be a change in the cases arriving, which is the default explanation and the one the stem rules out by holding the mix and volumes constant. It would be right if the utility had started seeing different faults, or the same faults in different proportions, from those the model was built on.
- B
Adversarial input, because outside parties are manipulating submissions
This reaches for manipulation with no opponent in the facts. Adversarial failure needs a party deliberately shaping inputs to move the output in their favour, and note that the estimates run low, which is the opposite of what anyone gaming a cost model would engineer. It would be right if, say, contractors had learned to word descriptions so as to inflate estimates.
- C
Edge case failure, because rare inputs fall outside the tested range
An edge case is a legitimate but unusual input the model was never really tested on. Here the failure runs across ordinary, well-represented faults and is consistent in one direction, which is a pattern, not an exception. It would be right if only the rare or extreme fault types were being mis-estimated.
- D
Concept drift, because the input-to-outcome relationship has moved
Correct: the inputs are stable and what those inputs imply about true cost has changed underneath the model.
Why that is the answer
Two different things can move after a model is built: the cases that arrive, and what those cases mean. Distribution shift is the first, concept drift is the second, and telling them apart decides what monitoring will catch the problem. Here the arriving population is explicitly unchanged while the relationship between a fault description and its real cost has moved with rates and prices, so input monitoring will show reassuring stability while the answers get steadily worse. Concept drift is caught only by measuring outcomes, comparing estimates against settled costs, which is the check the utility evidently does not run.
- 11Bias, Robustness and Testing ConceptsStandard
A vendor tells a Portuguese bank that its collections-prioritisation model equalises error rates across customer segments, keeps a score meaning the same thing in each segment, and flags each segment at the same rate. Confirmed default rates differ materially between the segments. What should the bank conclude?
- A
The claims hold only if the segments were sampled at equal size.
This mistakes an arithmetic impossibility for a sampling artefact that better balanced data would cure. Sample sizes affect how precisely you can estimate each quantity; they do not affect whether the three can co-exist, which turns on the base rates differing. It would be right if the conflict arose from unequal sampling rather than from unequal underlying default rates.
- B
The three claims cannot hold together where base rates differ.
Correct: with imperfect prediction and different base rates, equal error rates, equal score meaning and equal flag rates are mutually incompatible.
- C
The claims are credible given the vendor's larger validation set.
This treats volume of evidence as a proxy for quality, and no quantity of data can make three incompatible statements true at once. A large validation set on the vendor's own population also says nothing about how the model behaves on this bank's segments. It would be right only if the claims were compatible and the question were how confident you can be in the estimates.
- D
The claims need a subgroup breakdown before any assessment is made.
This applies the standard and usually sensible reflex, ask for more evidence, to a claim you can already reject on its face. Requesting breakdowns here delays the real conversation and implicitly concedes the claims might be true. It would be right if the vendor had claimed a single fairness criterion and you needed to verify it.
Why that is the answer
Where base rates genuinely differ between groups and no model predicts perfectly, equalising error rates, keeping the score calibrated so it means the same thing in each group, and flagging each group at the same rate cannot all be achieved together. This is arithmetic, not a matter of vendor effort or data quality, which is why a claim to have done all three tells you something about the vendor rather than about the model. The productive response is not to ask for more evidence but to ask which criterion was chosen, on what reasoning, and which customers absorb the error that choice leaves behind. Fairness work in practice is the defence of a choice among incompatible definitions.
- 12Incident Handling and Organisational StructuresStandard
A Portuguese bank's identity-verification model wrongly flagged eleven customers as suspected fraud; their accounts were frozen and each was told fraud was suspected. The triage lead scores scale 1, harm 4, reversibility 4 and exposure ongoing, then proposes a medium rating. Which rating and reasoning should stand?
- A
Medium, because the average of the four axis scores lands in mid-range.
This applies an averaging rule the scoring method deliberately rejects, letting a low scale score dilute the two axes that carry the seriousness here. Averaging is precisely how a small affected population conceals an irreversible harm. It would be right only if severity were defined as the mean of the axes, which would defeat the purpose of scoring them separately.
- B
Medium, because the accounts can be unfrozen and the charges refunded.
This counts only the financial layer as reversal. Restoring an account and refunding money does not withdraw an accusation of fraud already made to a named person, and it is that accusation the reversibility score of 4 is recording. It would be right if the only harm suffered were temporary loss of access to funds.
- C
Critical, because the top axis governs and the harm is irreversible.
Correct: severity takes the highest single axis, and an unretractable accusation with exposure still running is a critical event.
- D
Low, because eleven customers sits below the scale band for customers.
This scores scale alone, an approach usually inherited from IT incident scales built around outage size and users affected. On that scale almost every AI harm will look small, because AI events concentrate severe harm on few people. It would be right only if the number affected were the only axis that mattered.
Why that is the answer
Severity is set by the highest single axis rather than by an average, and the reason is visible in this case: eleven customers is a small number, and telling someone the bank suspects them of fraud cannot be taken back. Harm at 4 and reversibility at 4 are therefore the axes that govern, and ongoing exposure pushes the rating up rather than down. Averaging would produce a comfortable medium and, with it, a medium response tempo for an event that is still running and still doing permanent damage. If your scale routinely files AI incidents as medium, check whether it was designed for outages and is measuring the wrong thing.
About these items
These twelve items are written to the specification of the live CBA-AIG paper, and none of them will appear on one. Every item in the bank is reviewed by a named subject-matter expert and audited for answer cueing domain by domain.
