CBA-GAI · Technology
Specimen paper
Twelve examination items for the CBA Certified Generative AI Practitioner, with the answer and a rationale for every option.
Examiner’s note
This specimen shows you what the paper is like. Every item sets a short workplace scene and asks for the response a competent practitioner would give; nothing is tested by recall. The wrong options are all defensible-sounding, and each comes from an error somebody has genuinely made. The commonest way to lose marks here is to reach for a familiar remedy without asking what has failed: fine-tuning offered against a facts problem, added emphasis where the model has no source, another reviewer where the process design is at fault, or a small difference on a small sample read as a result. Expect four items you should settle quickly, six that turn on naming the mechanism, and two that need arithmetic or a judgement about evidence.
Items
12
Domains
6
Questions in the examination
80
- 01Generative Model FoundationsFoundational
A Japanese ceramics manufacturer asks an assistant, with no document attached, for the number of the standard governing lead release from glazes and the clause that sets the limit. The reply names both, flatly and without hedging. How should the quality manager read it?
- A
High risk: a standard reference is a precise string with little redundancy.
Correct: With no document supplied, the reference is produced as a likely continuation, and a near miss reads exactly like a hit.
- B
Low risk: standards of this kind are widely published and consistently cited.
This comes from reasoning about the topic when the question is about the string. Wide publication makes it likely the subject is well represented in training text; it does nothing for these particular digits, which have to be produced one token at a time. It would be sound only if the model held an index of standards it could look up and cite, which is not the mechanism.
- C
Low risk: an uncertain reference would have arrived hedged rather than flat.
This treats register as a confidence readout. The model writes its least supported sentence in the same voice as its best supported one, so hedging is a feature of the genre rather than a measurement of support. For this to be right, the system would need an internal estimate of how well founded the answer is and an instruction to express it, and it has neither.
- D
High risk: the number will be right and only the clause wording invented.
This is the partial-risk intuition: the hard identifier feels anchored while the surrounding prose feels free to drift. Both are guesses in the shape of facts, and the number, with no redundancy to absorb a slip, is the more fragile of the two. It would hold only if the standard number had been supplied in the request and merely the clause left to memory.
Why that is the answer
Identifiers and clause references are high-precision, low-redundancy strings: alter one character and the reference points somewhere else, or nowhere, while reading exactly as it did before. Ordinary prose survives small errors because the surrounding context carries the meaning, and a reference has no such context to lean on. With nothing attached to the request, the reply is generated from priors rather than read from a record, so a plausible near miss is indistinguishable from a hit. The control is to supply the standard, or to decline the question, rather than to judge the answer by how confidently it arrives.
- 02Generative Model FoundationsDemanding
Nile Delta Pharma prices the risk of assistant-drafted product leaflets as volume times error rate times cost per error. The errors most likely concern the current storage guideline, which the model states from its own memory rather than from a supplied source. Why does the figure understate the exposure?
- A
One wrong storage figure costs more to put right than the average error in the set.
This adjusts the cost-per-error term while leaving the shape of the calculation intact. It is a fair observation about severity, and it still treats each leaflet as carrying its own independent draw of good or bad luck. It would be the right answer if the errors genuinely were independent and had simply been priced too cheaply.
- B
One wrong belief about the guideline repeats across every leaflet drafted this way.
Correct: The defect is a single belief reproduced across the catalogue, so the failures are correlated and surface together rather than averaging out.
- C
The error rate came from a sample too small to give a usable average in the first place.
This reads the problem as one of statistical precision. A wider interval around the rate would matter if the underlying model of risk were sound, but here an independence assumption is being applied to events that move together, so a better-measured rate would still be multiplied in the wrong way. It would be the answer if the calculation were structurally right and only imprecise.
- D
The calculation leaves out what it costs to generate the leaflets in the first place.
This reaches for a missing cost line, which is a habit worth having in the wrong place here. Generation is close to free at leaflet volumes, and in any case it belongs to the business case rather than to the exposure. It would be the answer if the question asked what the programme costs rather than what it risks.
Why that is the answer
Volume times rate times cost assumes that errors are independent, drawn afresh for each item, so that good and bad luck average out over a catalogue. An error whose source is the model's own memory of a guideline does not behave that way: the belief is a single fixed thing, so every leaflet drafted by the same route carries the same defect, and they are discovered together. That changes the consequence as well as the count, because one correlated failure means one recall, one regulatory finding, one catalogue-wide correction. Where a class of claim comes from priors rather than from a supplied document, price the class rather than the average item.
- 03Prompt Design and IterationFoundational
A Ghanaian microfinance lender, Accra Thrift Finance, is deciding how many sample letters to embed in a prompt. Testing on a fixed set gives conformance of 24% at zero examples, 71% at three, 73% at five and 70% at eight, with reuse of example phrasing rising steadily across the runs. Which choice fits?
- A
Three, chosen to span the loan types the prompt will meet.
Correct: The curve flattens after three, so the remaining decision is which loan types those three examples cover.
- B
Eight, since reused phrasing is a matter an editor can fix.
This dismisses leakage as a tidy-up. Identical phrasing across hundreds of member letters is itself the defect, and what leaks from examples is not only phrasing but specifics, which travel as facts into letters they do not belong in. It would be defensible if eight examples bought a real gain in conformance, and on this evidence they buy slightly less than three.
- C
Five, since it records the highest conformance of the four.
This reads a two-point difference on one fixed set as a result, when two points is a couple of items and well inside ordinary noise. Choosing the top number in a table is how noise gets promoted into a decision that then stays in the prompt. It would be right only if the gap were larger than the set's margin, which nothing here shows.
- D
Zero, since the register can be described in more detail.
This assumes description substitutes for demonstration, when the measured difference between them is 47 points of conformance. A described register is interpreted a dozen ways; a shown one is not. Going without examples makes sense only where their specifics cannot be made unusable and the leakage cannot be tolerated at all.
Why that is the answer
Almost the whole benefit of examples sits between zero and three, and after that the conformance curve is flat while the costs of examples keep rising: phrase mimicry, borrowed specifics and a longer prompt competing with the clauses that carry the risk. When a count stops buying anything, the question changes from how many to which, because examples define the region the system generalises from and a set clustered on one loan type leaves a deficit everywhere it never looked. Read the table as a curve with a knee rather than as a league table with a winner.
- 04Prompt Design and IterationStandard
A Turkish tour operator, Anadolu Seyahat, finds that generated itinerary copy states opening hours and entry fees the supplied venue record does not mention. The digital manager proposes fine-tuning a model on five years of past itineraries. Why will that not address it?
- A
Training on past itineraries teaches how hours are phrased, not what they now are.
Correct: Training moves the form of the claim and leaves the value itself to memory, which is where the invention began.
- B
Training on past itineraries needs many more examples than five years can supply.
This reads a category error as a quantity shortfall. Hundreds to thousands of examples is an ordinary fine-tuning scale, so volume is not the obstacle, and a larger corpus of itineraries would only teach the same confident phrasing more thoroughly. It would be the answer if fine-tuning could install current facts and merely lacked the material.
- C
Training on past itineraries costs more per item than the present route does.
This converts a correctness question into a budget one, and it also has the economics backwards, since a fine-tuned route is typically cheaper per item and expensive as a fixed setup. A cheaper route to a wrong entry fee is not an improvement. Cost would decide only between two routes that both produced correct copy.
- D
Training on past itineraries requires a provider agreement not yet in place.
This mistakes a procedural obstacle for a reason. Sign the agreement tomorrow and the copy still states opening hours the venue record does not contain. A missing agreement is a real objection only where the approach would otherwise work.
Why that is the answer
Fine-tuning is a lever on form and disposition: how an itinerary sounds, how it is laid out, what it habitually mentions. It installs no current opening hour, and an opening hour is a fact that moves, so even a perfect training run is stale by the following season and carries no passage anyone can check. The actual cause is that the prompt asked for a claim its inputs never supplied, so the model produced the hours such a venue usually keeps. The lever for facts is in-context: supply the venue record, require the values to be taken from it, and forbid stating any value the record omits.
- 05Prompt Design and IterationStandard
An Indian pharmaceutical distributor, Sahyadri Pharma, extracts six fields from 1,500 archived supply agreements. Every field comes back filled, and the contracts database moves from 68% to 100% complete in an afternoon. Why should the operations director be alarmed?
- A
Fields the agreements never stated now hold plausible values that read as records.
Correct: A gap filled with the industry-typical value is indistinguishable from a value that was read, and the jump to complete is the signature of it.
- B
Fields the agreements stated in full may have kept the wrong wording on transfer.
This is an ordinary transcription worry brought to a structural problem. Wording errors sit on values that exist and can be caught by comparing the entry with the source, and they explain nothing about where the missing 32 per cent came from. It would be the main concern if completeness had stayed near 68 per cent.
- C
Fields the agreements stated ambiguously will have been tidied into a house format.
This names a genuine hazard, since normalising an ambiguous value destroys the ambiguity a specialist needs to see, and it describes a smaller defect on values that at least came from the page. It is worth controlling separately with a quoted phrase behind each value. It would be the strongest objection if nothing had been invented.
- D
Fields the agreements stated twice may have been merged into a single entry.
This is a data-cleaning artefact offered where the alarming event is fabrication. A merged duplicate leaves a visible, checkable trace in the source; an invented indemnity cap leaves none. It would be the answer if the extraction had reported fewer values than the agreements contain rather than more.
Why that is the answer
A completeness figure that jumps to 100 per cent is evidence about the extraction, not about the agreements. A system with no vocabulary for absence has only one move when a field is not stated, which is to supply the value such an agreement usually carries, and that value enters the database looking exactly like one that was read off the page. The harm compounds with time: three years on, somebody relies on the field with no way to separate record from guess. The controls are a permitted NOT STATED value, a separate UNCLEAR value for illegible material, a ban on inferring one field from another, and a quoted phrase behind every value returned.
- 06Retrieval and Grounded WorkflowsStandard
A Czech insurance broker samples 40 assistant answers to adviser queries, confirms that every one carries a policy-clause reference, and reports no defects to the board. Six of those answers are later found to state the wrong excess. What was wrong with the check?
- A
It confirmed a reference was present, not that it supported the statement
Correct: Presence of a citation is guaranteed by the instruction to cite, so the check tested format compliance and never tested support.
- B
It sampled only 40 answers, too few to establish a defect rate with confidence
This reads a validity problem as a precision problem. Four hundred answers checked the same way would have passed the same wrong excesses, because the test cannot see them at any sample size. Sample size becomes the live question only once the check is measuring the property you care about.
- C
It used a single reviewer, so a second reader was needed on each answer
This proposes reviewer capacity for a check no number of readers can rescue: a second person applying the same test would also see a reference and stop there. Adding a reader helps where verdicts require judgement and reviewers disagree, which is not what happened here.
- D
It ran after the answers were sent rather than before they left the desk
This treats timing as the fault, and moving the same test earlier catches the same nothing. Earlier checking matters when the check detects the defect and merely detects it too late, so this would be the answer if the reviewers had been reading the clauses all along.
Why that is the answer
A check is only worth what it actually tests. Requiring a citation and then confirming that a citation is present tests whether the instruction was obeyed, and every answer produced under that instruction passes it. The failure that survives is plausible attribution: a genuine clause attached to a claim it does not support, which reads well and defeats any test that looks for the presence of a reference rather than at its content. The control that works is to open the cited clause and read it against the sentence it is attached to, at a defined sampling rate, so that the reference is examined rather than counted.
- 07Retrieval and Grounded WorkflowsDemanding
A Malaysian hotel group's grounded assistant answers rate and group-booking questions from terms held separately for each of its five properties. Raising k from 5 to 15 lifted hit rate from 0.79 to 0.91, while answer accuracy fell from 0.70 to 0.63 and confidently wrong answers rose. What does this indicate?
- A
The added passages are neighbouring property terms that contradict the right one
Correct: The next nearest passages on this corpus are the other properties' terms, which read like the right passage and say something different.
- B
The context window is now exceeded, so the earliest passages are being dropped
This reaches for truncation whenever more text is supplied, which is a sound instinct in the wrong place: fifteen passages of booking terms sit comfortably inside a modern window, and material that never arrived would tend to depress hit rate rather than lift it. It would be worth investigating if the figures showed retrieval improving while the supplied text demonstrably failed to reach the model.
- C
The embedding step has degraded and now ranks unrelated passages too highly
This blames the retriever at the very moment the retrieval measure improved. Ranking that had gone wrong would show as a falling hit rate, not a rise from 0.79 to 0.91. It would be the answer if both columns had moved down together.
- D
Hit rate and accuracy were scored on different question sets, so both mislead
This is a methodological objection worth putting to any table of two measures, and the stem states that one gold set produced both columns. Raising it here abandons a diagnosis the figures fully support in favour of doubting the figures. It would be right if the two rates came from separate evaluations run at different times.
Why that is the answer
Hit rate and answer accuracy measure different things, and on a corpus of near-identical documents raising k moves them in opposite directions. More passages give more chances of including the right one, so hit rate climbs almost by construction. What arrives alongside are the next nearest neighbours, and where five properties keep separately worded versions of the same terms, those neighbours are confidently phrased, structurally identical and numerically different. The correct passage now sits in a crowd of plausible rivals with nothing marking which property each belongs to. Choose k on answer accuracy rather than hit rate, and label every passage with the property it governs.
- 08Output Evaluation and Quality ControlStandard
A Norwegian seafood exporter audits 60 randomly drawn generated product sheets against the source specification and finds no defects at all. The head of quality proposes reporting the process as defect free. What is the sound reading of that result?
- A
The true rate is under about 5 per cent, which is not the same as zero.
Correct: Zero defects in 60 random draws bounds the rate at roughly three divided by 60, and that bound is the honest report.
- B
The true rate is zero to the precision that the audit was designed to give.
This reads an observed count as a measured level, dressed in the language of precision. No finite sample resolves a rate to zero, and the phrase invites the board to hear the word zero and stop listening. It would be right only if every one of the sheets produced had been checked.
- C
The true rate is under about 1.7 per cent, one item being the smallest unit.
This confuses resolution with confidence: one item in 60 is the smallest defect the audit could have counted, which is not the same as the highest rate consistent with having counted none. The arithmetic is real, which is what makes it attractive, and it understates the bound by about a factor of three.
- D
The true rate cannot be bounded at all until some defect has been observed.
This assumes a zero count carries no information, when it carries a great deal: it rules out the high rates that would almost certainly have shown up in 60 draws. Refusing to bound the rate also leaves the exporter with nothing to report. It would be right if the 60 sheets had been chosen rather than drawn at random.
Why that is the answer
A clean sample bounds a rate; it does not measure one. The working rule of thumb puts the upper 95 per cent limit near three divided by the sample size, so 60 sheets with nothing found are consistent with a defect rate anywhere up to roughly 5 per cent, which across a year of product sheets is a substantial number of wrong statements reaching customers. Nothing was seen, and that is a fact about what 60 draws can see rather than about the process. If the exporter needs to claim a rate below 1 per cent, that costs several hundred items, or an automatic check applied to every sheet.
- 09Output Evaluation and Quality ControlStandard
A Singaporean logistics firm keeps a golden set of 80 customs enquiries for regression testing. To lift performance, someone adds the twelve hardest golden-set enquiries and their model answers to the prompt as worked examples. The next run scores 96 per cent, against 78 before. What has happened?
- A
The set now measures the system reproducing what the prompt showed it.
Correct: The system was shown twelve of the items and their answers before being scored on them, so the score reflects the prompt rather than the capability.
- B
The score is genuine: strong worked examples improved the assistant.
This trusts a score whose inputs the system has already been given, and it is tempting precisely because examples do lift conformance for real. The gain has to be demonstrated on items the prompt has never shown it. It would be right if the twelve worked examples had been drawn from enquiries held outside the evaluation pool.
- C
The set has aged, so those twelve enquiries no longer match live traffic.
This applies a genuine maintenance rule to the wrong symptom. Ageing shows as a slow divergence between set and live traffic, not as an eighteen-point jump in the single run after the prompt changed. It would be the diagnosis if the score had drifted downward over several quarters with the prompt untouched.
- D
The score is inflated because twelve of the eighty items count twice.
This reads contamination as an arithmetic slip in scoring. Nothing was counted twice: the twelve items were still scored once each, and they were answered from the prompt rather than from the system's handling of customs enquiries. It would be the answer if the twelve had been appended to the set instead of to the prompt.
Why that is the answer
A regression set is worth exactly what its independence is worth. Once its items and their answers appear in the prompt, the run on those items is a memory test: the number moves and nothing has changed for an enquiry that arrives tomorrow, which is the one thing the set existed to predict. The discipline this protects is that demonstration examples and evaluation items come from separate pools, and the evaluation pool stays unseen. When a score jumps sharply straight after a change to the prompt's examples, your first question should be what the set and the prompt now have in common.
- 10Multimodal Creation and EditingFoundational
Veldhuis en Roos, a Rotterdam architecture practice, is generating atmosphere images for a competition board. The prompt ends: 'Ground floor shopfronts, no text, no signage, no lettering anywhere.' Renders keep returning shopfronts with lettering across the glass. What is the most reliable fix?
- A
Move the prohibition to the front, where terms carry more weight.
This applies a true rule about term weighting to a clause whose mechanism is the problem. Early terms do carry more weight, which here means the concepts of text and signage are raised earlier and more strongly, so the change can make matters worse. It would be the right move for a positive term that was being under-weighted at the end of a long prompt.
- B
Extend it to cover numerals, house numbers and street names too.
This reads the failure as under-specification, as though the renders were returning lettering nobody had thought to exclude. Each addition names another kind of writing and strengthens the association the prompt is trying to break. It would be sound if prohibitions worked as filters, in which case a wider filter would catch more.
- C
Render a larger batch and keep only the frames without lettering.
This buys a way past the cause with selection, and it does produce usable frames, at the price of every discarded render and every minute spent reviewing them. The practice is back at the same starting point on the next competition board. It is the reasonable answer only when a fix is needed today and the prompt cannot be revised.
- D
Describe the glass as plain and unbroken, with no prohibition at all.
Correct: A positively specified surface gives the render a target that contains no writing, rather than a rule about writing to be interpreted.
Why that is the answer
A prompt to an image generator is a description of a target rather than a set of rules to obey. Naming text three times puts text firmly into that description, and the system has no dependable operator for not: it denoises towards images matching the words it was given, and those words include text, signage and lettering. Positive specification works because it removes the reason for writing to appear at all, describing a surface that would not carry any. The general habit to build is to say what should be in the frame rather than what should be kept out of it.
- 11Safe Use, Data Care and Work IntegrationFoundational
A partner at Trelane Fischer LLP in Halifax wants to summarise a client's unredacted merger documents. The firm's workspace agreement bars training on inputs, sets a thirty-day retention window and names the provider as a processor. The partner concludes the summary is permitted. What is the correct assessment?
- A
It is not permitted: the client's confidence binds regardless of tier
Correct: The duty runs to the client, who has no relationship with the provider, and nothing the firm agrees with a provider can discharge it.
- B
It is permitted, since a processing agreement covers material held for a client
This reads a data processing agreement as though it settled confidentiality. Such an agreement governs how the provider handles material on the firm's instructions; it says nothing about whether the firm was entitled to disclose that material in the first place. It would be right if the client had agreed to this route, which is what terms of engagement or specific instructions can provide.
- C
It is not permitted until retention for the tenant is reduced to zero
This reaches the right verdict by the wrong route, and the route matters because it implies a console setting could cure the objection. Zero retention still discloses the client's documents to a party the client never approved. It would be the answer if retention, rather than disclosure, were the thing the client had objected to.
- D
It is permitted once the file is marked confidential in the document system
This mistakes an internal classification for an external permission. A confidentiality label controls who inside the firm may open the file and has no effect whatever on what may be sent outside it. It would be relevant to a question about internal access, which is not the question the partner asked.
Why that is the answer
A workspace agreement is a promise a provider makes to your firm. The obligation in play here is a promise your firm made to its client, who is not a party to that agreement and consented to nothing. A contract between two parties cannot discharge a duty owed to a third, so the tier, the training setting and the retention window are all careful answers to a different question. What would change the position is the client's informed agreement, or engagement terms that already permit it. This is the most reliable trap in the area, because good procurement genuinely feels like permission.
- 12Safe Use, Data Care and Work IntegrationStandard
Torvia Rail's programme report claims PLN 310,000 of annual savings from document drafting. Headcount is unchanged, the freed hours have not been assigned to new work, and no budget line has fallen. What should the report say instead?
- A
Report the hours as capacity released, not as money saved
Correct: With no change to headcount, budget or workload, the hours are released capacity, and the report should say where they sit and what they are for.
- B
Report the saving net of the licence and training costs incurred
This refines the arithmetic of a category that is wrong, and a smaller false saving is still a false saving. Netting off costs is the right presentation once the hours have genuinely converted into money or into other output. As the report stands, it would leave the same claim in place at a lower figure.
- C
Report the saving once finance has verified the hourly rate used
This disputes an input to the calculation rather than the claim the calculation is making. A verified rate applied to hours nobody has reallocated produces a carefully verified fiction. It would be the right challenge if the hours had been redeployed and only their valuation were in doubt.
- D
Report the saving across three years so the benefit is visible
This spreads an unrealised benefit over a longer period, which restates the same claim more slowly and gives it three years to be found out. Multi-year reporting is appropriate where the benefit is real and its profile uneven. Here nothing has yet happened in any year.
Why that is the answer
Hours become money only when they turn into other output or into a smaller cost base. Headcount is unchanged, the freed hours are unassigned and no budget line has moved, so nothing in the accounts has altered and the PLN 310,000 describes an opportunity rather than a saving. The honest report states the capacity released, where it sits and what it is to be used for, which is also the version that survives the finance director's first question. Overstating the first year is how programmes lose the confidence of the people funding them, and the second year's claim is then discounted whether or not it is true.
About these items
These twelve items are written to the specification of the live CBA-GAI paper, and none of them will appear on one. Every item in the bank is reviewed by a named subject-matter expert and audited for answer cueing domain by domain.
