CBA-GAI · sample lesson
Chapter 1 · Free sample
Why the same prompt gives a different answer twice
4 min read
The model has to choose, and choosing well is not choosing the best
Given the text so far, the model produces a probability distribution over what comes next. Then it has to choose one option from it. If it always chose the single most probable option, output would be repetitive and often oddly flat, so most systems sample from the distribution instead, and settings exist that control how adventurously they sample.
These settings are usually described in terms of temperature, or a cut-off applied to cumulative probability, and the vocabulary varies between products. Do not memorise the names, because they will be different again in whatever your organisation buys next year. Memorise what the settings do, because that does not change, and the exam is written against the behaviour rather than the label.
Slide 1 of 9. The model has to choose, and choosing well is not choosing the best
The same lesson, in full
Given the text so far, the model produces a probability distribution over what comes next. It then has to choose. If it always chose the single most probable option, output would be repetitive and often oddly flat; so most systems sample from the distribution, and settings exist that control how adventurously they sample. These are usually described in terms of temperature, or a top-probability cut-off, and the vocabulary varies between products. Do not memorise the names. Memorise what they do.
Turning these settings down makes the model more likely to pick its highest-probability continuation. Output becomes more consistent between runs and more conventional. Turning them up flattens the distribution, so lower-probability continuations get a better chance: more variety, more surprising phrasing, and more excursions away from the safe centre.
Now the part that carries most of the exam questions in this section. These settings control variability, not accuracy. If the model's most probable continuation is a fabricated chapter, then turning the randomness down does not remove the fabrication; it makes the fabrication arrive reliably, in the same words, every time. Consistency reads as authority to a human reviewer, so a low-variability wrong answer is more likely to be believed than a high-variability one. What lower settings genuinely buy you is testability: a system that behaves the same way twice can be evaluated with fewer runs, and a change to the prompt can be attributed to the prompt rather than to chance. That is a real benefit for a publisher trying to build a controlled process. It is not a truth setting, and the exam will offer it to you as one.
Two further sources of variation matter and neither is a setting you control. Systems are updated: the model behind a service can change beneath a prompt you validated in March, and its behaviour on your specific edge cases can shift without any announcement that names your use case. And the surrounding software changes: retrieval indexes are rebuilt, standing instructions are edited by somebody else, tools are added. Validation is therefore a recurring cost, not a one-off. Any business case that counts prompt testing once has understated its own running costs.
Worked example 5: the same alt text prompt, five times
Lindenmoor's accessibility lead tested a draft alt text instruction on Figure 3.2 of a 2011 public health title: a bar chart of notified measles cases by UK region, with no time dimension on the chart and the highest bar in the North West. Five runs, at a mid-range randomness setting, produced:
| Run | Output | Defect |
|---|---|---|
| 1 | "Bar chart showing notified measles cases by UK region, highest in the North West." | None |
| 2 | "Bar chart comparing measles cases across six UK regions between 2008 and 2011." | Invented date range |
| 3 | "Chart of regional variation in measles notifications, highest North West, lowest South West." | None |
| 4 | "Histogram of measles case numbers by region." | Wrong chart type |
| 5 | "Bar chart showing a rise in measles cases across all regions." | Invented trend |
Three of five runs carried a defect, and notice that the defects are not random noise. Each one adds the thing such a chart usually has: a period, a trend, a familiar chart-type label. The model is filling the distribution's expectations for the genre, which is the same mechanism that produced the phantom chapter, appearing here in a task that looks trivially safe.
The testing arithmetic matters more than the individual defects. If a prompt produces a defect on a proportion p of runs, the chance of seeing at least one defect in n runs is 1 minus (1 minus p) to the power n. Turn that around and ask how many runs you need before a clean result means anything.
| True defect rate | Runs for a 90% chance of seeing it | Runs for a 95% chance |
|---|---|---|
| 20% | 11 | 14 |
| 10% | 22 | 29 |
| 5% | 45 | 59 |
| 2% | 114 | 149 |
The marketing team tested their blurb prompt once, liked what came back, and rolled it out across 600 titles a year. At a 10% defect rate, a single clean run had a 90% chance of occurring anyway. One good output is not evidence about a prompt; it is one sample from a distribution you have not characterised.
And be honest about the imprecision of small tests in the other direction too. Three defects in five runs gives a point estimate of 60% and a 95% confidence interval running from roughly 15% to 95%. That is not a measurement of a defect rate. It is sufficient evidence to stop and design a proper one, which is all a pilot ever needs to deliver.
The full contents
Every chapter and lesson of the CBA-GAI study material, with reading times and where the assessed workbooks fall.
