Sampling Strategies
How the model actually picks its next word
Once a trained model has worked out a probability for every word it could say next — a number showing how likely each one is — something still has to pick one. That choice is called sampling. The plainest choice, "greedy," always picks whichever word has the highest probability. A more natural-sounding choice, "top-p," first narrows the list down to the smallest group of words whose probabilities add up to some target like 90 per cent, and only then picks randomly from that smaller group, weighted by how likely each one is.
The code
import random
def greedy_sample(probabilities):
"""Always pick the single most likely word."""
return probabilities.index(max(probabilities))
def top_p_sample(probabilities, p=0.9):
"""Narrow down to the smallest group of words whose probabilities
add up to at least p, then pick randomly from that group."""
ranked = sorted(
range(len(probabilities)),
key=lambda i: probabilities[i],
reverse=True,
)
pool, running_total = [], 0.0
for i in ranked:
pool.append(i)
running_total += probabilities[i]
if running_total >= p:
break
weights = [probabilities[i] for i in pool]
return random.choices(pool, weights=weights, k=1)[0]
What each part does
greedy_sample takes the list of probabilities and returns the position of the biggest one. It never varies, which makes it predictable but also makes it repeat itself in longer answers.
top_p_sample first sorts every word's position by how likely it is, from most to least likely. It then walks down that sorted list, keeping a running total, and stops as soon as the group it has kept adds up to at least p (90 per cent, by default). Finally, random.choices picks one position from that smaller group, giving more likely words a bigger chance without always forcing the single most likely one.
Why this shape
Any accepted submission in this category takes one list of probability numbers (all adding up to 1) and returns a single whole number: the position of the word that was chosen. That fixed shape is what lets any sampling piece plug into any model's output without either side needing to know how the other one works inside.