A Simple Overview of LLM Sampling Methods

Almost all the dis­course around LLMs to­day is cen­tered on the mod­els them­selves. This model is bet­ter than that model which is bet­ter than this other model. It’s… ex­haust­ing. It’s also barely scratch­ing the sur­face of what makes an ef­fec­tive LLM-based sys­tem.

Today, I want to in­tro­duce you to an en­tirely new world. The world of sam­pling al­go­rithms. We’ll dis­cuss what they are, how some of them work, and the im­pact they can have on sys­tem per­for­mance.

What Is a Sampling Algorithm?

If you’ve re­ceived even a sim­ple ex­pla­na­tion of how au­tore­gres­sive lan­guage mod­els work, you likely al­ready have an in­tu­itive un­der­stand­ing of what a sam­pling al­go­rithm is. If you have not, do not worry. It is not ac­tu­ally as com­plex as it sounds.

When you give an LLM a piece of text, it does­n’t ac­tu­ally spit out more text. The act of “writing” as­so­ci­ated with LLMs when used in chat­bots or else­where comes from a higher or­der pro­gram that sim­ply uses an LLM. The LLM it­self is only ca­pa­ble of spit­ting out a prob­a­bil­ity dis­tri­b­u­tion.

What dis­tri­b­u­tion might that be? Given a se­quence of words, an LLM emits the prob­a­bil­ity dis­tri­b­u­tion of what word might come next.

I talk about this in more depth in my blog post about Markov Chains.

For ex­am­ple, let’s say we of­fer an LLM the phrase, “The cat sat on the ”. In re­turn, the LLM will give us a prob­a­bil­ity dis­tri­b­u­tion over all the pos­si­ble words that might come next.

Word Probability
hat 0.9
bat 0.05
ra­di­a­tor 0.04
… 0.01

In or­der to gen­er­ate text, all we need to do is choose among these pos­si­ble words and ap­pend it to our in­put. From there, we may re­peat the cy­cle all over again. If we do this enough times, we can get co­her­ent-sound­ing prose from the model.

Here’s the hard part: How we do choose a word from the prob­a­bil­ity dis­tri­b­u­tion? The an­swer: Use a sam­pling al­go­rithm.

Greedy Sampling.

Greedy sam­pling is the most ob­vi­ous so­lu­tion. Given a prob­a­bil­ity dis­tri­b­u­tion, a greedy sam­pling al­go­rithm will sim­ply choose the most likely word every time.

In our “The cat sat on the ” ex­am­ple, our greedy al­go­rithm would choose the most prob­a­ble word, “hat”. It’s sim­ple, el­e­gant, and it can some­times get the job done. But, as you may have guessed, it is not op­ti­mal.

The most ob­vi­ous prob­lem with greedy sam­pling ap­pears as soon as you gen­er­ate a mean­ing­ful amount of text. It’s filled with rep­e­ti­tions.

When I ask a small lan­guage model like SmolLM2:135M about Python, I get:

Python is a great lan­guage for be­gin­ners, as it’s easy to learn and has a large com­mu­nity of users. However, it’s worth not­ing that Python can be chal­leng­ing to learn, es­pe­cially for those with­out prior ex­pe­ri­ence with pro­gram­ming. But with the right re­sources and a solid foun­da­tion, Python can be a pow­er­ful tool for any­one look­ing to de­velop in­no­v­a­tive so­lu­tions.

Some of the key fea­tures of Python in­clude:

  • Python 3.x: Python 3.x is the lat­est ver­sion of Python, which is widely used in the in­dus­try.
  • Standard Library: Python has a large stan­dard li­brary that in­cludes a wide range of tools and func­tions for var­i­ous tasks.
  • Extensive Object-Oriented Programming (OOP): Python has a strong fo­cus on OOP, mak­ing it suit­able for build­ing com­plex ap­pli­ca­tions and sys­tems.
  • Cross-Platform: Python can run on mul­ti­ple plat­forms, in­clud­ing Windows, ma­cOS, Linux, and Android.
  • Cross-Platform: Python can run on mul­ti­ple plat­forms, in­clud­ing Windows, ma­cOS, Linux, and Android.

Notice any­thing?

Most ob­vi­ously, it re­peats it­self un­nec­es­sar­ily. It men­tions “Cross-Platform” twice, with the same de­scrip­tion text. It’s only a mi­nor an­noy­ance here, but rep­e­ti­tion like this can be dev­as­tat­ing for long-con­text rea­son­ing mod­els. Us­ing a sam­pling al­go­rithm that is “too greedy” is a com­mon rea­son for “looping” be­hav­ior in off-the-shelf rea­son­ing mod­els when they are set up in­cor­rectly.

Additionally, the model finds it­self too ad­her­ent to its own train­ing data. No­tice how it lists “Python 3.x” as a “feature” of the lan­guage? This is be­cause that par­tic­u­lar Python ver­sion num­ber ap­pears fre­quently in the data. That means it can get ranked at the top of the prob­a­bil­ity dis­tri­b­u­tion, even if it has noth­ing to do with the sur­round­ing con­text.

Okay, how can we fix these prob­lems?

Statistical Sampling

The sec­ond most ob­vi­ous sam­pling method is to sam­ple from the prob­a­bil­ity dis­tri­b­u­tion ran­domly. Let’s pull up that dis­tri­b­u­tion from ear­lier:

Word Probability
hat 0.9
bat 0.05
ra­di­a­tor 0.04
… 0.01

Our model sug­gests that the first word, “hat” is 18 times more likely to be the next word. In sta­tis­ti­cal sam­pling, we honor that. In­stead of only sam­pling the most likely word like in greedy sam­pling, in sta­tis­ti­cal sam­pling we sam­ple ran­domly, bi­as­ing our­selves to sam­ple the more likely words more of­ten.

The sim­plest al­go­rithm to ac­com­plish this is to:

  1. Generate a ran­dom num­ber rr be­tween 0 and 1.
  2. Iterate through the prob­a­bil­ity dis­tri­b­u­tion pip_i, start­ing with the most prob­a­ble word first.
  3. Once we en­counter a word pxp_x such that ∑n=1xpn>=r\sum_{n = 1}^{x} p_n >= r, we sam­ple it.

In prac­tice, this leads to more var­ied text. The model no longer needs to stick to the most likely word. In­stead, we can oc­ca­sion­ally dip down and emit the less likely (potentially more in­ter­est­ing) words.

In rea­son­ing mod­els, this al­lows the model to ex­plore more cre­ative so­lu­tions to prob­lems. Sta­tis­ti­cal sam­pling (and the vari­ants we ex­plore be­low) also mit­i­gate the rep­e­ti­tion we saw in the greedy al­go­rithm.

One of the prob­lems with sta­tis­tic sam­pling is that the ran­dom­ness can oc­ca­sion­ally pro­duce words that are en­tirely un­re­lated to the topic at hand.

By us­ing a rather ag­gres­sive sta­tis­ti­cal sam­pler with the same model as be­fore, we get some­thing that seems fine at first, but de­grades rather quickly.

Python is an in­ter­preted, dy­nam­i­cally-typed pro­gram­ming lan­guage cre­ated by Guido van Rossum. Its core fea­tures in­clude syn­tax for com­pact read­able code, ease of use with vast li­braries and mod­ules, good per­for­mance, ease of de­vel­op­ment with con­ve­nient de­vel­op­ment en­vi­ron­ments, user-friendly GUI, seam­less in­ter­ac­tions with on­line ser­vices, se­cure web cod­ing, and data­bases for seam­less busi­ness pro­duc­tiv­ity ap­pli­ca­tions.

What other fac­tors dis­tin­guishes python be­sides sim­ple gram­mar have been uti­lized to in­ter­act with is­sues such as sched­ul­ing fail­ure re­quest for your ma­chine ser­vice in just sec­onds be­cause the load fol­low­ing how the role of do­main guys emaild Osea.. I Hate syn­tax hard w

Programs that use LLMs with sta­tis­ti­cal sam­plers de­grade at longer con­text sizes partly for this rea­son. Over time, the er­ror in the ran­dom­ness builds un­til the model can no longer pro­duce log­i­cal text.

How can we mit­i­gate this new prob­lem?

Nucleus Sampling

Nucleus sam­pling is just like sta­tis­ti­cal sam­pling, but with a de­fined limit to how un­likely a word is al­lowed to be.

To il­lus­trate the process, let’s set a limit (0.95) and go through each step.

Word Probability
hat 0.9
bat 0.05
ra­di­a­tor 0.04
… 0.01

Before we can do any­thing, we need to pro­duce our “nucleus”. To do this, we step through each item in our prob­a­bil­ity dis­tri­b­u­tion. While do­ing so, we keep track of the cu­mu­la­tive prob­a­bil­ity, just like in sta­tis­ti­cal sam­pling. If we en­counter any words whose cu­mu­la­tive prob­a­bil­ity is greater than our limit (that is, ∑n=1xpn>0.95\sum_{n = 1}^{x} p_n > 0.95) we dis­card it.

After that, we are left with a new list of “probabilities”:

Word Probability
hat 0.9
bat 0.05

Except, these are not quite prob­a­bil­i­ties. After all, the do not add up to 1.0. Let’s fix that by ap­ply­ing a soft­max to the col­lec­tion. I will not go into great de­tail of how that works. To put it sim­ply, a soft­max turns a col­lec­tion of num­bers into a proper prob­a­bil­ity dis­tri­b­u­tion that sums to 1.0.

Word Probability
hat 0.7
bat 0.3

From here, we just ap­ply nor­mal sta­tis­ti­cal sam­pling to our new prob­a­bil­ity dis­tri­b­u­tion.

To sum it up, nu­cleus sam­pling the same as sta­tis­tic sam­pling, but we re­move all the most un­likely words from the dis­tri­b­u­tion first.

Nucleus sam­pling mit­i­gates all the down­sides of the first two al­go­rithms, while re­main­ing per­for­mant. Text is gen­er­ally free of un­nec­es­sary rep­e­ti­tion, and it avoids the ran­dom in­ser­tion of low-prob­a­bil­ity words.

Conclusion

There are a few more salient sam­pling al­go­rithms that per­haps should be in­cluded in to­day’s post, like DRY sam­pling, but un­for­tu­nately I am out of time. I’ve cov­ered the more com­mon ones to­day. Most com­mer­cial apps use nu­cleus sam­pling. Some more niche apps use a kind of greedy sam­pling. Know­ing which to use is an im­por­tant de­ci­sion. Make it wisely.

Published October 9, 2026 at 6:31 PM

Proofread by Harper.

Comments