LLM sampling stunts

Inference-time-scaling considered unreasonably accurate

2026-07-30 — 2026-09-30

quality 8.1

Wherein Sequential Monte Carlo Methods Are Applied to LLM Inference, \(N\) Candidate Trajectories Are Reweighted by Constraints, and the Resulting Particle-Like Filtering Is Contrasted Against the Manual, Lineage-Collapsing Technique Known as Looming.

compsci
language
machine learning
meta learning
Monte Carlo
neural nets
NLP
particle
state space models
stochastic processes
time series
Figure 1

I’ve noticed a number of interesting sampling tricks for LLMs that seem related. The category in my mind is something like sampling from LLMs in ways that resemble classic particle methods, i.e., using the kinematics of multiple hypothetical trajectories to inform the next step(s). So! A new notebook!

1 Sequential Monte Carlo Steering for LLMs

The thing that first drew my attention was the Sequential Monte Carlo (SMC) steering trick for LLMs (Lew et al. 2023): instead of decoding one continuation—or keeping a fixed beam—we keep a small population of \(N\) candidate continuations, score them against a desired constraint or reward, and then repeatedly resample promising candidates as generation proceeds. This can enforce constraints or steer semantic properties at inference time without retraining the base model, and it should look reminiscent of particle filtering except the likelihood is a dog’s breakfast and generally intractable in an annoying way.

Relatedly, LLaMPPL (Lew et al. 2023) claims to be a Probabilistic Programming Language for LLMs based on this idea. It uses Sequential Monte Carlo (SMC) methods to perform probabilistic inference over the outputs of LLMs, which seems logical for prefix sampling; I’m a bit confused about how it works for infill sampling, although they claim both.

Basic idea: We treat a completion \(y = (y_1,\ldots,y_T)\) as a trajectory sampled from the stochastic base model \(p_\theta(y\mid x)\), assuming that it comes from something similar to, but not the same as, the target distribution that we want to hit,

\[ \pi(y\mid x) \propto p_\theta(y\mid x)\,R(y) \]

where \(R(y)\) is a score—e.g., “is valid JSON,” “satisfies this regex/grammar,” “answers both prompts,” “is non-toxic,” or “meets a programmatic verifier.”

We…

  1. Maintain \(N\) partial completions.
  2. Sample the next token (or tokens) for each from the “base” LLM (which in SD-SMC might be a different model entirely).
  3. Give each continuation an incremental weight based on how well it performs under the constraint.
  4. When weights become concentrated, resample: copy high-weight trajectories and discard weak ones.
  5. Continue until completion; return a high-weight sample or sample from the final weighted population.

The distinction from simple rejection sampling is that SMC doesn’t wait until the entire response is finished to discover failure. We can prune bad paths early and concentrate compute on promising ones.

The “trick” is worthwhile because sometimes we can write a prefix-aware potential—a function that scores a partial answer, not merely a final answer.

Examples:

  • Strict JSON: incremental parses can bail early from impossible prefixes
  • Prompt intersection: Likelihood or task score under each of several prompts
  • Other stuff that doesn’t seem very practical to me.

For hard constraints, a particle that makes the remaining task impossible gets weight zero. For soft preferences, the weight rises smoothly to interpolate between “worse” and “better” paths.

The original SMC-steering formulation interprets constrained language generation as posterior inference in a discrete sequence model. It reports capabilities including infilling, syntactic constraints, and “prompt intersection,” at a computational cost comparable to beam search.

LLM vibe-coded pseudo-code example:

particles = [("", 0.0, cache_i) for i in range(N)]  # text, log_weight, KV cache

for t in range(max_new_tokens):
    proposed = []

    for text, logw, cache in particles:
        token, logp, cache2 = sample_next_token(model, prompt + text, cache)
        text2 = text + decode(token)

        # potential should assess the prefix / remaining feasibility
        delta = log_potential(text2, t + 1) - log_potential(text, t)
        proposed.append((text2, logw + delta, cache2))

    particles = proposed
    weights = softmax([logw for _, logw, _ in particles])

    if effective_sample_size(weights) < N * 0.5:
        particles = systematic_resample(particles, weights)
        particles = [(text, 0.0, cache) for text, _, cache in particles]

answer = select_or_sample_final(particles)

A common formulation uses a sequence of potentials \(\phi_t(y_{1:t})\), with the incremental correction:

\[ \log w_t = \log w_{t-1} + \log \phi_t(y_{1:t}) - \log \phi_{t-1}(y_{1:t-1}) \]

That incremental difference prevents us from repeatedly counting the same “goodness” evidence at every token. In practice, this is not the true target potential, because that is almost always super weird in token space.

There are a few adjacent “SMC for LLMs” ideas that have a similar shape.

  • SMC steering: use particle filtering to steer toward constraints or a reward at inference time, as discussed above.
  • SMC speculative decoding (SMC-SD): drafts from a cheaper model, score blocks under an expensive target model, then reweights/resamples rather than rejecting drafts token-by-token. The aim is faster inference while retaining close target-model behaviour (Emara et al. 2026)
  • Self-consistency: sample many complete reasoning trajectories and majority-vote the answer. This is not sequential resampling, but people sometimes loosely describe it as a particle-like trick. This one has the shape of, and relevance to, Reasoning LLMs where it is called maj@\(k\).
  • Power-SMC: sample from a sharpened sequence distribution \(p_\theta(y\mid x)^\alpha\), using parallel particles and token-level importance weights/resampling. It is positioned as a low-latency, training-free alternative to serial Metropolis–Hastings-style sampling (Azizi et al. 2026) of (in the model’s opinion) higher likelihood completions. NGL I cannot even work out what this one is supposed to do.

2 Looming

I have just been schooled on the prior art: the cyborgists have been doing population-based steering of language models since 2020 under the name looming. HT Theia Vogel.

Loom is Janus’s tree-based writing interface to base models, which is to say models that continue a text rather than answer a chat turn. The original is pyloom; there are descendants such as Loomsidian and Conjecture’s Bonsai, more on the cyborgism wiki. The procedure, in Janus’s description, is “manual iterative rejection sampling”. The operator (loomer, let us say) generates \(N\) continuations from the current node, a paragraph or less each, reads them, picks the one we like, and generates \(N\) more from there. Janus recommends \(N\) from 5 to over 100, depending on how fussy the passage is and the loomer’s patience. The interface keeps every rejected branch, so we can go back into history and replay to get different outcomes. (Similarities to fan-out in mathematical proving left as an exercise).

In the vocabulary of this notebook, this is a particle method with a human defining the selection process instead of a potential. The proposal is generated by a base model \(p_\theta\), as in the bootstrap version of SMC steering. The weight function is “black box”; i.e. it is whatever goes on in the loomer’s head. In terms of the pseudo-code above, we replace systematic_resample with “\(N\) copies of the one I clicked on”.

Damn but that resembles SMC. There are a few differences however.

One lineage. SMC in the default forward mode resamples \(N\) survivors from \(N\) candidates and ideally keeps several distinct ancestries alive, because collapsing onto a single ancestor is indicative of failure for most purposes. Looming collapses onto one ancestor at every step, on purpose. This changes what is being sampled, as a statistical object. Split the text into blocks \(b_1,\dots,b_K\) between branch points and suppose, generously, that the loomer picks candidate \(i\) with some Boltzmann-rational preference \(g_k(b_{1:k}^{(i)})\). For large \(N\), looming draws from the locally normalized distribution

\[ q(y) = \prod_{k=1}^K \frac{p_\theta(b_k\mid b_{<k})\,g_k(b_{1:k})}{Z_k(b_{<k})}, \qquad Z_k(b_{<k}) = \mathbb{E}_{p_\theta}\left[g_k(b_{1:k})\mid b_{<k}\right], \]

whereas SMC with the same incremental weights targets the globally normalized

\[ \pi(y) \propto \prod_{k=1}^K p_\theta(b_k\mid b_{<k})\,g_k(b_{1:k}). \]

The two differ by the factor \(\prod_k Z_k(b_{<k})\), which depends on the path. \(Z_k(b_{<k})\) measures how good the options were at step \(k\). SMC ‘remembers’ that information, in that a particle whose children are mostly duds ends up with few descendants. A loomed lineage cannot remember it, because there is no rival prefix to lose out to. So looming can walk into a prefix where only 1 continuation in 100 is any good; it doesn’t, by sampling alone, help us find regions that don’t suck. This ‘distortion’ is what token-masking approaches to constrained decoding suffer, and removing it is the stated reason for using SMC in the first place (Lew et al. 2023; Loula et al. 2025).

There is an interesting special case where the distortion vanishes. If \(g_k = \psi_k/\psi_{k-1}\) with \(\psi_k(b_{1:k}) = \mathbb{E}_{p_\theta}[R(y)\mid b_{1:k}]\), the expected final score given the prefix, then every \(Z_k = 1\) and \(q = \pi\). That \(\psi_k\) is the optimal twist function of twisted SMC, which Zhao et al. (2024) try to learn. A loomer is asked to be the twist function: to judge a paragraph by where it is likely to lead, not by how it reads now. I do not know how well people do this, but I suspect we are not terrible at it, insofar as we solve problems like this every time we flap our gobs. The invective that streams from my cake hole is by no means globally optimal.

No stated target. SMC steering starts from explicitly given \(R(y)\), and the challenge is finding prefix potentials that approximate it. A loomer has only the prefix potential, and it is not required to be consistent from one step to the next. Janus says they expect whatever purpose they started with to itself mutate/branch along the way. In that sense there is no posterior for the procedure to be “wrong” about, and the guarantees that make SMC attractive to statisticians (consistency, unbiased estimates of the normalizing constant) are not especially relevant to cyborgists. As such, maybe we should think of looming as closer to search than to sampling. There turns out to be some mileage in this perspective, which we come back to in a moment.

Janus attempts to quantify such curation: Choosing 1 of \(N\) equiprobable completions applies \(\log_2 N\) bits of optimization, and bits add across choices. AFAICT the particle-filter equivalent is \(\log_2(N/\mathrm{ESS})\) (Effective Sample Size) bits per resampling, since \(N/\mathrm{ESS}\) is the factor by which the weights concentrate the population. The ESS < N/2 rule in the pseudo-code says “resample once we have accumulated about 1 bit”. Loaming with \(N=8\) takes ESS to 1 and injects 3 bits of information every paragraph.

The tree is kept. SMC in the default forward mode is forward-only; a discarded particle is gone, and the history of the survivors is of interest only as a diagnostic. A loom stores the whole genealogy and extensions over the default allow us to “rewind” and “replay”, walking back up the tree and back down to alternative limbs, branching again from an earlier node when the current line goes nowhere fun. Backtracking is, as such, an interactive, manual fix for the greediness of looming. It makes looming look more like tree search than like filtering.

Branch points follow the model’s uncertainty. SMC resamples when the weights degenerate (concentrating on a small number of options), which is to say when the constraint has said something informative. Janus’s adaptive branching puts branch points in according to their idiosyncratic logic, e.g. where the model’s own next-token distribution is spread out, or where a particular sampled token had low probability. One argument for this is that forcing a branch (do loomers force them to be distinct?) where the model is 99% sure of the next token only produces incoherent alternatives. I assume that some hip flavour of adaptive SMC does such things, but I have not seen it.

There is some work to unify looming and SMC. JD Pressman’s MiniHF is one galaxy-brained example. Pressman argues that because picking branches by hand is tiring, a loomer will want to distill their judgement into a reward model to do it for them. MiniHF’s Weave algorithm does that, running Monte Carlo tree search against an evaluator trained to approximate the loomer’s preferences, and by the author’s account it ends up close to Tree of Thoughts (Yao et al. 2023). Once the human has been replaced by a learned evaluator we seem to be doing something like SMC steering with a learned potential, give or take the choice of search algorithm. c.f. InFerActive (Hwangbo et al. 2025) which uses SMC with adaptive resampling to pre-populate the tree that a human evaluator then explores.

If we wanted to take the parallel totally seriously, we could combine them completely: devise a loom that keeps several lineages alive and asks the operator for rough weights on a handful of candidates each step. That would be a particle filter with a human likelihood. It would cost more reading per step but OTOH would buy back the global normalization, which is to say, estimate the partition function. Or is that taking the analogy too far? If, as Janus says, the loop is more exploratory than inferential, does it even make sense to worry about global normalization? Anyway, that is a side quest for someone with more side quest time than I.

3 References

Azizi, Potraghloo, Ahmadi, et al. 2026. “Power-SMC: Low-Latency Sequence-Level Power Sampling for Training-Free LLM Reasoning.”
Emara, da Costa, Chang, et al. 2026. “Faster LLM Inference via Sequential Monte Carlo.”
Hwangbo, Lee, Jeon, et al. 2025. “InFerActive: Interactive Tree-Based Exploration of LLM Sampling for Safety Evaluation.”
Lew, Zhi-Xuan, Grand, et al. 2023. “Sequential Monte Carlo Steering of Large Language Models Using Probabilistic Programs.”
Loula, LeBrun, Du, et al. 2025. “Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo.”
Wang, Wei, Schuurmans, et al. 2023. “Self-Consistency Improves Chain of Thought Reasoning in Language Models.”
Yao, Yu, Zhao, et al. 2023. “Tree of Thoughts: Deliberate Problem Solving with Large Language Models.”
Zhao, Brekelmans, Makhzani, et al. 2024. “Probabilistic Inference in Language Models via Twisted Sequential Monte Carlo.” In Proceedings of the 41st International Conference on Machine Learning.