Popular searches
//

Diffusion Language Models: Beyond Left-to-Right Generation

9.10.2026 | 11 minutes reading time

You might think every large language model (LLM) generates the same way: one token, left to right, no going back. That constraint is so fundamental to how we think about LLMs that it's easy to forget: it was a choice.

So what would an alternative look like? Think about how you write code: you sketch a signature, jump to the call-site, rename a parameter, then come back and fill in the body. Diffusion language models (DLMs) work a lot like that: they start with a whole sequence of placeholders and refine it, filling many positions per step and revising earlier guesses when later context proves them wrong.

This article explains the paradigm, where it stands today, and where it could be useful.

Autoregressive models: a brief refresher

All current state-of-the-art LLMs generate text left to right. The formal way to say this is that they work autoregressively: one token comes out at a time. For each token, the model runs a full forward pass through all of its parameters, often billions of them, then appends the result and starts over for the next one.

This approach has two flaws.

First, inefficiency. When writing a class definition or a common function pattern, a whole sequence of upcoming tokens is often highly predictable from what came before. Yet, the model is still forced to produce them one by one, spending massive compute on every single step regardless of how obvious the next word is.

Second, the inability to self-correct. Because generation only moves forward, the model cannot go back and revise earlier text if the newly unfolding logic means a previous choice should have been different. The attention mechanism in these models only looks backward (causal attention). Information only flows from the past to the next token, locking in the generated context without any way to adapt it in reverse.

To be clear, this paradigm works remarkably well, as we all experience in our everyday work. And there are fixes for both issues: speculative decoding and the KV cache make inference faster, and an agentic loop can run multiple times to find problems and correct them.

Still, in this fast-moving field it is always worth looking left and right to see what else exists. The most promising ones for me are DLMs.

How diffusion language models work

Diffusion first became popular in image generation, where it powers Stable Diffusion, Midjourney, and DALL-E. These models start from a noisy image and add more meaning at each step, iteratively "denoising" a pixel grid until the final image emerges.

diffusion_process.png

The principle they use can be applied not only to pixels but also to text. The underlying model is almost the same as in an autoregressive LLM; what changes is how it's trained and how it generates. Instead of generating tokens one by one from left to right, a diffusion language model starts with a sequence of placeholders. Depending on the method, these can be mask tokens or random tokens sampled from the vocabulary. Either way, the starting point is essentially meaningless text. It is the equivalent of the noisy image at the start of Stable Diffusion, only in token space.

Then comes the prompt. Say we ask the model to write a function that checks whether a number is even. The model sees the prompt followed by a row of masks and predicts every masked position at once. So why not stop there?

Because on the first attempt the model has nothing to work with: every position is guessed in the presence of nothing but other guesses. We need a way to let the good parts survive and the weak parts be reconsidered, and that is what remasking does. One step looks like this:

  • Feed the current sequence to the model: the prompt, the masks, and whatever has already been filled in.
  • The model predicts a token for every masked position, each with a confidence. Keep the best ones, remask the rest, so weak guesses are thrown away before they become part of the sequence.
  • Repeat until no masks are left.

dlm-run.png

Notice how much control this gives us compared to an autoregressive model. We can keep everything from the first prediction and be done in one step, resulting in a very fast model, but at the cost of lower quality. Alternatively, we can keep only the single best token each time and take as many steps as there are tokens. This creates a slower, more autoregressive-like model, but with the distinct advantage that all tokens can inform each other using context from both the past and the future (bidirectional attention). Ultimately, DLMs give us the ability to dynamically steer the balance between output quality and inference time.

Why diffusion language models are worth looking at

DLMs are no longer an academic niche. Google has Gemini Diffusion, Apple has DiffuCoder, ByteDance has Seed-Diffusion. From the smaller labs there is Mercury by Inception Labs and Celeris-1, which came out in July 2026.

The first systematic study of diffusion models on code (Li et al. 2025) identifies two core advantages over autoregressive models.

Multi-token prediction. A diffusion model commits multiple tokens per forward pass instead of one. On Artificial Analysis, diffusion models are currently at the top of the speed benchmark. Celeris-1 is the fastest model with 1523 token/s, followed by Mercury-2.5 with 690 token/s. For comparison, the fastest autoregressive model in the same benchmark reaches 227 token/s. That makes these models interesting for agentic systems where response time is the binding constraint, like voice agents.

artificial_analysis_DLMs.png

Flexible generation order. Diffusion models can generate tokens in any order, not just left to right. The model sees the full sequence at each step, including later positions, so for coding it can let the end of a function inform the beginning. An autoregressive model can't do that: once it has written a function signature, it has to live with it, even if the body later calls for a different parameter. This also makes diffusion models a natural fit for filling a gap in existing code, where the new lines have to fit both what comes before and what comes after. The same applies to structured output. Since the model sees the whole structure at every step, it can reconsider a weak guess for a bracket or a field in a JSON object or tool call before the output is final.

Current limitations and directions

While this all sounds promising, why is diffusion not the dominant paradigm for language models today?

Quality. The obvious answer is that DLMs don't perform nearly as well as their autoregressive counterparts. But that's not the complete picture. At comparable model sizes, DLMs actually show competitive performance on several benchmarks (LLaDA, Li et al.). Compared to state-of-the-art models they still fall short, but that might change with bigger scale, more research and more compute.

It is a young field. Autoregressive generation came first, and for years nearly all research investment went into making it better, so the training recipes and accumulated community knowledge that AR models enjoy are still being worked out for diffusion. But you can see an acceleration in this direction: DLM papers now show up across NeurIPS, ICLR and ICML. NeurIPS is hosting two dedicated DLM workshops in December, the first at the conference, and vLLM added native diffusion support in June 2026.

Fixed output length. As you might have noticed in the example, the model needs to know upfront how many tokens to generate. That is not always practical for code or text where the output length is hard to predict. Generate too few tokens and the function is incomplete. Generate too many and the model fills the extra space with redundant comments. Block Diffusion (Arriola et al. 2025) addresses this naturally by combining both paradigms: autoregressive generation at the block level, diffusion within each block. You just keep adding blocks until the output is done, and the model gets the left-to-right coherence of autoregressive generation across blocks and the parallel flexibility of diffusion within them.

Bildschirmfoto 2026-09-25 um 13.32.26.png source: https://arxiv.org/pdf/2503.09573

Hands On

Calling a DLM via API

Next I will show you how to use such a model. And surprisingly almost nothing changes compared to autoregressive models. Inception and most of the other providers expose OpenAI-compatible endpoints, so if you’re already using the OpenAI API, it’s usually just a different base URL and model name. In the next snippet I use Inception’s SDK and its model mercury-2.5 that I already mentioned in the speed benchmark.

Output:

As you can see, the model needed only three denoising steps to produce the final text. An autoregressive model would have needed 44 forward passes, one for every token.

DLMs as decision models

A few weeks ago TypeSafe launched JEV, a decision model (full breakdown here if you want to go deeper). In short: instead of free text, a decision model answers typed questions (yes/no, pick one of several options, a score on a scale) and returns a probability for every answer.

My immediate thought when I heard about it: I would bet DLMs are great at this task, because of their speed and bidirectional attention. The answer is short and structured, and the model sees the whole structure at every step. So as a second hands-on, I show you how to use a DLM as a decision model with the TypeSafe client.

Install the SDK (uv add typesafe-sdk) and point the client at a diffusion model. Here I use celeris-1-decision:

Then you describe the situation as state and ask your typed questions. A Noul is a yes/no, a Choice picks one of several options, a Score places the answer on a scale:

Output:

How good are they?

I also ran an evaluation on the public benchmark Typed Decisions (400 cases, 2,000 decisions) to get a feeling for how good diffusion models are at this kind of task. Next to celeris-1-decision I tested mercury-decide, the decision model from Inception. They performed quite well: both stayed close to the other decision models in accuracy and had the lowest latencies in the table. celeris-1-decision even matched Jev while answering in about a third of the time. It is only one benchmark, but it shows the potential of these models and the role they could play in this new wave of fast decision models.

ModelAccuracy ↑p50 Latency
meraGPT Decider 10.768526 ms
Liquid AI d10.742525 ms
celeris-1-decision (diffusion)0.739243 ms
TypeSafe Jev 1.13.00.727710 ms
mercury-decide (diffusion)0.715409 ms

You can find the full code, including the benchmark and a more detailed report, in the companion repository.

Conclusion

DLMs are not going to replace autoregressive models tomorrow, and it is not clear that they ever will. What they offer is a different approach: they refine a whole sequence in parallel instead of writing one token after the other, which gives you many tokens per forward pass, bidirectional context, and a dial between quality and latency. The field is still young, with a lot of room for optimization.

Where I see their potential is next to the frontier rather than at it: tasks where speed matters more than peak intelligence, or where the output has to hold together as a whole. Fast, typed decisions are one example, as the benchmark showed, alongside code edits, voice agents and short agent steps.

I would not be surprised if diffusion models matter noticeably more a year from now than they do today.

//

More articles in this subject area

Discover exciting further topics and let the codecentric world inspire you.