---
title: Performance
description: "How we measure the Workflow SDK runtime's performance, the metrics we track, and the latest results."
docs_index: /llms.txt
lastUpdated: 2026-10-08
type: reference
summary: The benchmarks we run on the Workflow SDK runtime, how each one is measured, and the latest results.
related:
  - /docs/whats-new
  - /docs/how-it-works/event-sourcing
  - /docs/foundations/hooks
  - /docs/foundations/streaming
---

> For an index of all documentation, see [/llms.txt](/llms.txt).

We benchmark the Workflow SDK runtime on Vercel at each point where a workflow waits on it: starting a run, moving from one step to the next, fanning out, resuming from a hook, and delivering a stream. This page explains what each benchmark measures and how, and then shows the [latest results](#latest-results), which compare a new release with an earlier one.

## What we measure

Each benchmark starts a clock at one line of code and stops it at another. The work inside a step is not counted, so the numbers measure the runtime itself. Choose a metric below, then play or scrub through a run to see when its clock runs.

### Time to first step (TTFS)

How long it takes from your app calling [`start()`](/docs/api-reference/workflow-api/start) until the first line of the workflow's first step runs. This covers creating and queuing the run, running the workflow code up to its first step, and starting that step. It is measured on a five-step workflow. See [Steps and long runs](#steps-and-long-runs) for the latest results.

### Step-to-step overhead (STSO)

How long it takes from one step finishing until the next step starts: recording the step's result, resuming the workflow, and starting the next step. It is measured on the same five-step workflow, and at steps 1 to 20, 101 to 120, and 1,001 to 1,020 of a 1,020-step workflow, which shows whether the overhead grows as the run's [event log](/docs/how-it-works/event-sourcing) grows. See [Steps and long runs](#steps-and-long-runs) for the latest results.

### Fan-out (TTFS, TTLS, join)

A workflow starts 64 one-step branches at once with `Promise.all`. The runtime queues the branches in batches, so they start close together and in no fixed order. The first branch (TTFS) and the slowest branch to start, or time to last step (TTLS), are timed from `start()` until that branch's first line runs. The join is timed from the slowest branch finishing until the code after `Promise.all` runs. See [Fan-out](#fan-out) for the latest results.

### Time to resume (TTR)

A workflow creates a [hook](/docs/foundations/hooks) and waits on it for 1 s, the way an agent waits for a person's approval, and then your app calls `resumeHook()`. Time to resume runs from that call until the first line of the workflow's next step runs. Each run resumes five times. See [Resuming a workflow](#resuming-a-workflow) for the latest results.

### Chunk trip time (CTT)

One step replays a recorded AI agent's token stream, 2,593 chunks over 52.4 s, in real time to a [namespaced stream](/docs/foundations/streaming#namespaced-streams) with [`getWritable()`](/docs/api-reference/workflow/get-writable), and a second step in the same run reads it with `getRun(runId).getReadable()`. Chunk trip time runs from the writer writing a chunk until the reader receives it. The same benchmark records how long the whole response takes to reach the reader and how many runs had a chunk delayed by over 1 s. See [Streams](#streams) for the latest results.

## Methodology

- **Platform**: Each version runs the same benchmark app on Vercel production in `iad1`, against the managed [Vercel World](/worlds/vercel), on its own deployment. The versions are measured at the same time. Before measuring, each deployment runs two short workflows and one fan-out that are not counted.
- **Workload**: Every step takes 1 KB of JSON in and out and does 100 ms of simulated work, which the overhead measurements exclude.
- **Runs**: Each version runs every benchmark several times, and each result pools the samples from all of its runs. The run counts are listed with the latest results.
- **Percentiles**: Each percentile pools every sample across a version's runs and uses the nearest-rank method. Half of the samples are at or below the p50 (the median), three quarters at or below the p75, and all but the slowest 1% at or below the p99. With 100 or fewer samples, a p99 is the slowest sample.

### Limits

- The benchmarks run in one region, against one World. Other regions, Worlds, and workloads will see different numbers.
- The versions share the platform while they are measured. Any exception is noted with the results.
- The explainer and the animations are schematic: they show where each clock starts and stops, not real timings.

## Latest results

> **Latest results:** Workflow SDK v5.2.0 compared with v4.8.12, measured October 8, 2026. We run these benchmarks regularly and publish updated results here.

At p75, v5.2.0 is as fast as or faster than v4.8.12 on every benchmark. The largest change is in long runs: 1,000 steps into a run, v5.2.0 adds 70 ms between steps where v4.8.12 adds 1,748 ms.

Both versions ran at the same time and shared the platform. The exception is the 1,020-step benchmark on v4: one of its five runs ran alongside v5.2.0's, and the other four ran on their own afterwards, the same day with the same configuration. Each version ran the sequential workflow 10 times, the agent stream 10 times, the fan-out 25 times, the resume benchmark 25 times, and the 1,020-step run 5 times.

![Dot plot of how many times faster the newer version is than the older one at p75 on each benchmark, on a log scale. The full values are in the results table below.](/performance/v5.2.0/01-overview-light.png)

### Steps and long runs

At p75, a five-step workflow reaches its first step in 172 ms on v5.2.0 and 418 ms on v4.8.12, and v5.2.0 adds 109 ms between steps where v4.8.12 adds 618 ms.

Through a 1,020-step run, step-to-step overhead on v4.8.12 grows from 622 ms in steps 1 to 20 to 1,748 ms in steps 1,001 to 1,020, at p75. On v5.2.0 it is 110 ms and 70 ms. The median run takes 21.0 min on v4.8.12 and 3.0 min on v5.2.0.

![Three charts comparing the two versions: time to first step and step-to-step overhead for a five-step workflow, step-to-step overhead through a 1,020-step run at p75, and the time to finish all 1,020 steps. The full values are in the results table below.](/performance/v5.2.0/02-sequential-light.png)

### Fan-out

At p75, v5.2.0 starts the first of 64 branches in 360 ms against 478 ms on v4.8.12, then starts the slowest branch at about the same time (996 ms against 972 ms) and joins in 448 ms against 464 ms.

![Dumbbell chart of a 64-branch fan-out at p75, comparing the two versions for the first branch to start, the slowest branch to start, and the join. The full values are in the results table below.](/performance/v5.2.0/03-fanout-light.png)

### Resuming a workflow

At p75, a workflow waiting on a hook resumes in 425 ms on v5.2.0 against 699 ms on v4.8.12.

![Dumbbell chart of time to resume after 1 s idle at p50 and p75, comparing the two versions. The full values are in the results table below.](/performance/v5.2.0/04-resume-light.png)

### Streams

The agent produced its 2,593-chunk response over 52.4 s. On v4.8.12 the writer step manages about 12 chunks per second, so the response takes 226.1 s to reach the reader. v5.2.0 writes streams over a WebSocket by default, keeps up at about 49 chunks per second, and delivers the response in 53.5 s.

At p75, a chunk reaches the reader in 102 ms on v5.2.0 against 172 ms on v4.8.12. v5.2.0 had a chunk delayed by over 1 s in 0 of 10 runs, against 6 of 10 on v4.8.12.

On the Vercel World, [`WORKFLOW_STREAMS_TRANSPORT`](/docs/configuration/worlds#workflow_streams_transport) sets the transport for stream writes.

![Three charts of a replayed agent stream comparing the two versions: the time to deliver the whole response against the recording's length, chunk trip time at p50 and p75, and the number of runs with a chunk delayed by over 1 s. The full values are in the results table below.](/performance/v5.2.0/05-streams-light.png)

### All results

Every measured value from the latest results. Percentiles follow the [methodology](#methodology) above.

| Metric                                          |  v4.8.12 |  v5.2.0 |       Change | Samples per version |
| ----------------------------------------------- | -------: | ------: | -----------: | ------------------: |
| **Sequential workflow, 5 steps, 10 runs**       |          |         |              |                     |
| Time to first step, p75                         |   418 ms |  172 ms |  2.4× faster |             10 runs |
| Time to first step, slowest run                 |   451 ms |  253 ms |  1.8× faster |             10 runs |
| Step-to-step overhead, p75                      |   618 ms |  109 ms |  5.7× faster |             40 gaps |
| Step-to-step overhead, p99                      |   777 ms |  327 ms |  2.4× faster |             40 gaps |
| **Long run, 1,020 steps, 5 runs**               |          |         |              |                     |
| Step overhead, steps 1 to 20, p75               |   622 ms |  110 ms |  5.7× faster |             95 gaps |
| Step overhead, steps 101 to 120, p75            |   719 ms |   74 ms |  9.7× faster |             95 gaps |
| Step overhead, steps 1,001 to 1,020, p75        | 1,748 ms |   70 ms | 25.0× faster |             95 gaps |
| Whole run, median wall time                     | 21.0 min | 3.0 min |  7.0× faster |              5 runs |
| **Fan-out, 64 parallel branches, 25 runs**      |          |         |              |                     |
| First branch starts, p75                        |   478 ms |  360 ms |  1.3× faster |             25 runs |
| Slowest branch to start, p75                    |   972 ms |  996 ms |    No change |             25 runs |
| Join, p75                                       |   464 ms |  448 ms |    No change |             25 runs |
| **Resume after 1 s idle, 25 runs of 5 resumes** |          |         |              |                     |
| Time to resume, p50                             |   619 ms |  383 ms |  1.6× faster |         125 resumes |
| Time to resume, p75                             |   699 ms |  425 ms |  1.6× faster |         125 resumes |
| **Agent stream, 52.4 s recording, 10 runs**     |          |         |              |                     |
| Delivery, median wall time                      |  226.1 s |  53.5 s |  4.2× faster |             10 runs |
| Chunk trip time, p75                            |   172 ms |  102 ms |  1.7× faster |       25,930 chunks |

---

For a semantic overview of all documentation, see [/sitemap.md](/sitemap.md)

For an index of all available documentation, see [/llms.txt](/llms.txt)

For agent-facing discovery, including API and MCP surfaces, see [/agents.md](/agents.md)