Tensor Parallelism vs. Pipeline Parallelism vs. Layer Splitting: Comparing 3 Methods for Running LLMs on Multi-GPU Systems

Tensor Parallelism vs. Pipeline Parallelism vs. Layer Splitting: Comparing 3 Methods for Running LLMs on Multi-GPU Systems

Parallelization methods for multi-GPU inference are ways of deciding where to split a model and when the GPUs communicate with each other, in order to run an LLM that doesn't fit on a single GPU, or that runs too slowly on one GPU, across multiple GPUs. The representative methods are three: tensor parallelism (TP), pipeline parallelism (PP), and the layer-splitting approach used by tools such as Ollama.

All three share the characteristic of "using multiple GPUs," but they differ considerably in what gets faster, what kind of connection is required, and what usage patterns they suit. The phenomenon where adding a second GPU doesn't actually speed up the response can also be explained once you understand these differences.

This article is aimed at technical staff within companies who are considering local LLMs or on-premises GPU servers. It organizes how the three methods work using a comparison table, and explains their suitability for PCIe versus NVLink, as well as how to choose between vLLM and Ollama.

What's the Difference Between Tensor Parallelism, Pipeline Parallelism, and Layer Splitting?

What's the Difference Between Tensor Parallelism, Pipeline Parallelism, and Layer Splitting?

The difference between the three methods lies in whether the model is split "within a layer" or "between layers," and in how the GPUs are made to work after the split. We'll first grasp the overall picture using a table, and then look at why these differences arise.

Comparison Table of the 3 Methods

To state the conclusion up front: if you want to speed up the response for a single request, tensor parallelism is the starting point; if you want to handle many requests simultaneously, pipeline parallelism is the starting point; and if you first need to simply get a model running that doesn't fit on one GPU, layer splitting is the starting point.

Comparison axisTensor Parallelism (TP)Pipeline Parallelism (PP)Layer Splitting
Unit of divisionWeight matrices within a layerA contiguous group of layers (a stage)Layers (each GPU is assigned certain layers)
Inter-GPU communicationResults are aggregated across all GPUs for every layer (all-reduce)Handed off to the neighboring GPU at stage boundariesHanded off to the next GPU at the assigned boundary
Amount/frequency of communicationHighLowLow
Response time for a single requestBecomes fasterDoes not become fasterDoes not become faster
Simultaneous handling of many requestsImprovesImproves (assumes a continuous stream of requests)Not the primary purpose
Suitable connectionNVLink / NVSwitchManageable even over PCIeRarely problematic even over PCIe
Main implementation examplesvLLM's --tensor-parallel-sizevLLM's --pipeline-parallel-sizeOllama, llama.cpp defaults

The row in the table that best captures the character of the three methods is "Response time for a single request." With methods that split between layers, no matter how many GPUs you add, a single request still has to pass through the GPUs in the order of the layers. Only tensor parallelism, which splits within a layer, allows all GPUs to work simultaneously on the computation of a single layer.

In exchange, tensor parallelism requires far more frequent communication. Which method is faster depends on the speed of the connection between GPUs, and on whether the system is being used by a single person or by many people at once.

What Creates the Difference: "Where to Split" and "When to Communicate"

An LLM has a structure in which dozens of identically shaped layers (Transformer blocks) are stacked on top of each other. Each time a single token of text is generated, the input passes through these layers in order, from the first to the last. It's easier to make sense of the parallelization methods if you think of them as differing in which direction this stack is cut.

Cutting the stacked layers between layers is what pipeline parallelism and layer splitting do. GPU0 might be assigned layers 1 through 20, and GPU1 layers 21 through 40, for example, so the GPUs only need to exchange data at the boundary where responsibility switches. In exchange, at the moment a given token is being computed, only the GPU responsible for the layer that token is currently in is actually working.

The other way to cut is tensor parallelism, which splits the contents of each layer. Every GPU is responsible for all of the layers, each taking on a portion of the computation for each layer. This lets everyone work simultaneously, but they must bring their partial results together before they can move on to the next layer.

To use an analogy: layer splitting and pipeline parallelism are like an assembly line where each worker is responsible for a different step, while tensor parallelism is like having everyone share a single step and then gather together to compare answers each time that step finishes. Which is faster depends on how long it takes to gather together — that is, on the speed of the connection between GPUs.

How Does Tensor Parallelism (TP) Work, and Where Does It Speed Things Up?

How Does Tensor Parallelism (TP) Work, and Where Does It Speed Things Up?

Tensor parallelism is a method in which the computation of a single layer itself is divided among multiple GPUs. While it can shorten the response time for a single request, it is heavily affected by the speed of the connection between GPUs.

Splitting Matrices Within Layers and Summing Results Per Layer

At the heart of the computation performed in each layer of an LLM is large matrix multiplication. In tensor parallelism, this weight matrix is sliced along the column or row direction and distributed across GPUs, with each one computing its assigned portion simultaneously. With TP=4, each GPU holds roughly one-quarter of each layer's weights.

Computing just the assigned portion doesn't complete the layer's output. The partial results from all GPUs must be summed together and the same result redistributed to everyone—this communication is called all-reduce. The Megatron-LM paper, which presents a representative design for tensor parallelism, explains a configuration in which, for each Transformer layer, all-reduce occurs twice during the forward computation (after the attention mechanism and after the fully connected layer). Megatron-LM paper

If a model has dozens of layers, then for every token generated, all GPUs must synchronize twice that many times. Even so, a single request becomes faster because each GPU only needs to read out the weights for its own assigned portion. During the stage of generating one token at a time, the time spent reading weights from memory tends to matter more than the computation itself, and this reading can be shared across multiple GPUs.

In vLLM, this is specified as --tensor-parallel-size 4. The commonly seen "TP=4" refers to this setting.

Why Gains Are Limited Over PCIe Connections

All-reduce is a communication operation that cannot proceed until everyone's results are ready. Since it occurs every layer, even if the wait time per occurrence is small, it accumulates and directly affects response time. For this reason, the effectiveness of tensor parallelism depends heavily on the speed of the interconnect between GPUs.

For reference, NVIDIA publishes the interconnect bandwidth for its data center GPU H100 (SXM version) as 900GB/s over NVLink and 128GB/s over PCIe Gen5. NVIDIA H100 specifications Actual communication performance varies by configuration, but this shows there can be a large difference in interconnect bandwidth.

The official vLLM documentation also advises that on nodes with GPUs lacking NVLink, such as the L40S, using pipeline parallelism instead of tensor parallelism yields higher throughput and less communication overhead. vLLM official documentation

In configurations with multiple workstation or gaming GPUs installed, it's common for the GPUs to be connected only via PCIe. Tensor parallelism will still work in such environments, but it won't necessarily speed up in proportion to the number of GPUs. Results vary depending on model size and the number of concurrent requests, so the safest approach is to measure the response time for a single request in your own environment before deciding whether to adopt it.

How Does Pipeline Parallelism (PP) Work, and What Is It Suited For?

How Does Pipeline Parallelism (PP) Work, and What Is It Suited For?

Pipeline parallelism divides the model into groups of layers (stages) and processes them sequentially, like a factory conveyor belt. This involves less communication and is easier to handle over PCIe, but it doesn't speed up the response for a single request.

Streaming Groups of Layers in Sequence to Gain Throughput

For example, if an 80-layer model is pipeline-parallelized across 4 GPUs, it's divided into 4 stages: layers 1–20, 21–40, 41–60, and 61–80. Each GPU, after finishing the computation for its own stage, simply passes the intermediate results to the next GPU. There's no need for all-reduce to aggregate results across everyone; communication is limited to sending data to the neighboring GPU at the stage boundaries. This is why it's easy to handle even over PCIe.

This approach proves its strength when requests arrive one after another. While GPU1 is computing the continuation of request A, GPU0 can start on the first half of the next request, B. As long as requests keep flowing in, no stage sits idle, and overall throughput (the amount processed in a given time) increases.

Conversely, when the flow is interrupted, waiting occurs. The periods at the start and end of the pipeline when some GPUs are idle are called "bubbles." The GPipe paper, which proposed pipeline parallelism for training, reduces this waiting by splitting a mini-batch into smaller micro-batches and streaming them through. GPipe paper The same idea applies to inference servers: the more requests that can be lined up and streamed through, the greater the efficiency.

A Single Request Alone Won't Get Faster

What needs attention with pipeline parallelism is the speed when processing just a single request. Request A has no choice but to pass through stages 1 through 4 in order, and at any given moment, only one GPU is computing on A. Compared to a case where a single GPU of the same performance could compute all layers, the time taken for computation is roughly the same, and may even be slightly longer due to the handoffs between stages.

This is in contrast to tensor parallelism, which can shorten the response time for a single request. If the usage is light, such as a few people inside a company using it for chat with few concurrent requests, the throughput advantage of pipeline parallelism is barely realized. The later stages end up spending a long time waiting for the computation of the earlier stages to finish.

In other words, pipeline parallelism is not a method for "making one person faster" but one for "handling many people without stopping." It gains value in usage patterns where requests overlap, such as serving as an internal API called by many departments, or batch-processing documents overnight. If you want to speed up the response for a chat used by a single person, it's a more direct path to consider tensor parallelism or to choose a model small enough to fit on a single GPU.

What Is Layer Splitting (Ollama's Default) For?

What Is Layer Splitting (Ollama's Default) For?

Layer splitting is a method that assigns layers to each GPU and passes processing along in order. It's easier to understand not as a method for making things faster, but as a method for running models that don't fit on a single GPU by combining VRAM across GPUs.

The Goal Is Fitting VRAM, Not Speed

With layer splitting, you simply assign responsibility like "this layer goes to GPU0, and from the next layer onward it's GPU1," and the computation when generating tokens in a single conversation is almost entirely sequential. Once GPU0 finishes its assigned portion, it hands off to GPU1, during which GPU0 waits for its next turn. Even with multiple GPUs, you can hardly expect a single response to get faster.

In exchange, you can combine the VRAM of each GPU. For example, with two 24GB GPUs, you can split weights and the KV cache that wouldn't fit on one card across the two. Since communication only happens at the boundaries between assigned portions, this method tends not to cause problems even over PCIe connections.

On the surface, this looks a lot like pipeline parallelism. The difference lies in the purpose: pipeline parallelism is a method premised on keeping multiple requests flowing to keep the stages filled, while layer splitting is primarily a placement method aimed at simply getting the model to load.

In llama.cpp, the default for --split-mode is layer, which splits layers and the KV cache across GPUs. llama.cpp's llama-server documentation It also incorporates a mechanism that breaks up large batches, such as long prompts, into smaller pieces and overlaps processing across GPUs. The change that added pipeline parallelism to llama.cpp The room for speedup lies in this batch processing portion; the part that generates one token at a time proceeds strictly in order.

If the reason a model doesn't fit on one GPU is its size, it's also worth checking whether quantization can reduce the required memory before adding more GPUs.

Ollama Prioritizes "Fit on One GPU If Possible"

In a multi-GPU environment, Ollama doesn't always distribute a model across GPUs. According to the official FAQ, when Ollama loads a model, it compares the required VRAM against the available capacity, and if the model fits entirely on any single GPU, it loads it onto that GPU alone. This is because it reduces the data passing through the PCI bus during inference, which is usually the fastest approach. Only when the model doesn't fit on a single GPU does Ollama distribute it across all available GPUs. Ollama's FAQ

In other words, Ollama uses layer splitting mainly when "it doesn't fit on one card." If you add a second GPU but GPU utilization doesn't change, it may simply be that the model already fits on a single card.

If you always want to distribute across all GPUs, you can enable the environment variable OLLAMA_SCHED_SPREAD. In Ollama's source code, this is defined as a setting that "always schedules models across all GPUs." Ollama's environment variable definitions However, distributing the model doesn't make a single response faster—as the official FAQ explains, a model that fits on one card is typically faster when loaded onto that single card.

If you want to speed up the response for a single request using multiple GPUs, you'll need to consider an inference server that supports tensor parallelism, such as vLLM.

Which Should You Choose for Your Environment?

Which Should You Choose for Your Environment?

When choosing, it helps to think in this order: "Does it fit on one GPU?", "Which matters more, single-request speed or concurrent throughput?", and "What kind of connection is between the GPUs?"

Order of Decision-Making

The first thing to check is whether the model fits on a single GPU. If it fits, the basic approach is not to split it. If you have multiple GPUs, launching the same model separately on each GPU and distributing requests across them is a more straightforward way to increase overall throughput.

If it doesn't fit, the next step is to decide on your goal.

  1. You want to speed up the response for each individual user: consider tensor parallelism. However, this assumes a fast interconnect such as NVLink
  2. You want to handle many requests concurrently: consider pipeline parallelism. This tends to be effective even over PCIe connections
  3. Just getting it to work is enough, and you have few users: layer splitting is sufficient. With Ollama, this is the default setup with no configuration needed

Finally, check the interconnect. For servers connected via NVLink or NVSwitch, tensor parallelism becomes the first candidate; for setups with only PCIe, pipeline parallelism or layer splitting is more realistic. For models too large to fit on a single node, vLLM's official documentation shows a method of combining tensor parallelism within a node with pipeline parallelism across nodes. vLLM's official documentation

Even the outcome of "I added more GPUs but it didn't get faster" becomes understandable when you review the combination of goal and method in this order.

Configuration Examples by Use Case

When applying this to specific use cases, the following combinations serve as a starting point.

For testing or small-scale use with a workstation equipped with two PCIe-connected GPUs, it's easiest to start by running Ollama's layer splitting. If response slowness is a concern, trying a quantized model or a model small enough to fit on a single GPU tends to be more effective than adding more GPUs.

If you're providing an API to multiple internal departments and have a GPU server connected via NVLink, vLLM's tensor parallelism is a strong option. It can shorten the response time for a single request while also handling concurrent processing.

If you have a server with multiple GPUs but no NVLink, and there are many concurrent requests, it's worth trying vLLM's pipeline parallelism. As mentioned earlier, vLLM's official documentation also recommends this approach for nodes without NVLink.

Regardless of configuration, the most reliable approach is to measure performance under your own workload. Check both the response time for a single request and how many requests can be handled concurrently, using prompt lengths close to those used in production. For guidance on how to position local LLMs relative to cloud APIs, see Local LLM / SLM Deployment Comparison for a detailed discussion.

What Are Common Mistakes in Multi-GPU Inference?

What Are Common Mistakes in Multi-GPU Inference?

Many failures occur when the characteristics of a given approach don't match the intended goal. Here are two representative cases and how to avoid them.

Assuming Adding GPUs Speeds Up Responses

The first thing to keep in mind is the assumption that adding more GPUs—two, then four—should speed up a single response proportionally. As we've seen, with layer splitting and pipeline parallelism, the computation for a single request proceeds in sequence, so increasing the number of GPUs barely changes the response time. Even with tensor parallelism, PCIe connections can introduce increased communication wait times, meaning the benefit may not scale with the number of GPUs.

To avoid this misunderstanding, decide numerically what you want to speed up before adding GPUs. Whether it's the response time for a single request or the number of requests processed per second determines which approach to choose and how to measure it.

When measuring, keep the same prompt, the same output length, and the same number of concurrent connections before and after adding GPUs. For response time, separating the time until the first character appears from the subsequent generation speed makes it easier to see what actually changed. If the results differ from expectations, first reconsider the combination of approach and goal—adding more GPUs should be the last resort.

Mismatch Between GPU Count and Tensor Parallelism Degree

Tensor parallelism also comes with the constraint that you can't freely choose the number of GPUs. This is because the attention mechanism is divided into multiple "heads," and tensor parallelism distributes these heads evenly across GPUs. In vLLM, if the total number of heads isn't divisible by the tensor parallel size, it fails to start and produces the error Total number of attention heads (X) must be divisible by tensor parallel size (Y). See the relevant vLLM discussion

You might think, "I have 3 GPUs, so let's set TP=3," and run into this wall. For models with a head count like 32 or 64, which isn't divisible by 3, TP=3 won't work.

There are two ways to handle this. One is to use a divisible number such as TP=2, and allocate the remaining GPU to another purpose. The other is to use pipeline parallelism, which splits by layer and isn't constrained by head count. Checking the head count of the model you plan to use in its configuration file before buying additional GPUs can help you avoid wasted investment.

Frequently Asked Questions About Multi-GPU Inference

Frequently Asked Questions About Multi-GPU Inference

Finally, let's address some questions that commonly come up when considering a configuration.

Q1: Can Tensor Parallelism Be Used with 2 PCIe-Connected GPUs?

can be used. Inference servers like vLLM can be launched with tensor parallelism even in configurations without NVLink. The issue is the magnitude of the effect: since the per-layer all-reduce goes through PCIe, communication wait time increases.

A single request's response may become somewhat faster, but for use cases with many concurrent requests, as the official vLLM documentation recommends, pipeline parallelism tends to achieve higher throughput. Which one fits better is best determined by actually testing both with your specific model and workload.

Q2: Why Doesn't Adding a Second GPU Speed Up Ollama?

Two reasons are possible. One is that the model fits on a single GPU, and Ollama is simply running it on just that one GPU. The other is that even if it's distributed across two GPUs, since it's layer splitting, the computation for a single response proceeds sequentially. Neither case is a malfunction—both are Ollama behaving exactly as designed.

Q3: Can Tensor Parallelism and Pipeline Parallelism Be Combined?

They can be combined. In vLLM, both settings can be used at the same time, and the official documentation includes an example like --tensor-parallel-size 4 --pipeline-parallel-size 2, which connects tensor parallelism across groups of 4 GPUs with a 2-stage pipeline. In this case, the total number of GPUs used is 8.

A typical setup uses tensor parallelism within a single server connected via NVLink, and pipeline parallelism across the slower paths connecting servers to each other. The idea is to assign communication-heavy processing to the fast path and communication-light processing to the slow path. For a model that fits within a single server, it's sufficient to start with just one of the two approaches and only consider combining them once that proves insufficient.

Summary: Choose a Method Based on Speed, Concurrency, or Fitting the Model

Summary: Choose a Method Based on Speed, Concurrency, or Fitting the Model

Tensor parallelism splits the inside of a layer and computes it simultaneously across all GPUs, shortening the response time for a single request. However, since this generates an all-reduce for every layer, it presupposes a fast connection like NVLink. Pipeline parallelism is a method that flows groups of layers in sequence; it involves less communication, is easier to handle even over PCIe, and demonstrates its strength the more requests you keep flowing through it. Layer splitting is simply a placement method that allocates layers; its purpose is not speed but rather fitting a large model by pooling VRAM together.

Even in multi-GPU environments, Ollama loads a model onto a single GPU if it fits, and uses layer splitting to distribute it across all GPUs when it doesn't fit. If adding GPUs doesn't speed things up, start by suspecting this mechanism.

When deciding on a configuration, check things in this order: whether the model fits on a single GPU, whether you want to speed up a single person's response or handle a large number of people, and what kind of connection exists between the GPUs. If you're considering introducing a local LLM in-house or configuring a GPU server, feel free to contact us as well.

Author & Supervisor

Yusuke Ishihara

Yusuke Ishihara

Started programming at age 13 with MSX. After graduating from Musashi University, worked on large-scale system development including airline core systems and Japan's first Windows server hosting/VPS infrastructure. Co-founded Site Engine Inc. in 2008. Founded Unimon Inc. in 2010 and Enison Inc. in 2025, leading development of business systems, NLP, and platform solutions. Currently focuses on product development and AI/DX initiatives leveraging generative AI and large language models (LLMs).