What Is TensorRT? NVIDIA's Inference SDK Explained

NVIDIA TensorRT is an SDK for optimizing and running deep learning inference on NVIDIA GPUs — it takes a model you already trained in PyTorch, ONNX, or another framework, compiles it into a GPU-specific "engine," and runs that engine with lower latency and higher throughput than the original framework alone. It isn't a training framework and it isn't a replacement for CUDA; it's a compiler and runtime that sits on top of CUDA, and it's the same technology NVIDIA uses to accelerate everything from a single Jetson board to a rack of H100s. This guide explains what's actually inside TensorRT, how a model gets optimized, what TensorRT-LLM adds for language models, and where the whole stack runs in practice.
- What TensorRT actually is
- How TensorRT optimizes a model: build phase and runtime phase
- Getting a model into TensorRT: ONNX, Torch-TensorRT, and Model Optimizer
- TensorRT-LLM: TensorRT for large language models
- Where TensorRT runs: from Jetson to data-center GPUs
- Frequently asked questions about TensorRT
What TensorRT actually is
According to NVIDIA's own developer documentation, TensorRT is "an ecosystem of tools for developers to achieve high-performance deep learning inference," made up of inference compilers, runtimes, and model optimizations. In current usage, "TensorRT" refers to a family of related products rather than one program: the core TensorRT compiler and runtime (for CNNs, diffusion models, transformers and other non-LLM networks), TensorRT-LLM (a separate library purpose-built for large language models), TensorRT Model Optimizer (quantization, pruning, and distillation), TensorRT for RTX (a lightweight build for consumer RTX GPUs), and TensorRT Cloud (a hosted service for generating optimized engines, currently limited to select partners).
The core idea that ties them together: TensorRT never trains a model. It takes a network that has already been trained somewhere else, analyzes its graph, and rebuilds it into a form that runs as fast as possible on a specific NVIDIA GPU. NVIDIA's developer page states that this can speed up inference by up to 36x compared to CPU-only platforms — a vendor figure, not an independent benchmark, but it illustrates the order of magnitude TensorRT targets: it's optimizing an already-trained model for deployment, not changing what the model computes.
How TensorRT optimizes a model: build phase and runtime phase
NVIDIA's architecture documentation describes TensorRT working in two distinct phases. In the build phase, a builder compiles the network and, for every layer, selects the fastest available kernel (a low-level GPU function) for the target GPU; the output is a serialized binary called an engine (also called a plan file). In the runtime phase, a lightweight runtime loads that engine into an application and executes it on the GPU. Because the expensive analysis happens once, at build time, the runtime itself only has to load a pre-optimized plan and run it — that split is where most of the latency and throughput gains come from.
Inside the build phase, NVIDIA's developer page names three specific optimization techniques: quantization, layer and tensor fusion, and kernel tuning.
Precision: from FP32 down to INT4
TensorRT can run a model at several numeric precisions, and picking a lower one is one of the main levers for speed. Per the current TensorRT documentation, the SDK supports mixed precision across FP32, FP16, BF16, FP8, INT8, FP4, and INT4, and TensorRT Model Optimizer adds post-training quantization and quantization-aware training — including techniques such as AWQ — to compress a model into a lower precision with minimal accuracy loss. That's a wider menu than the FP16/INT8 pairing many older explainers still describe; FP8 and FP4 support in particular is tied to newer GPU generations (Hopper and Blackwell), so which precisions are actually usable depends on the target hardware.
Layer and tensor fusion, kernel tuning
"Fusion" means combining multiple operations in the network graph — for example a convolution followed by a bias add and an activation — into a single GPU kernel, so the GPU does one pass over memory instead of several. "Kernel tuning" is the build-time step where TensorRT benchmarks several candidate GPU implementations of a layer on the actual target device and keeps whichever one measures fastest. Both are things the builder decides automatically; a developer doesn't hand-pick fusions or kernels.
Getting a model into TensorRT: ONNX, Torch-TensorRT, and Model Optimizer
TensorRT doesn't read PyTorch or TensorFlow code directly — a model needs to reach it through one of a small number of supported paths.
The primary path, per NVIDIA's architecture overview, is the ONNX interchange format: TensorRT ships its own ONNX parser, and the documented recommendation is to export a model to the latest supported ONNX opset from whichever framework trained it (PyTorch, TensorFlow, JAX, and others), then parse it with that parser. NVIDIA also recommends running Polygraphy's constant folding on the exported ONNX graph first, since that step resolves many of the conversion errors the parser would otherwise hit.
The second path is Torch-TensorRT, described on its own documentation site as compiling PyTorch models for NVIDIA GPUs using TensorRT "with minimal code changes." It supports two workflows: just-in-time compilation through torch.compile (pass backend="tensorrt" and the model compiles on first call), and ahead-of-time export through torch.export, which serializes an optimized module you can later load in Python or in a C++ application via LibTorch. Rather than converting the whole model, Torch-TensorRT accelerates the subgraphs it can convert and lets PyTorch continue to execute the rest natively — which is why NVIDIA can describe an ONNX export as an all-or-nothing conversion, and Torch-TensorRT as a partial, in-place one.
Sitting alongside both import paths is TensorRT Model Optimizer, a separate library for quantization, pruning, speculative-decoding preparation, sparsity, and distillation that compresses a model before it goes into TensorRT or TensorRT-LLM. It replaces NVIDIA's older, deprecated PyTorch and TensorFlow quantization toolkits, and — per NVIDIA's documentation — TensorFlow models need to be exported to ONNX first before Model Optimizer or TensorRT can work with them at all.
TensorRT-LLM: TensorRT for large language models
Standard TensorRT is a general-purpose inference compiler; it isn't the tool NVIDIA points developers toward for serving large language models. That's TensorRT-LLM, described in its own documentation as "NVIDIA's comprehensive open-source library for accelerating and optimizing inference performance of the latest large language models (LLMs) on NVIDIA GPUs," available free on GitHub.
A few things separate it from the base TensorRT workflow described above:
- It exposes a high-level Python LLM API rather than requiring a manual ONNX or Torch-TensorRT conversion, and its architecture is described as "PyTorch-native," letting developers extend or customize supported models with ordinary PyTorch code.
- It targets the specific bottlenecks of LLM serving: in-flight batching and paged attention to keep GPUs busy across many concurrent requests, plus built-in support for tensor, pipeline, and expert parallelism across multiple GPUs and nodes.
- Its quantization is LLM-specific: FP8 support on NVIDIA H100-class GPUs and later, which TensorRT-LLM's documentation says can double throughput and halve memory use versus 16-bit floating point "with minimal impact on model accuracy," and FP4 support on B200-class GPUs for loading model weights in that even lower precision.
The practical distinction for anyone choosing between them: reach for base TensorRT (via ONNX or Torch-TensorRT) when deploying CNNs, diffusion models, or general transformer architectures; reach for TensorRT-LLM specifically when the workload is serving a large language model in production.
Where TensorRT runs: from Jetson to data-center GPUs
TensorRT isn't tied to one class of hardware. NVIDIA's developer page lays out separate, purpose-built entry points depending on the deployment target:
| Deployment target | Hardware examples | What NVIDIA points you to |
|---|---|---|
| Data-center LLM serving | GB100, H100, A100-class GPUs | TensorRT-LLM |
| Data-center / embedded / edge, non-LLM models | Data-center GPUs, NVIDIA DRIVE AGX (automotive), Jetson and NVIDIA IGX (edge) | Base TensorRT SDK |
| Consumer RTX laptops and desktops | GeForce RTX, RTX Pro GPUs | TensorRT for RTX |
| Model compression before deployment | Data-center GPUs (GB100, H100, etc.) | TensorRT Model Optimizer |
On the edge side, the base TensorRT SDK is what runs on NVIDIA Jetson boards like the Orin Nano — our Jetson Orin Nano guide covers getting JetPack and TensorRT running on that specific board, so this piece won't repeat that setup here. On the other end of the spectrum, NVIDIA's own materials describe TensorRT-LLM running on data-center GPUs to serve models in production, the same category of hardware compared in our guide to cloud GPU providers. In between, TensorRT for RTX targets local workstation and desktop hardware — relevant if you're evaluating which GPU to buy for deep learning or planning out a GPU workstation build. NVIDIA also describes TensorRT for RTX as building inference engines directly on the target PC in 15–30 seconds with a total library footprint under 200 MB, which is a different engineering trade-off than the data-center build phase described above: on RTX, the "build once, deploy to identical hardware" model gives way to compiling per-device, since consumer GPUs vary far more than a fleet of identical data-center cards. The broader trade-off between running inference locally versus in the cloud is the subject of our edge AI vs. cloud AI guide.
Frequently asked questions about TensorRT
Is TensorRT free?
The core TensorRT SDK is a free download from the NVIDIA Developer site, and TensorRT-LLM and TensorRT Model Optimizer are open-source and free to use (on GitHub and NVIDIA's PyPI, respectively). Separately, NVIDIA also offers NVIDIA AI Enterprise, a paid production support subscription with a 90-day trial license — that's an optional support layer on top of the free software, not a license fee for TensorRT itself.
What is the difference between CUDA and TensorRT?
CUDA is NVIDIA's general-purpose parallel computing platform and programming model — the foundation that lets code run on the GPU at all. TensorRT is built on top of CUDA: it's a higher-level SDK specifically for compiling and running deep learning inference, handling things like layer fusion, precision selection, and kernel tuning that you would otherwise have to write CUDA code for by hand.
What is TensorRT vs TensorFlow?
They operate at different stages of the pipeline. TensorFlow is a framework for building and training models. TensorRT is an inference optimizer and runtime for models that have already been trained — it doesn't train anything. To move a TensorFlow (or Keras) model into TensorRT, NVIDIA's documented path is to export it to ONNX first, using a tool like tf2onnx, and then parse that ONNX graph with TensorRT.
What is the difference between TensorRT and TensorRT-LLM?
Base TensorRT is a general-purpose deep learning inference compiler used for CNNs, diffusion models, and general transformer architectures, typically reached through an ONNX export or Torch-TensorRT. TensorRT-LLM is a separate, open-source library purpose-built for large language models, with its own high-level Python API, in-flight batching and paged attention for serving many requests at once, and LLM-specific FP8/FP4 quantization on supported GPUs.
Do I need to convert my model to ONNX to use TensorRT?
Not necessarily. ONNX plus TensorRT's ONNX parser is the primary and most broadly supported path, covering models exported from PyTorch, TensorFlow, JAX, and other frameworks. But if you're already working in PyTorch, Torch-TensorRT can compile a model directly through torch.compile or torch.export without an ONNX export step in between, accelerating whatever subgraphs it supports while PyTorch runs the rest.
Does TensorRT run on Jetson devices?
Yes. NVIDIA lists Jetson and NVIDIA IGX alongside data-center GPUs and NVIDIA DRIVE AGX (automotive) as target platforms for the base TensorRT SDK, and separately notes that MATLAB's GPU Coder generates TensorRT inference engines for Jetson, DRIVE, and data-center platforms. See our Jetson Orin Nano guide for setup on that specific board.
Recommended: