What Are Embeddings? Text, Image, and Multimodal Vectors

An embedding is a list of numbers — a vector — that represents a piece of data (a word, a sentence, an image) so that items with similar meaning end up close together in that numerical space. Instead of comparing raw text or pixels, a model compares vectors, which is what makes search, clustering, and recommendation systems work on meaning rather than exact keyword matches. This guide covers how embeddings are built, how text, image, and multimodal embeddings differ, how "similar" gets measured mathematically, and the dimensionality trade-offs that decide how much an embedding costs to store and search.
What is an embedding?
Formally, a word (or sentence, or image) embedding is a real-valued vector that encodes an item's meaning so that items closer together in the vector space are expected to be more similar. Embedding models learn this mapping from data — through neural networks, dimensionality reduction on co-occurrence statistics, or probabilistic models — rather than having it hand-built.
It helps to see what embeddings replace. A one-hot encoding represents each item in a vocabulary as a vector of mostly zeros with a single 1, and that scheme has three concrete problems: the vector length equals the size of the vocabulary (tens of thousands of entries for a real dataset), the number of weights a downstream model has to learn scales with that length, and — most importantly — a one-hot vector for "hot dog" is exactly as different from "shawarma" as it is from "salad." There's no built-in notion of similarity. An embedding model learns a dense, lower-dimensional vector instead, one where the geometry of the space captures relationships the one-hot representation couldn't: closer for "hot dog" and "shawarma," farther for "hot dog" and "salad."
OpenAI's own framing is simpler: an embedding is a vector of floating point numbers, and the distance between two vectors measures how related they are — small distance, high relatedness; large distance, low relatedness. That single property is the entire basis for using embeddings in search, clustering, recommendations, anomaly detection, and classification.
From word vectors to contextual embeddings
The idea behind embeddings is older than deep learning. Linguist John Rupert Firth's 1957 observation that "a word is characterized by the company it keeps" is the founding idea of distributional semantics: you can infer what a word means from the words that tend to appear around it. Early implementations of that idea — vector space models for information retrieval, then latent semantic analysis in the late 1980s — produced sparse, high-dimensional vectors from word co-occurrence statistics.
Static embeddings: one vector per word
The shift toward the embeddings used today came in 2013, when a team at Google led by Tomas Mikolov released word2vec, a toolkit that trained dense word vectors far faster than earlier approaches and made embeddings practical at scale. Word2vec and its contemporaries (like GloVe) produce static embeddings: each word gets exactly one vector, learned once from a training corpus and reused everywhere that word appears. The word "bank" gets a single vector whether the sentence is about a river or a loan — a limitation that static embeddings simply can't resolve.
Contextual embeddings: the vector depends on the sentence
BERT (Bidirectional Encoder Representations from Transformers), introduced by Devlin et al. in 2018, addressed exactly that limitation. Instead of one fixed vector per word, BERT computes a representation for each token by jointly conditioning on the words to its left and right in every layer of the network, so the same word gets a different vector depending on what surrounds it. That architecture — pre-trained on unlabeled text, then fine-tuned with a single added output layer — pushed state-of-the-art results across eleven NLP tasks at the time, including an F1 score of 93.2 on the SQuAD v1.1 question-answering benchmark. Modern text embedding models, including the ones behind most current semantic search and retrieval systems, build on this contextual approach: the embedding of a sentence reflects the whole sentence, not just a lookup of its individual words.
Text, image, and multimodal embeddings
Text embeddings
A text embedding model takes a string — a word, a sentence, a full document up to some token limit — and returns a fixed-length vector. OpenAI's current models, for example, accept inputs up to 8,192 tokens and return a 1,536-number vector by default for the smaller model or a 3,072-number vector for the larger one. Under the hood, generating one means sending the text to the model's API and getting back a list of floating-point numbers, which is then typically stored so it can be compared against other vectors later.
Image embeddings
Image embedding models follow the same principle with a different input: a vision model processes pixels and outputs a vector positioned so that visually or semantically similar images land near each other — two photos of golden retrievers close together, a photo of a golden retriever and a spreadsheet far apart. On their own, image embeddings support reverse image search and visual clustering, but they don't understand text queries — that requires a model trained jointly on both.
Multimodal embeddings
Multimodal embeddings put text and images (or other modalities) into the same vector space, so a text query and an image can be compared directly. CLIP, published by OpenAI researchers in 2021, is the model that popularized this approach: it was trained on 400 million (image, text) pairs collected from the internet, with the pre-training task simply being to predict which caption goes with which image. The result is a model that transfers to new tasks without task-specific training — the paper reports it matching the original ResNet-50's accuracy on ImageNet in a zero-shot setting, without using any of that model's 1.28 million training images. That joint-space property is what powers "search images by typing a sentence" and is one of the building blocks behind broader multimodal AI systems; for a closer look at the model architectures that make this possible, see our guide to multimodal AI models.
How similarity between embeddings is measured
Once two items are vectors, "how similar are they" becomes a geometry question. Three measures cover almost every use case:
| Measure | What it captures | Formula | As similarity increases |
|---|---|---|---|
| Cosine similarity | The angle between two vectors, ignoring their length | (aᵀb) / (‖a‖·‖b‖) | Value increases toward 1 |
| Dot product | Angle and magnitude — vectors with larger length score higher | a₁b₁ + a₂b₂ + … + aₙbₙ | Value increases (also increases with vector length) |
| Euclidean distance | Straight-line distance between the two vector endpoints | √((a₁−b₁)² + … + (aₙ−bₙ)²) | Value decreases |
The choice matters less than it looks once vectors are normalized to unit length: after scaling both vectors so ‖a‖ = ‖b‖ = 1, cosine similarity, dot product, and (squared) Euclidean distance all become proportional to the same quantity, the cosine of the angle between them — so they rank results identically. That's exactly why OpenAI's embedding models return normalized vectors: cosine similarity between a query and a set of candidate documents (the same operation used for semantic search) reduces to a plain dot product, which is cheaper to compute at scale. The one case where the choice still matters is when vectors are not normalized and their length is meaningful — for example if longer or more "popular" items naturally produce larger-magnitude vectors, dot product will favor them in a way cosine similarity won't.
What embeddings are used for
The same underlying operation — turn items into vectors, then measure distance — supports a handful of distinct jobs:
Semantic search and retrieval
A query is embedded with the same model used for the documents, then compared against every stored document vector; the closest matches are returned regardless of whether they share exact keywords with the query. This is also the retrieval step inside retrieval-augmented generation pipelines, where the retrieved passages are handed to a language model as context.
Clustering and anomaly detection
Grouping vectors that land near each other in the embedding space clusters items by similarity without predefined categories; items that sit far from every cluster are flagged as outliers — the same distance measures above, applied to find what doesn't fit rather than what matches.
Recommendations
If a user has engaged with an item, the nearest neighbors of that item's embedding are natural candidates to recommend — the same mechanism behind "customers also liked," applied to whatever the embedding represents.
Classification
A piece of text can be classified by comparing its embedding against the embeddings of a fixed set of labels and assigning whichever is closest, without training a dedicated classifier for every new label set.
Dimensionality and cost trade-offs
Every embedding has a fixed number of dimensions, and that number is a direct trade-off between quality and cost. Larger vectors generally score better on benchmarks but cost more to store, move, and compare at scale — storing and searching millions of 3,072-number vectors consumes meaningfully more memory and compute than the same count of 256-number vectors. OpenAI's own comparison makes the trade-off concrete: on the MTEB benchmark, its smaller current-generation model scores 62.3% at 1,536 dimensions, its larger model scores 64.6% at 3,072 dimensions, and the previous-generation model scores 61.0% at 1,536 dimensions — and, notably, the larger current model can be shortened to just 256 dimensions and still outperform the full, unshortened 1,536-dimension previous-generation model. That shortening is exposed as a dimensions parameter at generation time, letting a team pick a smaller, cheaper vector without retraining anything; if dimensions are cut manually after the fact instead, the resulting vector needs to be renormalized before similarity comparisons will behave correctly.
In practice this means dimensionality is a deployment decision, not just a modeling one: a system computing and comparing embeddings at real scale is also a compute-provisioning problem, which is where choices about cloud GPU capacity or a dedicated GPU for deep learning workloads start to matter alongside the embedding model itself.
Storing and searching embeddings at scale
Generating embeddings is only half the system. Comparing a query vector against a handful of documents is trivial; comparing it against millions of them one-by-one is not, which is why production systems store vectors in a purpose-built vector database that indexes them for fast approximate nearest-neighbor search instead of a brute-force scan. That storage and indexing layer is a large enough topic on its own that it's covered separately — this guide focuses on what an embedding is and how to reason about it, not on which database to run it in.
Frequently asked questions about embeddings
What are embeddings in AI?
In AI, an embedding is a vector of real numbers that represents a piece of data — a word, sentence, image, or other item — positioned so that items with similar meaning end up close together in that vector space. Models compare these vectors mathematically instead of comparing raw text or pixels directly, which is what makes semantic search, clustering, and recommendation systems possible.
What is the meaning of "embedding" as opposed to just encoding data?
Encoding is simply choosing an initial numerical representation for data — one-hot encoding, for instance, assigns each item an arbitrary vector of zeros with a single 1, with no relationship between the codes for similar items. An embedding is a learned, dense representation instead, where a model has organized the space so that distance reflects actual similarity in meaning, not just distinctness.
What is an example of an embedding?
A short piece of text sent to an embedding API comes back as a plain list of floating-point numbers — for example, OpenAI's embeddings API returns values like -0.006929283495992422, -0.005336422007530928, ..., with 1,536 or 3,072 such numbers depending on the model. That list is the embedding; two lists that are close together (by cosine similarity or another distance measure) represent text with related meaning.
What are embeddings in ChatGPT?
ChatGPT's chat interface doesn't expose embeddings directly to users. What people usually mean by "embeddings in ChatGPT" is OpenAI's separate embeddings API (the text-embedding-3-small and text-embedding-3-large models), which developers use to build search, retrieval, and recommendation features around chat-based products. Internally, transformer language models also use token and positional embeddings as an input representation, but that's a distinct, lower-level use of the term from the embeddings API most builders interact with.
Are embeddings the same regardless of context?
No, and the difference matters. Static embeddings, like the ones produced by word2vec, assign a single fixed vector to each word regardless of how it's used. Contextual embeddings, the approach introduced by BERT and used by modern text embedding models, compute a different vector for the same word depending on the surrounding text, which is why they capture ambiguous words and nuanced meaning far better than static approaches.
What's the difference between a vector and an embedding?
Every embedding is a vector, but not every vector is an embedding. A vector is simply an ordered list of numbers; an embedding is a vector specifically produced by a model so that its position in space encodes semantic meaning — the geometric relationship between embeddings is meaningful, while an arbitrary vector's coordinates may not be.
Recommended: