Vector embeddings
Vector embeddings represent data (ex. text) as lists of numbers (vectors). Vectors allow you to query for documents basing on their similarity to other documents or to specific vector embeddings (see Vector indexes and Query similar documents).
You can create an embedding field with the @embedding directive on a collection field of type [Float32!]. Fields marked with @embedding become a vector mirror of the document's fields: when a document is added or updated, the corresponding vector embedding is regenerated to match the new content.
The @embedding directive is useful to entirely delegate the generation of embeddings to the database. If you plan to manually generate vector embeddings, use a field of type [Float32!] and don't mark it with the @embedding directive – you'll still be able to create vector indexes on it and run similarity queries.
Syntax
@embedding(
fields: [String]!,
provider: String!,
model: String!,
url: String
)
fields– Collection fields to create embeddings of.
Supported field types are:Float32,Float64,Int,String.provider– Embedding provider.
Supported values:ollama,openai.model– Embedding model (ex.embeddinggemma,text-embedding-3-small).url– (Optional) URL of provider's API.
Default:https://api.openai.com/v1foropenai;http://localhost:11434/apiforollama.
If fields contains several entries, embeddings will be generated even if a document lacks value for some of the fields.
Providers
Ollama
Use the ollama provider to generate embeddings with Ollama. The url field is only needed if you use Ollama models hosted elsewhere than the default http://localhost:11434/api.
type Book {
title: String
plot: String
about_v: [Float32!] @embedding(
fields: ["title", "plot"],
provider: "ollama",
model: "embeddinggemma"
)
}
The selected embedding model must be available in the Ollama instance ahead of usage.
For example, to install embeddinggemma:
ollama pull embeddinggemma
OpenAI
Use the openai provider to generate embeddings using OpenAI. The url field is only needed if you use OpenAI models hosted elsewhere than the default https://api.openai.com/v1.
Provide your OpenAI API key via the environment variable OPENAI_API_KEY. The variable must be defined in the environment in which defradb runs.
type Book {
title: String
plot: String
about_v: [Float32!] @embedding(
fields: ["title", "plot"],
provider: "openai",
model: "text-embedding-3-small"
)
}
Storing embedding vectors
An embedding field is a vector mirror of the fields it encodes:
- when a new document is created, the embedding field gets populated with a vector encoding the content fields
- when a document is updated, the embedding field is regenerated to account for changes in the content fields
The embedding is generated even if a document lacks value for one or more of the embedding fields.
mutation {
add_Book(input: {
title: "Infinite Jest"
}) {
title
plot
about_v
}
}
{
"data": {
"add_Book": [
{
"about_v": [
-0.14905636,
-0.012006985,
0.030906163,
...
(768 entries)
],
"plot": "",
"title": "Infinite Jest"
}
]
}
}
Explicit values
You can also set an embedding value to an explicit vector value. However, the vector value will be overwritten if the document is updated, as the database will sync the source content fields with the embedding value.
mutation {
update_Book(
filter: { title: { _eq: "Infinite Jest" } },
input: { about_v: [1,2,3] }
) {
title
about_v
}
}