Connecting OpenAI embeddings to vector databases
What You’ll Need
- n8n Cloud or self-hosted n8n
- Hetzner VPS or Contabo VPS for hosting
- DigitalOcean as alternative
- OpenAI API key (for embeddings)
- A vector database (Pinecone, Weaviate, Milvus, or Qdrant)
- Docker and Docker Compose for containerization
- Node.js 18+ if running workflows locally
Table of Contents
- Understanding Vector Embeddings and Why They Matter
- Setting Up Your Vector Database Infrastructure
- Configuring OpenAI Embeddings in n8n
- Building Your First Embedding Pipeline
- Storing and Retrieving Embeddings at Scale
- Production Deployment Considerations
- Getting Started
I’ve spent the last three years building AI-powered search and retrieval systems, and the most transformative shift in that work came when I stopped thinking about databases as simple key-value stores and started treating them as semantic search engines. Vector databases paired with OpenAI embeddings are the backbone of modern RAG (Retrieval-Augmented Generation) systems, recommendation engines, and intelligent content discovery platforms.
In this guide, I’ll walk you through connecting OpenAI embeddings to vector databases using n8n, and I’ll give you the complete configuration to move from experimentation to production.
Understanding Vector Embeddings and Why They Matter
Before we wire up any infrastructure, let’s be clear about what we’re actually doing. When you send text to OpenAI’s embedding API, it returns a 1536-dimensional vector representing the semantic meaning of that text. This isn’t a hash or a fingerprint. It’s a point in high-dimensional space where semantically similar texts are close to each other.
Traditional databases use exact matching. You search for “cat” and get results with “cat” in them. Vector databases use similarity matching. You search for “feline pet” and get results for “cat,” “kitten,” and “tabby” because they’re all close in vector space. This is why RAG systems using vector embeddings can understand context, nuance, and intent in ways keyword search cannot.
The workflow is straightforward: text goes in, embeddings come out, and those embeddings get stored in a vector database optimized for similarity search. When you query later, your query text gets embedded with the same model, and the database returns the most similar vectors, which correspond to your original documents.
Setting Up Your Vector Database Infrastructure
Your vector database is where embeddings live. I recommend starting with one of these options, depending on your scale and preferences:
Pinecone handles managed infrastructure and scales effortlessly, but costs money immediately. Weaviate runs self-hosted with Docker and gives you full control. Milvus is open-source and runs on any Kubernetes cluster. Qdrant is newer but exceptionally fast and has great Python and JavaScript SDKs.
For this guide, I’m going to use Weaviate because it’s production-grade, self-hostable on a Hetzner VPS, and handles 1536-dimensional vectors without breaking a sweat.
Start by spinning up a VPS with at least 4GB of RAM. Here’s the Docker Compose configuration to run Weaviate:
version: '3.8'
services:
weaviate:
image: semitechnologies/weaviate:latest
restart: always
ports:
- "8080:8080"
- "50051:50051"
environment:
QUERY_DEFAULTS_LIMIT: 25
AUTHENTICATION_ANONYMOUS_ACCESS_ENABLED: 'true'
PERSISTENCE_DATA_PATH: /var/lib/weaviate
DEFAULT_VECTORIZER_MODULE: none
ENABLE_MODULES: 'text2vec-openai'
OPENAI_APIKEY: '${OPENAI_API_KEY}'
volumes:
- weaviate_data:/var/lib/weaviate
command:
- --host
- "0.0.0.0"
- --port
- "8080"
- --scheme
- http
volumes:
weaviate_data:
driver: local
Save this as docker-compose.yml, then run:
export OPENAI_API_KEY="your-openai-api-key-here"
docker compose up -d
Weaviate will be accessible at http://localhost:8080. You can test it immediately:
curl http://localhost:8080/v1/.well-known/ready
If you’re deploying to Hetzner VPS, you’ll want to restrict container resource usage to prevent one service from consuming all memory and crashing other processes. I cover the specifics in my guide on how to restrict Docker container resource usage, but the core addition to your docker-compose.yml is:
services:
weaviate:
# ... existing config ...
deploy:
resources:
limits:
cpus: '2'
memory: 3G
reservations:
cpus: '1'
memory: 1G
Now Weaviate is bounded and won’t destroy your server.
💡 Fast-Track Your Project: Don’t want to configure this yourself? I build custom n8n pipelines and bots. Message me with code SYS3-HUGO.
Configuring OpenAI Embeddings in n8n
n8n has first-class support for both OpenAI and vector databases. If you’re using n8n Cloud, you can skip infrastructure setup entirely. If you’re self-hosting on Hetzner VPS or Contabo VPS, you’ll deploy n8n in the same Docker environment as Weaviate.
Here’s a self-hosted setup for both:
version: '3.8'
services:
postgres:
image: postgres:15-alpine
restart: always
environment:
POSTGRES_DB: n8n
POSTGRES_USER: n8n
POSTGRES_PASSWORD: ${DB_PASSWORD}
volumes:
- postgres_data:/var/lib/postgresql/data
healthcheck:
test: ['CMD-SHELL', 'pg_isready -U n8n']
interval: 10s
timeout: 5s
retries: 5
n8n:
image: n8nio/n8n:latest
restart: always
ports:
- "5678:5678"
environment:
DB_TYPE: postgresdb
DB_POSTGRESDB_HOST: postgres
DB_POSTGRESDB_PORT: 5432
DB_POSTGRESDB_DATABASE: n8n
DB_POSTGRESDB_USER: n8n
DB_POSTGRESDB_PASSWORD: ${DB_PASSWORD}
N8N_HOST: n8n.yourdomain.com
N8N_PORT: 5678
N8N_PROTOCOL: https
NODE_ENV: production
WEBHOOK_URL: https://n8n.yourdomain.com/
depends_on:
postgres:
condition: service_healthy
volumes:
- n8n_data:/home/node/.n8n
deploy:
resources:
limits:
cpus: '2'
memory: 2G
weaviate:
image: semitechnologies/weaviate:latest
restart: always
ports:
- "8080:8080"
environment:
QUERY_DEFAULTS_LIMIT: 25
AUTHENTICATION_ANONYMOUS_ACCESS_ENABLED: 'true'
PERSISTENCE_DATA_PATH: /var/lib/weaviate
DEFAULT_VECTORIZER_MODULE: none
ENABLE_MODULES: 'text2vec-openai'
OPENAI_APIKEY: '${OPENAI_API_KEY}'
volumes:
- weaviate_data:/var/lib/weaviate
deploy:
resources:
limits:
cpus: '2'
memory: 3G
volumes:
postgres_data:
n8n_data:
weaviate_data:
For HTTPS on a public domain, layer Traefik with Docker Compose in front of n8n. That guide shows you exactly how to set up automatic SSL certificates and reverse proxy routing, which is essential for production workflows.
Start everything:
export DB_PASSWORD="your-secure-password"
export OPENAI_API_KEY="your-openai-api-key"
docker compose up -d
Access n8n at port 5678, create your account, and you’re ready to build workflows.
Building Your First Embedding Pipeline
Inside n8n, create a new workflow. We’ll build something practical: a system that takes incoming text (maybe from an HTTP trigger, a CRM, or an email), generates embeddings, and stores them in Weaviate.
Here’s the step-by-step workflow logic:
Step 1: HTTP Trigger (Webhook)
Add an HTTP node with POST method. This becomes your endpoint for ingesting documents. The incoming JSON will look like:
{
"text": "The quick brown fox jumps over the lazy dog",
"document_id": "doc_12345",
"metadata": {
"source": "blog",
"author": "john"
}
}
Step 2: OpenAI Embedding Node
Add the OpenAI node. Select the embeddings model. Configure it as:
Model: text-embedding-3-small (cheaper and faster) or text-embedding-3-large (higher quality)
Input: {{ $json.text }}
The output is a 1536-dimensional vector. n8n handles the array serialization automatically, so you get clean JSON back.
Step 3: Weaviate Vector Database Node
Add a Weaviate node. Connect to your instance:
Host: http://weaviate:8080 (if in same Docker network) or http://your-ip:8080
API Key: Leave blank if authentication is disabled
Configure it to create or update a Weaviate class:
{
"class": "Document",
"description": "Documents with OpenAI embeddings",
"vectorizer": "none",
"properties": [
{
"name": "text",
"dataType": ["text"],
"description": "The original text"
},
{
"name": "document_id",
"dataType": ["text"],
"description": "Unique document identifier"
},
{
"name": "source",
"dataType": ["text"],
"description": "Source of the document"
},
{
"name": "author",
"dataType": ["text"],
"description": "Author of the document"
}
]
}
This schema setup happens once. After that, each insert operation sends:
{
"class": "Document",
"vector": {{ $json.data.embedding }},
"properties": {
"text": "{{ $json.text }}",
"document_id": "{{ $json.document_id }}",
"source": "{{ $json.metadata.source }}",
"author": "{{ $json.metadata.author }}"
}
}
The vector field gets your embedding array, and properties hold the metadata and original text.
Step 4: Response Node
Return confirmation to the caller:
{
"success": true,
"document_id": "{{ $json.document_id }}",
"embedding_dimension": 1536,
"stored_in": "weaviate"
}
That workflow is now live. POST JSON to it, and documents get embedded and stored. Test it:
curl -X POST http://localhost:5678/webhook/my-embedding-webhook \
-H "Content-Type: application/json" \
-d '{
"text": "Machine learning models need massive amounts of training data",
"document_id": "ml_doc_001",
"metadata": {
"source": "research",
"author": "alice"
}
}'
Storing and Retrieving Embeddings at Scale
Once documents are embedded and stored, the real magic happens during retrieval. Create a second workflow for semantic search. This workflow accepts a query string and returns the most similar documents.
Query Workflow:
Step 1: HTTP POST Trigger
Accept:
{
"query": "What is machine learning?",
"limit": 5
}
Step 2: OpenAI Embedding for Query
Use the same embedding model on the query text:
Input: {{ $json.query }}
Step 3: Weaviate Similarity Search
Use the Weaviate node with a GraphQL query:
{
Get {
Document(
nearVector: {
vector: {{ JSON.stringify($json.data.embedding) }}
}
limit: {{ $json.limit || 5 }}
) {
text
document_id
source
author
_additional {
distance
}
}
}
}
The distance metric tells you how similar each result is to your query (0 is identical, 1 is completely different).
Step 4: Response Formatter
Transform the Weaviate results:
{
"query": "{{ $json.query }}",
"results": {{ $json.data.Get.Document.map(item => ({
text: item.text,
document_id: item.document_id,
source: item.source,
author: item.author,
similarity_score: (1 - item._additional.distance).toFixed(3)
})) }}
}
Now you have a semantic search API. Test it:
curl -X POST http://localhost:5678/webhook/semantic-search \
-H "Content-Type: application/json" \
-d '{
"query": "deep learning algorithms",
"limit": 3
}'
You’ll get back documents sorted by semantic similarity, not keyword matching.
For high-traffic scenarios where reliability matters, consider adding resilient webhook endpoints with Redis queues. That guide shows how to buffer requests in Redis, process them asynchronously, and ensure no embeddings are lost during spikes.
Production Deployment Considerations
Before going live, you need to think about scale, cost, and reliability.
Embedding costs add up fast. OpenAI charges per token. At 1.35 million tokens per $1 USD, you’re fine for thousands of documents, but millions get expensive. Cache embeddings aggressively. If you’re embedding the same text twice, you’re wasting money.
Rate limiting is critical. OpenAI allows 3,500 requests per minute on the free tier, more on paid accounts. If you’re ingesting high volumes, add rate limiting inside n8n:
{
"execution_time_ms": 300,
"requests_per_minute": 3000
}
The OpenAI node respects this. Overflow requests queue automatically.
Vector database size matters. Weaviate stores the original text plus the 1536-dimensional vector (roughly 6KB per document). 1 million documents is about 6GB. Plan accordingly. If you’re on a Hetzner VPS with 40GB storage, you can handle about 6.5 million documents.
Backup your embeddings. Docker volumes are ephemeral if not configured with persistent storage. Use the volume mount strategy I showed earlier, or back up to S3 regularly:
docker exec weaviate-container tar czf - /var/lib/weaviate | aws s3 cp - s3://my-bucket/weaviate-backup-$(date +%s).tar.gz
Monitor everything. Watch Weaviate’s query latency with Prometheus, track OpenAI costs, and alert on failed embeddings. n8n has built-in logging and error handling, but you should export execution logs to a time-series database for analysis.
Getting Started
Everything you need is here. If you’re starting from scratch:
- Provision a VPS on Hetzner VPS or DigitalOcean. A 4GB starter plan is sufficient.
- Deploy the Docker Compose stack with Weaviate and n8n Cloud or self-hosted n8n.
- Create your embedding workflow.
- Create your search workflow.
- Test with real data.
- Add monitoring and backups.
The patterns I’ve shown here scale to billions of embeddings. I’ve deployed variants of this architecture for recommendation engines, customer support chatbots, and content moderation systems. The principles stay the same: embed consistently, store efficiently, query semantically.
Outsource Your Automation
Don’t have time? I build production n8n workflows, WhatsApp bots, and fully automated YouTube Shorts pipelines. Hire me on Fiverr, mention SYS3-HUGO for priority. Or DM at chasebot.online.
Want to automate this yourself?
Start with n8n Cloud (free tier available) or self-host on a Hetzner VPS for full control.
Want this engine running on your own VPS?
This blog publishes itself — daily, unattended, on free API tiers. The full engine, Hugo theme, and setup guide are available as System 3.
Get System 3