Connecting OpenAI embeddings to vector databases

Connecting OpenAI embeddings to vector databases

What You’ll Need

  • n8n Cloud or self-hosted n8n
  • Hetzner VPS or Contabo VPS for hosting
  • DigitalOcean as alternative
  • OpenAI API key (for embeddings)
  • A vector database (Pinecone, Weaviate, Milvus, or Qdrant)
  • Docker and Docker Compose for containerization
  • Node.js 18+ if running workflows locally

Table of Contents

  1. Understanding Vector Embeddings and Why They Matter
  2. Setting Up Your Vector Database Infrastructure
  3. Configuring OpenAI Embeddings in n8n
  4. Building Your First Embedding Pipeline
  5. Storing and Retrieving Embeddings at Scale
  6. Production Deployment Considerations
  7. Getting Started

I’ve spent the last three years building AI-powered search and retrieval systems, and the most transformative shift in that work came when I stopped thinking about databases as simple key-value stores and started treating them as semantic search engines. Vector databases paired with OpenAI embeddings are the backbone of modern RAG (Retrieval-Augmented Generation) systems, recommendation engines, and intelligent content discovery platforms.

In this guide, I’ll walk you through connecting OpenAI embeddings to vector databases using n8n, and I’ll give you the complete configuration to move from experimentation to production.

Understanding Vector Embeddings and Why They Matter

Before we wire up any infrastructure, let’s be clear about what we’re actually doing. When you send text to OpenAI’s embedding API, it returns a 1536-dimensional vector representing the semantic meaning of that text. This isn’t a hash or a fingerprint. It’s a point in high-dimensional space where semantically similar texts are close to each other.

Traditional databases use exact matching. You search for “cat” and get results with “cat” in them. Vector databases use similarity matching. You search for “feline pet” and get results for “cat,” “kitten,” and “tabby” because they’re all close in vector space. This is why RAG systems using vector embeddings can understand context, nuance, and intent in ways keyword search cannot.

The workflow is straightforward: text goes in, embeddings come out, and those embeddings get stored in a vector database optimized for similarity search. When you query later, your query text gets embedded with the same model, and the database returns the most similar vectors, which correspond to your original documents.

Setting Up Your Vector Database Infrastructure

Your vector database is where embeddings live. I recommend starting with one of these options, depending on your scale and preferences:

Pinecone handles managed infrastructure and scales effortlessly, but costs money immediately. Weaviate runs self-hosted with Docker and gives you full control. Milvus is open-source and runs on any Kubernetes cluster. Qdrant is newer but exceptionally fast and has great Python and JavaScript SDKs.

For this guide, I’m going to use Weaviate because it’s production-grade, self-hostable on a Hetzner VPS, and handles 1536-dimensional vectors without breaking a sweat.

Start by spinning up a VPS with at least 4GB of RAM. Here’s the Docker Compose configuration to run Weaviate:

version: '3.8'

services:
  weaviate:
    image: semitechnologies/weaviate:latest
    restart: always
    ports:
      - "8080:8080"
      - "50051:50051"
    environment:
      QUERY_DEFAULTS_LIMIT: 25
      AUTHENTICATION_ANONYMOUS_ACCESS_ENABLED: 'true'
      PERSISTENCE_DATA_PATH: /var/lib/weaviate
      DEFAULT_VECTORIZER_MODULE: none
      ENABLE_MODULES: 'text2vec-openai'
      OPENAI_APIKEY: '${OPENAI_API_KEY}'
    volumes:
      - weaviate_data:/var/lib/weaviate
    command:
      - --host
      - "0.0.0.0"
      - --port
      - "8080"
      - --scheme
      - http

volumes:
  weaviate_data:
    driver: local

Save this as docker-compose.yml, then run:

export OPENAI_API_KEY="your-openai-api-key-here"
docker compose up -d

Weaviate will be accessible at http://localhost:8080. You can test it immediately:

curl http://localhost:8080/v1/.well-known/ready

If you’re deploying to Hetzner VPS, you’ll want to restrict container resource usage to prevent one service from consuming all memory and crashing other processes. I cover the specifics in my guide on how to restrict Docker container resource usage, but the core addition to your docker-compose.yml is:

services:
  weaviate:
    # ... existing config ...
    deploy:
      resources:
        limits:
          cpus: '2'
          memory: 3G
        reservations:
          cpus: '1'
          memory: 1G

Now Weaviate is bounded and won’t destroy your server.

💡 Fast-Track Your Project: Don’t want to configure this yourself? I build custom n8n pipelines and bots. Message me with code SYS3-HUGO.

Configuring OpenAI Embeddings in n8n

n8n has first-class support for both OpenAI and vector databases. If you’re using n8n Cloud, you can skip infrastructure setup entirely. If you’re self-hosting on Hetzner VPS or Contabo VPS, you’ll deploy n8n in the same Docker environment as Weaviate.

Here’s a self-hosted setup for both:

version: '3.8'

services:
  postgres:
    image: postgres:15-alpine
    restart: always
    environment:
      POSTGRES_DB: n8n
      POSTGRES_USER: n8n
      POSTGRES_PASSWORD: ${DB_PASSWORD}
    volumes:
      - postgres_data:/var/lib/postgresql/data
    healthcheck:
      test: ['CMD-SHELL', 'pg_isready -U n8n']
      interval: 10s
      timeout: 5s
      retries: 5

  n8n:
    image: n8nio/n8n:latest
    restart: always
    ports:
      - "5678:5678"
    environment:
      DB_TYPE: postgresdb
      DB_POSTGRESDB_HOST: postgres
      DB_POSTGRESDB_PORT: 5432
      DB_POSTGRESDB_DATABASE: n8n
      DB_POSTGRESDB_USER: n8n
      DB_POSTGRESDB_PASSWORD: ${DB_PASSWORD}
      N8N_HOST: n8n.yourdomain.com
      N8N_PORT: 5678
      N8N_PROTOCOL: https
      NODE_ENV: production
      WEBHOOK_URL: https://n8n.yourdomain.com/
    depends_on:
      postgres:
        condition: service_healthy
    volumes:
      - n8n_data:/home/node/.n8n
    deploy:
      resources:
        limits:
          cpus: '2'
          memory: 2G

  weaviate:
    image: semitechnologies/weaviate:latest
    restart: always
    ports:
      - "8080:8080"
    environment:
      QUERY_DEFAULTS_LIMIT: 25
      AUTHENTICATION_ANONYMOUS_ACCESS_ENABLED: 'true'
      PERSISTENCE_DATA_PATH: /var/lib/weaviate
      DEFAULT_VECTORIZER_MODULE: none
      ENABLE_MODULES: 'text2vec-openai'
      OPENAI_APIKEY: '${OPENAI_API_KEY}'
    volumes:
      - weaviate_data:/var/lib/weaviate
    deploy:
      resources:
        limits:
          cpus: '2'
          memory: 3G

volumes:
  postgres_data:
  n8n_data:
  weaviate_data:

For HTTPS on a public domain, layer Traefik with Docker Compose in front of n8n. That guide shows you exactly how to set up automatic SSL certificates and reverse proxy routing, which is essential for production workflows.

Start everything:

export DB_PASSWORD="your-secure-password"
export OPENAI_API_KEY="your-openai-api-key"
docker compose up -d

Access n8n at port 5678, create your account, and you’re ready to build workflows.

Building Your First Embedding Pipeline

Inside n8n, create a new workflow. We’ll build something practical: a system that takes incoming text (maybe from an HTTP trigger, a CRM, or an email), generates embeddings, and stores them in Weaviate.

Here’s the step-by-step workflow logic:

Step 1: HTTP Trigger (Webhook)

Add an HTTP node with POST method. This becomes your endpoint for ingesting documents. The incoming JSON will look like:

{
  "text": "The quick brown fox jumps over the lazy dog",
  "document_id": "doc_12345",
  "metadata": {
    "source": "blog",
    "author": "john"
  }
}

Step 2: OpenAI Embedding Node

Add the OpenAI node. Select the embeddings model. Configure it as:

Model: text-embedding-3-small (cheaper and faster) or text-embedding-3-large (higher quality)
Input: {{ $json.text }}

The output is a 1536-dimensional vector. n8n handles the array serialization automatically, so you get clean JSON back.

Step 3: Weaviate Vector Database Node

Add a Weaviate node. Connect to your instance:

Host: http://weaviate:8080 (if in same Docker network) or http://your-ip:8080
API Key: Leave blank if authentication is disabled

Configure it to create or update a Weaviate class:

{
  "class": "Document",
  "description": "Documents with OpenAI embeddings",
  "vectorizer": "none",
  "properties": [
    {
      "name": "text",
      "dataType": ["text"],
      "description": "The original text"
    },
    {
      "name": "document_id",
      "dataType": ["text"],
      "description": "Unique document identifier"
    },
    {
      "name": "source",
      "dataType": ["text"],
      "description": "Source of the document"
    },
    {
      "name": "author",
      "dataType": ["text"],
      "description": "Author of the document"
    }
  ]
}

This schema setup happens once. After that, each insert operation sends:

{
  "class": "Document",
  "vector": {{ $json.data.embedding }},
  "properties": {
    "text": "{{ $json.text }}",
    "document_id": "{{ $json.document_id }}",
    "source": "{{ $json.metadata.source }}",
    "author": "{{ $json.metadata.author }}"
  }
}

The vector field gets your embedding array, and properties hold the metadata and original text.

Step 4: Response Node

Return confirmation to the caller:

{
  "success": true,
  "document_id": "{{ $json.document_id }}",
  "embedding_dimension": 1536,
  "stored_in": "weaviate"
}

That workflow is now live. POST JSON to it, and documents get embedded and stored. Test it:

curl -X POST http://localhost:5678/webhook/my-embedding-webhook \
  -H "Content-Type: application/json" \
  -d '{
    "text": "Machine learning models need massive amounts of training data",
    "document_id": "ml_doc_001",
    "metadata": {
      "source": "research",
      "author": "alice"
    }
  }'

Storing and Retrieving Embeddings at Scale

Once documents are embedded and stored, the real magic happens during retrieval. Create a second workflow for semantic search. This workflow accepts a query string and returns the most similar documents.

Query Workflow:

Step 1: HTTP POST Trigger

Accept:

{
  "query": "What is machine learning?",
  "limit": 5
}

Step 2: OpenAI Embedding for Query

Use the same embedding model on the query text:

Input: {{ $json.query }}

Step 3: Weaviate Similarity Search

Use the Weaviate node with a GraphQL query:

{
  Get {
    Document(
      nearVector: {
        vector: {{ JSON.stringify($json.data.embedding) }}
      }
      limit: {{ $json.limit || 5 }}
    ) {
      text
      document_id
      source
      author
      _additional {
        distance
      }
    }
  }
}

The distance metric tells you how similar each result is to your query (0 is identical, 1 is completely different).

Step 4: Response Formatter

Transform the Weaviate results:

{
  "query": "{{ $json.query }}",
  "results": {{ $json.data.Get.Document.map(item => ({
    text: item.text,
    document_id: item.document_id,
    source: item.source,
    author: item.author,
    similarity_score: (1 - item._additional.distance).toFixed(3)
  })) }}
}

Now you have a semantic search API. Test it:

curl -X POST http://localhost:5678/webhook/semantic-search \
  -H "Content-Type: application/json" \
  -d '{
    "query": "deep learning algorithms",
    "limit": 3
  }'

You’ll get back documents sorted by semantic similarity, not keyword matching.

For high-traffic scenarios where reliability matters, consider adding resilient webhook endpoints with Redis queues. That guide shows how to buffer requests in Redis, process them asynchronously, and ensure no embeddings are lost during spikes.

Production Deployment Considerations

Before going live, you need to think about scale, cost, and reliability.

Embedding costs add up fast. OpenAI charges per token. At 1.35 million tokens per $1 USD, you’re fine for thousands of documents, but millions get expensive. Cache embeddings aggressively. If you’re embedding the same text twice, you’re wasting money.

Rate limiting is critical. OpenAI allows 3,500 requests per minute on the free tier, more on paid accounts. If you’re ingesting high volumes, add rate limiting inside n8n:

{
  "execution_time_ms": 300,
  "requests_per_minute": 3000
}

The OpenAI node respects this. Overflow requests queue automatically.

Vector database size matters. Weaviate stores the original text plus the 1536-dimensional vector (roughly 6KB per document). 1 million documents is about 6GB. Plan accordingly. If you’re on a Hetzner VPS with 40GB storage, you can handle about 6.5 million documents.

Backup your embeddings. Docker volumes are ephemeral if not configured with persistent storage. Use the volume mount strategy I showed earlier, or back up to S3 regularly:

docker exec weaviate-container tar czf - /var/lib/weaviate | aws s3 cp - s3://my-bucket/weaviate-backup-$(date +%s).tar.gz

Monitor everything. Watch Weaviate’s query latency with Prometheus, track OpenAI costs, and alert on failed embeddings. n8n has built-in logging and error handling, but you should export execution logs to a time-series database for analysis.

Getting Started

Everything you need is here. If you’re starting from scratch:

  1. Provision a VPS on Hetzner VPS or DigitalOcean. A 4GB starter plan is sufficient.
  2. Deploy the Docker Compose stack with Weaviate and n8n Cloud or self-hosted n8n.
  3. Create your embedding workflow.
  4. Create your search workflow.
  5. Test with real data.
  6. Add monitoring and backups.

The patterns I’ve shown here scale to billions of embeddings. I’ve deployed variants of this architecture for recommendation engines, customer support chatbots, and content moderation systems. The principles stay the same: embed consistently, store efficiently, query semantically.

Outsource Your Automation

Don’t have time? I build production n8n workflows, WhatsApp bots, and fully automated YouTube Shorts pipelines. Hire me on Fiverr, mention SYS3-HUGO for priority. Or DM at chasebot.online.

Want to automate this yourself?

Start with n8n Cloud (free tier available) or self-host on a Hetzner VPS for full control.

Want this engine running on your own VPS?

This blog publishes itself — daily, unattended, on free API tiers. The full engine, Hugo theme, and setup guide are available as System 3.

Get System 3
system online