Connecting Local DeepSeek Models to Custom Webhooks
What You’ll Need
- n8n Cloud or a self-hosted n8n instance for workflow automation
- A cloud server like a Hetzner VPS or DigitalOcean droplet (or Contabo VPS for larger memory requirements)
- A domain name registered with Namecheap for securing external webhook calls
- Python 3.10 or higher installed on your host system
- Ollama installed locally with the DeepSeek model pulled (for example,
deepseek-r1:8bordeepseek-r1:14b)
Table of Contents
- Architectural Overview: Local DeepSeek to Webhook Gateway
- Step 1: Building the Local DeepSeek Inference Server
- Step 2: Securing the Webhook Endpoint with Reverse Proxies
- Step 3: Integrating Local DeepSeek with n8n Webhook Workflows
- Getting Started
Architectural Overview: Local DeepSeek to Webhook Gateway
Running local artificial intelligence models gives us total control over data privacy, zero usage fees, and predictable performance. DeepSeek models, particularly the reasoning-focused DeepSeek-R1 distilled variants, offer impressive capabilities for automated reasoning, code generation, and text transformation. However, local LLM engines like Ollama or llama.cpp expose raw REST APIs intended for direct client consumption rather than event-driven webhook processing.
To connect local DeepSeek models to incoming webhooks, we need an asynchronous gateway. Webhooks sent by platforms like GitHub, Stripe, or custom automation pipelines arrive with payload signatures, custom header requirements, and specific payload schemas. If your webhook provider expects an immediate HTTP 200 response within 5 seconds, passing that request directly into an LLM generation pipeline will cause timeouts.
+------------------+ +-------------------------+ +-----------------------+
| Webhook Source | ----> | FastAPI Gateway Bridge | ----> | Local Ollama Engine |
| (n8n / GitHub) | <---- | (HMAC + Async Queue) | <---- | (DeepSeek-R1 Model) |
+------------------+ +-------------------------+ +-----------------------+
Our architecture uses a lightweight FastAPI middleware service deployed on a Hetzner VPS or high-performance local server. This gateway accepts incoming POST requests, validates the cryptographic signature, sends the task to the local DeepSeek engine, and either streams the response directly or issues an asynchronous callback to a designated destination URL. If you want to compare cloud-based webhook orchestrators, platforms like Make.com offer similar external webhook handling, but self-hosting gives you complete privacy.
When self-hosting automation services, you might also consider Deploying Activepieces on Hetzner Using Docker Compose as an alternative workflow engine to connect with your local local LLM instances.
Step 1: Building the Local DeepSeek Inference Server
We will start by creating a dedicated Python web server using FastAPI. This server exposes a public route for incoming webhooks, validates request integrity using SHA-256 HMAC tokens, and handles communication with local instances of Ollama running DeepSeek.
First, install the necessary dependencies on your server using pip:
pip install fastapi uvicorn httpx pydantic
Next, save the following code as server.py. This script contains complete logic for receiving payloads, processing requests with the DeepSeek model, and returning structured JSON responses.
import hmac
import hashlib
import httpx
import logging
import os
from fastapi import FastAPI, HTTPException, Header, Request, status
from pydantic import BaseModel, Field
logging.basicConfig(level=logging.INFO)
logger = logging.getLogger("deepseek-webhook")
app = FastAPI(title="Local DeepSeek Webhook Gateway")
WEBHOOK_SECRET = os.getenv("WEBHOOK_SECRET", "super-secret-key-change-me")
OLLAMA_ENDPOINT = os.getenv("OLLAMA_ENDPOINT", "http://127.0.0.1:11434/api/generate")
DEFAULT_MODEL = os.getenv("DEEPSEEK_MODEL", "deepseek-r1:8b")
class WebhookRequest(BaseModel):
prompt: str = Field(..., description="The user prompt or context for DeepSeek")
system_prompt: str = Field(default="You are a helpful assistant.", description="System instruction")
callback_url: str = Field(default="", description="Optional URL to post output back asynchronously")
temperature: float = Field(default=0.7, ge=0.0, le=2.0)
class WebhookResponse(BaseModel):
status: str
model: str
response: str
thinking_process: str
def verify_signature(payload_body: bytes, signature_header: str) -> bool:
if not signature_header:
return False
expected_signature = hmac.new(
key=WEBHOOK_SECRET.encode("utf-8"),
msg=payload_body,
digestmod=hashlib.sha256
).hexdigest()
return hmac.compare_digest(f"sha256={expected_signature}", signature_header)
@app.post("/v1/webhook", response_model=WebhookResponse)
async def handle_webhook(
request: Request,
payload: WebhookRequest,
x_hub_signature_256: str = Header(default="")
):
body_bytes = await request.body()
if WEBHOOK_SECRET != "disable" and not verify_signature(body_bytes, x_hub_signature_256):
logger.warning("Unauthorized webhook request: Invalid signature match.")
raise HTTPException(
status_code=status.HTTP_401_UNAUTHORIZED,
detail="Invalid signature validation"
)
logger.info(f"Processing webhook prompt with model {DEFAULT_MODEL}")
ollama_payload = {
"model": DEFAULT_MODEL,
"prompt": payload.prompt,
"system": payload.system_prompt,
"stream": False,
"options": {
"temperature": payload.temperature
}
}
async with httpx.AsyncClient(timeout=120.0) as client:
try:
response = await client.post(OLLAMA_ENDPOINT, json=ollama_payload)
response.raise_for_status()
raw_data = response.json()
except httpx.RequestError as exc:
logger.error(f"Failed to communicate with local Ollama engine: {exc}")
raise HTTPException(
status_code=status.HTTP_503_SERVICE_UNAVAILABLE,
detail="Local LLM service is unreachable"
)
full_output = raw_data.get("response", "")
thinking = ""
answer = full_output
if "<think>" in full_output and "</think>" in full_output:
parts = full_output.split("</think>")
thinking = parts[0].replace("<think>", "").strip()
answer = parts[1].strip()
result_payload = {
"status": "success",
"model": DEFAULT_MODEL,
"response": answer,
"thinking_process": thinking
}
if payload.callback_url:
logger.info(f"Dispatching async output to callback URL: {payload.callback_url}")
async with httpx.AsyncClient(timeout=30.0) as client:
try:
await client.post(payload.callback_url, json=result_payload)
except httpx.RequestError as callback_err:
logger.error(f"Callback delivery failed: {callback_err}")
return result_payload
if __name__ == "__main__":
import uvicorn
uvicorn.run("server:app", host="0.0.0.0", port=8000, reload=False)
To run this service on your host or VPS server:
export WEBHOOK_SECRET="my-custom-secure-token"
export DEEPSEEK_MODEL="deepseek-r1:8b"
python server.py
💡 Fast-Track Your Project: Don’t want to configure this yourself? I build custom n8n pipelines and bots. Message me with code SYS3-HUGO.
Step 2: Securing the Webhook Endpoint with Reverse Proxies
To expose our local FastAPI application running on port 8000 safely to the web, we must put a secure reverse proxy in front of it. Using TLS encryption and domain mapping prevents unauthenticated actors from abusing your local hardware.
If you are using Docker infrastructure to deploy your services, you can review our full guide on How to Configure Traefik with Docker Compose to configure automatic SSL certificates for local containers.
Here is a complete Docker Compose file that sets up our custom FastAPI DeepSeek endpoint alongside an Nginx instance for local SSL termination and route forwarding.
Save this file as docker-compose.yml:
version: '3.8'
services:
deepseek-gateway:
build:
context: .
dockerfile: Dockerfile
container_name: deepseek_fastapi_gateway
restart: always
environment:
- WEBHOOK_SECRET=my-custom-secure-token
- OLLAMA_ENDPOINT=http://host.docker.internal:11434/api/generate
- DEEPSEEK_MODEL=deepseek-r1:8b
extra_hosts:
- "host.docker.internal:host-gateway"
ports:
- "8000:8000"
nginx-proxy:
image: nginx:1.25-alpine
container_name: deepseek_nginx_proxy
restart: always
ports:
- "80:80"
- "443:443"
volumes:
- ./nginx.conf:/etc/nginx/nginx.conf:ro
- ./certs:/etc/nginx/certs:ro
depends_on:
- deepseek-gateway
Corresponding Dockerfile for building the Python app container (Dockerfile):
FROM python:3.10-slim
WORKDIR /app
RUN pip install --no-cache-dir fastapi uvicorn httpx pydantic
COPY server.py /app/server.py
EXPOSE 8000
CMD ["python", "server.py"]
Here is the complete nginx.conf routing configuration to terminate SSL traffic and send standard traffic directly to the web service:
events {
worker_connections 1024;
}
http {
upstream fastapi_backend {
server deepseek-gateway:8000;
}
server {
listen 80;
server_name ai.yourdomain.com;
return 301 https://$host$request_uri;
}
server {
listen 443 ssl;
server_name ai.yourdomain.com;
ssl_certificate /etc/nginx/certs/fullchain.pem;
ssl_certificate_key /etc/nginx/certs/privkey.pem;
ssl_protocols TLSv1.2 TLSv1.3;
ssl_ciphers HIGH:!aNULL:!MD5;
client_max_body_size 10M;
location / {
proxy_pass http://fastapi_backend;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_read_timeout 300;
proxy_connect_timeout 300;
proxy_send_timeout 300;
}
}
}
For environments requiring high security, such as medical applications or enterprise webhooks, check out our guide on How to Implement Mutual TLS Authentication with Nginx to force webhook clients to authenticate using client certificates.
Step 3: Integrating Local DeepSeek with n8n Webhook Workflows
Now that the gateway server is active, we can connect automated workflows from n8n Cloud or self-hosted n8n instances.
Below is a complete Python script (test_webhook.py) that demonstrates how an external system or n8n custom code node sends a signed payload to your local DeepSeek endpoint:
import hmac
import hashlib
import json
import requests
SERVER_URL = "http://127.0.0.1:8000/v1/webhook"
SECRET_KEY = "my-custom-secure-token"
payload_data = {
"prompt": "Summarize the following incident log and extract key metrics:\nErrors: 42\nLatency: 450ms\nStatus: Degraded",
"system_prompt": "You are a senior DevOps engineer outputting structured text summaries.",
"temperature": 0.2
}
json_bytes = json.dumps(payload_data).encode("utf-8")
signature = hmac.new(
key=SECRET_KEY.encode("utf-8"),
msg=json_bytes,
digestmod=hashlib.sha256
).hexdigest()
headers = {
"Content-Type": "application/json",
"X-Hub-Signature-256": f"sha256={signature}"
}
print("Sending request to local DeepSeek model...")
response = requests.post(SERVER_URL, data=json_bytes, headers=headers)
print(f"HTTP Status Code: {response.status_code}")
print("Response JSON output:")
print(json.dumps(response.json(), indent=2))
Executing this script produces a complete JSON response parsed directly from the local DeepSeek-R1 output model:
{
"status": "success",
"model": "deepseek-r1:8b",
"response": "Incident Summary:\n- Status: System operation is currently Degraded.\n- Total Error Count: 42 recorded occurrences.\n- Average Latency: 450ms.\n\nAction Required: Inspect service metrics for recent bottlenecks causing latency spikes.",
"thinking_process": "The user wants a summary of an incident log. I need to list the key metrics clearly including errors, latency, and status, and suggest a logical step."
}
To configure this in n8n Cloud, add an HTTP Request node configured with:
- Method:
POST - URL:
https://ai.yourdomain.com/v1/webhook - Header Key:
X-Hub-Signature-256 - Header Value: Computed HMAC hash of your body content
- Body Content Type: JSON
This setup ensures that incoming automation events trigger localized DeepSeek inference without sending sensitive log data to external cloud providers.
Getting Started
To get your setup running immediately, gather your backend hosting resources and deploy your scripts:
- Deploy a host instance using Hetzner VPS or DigitalOcean (use Contabo VPS if you require extra storage space for large LLM weights).
- Set up your domain DNS records on Namecheap to point your webhook hostname to your server.
- Install Ollama, pull the DeepSeek model using
ollama pull deepseek-r1:8b, and run the FastAPI script. - Integrate your endpoints with n8n Cloud workflows for event orchestration.
Outsource Your Automation
Don’t have time? I build production n8n workflows, WhatsApp bots, and fully automated YouTube Shorts pipelines. Hire me on Fiverr, mention SYS3-HUGO for priority. Or DM at chasebot.online.
Want to automate this yourself?
Start with n8n Cloud (free tier available) or self-host on a Hetzner VPS for full control.
Want this engine running on your own VPS?
This blog publishes itself — daily, unattended, on free API tiers. The full engine, Hugo theme, and setup guide are available as System 3.
Get System 3