Local LLMs (Ollama, LM Studio)
The AI generate endpoint supports a special provider: "local" that points the engine at any OpenAI-compatible chat-completions server you host yourself. The server still does all the heavy lifting (QR rendering, ISO 18004 validation, scan guarantee) — only the language model that proposes the design lives on your machine.
When to use it
- Privacy-first stacks. Prompts never leave your network on their way to the LLM. The Signet API only ever sees the resulting JSON config, not the natural-language description you typed.
- Cost control. You're already running Ollama / LM Studio for other workloads and want to reuse the model you've paid for / downloaded.
- Custom-tuned models. You've fine-tuned a small open model on QR design prompts and want to point Signet at it.
Not a good fit for:
- SaaS callers that try to reach
localhoston the API server — the Signet API runs on Azure Functions and cannot reach your laptop directly. You need a public tunnel (see below). - Hosted Ollama clouds you don't control — for those, use the relevant native provider (
openai,groq, etc.) instead.
Setup
1. Start your local server
Ollama:
ollama serve
# in another shell, pull a model
ollama pull llama3.3
Ollama exposes an OpenAI-compatible API at http://localhost:11434/v1/chat/completions out of the box.
LM Studio:
Load a model in the GUI, open the "Local Server" tab, click "Start Server". Default endpoint is http://localhost:1234/v1.
vLLM / llama.cpp server: Both expose /v1/chat/completions natively. Use whatever host / port you started them on.
2. Make your server publicly reachable
Azure Functions cannot reach localhost on your machine, so you need a tunnel. Two zero-config options:
# ngrok
ngrok http 11434
# → https://your-tunnel.ngrok.io
# Cloudflare Tunnel
cloudflared tunnel --url http://localhost:11434
# → https://random-words.trycloudflare.com
Use the resulting public HTTPS URL as localEndpoint. The API rejects RFC1918 / link-local IPs (10.x, 172.16-31.x, 192.168.x, 169.254.x) to avoid SSRF — only public addresses or localhost (for local dev of the API itself) are accepted.
3. Call the API
curl -X POST https://signetqr-core.azurewebsites.net/api/qr/ai/generate \
-H "Content-Type: application/json" \
-H "X-RapidAPI-Proxy-Secret: $RAPIDAPI_PROXY_SECRET" \
-d '{
"provider": "local",
"localEndpoint": "https://your-tunnel.ngrok.io/v1",
"model": "llama3.3",
"apiKey": "ollama",
"prompt": "Tech startup QR, neon cyan and magenta on near-black, hexagonal modules with subtle glow",
"content": "https://example.com"
}'
Notes:
apiKeyis required by the field validator but its value is irrelevant for most local servers. Ollama accepts any non-empty string; pass"ollama"by convention. LM Studio also ignores it unless you've explicitly configured an API key.modelis required for local — there is no server-side default. Pass exactly the model name your server exposes (llama3.3,qwen2.5:14b,mistral, etc.).localEndpointcan be either the base URL (https://host/v1) or the full path (https://host/v1/chat/completions) — the API normalizes both.- JSON-mode (
response_format: { type: "json_object" }) is not forwarded to local servers. Some runtimes don't support it and would error. The system prompt is strict enough that compliant models still return clean JSON. If your model wraps the JSON in prose, lowercreativitytoward0.0or pick a more instruction-tuned model.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
400 localEndpoint must be a publicly reachable address | You passed a 10.x / 192.168.x / 127.0.0.1 IP from the public internet | Use a tunnel (ngrok, Cloudflare Tunnel) and pass the public URL |
400 For provider='local' the 'llmModel' field is required | You forgot model | Set model to whatever your local server exposes |
500 AI generated an invalid configuration | The local model returned malformed JSON | Try a stronger model (≥ 7B params) or lower creativity |
502 / connection refused from the AI service | Tunnel down / server not running | Restart ollama serve and re-run cloudflared tunnel |
Security notes
- The Signet API never logs the contents of
localEndpointorapiKey. It does log the provider name (local) and the model identifier for billing telemetry. - Treat your tunnel URL as a secret in the same tier as an API key — anyone who has it can call your local model on your hardware.
- For production stacks where the privacy guarantee matters, prefer Cloudflare Tunnel over ngrok: ngrok URLs are guessable and the free tier has no auth.