Blog
RAGLLMAgentsKafkaKubernetes

Quby: RAG, text agents, and voice agents on one stack

·11 min read

This is a study note, not a product pitch. It documents how I built and wired Quby, a multi-tenant lab for customer-facing AI: retrieval-augmented generation (RAG), text agents across channels, and voice agents for live support.

Where to try it (study links, not a sales page):

  • Voice agents only: https://demo.quby.com.br — pick a model in the catalog and start a browser call.
  • Product landing: https://quby.com.br — commercial context and form for a full Quby demo.
  • Signed-in product: app.quby.com.br — dashboard, knowledge upload, and multichannel inbox.
The goal of the study is practical: one knowledge base per organization, one semantic search path reused everywhere, and orchestration that keeps business rules in the backend while the LLM handles language.

What Quby is trying to prove

Quby bundles three ideas that usually get split across repos:

1. Grounded answers — answers should cite org-specific PDFs, price lists, FAQs, and process docs, not generic model memory. 2. Agents with guardrails — qualification, handoff to humans, and routing should survive model swaps and channel changes. 3. Same brain on chat and voice — WhatsApp, webchat, email, SMS, and phone should share leads, conversation history, and RAG where possible.

The stack runs on Kubernetes: PostgreSQL holds transactional data, conversation history, and pgvector embeddings for RAG. Apache Kafka inside the cluster carries asynchronous work (follow-ups, outbound notifications, enrichment jobs, integration events) so HTTP and voice paths stay fast. Clerk handles multi-tenant auth; OpenAI powers embeddings and chat. Voice uses Vapi with server-side tools on the same RAG layer as text. Observability goes through New Relic and Datadog (traces, metrics, and call-path visibility across webhooks and workers).

Architecture at a glance

Channels (webchat, WhatsApp worker, email, SMS, Vapi voice)
        │
        ▼
Route handlers / webhooks  ──►  conversations + messages (PostgreSQL)
        │
        ├── sync path: /api/ai/process (orchestrator)
        │      ├── conversation state machine
        │      ├── optional OpenAI Assistants per org
        │      ├── RAG: buildKnowledgeContextBlock → prompt
        │      └── handoff / routing (rules + LLM assist)
        │
        └── async path: publish to Kafka (Kubernetes)
               ├── follow-up & nurture jobs
               ├── outbound SMS / WhatsApp / email dispatch
               └── enrichment & integration side effects
        │
        ▼
Leads, quotes, inbox UI (dashboard)
        │
        ▼
New Relic + Datadog (traces, metrics, webhook latency)

Multi-tenancy: every row is scoped by organization_id. Knowledge chunks, assistants, and channel credentials are per org. That is what makes the study reusable as a pattern, not a single-tenant chatbot.

RAG: ingest, index, retrieve

The RAG module (lib/org-knowledge.ts) is intentionally small and explicit.

Ingest

  • Operators upload PDF, Markdown, or plain text via POST /api/knowledge/upload.
  • PDF text extraction uses unpdf on the server (avoids browser-only PDF APIs in serverless).
  • Raw text is stored on org_knowledge_documents with an indexing status (processing → ready or error).

Chunking and embeddings

  • Text is split into chunks (~2.4k characters) with paragraph-aware breaks and ~200 characters overlap.
  • Embeddings use OpenAI text-embedding-3-small (1536 dimensions).
  • Vectors land in org_knowledge_chunks, keyed by document and chunk index.
  • Re-index deletes old chunks for that document first (idempotent updates).

Retrieval

  • Queries are embedded with the same model.
  • Similarity search runs in PostgreSQL (pgvector + a scoped SQL function such as search_org_knowledge) with match_count and min_similarity (default 0.55).
  • Hits carry document title, source type, content, and similarity score.

Prompt assembly

buildKnowledgeContextBlock turns hits into a bounded block (default ~3k characters) prefixed as internal company sources. That block is injected into:

  • the main AI orchestrator (/api/ai/process),
  • tenant webchat streaming (lib/webchat-tenant-llm.ts),
  • quote generation and other flows that need catalog or policy context.
This is classic RAG: retrieve first, then generate. The study emphasis is one retrieval function, many callers.

Async work on Kafka (Kubernetes)

Not every side effect belongs in the request that answered the customer. Cron-style follow-ups, proactive outreach, and heavy integration steps are modeled as events on Kafka, consumed by workers in the same Kubernetes cluster as the API.

Typical pattern in the study:

1. The synchronous path commits intent in PostgreSQL (lead updated, quote draft ready, conversation state advanced). 2. The API or a scheduler publishes a message to a topic (org-scoped key when possible). 3. A consumer applies the work idempotently: send outbound message, call ERP webhook, run a second LLM pass, etc.

That split keeps webchat and Vapi tool handlers within tight timeouts while still guaranteeing at-least-once processing with retries on the consumer side. It also mirrors how you would scale agents in production: scale HTTP replicas for traffic, scale consumers for backlog.

Observability (New Relic and Datadog)

Agents fail in boring ways: slow retrieval, duplicate cron runs, webhook auth mismatches, Kafka lag. The study wires New Relic and Datadog as complementary layers:

  • Traces on /api/ai/process, knowledge upload, and Vapi tool endpoints (knowledge-search, ticket creation).
  • Metrics on consumer lag, embedding batch duration, and voice tool error rates.
  • Logs correlated by organization_id and conversation id where privacy policy allows.
Neither tool replaces domain logging inside the app; they make regressions visible when you change prompts, RAG thresholds, or broker topology.

Text agents: state machine plus LLM

A common failure mode is letting the model decide the funnel (“ask for budget now” vs “ask for phone”). Quby inverts that.

lib/conversation-state-machine.ts defines profiles (dealer, quby_product, tenant) and states such as IDENTIFY_INTENT, COLLECT_BUDGET, CREATE_LEAD, HANDOFF_OR_CLOSE. The backend decides the next state from user text and collected fields. The LLM’s job is to phrase the current step naturally and to use RAG when the user asks domain-specific questions.

The orchestrator POST /api/ai/process ties it together:

  • loads conversation state and lead context,
  • evaluates handoff and branch routing (rules plus optional AI-assisted decisions),
  • pulls RAG and cross-channel history blocks into the prompt,
  • calls either Chat Completions or, when configured, runOpenAIAssistantTurn for org-specific OpenAI Assistants.
Autonomous layers (lib/ai-autonomous-decisions.ts) sit on top for “should we transfer to a human?” and “which branch owns this lead?” without replacing the state machine.

Study takeaway: agents here are workflows with LLM skin, not a single prompt pretending to be a CRM.

Voice agents: Vapi plus the same RAG

Phone support adds latency and noisier input. The study integrates Vapi so the voice model can call back into Quby over HTTPS.

Lifecycle webhooks

POST /api/webhook/vapi receives call events (status, transcript, end-of-call reports) for logging and inbox continuity.

Tool: search_knowledge

During a call, the assistant can invoke search_knowledge. Vapi posts to POST /api/webhook/vapi/knowledge-search, which:

  • authenticates with VAPI_WEBHOOK_SECRET,
  • resolves the organization from call metadata, vapi_call_id, or vapi_assistant_id,
  • runs searchOrgKnowledge (same PostgreSQL retrieval as chat),
  • returns a short formatted result (≤800 characters) so the voice model summarizes in one or two sentences.
The tool schema and system-prompt snippet live in lib/vapi-knowledge-search-tool.ts and are auto-attached when provisioning assistants (lib/vapi-auto-provision.ts).

Tool: create lead

POST /api/webhook/vapi/ticket handles create_quby_lead (and legacy ticket names) so a qualified call opens a lead in the same pipeline as webchat.

Study takeaway: voice is not a separate knowledge base; it is RAG exposed as a function the realtime model can call under pressure.

Public voice demo catalog

The study ships a dedicated voice sandbox at demo.quby.com.br. The page is scoped to voice agents only: select a persona on the left, open the call panel, and talk in the browser. It is the fastest way to feel latency, tool use, and persona prompts without signing into the product.

| Agent | Role in the study | |--------|-------------------| | Atendimento (Bob) | Dealership-style support: sales, rental, service, and parts | | Status de frota | Internal fleet ops: collect data and update asset status | | Renovação de locação (Bob) | Executive renewal flow: lease extension and equipment buyout | | Prospecção ativa (Sarah) | Outbound-style qualification: propose equipment and capture intent |

For the full platform (webchat embed, inbox, knowledge upload, org setup), the landing at quby.com.br points to the commercial demo request; the voice URL above is intentionally narrow so readers can test agents in isolation.

Multichannel inbox

Quby’s product pillars (Connect / Automate / Integrate) map to code paths:

| Channel | Entry | Notes | |--------|--------|--------| | Webchat | /api/ai/webchat/stream, public widget | Tenant LLM + RAG on the client site embed | | WhatsApp | worker + /api/webhook/whatsapp-wwebjs | QR-linked worker; unified lead thread | | Email | Gmail OAuth + push handlers | Inbound into same conversations | | SMS | Twilio webhook | Outbound/inbound tied to org numbers | | Voice | Vapi | Tools for RAG + lead creation |

lib/multichannel-ai-autoresponse.ts and lib/cross-channel-context.ts keep autoresponses and history coherent when the same person writes on chat and later calls.

Operating the study safely

Points I kept explicit while building:

  • Secrets stay server-side — no service role or webhook secrets in Client Components.
  • Ground when unsure — voice and chat prompts instruct the model to offer human follow-up when RAG returns nothing useful.
  • Human in the loop for money — quote generation produces drafts; sending and pricing commits stay reviewable in the dashboard.
  • Per-org isolation — vector search and row access always filter by organization_id.

What to try on the demos

1. Voice: open https://demo.quby.com.br, choose Atendimento or Prospecção ativa, and ask a question that should hit search_knowledge (price, warranty, process). Listen for a short answer, not a document read aloud. 2. Full product story: skim https://quby.com.br if you want positioning and the form for a guided demo. 3. Operators: sign in on app.quby.com.br to upload PDFs to the org knowledge base and see the same RAG power text channels in the inbox.

Suggested mental checks:

  • Same org boundary in RAG whether the user typed in webchat or spoke on a Vapi call.
  • Lead creation after qualification on voice (create_quby_lead) should match the webchat funnel philosophy even when the UI differs.

Closing

Quby, as documented here, is a study platform for RAG plus agent orchestration plus voice tools in one multi-tenant stack. The interesting part is not “call OpenAI once”; it is shared retrieval, deterministic funnel state, and channel adapters that all honor the same org boundary.

If you are designing something similar, start with searchOrgKnowledge and buildKnowledgeContextBlock as the single source of truth, then hang chat, voice tools, and quote drafts off that spine.