Quby: RAG, text agents, and voice agents on one stack
This is a study note, not a product pitch. It documents how I built and wired Quby, a multi-tenant lab for customer-facing AI: retrieval-augmented generation (RAG), text agents across channels, and voice agents for live support.
Where to try it (study links, not a sales page):
- Voice agents only: https://demo.quby.com.br — pick a model in the catalog and start a browser call.
- Product landing: https://quby.com.br — commercial context and form for a full Quby demo.
- Signed-in product:
app.quby.com.br— dashboard, knowledge upload, and multichannel inbox.
What Quby is trying to prove
Quby bundles three ideas that usually get split across repos:
1. Grounded answers — answers should cite org-specific PDFs, price lists, FAQs, and process docs, not generic model memory. 2. Agents with guardrails — qualification, handoff to humans, and routing should survive model swaps and channel changes. 3. Same brain on chat and voice — WhatsApp, webchat, email, SMS, and phone should share leads, conversation history, and RAG where possible.
The stack runs on Kubernetes: PostgreSQL holds transactional data, conversation history, and pgvector embeddings for RAG. Apache Kafka inside the cluster carries asynchronous work (follow-ups, outbound notifications, enrichment jobs, integration events) so HTTP and voice paths stay fast. Clerk handles multi-tenant auth; OpenAI powers embeddings and chat. Voice uses Vapi with server-side tools on the same RAG layer as text. Observability goes through New Relic and Datadog (traces, metrics, and call-path visibility across webhooks and workers).
Architecture at a glance
Channels (webchat, WhatsApp worker, email, SMS, Vapi voice)
│
▼
Route handlers / webhooks ──► conversations + messages (PostgreSQL)
│
├── sync path: /api/ai/process (orchestrator)
│ ├── conversation state machine
│ ├── optional OpenAI Assistants per org
│ ├── RAG: buildKnowledgeContextBlock → prompt
│ └── handoff / routing (rules + LLM assist)
│
└── async path: publish to Kafka (Kubernetes)
├── follow-up & nurture jobs
├── outbound SMS / WhatsApp / email dispatch
└── enrichment & integration side effects
│
▼
Leads, quotes, inbox UI (dashboard)
│
▼
New Relic + Datadog (traces, metrics, webhook latency)Multi-tenancy: every row is scoped by organization_id. Knowledge chunks, assistants, and channel credentials are per org. That is what makes the study reusable as a pattern, not a single-tenant chatbot.
RAG: ingest, index, retrieve
The RAG module (lib/org-knowledge.ts) is intentionally small and explicit.
Ingest
- Operators upload PDF, Markdown, or plain text via
POST /api/knowledge/upload. - PDF text extraction uses
unpdfon the server (avoids browser-only PDF APIs in serverless). - Raw text is stored on
org_knowledge_documentswith an indexing status (processing→readyorerror).
Chunking and embeddings
- Text is split into chunks (~2.4k characters) with paragraph-aware breaks and ~200 characters overlap.
- Embeddings use OpenAI
text-embedding-3-small(1536 dimensions). - Vectors land in
org_knowledge_chunks, keyed by document and chunk index. - Re-index deletes old chunks for that document first (idempotent updates).
Retrieval
- Queries are embedded with the same model.
- Similarity search runs in PostgreSQL (pgvector + a scoped SQL function such as
search_org_knowledge) withmatch_countandmin_similarity(default 0.55). - Hits carry document title, source type, content, and similarity score.
Prompt assembly
buildKnowledgeContextBlock turns hits into a bounded block (default ~3k characters) prefixed as internal company sources. That block is injected into:
- the main AI orchestrator (
/api/ai/process), - tenant webchat streaming (
lib/webchat-tenant-llm.ts), - quote generation and other flows that need catalog or policy context.
Async work on Kafka (Kubernetes)
Not every side effect belongs in the request that answered the customer. Cron-style follow-ups, proactive outreach, and heavy integration steps are modeled as events on Kafka, consumed by workers in the same Kubernetes cluster as the API.
Typical pattern in the study:
1. The synchronous path commits intent in PostgreSQL (lead updated, quote draft ready, conversation state advanced). 2. The API or a scheduler publishes a message to a topic (org-scoped key when possible). 3. A consumer applies the work idempotently: send outbound message, call ERP webhook, run a second LLM pass, etc.
That split keeps webchat and Vapi tool handlers within tight timeouts while still guaranteeing at-least-once processing with retries on the consumer side. It also mirrors how you would scale agents in production: scale HTTP replicas for traffic, scale consumers for backlog.
Observability (New Relic and Datadog)
Agents fail in boring ways: slow retrieval, duplicate cron runs, webhook auth mismatches, Kafka lag. The study wires New Relic and Datadog as complementary layers:
- Traces on
/api/ai/process, knowledge upload, and Vapi tool endpoints (knowledge-search, ticket creation). - Metrics on consumer lag, embedding batch duration, and voice tool error rates.
- Logs correlated by
organization_idand conversation id where privacy policy allows.
Text agents: state machine plus LLM
A common failure mode is letting the model decide the funnel (“ask for budget now” vs “ask for phone”). Quby inverts that.
lib/conversation-state-machine.ts defines profiles (dealer, quby_product, tenant) and states such as IDENTIFY_INTENT, COLLECT_BUDGET, CREATE_LEAD, HANDOFF_OR_CLOSE. The backend decides the next state from user text and collected fields. The LLM’s job is to phrase the current step naturally and to use RAG when the user asks domain-specific questions.
The orchestrator POST /api/ai/process ties it together:
- loads conversation state and lead context,
- evaluates handoff and branch routing (rules plus optional AI-assisted decisions),
- pulls RAG and cross-channel history blocks into the prompt,
- calls either Chat Completions or, when configured,
runOpenAIAssistantTurnfor org-specific OpenAI Assistants.
lib/ai-autonomous-decisions.ts) sit on top for “should we transfer to a human?” and “which branch owns this lead?” without replacing the state machine.Study takeaway: agents here are workflows with LLM skin, not a single prompt pretending to be a CRM.
Voice agents: Vapi plus the same RAG
Phone support adds latency and noisier input. The study integrates Vapi so the voice model can call back into Quby over HTTPS.
Lifecycle webhooks
POST /api/webhook/vapi receives call events (status, transcript, end-of-call reports) for logging and inbox continuity.
Tool: search_knowledge
During a call, the assistant can invoke search_knowledge. Vapi posts to POST /api/webhook/vapi/knowledge-search, which:
- authenticates with
VAPI_WEBHOOK_SECRET, - resolves the organization from call metadata,
vapi_call_id, orvapi_assistant_id, - runs
searchOrgKnowledge(same PostgreSQL retrieval as chat), - returns a short formatted result (≤800 characters) so the voice model summarizes in one or two sentences.
lib/vapi-knowledge-search-tool.ts and are auto-attached when provisioning assistants (lib/vapi-auto-provision.ts).Tool: create lead
POST /api/webhook/vapi/ticket handles create_quby_lead (and legacy ticket names) so a qualified call opens a lead in the same pipeline as webchat.
Study takeaway: voice is not a separate knowledge base; it is RAG exposed as a function the realtime model can call under pressure.
Public voice demo catalog
The study ships a dedicated voice sandbox at demo.quby.com.br. The page is scoped to voice agents only: select a persona on the left, open the call panel, and talk in the browser. It is the fastest way to feel latency, tool use, and persona prompts without signing into the product.
| Agent | Role in the study | |--------|-------------------| | Atendimento (Bob) | Dealership-style support: sales, rental, service, and parts | | Status de frota | Internal fleet ops: collect data and update asset status | | Renovação de locação (Bob) | Executive renewal flow: lease extension and equipment buyout | | Prospecção ativa (Sarah) | Outbound-style qualification: propose equipment and capture intent |
For the full platform (webchat embed, inbox, knowledge upload, org setup), the landing at quby.com.br points to the commercial demo request; the voice URL above is intentionally narrow so readers can test agents in isolation.
Multichannel inbox
Quby’s product pillars (Connect / Automate / Integrate) map to code paths:
| Channel | Entry | Notes |
|--------|--------|--------|
| Webchat | /api/ai/webchat/stream, public widget | Tenant LLM + RAG on the client site embed |
| WhatsApp | worker + /api/webhook/whatsapp-wwebjs | QR-linked worker; unified lead thread |
| Email | Gmail OAuth + push handlers | Inbound into same conversations |
| SMS | Twilio webhook | Outbound/inbound tied to org numbers |
| Voice | Vapi | Tools for RAG + lead creation |
lib/multichannel-ai-autoresponse.ts and lib/cross-channel-context.ts keep autoresponses and history coherent when the same person writes on chat and later calls.
Operating the study safely
Points I kept explicit while building:
- Secrets stay server-side — no service role or webhook secrets in Client Components.
- Ground when unsure — voice and chat prompts instruct the model to offer human follow-up when RAG returns nothing useful.
- Human in the loop for money — quote generation produces drafts; sending and pricing commits stay reviewable in the dashboard.
- Per-org isolation — vector search and row access always filter by
organization_id.
What to try on the demos
1. Voice: open https://demo.quby.com.br, choose Atendimento or Prospecção ativa, and ask a question that should hit search_knowledge (price, warranty, process). Listen for a short answer, not a document read aloud.
2. Full product story: skim https://quby.com.br if you want positioning and the form for a guided demo.
3. Operators: sign in on app.quby.com.br to upload PDFs to the org knowledge base and see the same RAG power text channels in the inbox.
Suggested mental checks:
- Same org boundary in RAG whether the user typed in webchat or spoke on a Vapi call.
- Lead creation after qualification on voice (
create_quby_lead) should match the webchat funnel philosophy even when the UI differs.
Closing
Quby, as documented here, is a study platform for RAG plus agent orchestration plus voice tools in one multi-tenant stack. The interesting part is not “call OpenAI once”; it is shared retrieval, deterministic funnel state, and channel adapters that all honor the same org boundary.
If you are designing something similar, start with searchOrgKnowledge and buildKnowledgeContextBlock as the single source of truth, then hang chat, voice tools, and quote drafts off that spine.