Support emails become tickets that are auto-classified, summarized, and either auto-resolved by AI — grounded in a real retrieval-augmented knowledge base — or left for an agent, who gets AI summaries and AI-polished replies that stay in their own voice, all backed by a full management interface for agents and admins.
🔗 Live Demo — demo credentials are shown right on the login page
Want to see the real pipeline end to end? Email support@shahzadtariq.com or deskwise.support@shahzadtariq.com with a support-style question and it'll land as a real ticket within seconds — auto-resolved with a reply if the seeded knowledge base covers it, left open for an agent if not. (It's a shared public demo, so anything sent in is visible to anyone using the demo login above — don't send anything private.)
- Overview
- Screenshots
- Features
- Architecture
- Tech Stack
- Security
- Getting Started
- Project Structure
- Development Process
- Roadmap
- License
Support teams drown in repetitive email traffic — the same shipping, returns, and account questions, answered one at a time. Deskwise turns inbound support email into a structured ticket queue, then uses AI to lighten the manual load at every step: the moment a ticket arrives, it's checked against a retrieval-augmented knowledge base and resolved automatically — with a real, grounded email reply — if the knowledge base confidently answers it; everything else is auto-classified and left for an agent, who gets one-click AI summaries of long threads and AI-polished replies that stay in their own voice.
It's a full-stack TypeScript monorepo: an Express API backed by Postgres/Prisma, a React admin/agent interface, and an async job pipeline (pg-boss) that keeps every LLM and vector-database call off the request path.
The branded React Email template every outbound reply — human agent or AI auto-resolution — renders through before being sent via SES.
Ticket Management
- Paginated, sortable, filterable ticket list (by subject, status, category)
- Ticket detail view with a threaded reply history
- Assignment to agents, status transitions (Open → Resolved → Closed)
- Inbound-email webhook intake — automatically dedupes into an existing open ticket for the same sender + subject instead of creating duplicates. Fed by a real AWS-native pipeline, not just manual/test calls: an SES receipt rule on a live domain routes incoming mail through S3 and a Lambda that parses the raw MIME and posts to the secret-protected webhook. Try it yourself — email addresses are at the top of this page, under Live Demo
- Outbound email sending via AWS SES, rendered through a branded React Email template — every agent reply and AI auto-resolution reply is emailed to the customer, off the request path via a pg-boss job
AI-Powered Features (all via LangChain, never a provider SDK called directly — see Architecture)
- 🤖 Automatic resolution via RAG — every new ticket is checked against the knowledge base the instant it arrives; if the retrieved excerpts fully and confidently answer it, a complete, ready-to-send email reply is generated and posted automatically (signed off "Customer Support") and the ticket is marked resolved — without ever landing in an agent's queue. Anything the knowledge base can't answer is left untouched and open for a human — see How the RAG pipeline works
- 🏷️ Auto-classification — every inbound ticket is classified (general question / technical question / refund request) asynchronously, without blocking the webhook response
- 📝 AI summaries — one click to summarize a long ticket + reply thread for an agent picking it up cold
- ✍️ AI reply polish — improves an agent's draft (grammar, clarity, tone) while preserving their own wording and intent
- 📚 Retrieval-augmented knowledge base — admins upload PDF/DOCX/TXT/MD policy docs; they're chunked, embedded, and stored in Pinecone, powering the automatic resolution above instead of the model hallucinating an answer
User Management & Auth
- Session-based authentication (no public sign-up — admin-provisioned accounts only)
- Admin/Agent roles with route- and API-level authorization
- Soft-delete with session revocation, not destructive hard deletes
Deskwise is a Bun workspace monorepo with three packages: packages/server (Express 5 API), packages/client (React 19 SPA), and packages/core (Zod schemas and const-object enums shared by both, so the client and server can never drift on validation rules or status values).
The core architectural rule: any request that would call an LLM or a vector database never blocks the HTTP response. Ticket classification, automatic resolution, and knowledge-base ingestion are all handed off to pg-boss (a Postgres-backed job queue — no separate Redis/SQS infra needed) and processed by background workers, so a slow OpenAI call or a Pinecone hiccup can never make an API request hang.
browser (React client)
|
v HTTPS + session cookie
Express API (routes)
|
does this request only need
the database?
|
+--------+--------+
| |
yes no — also kick off a
| background job (the API
v responds right away,
read/write never waits for these)
PostgreSQL |
(tickets, users, v
replies, docs) pg-boss job queue
(just rows in the same
Postgres database —
no separate Redis/SQS)
|
+------------+-------+-------+------------+
| | | |
v v v v
classify- auto-resolve- ingest- send-reply-
ticket job ticket job document job email job
| | | |
v v v v
LangChain LangChain LangChain AWS SES
+ OpenAI + Pinecone + Pinecone (deliver
(pick a (search the + S3/disk the reply
category) KB, maybe (extract, email)
draft a chunk, embed,
reply) store)
(The exact branching inside each of those four jobs — new → processing → resolved/open, chunk → embed → upsert, and so on — is broken out in its own simple diagram further down.)
Key decisions worth knowing about:
- LangChain-only AI access. Every AI feature — chat completions and embeddings alike — goes through LangChain's abstractions (
@langchain/core,@langchain/openai,@langchain/pinecone), never a provider SDK called directly. This is a deliberate standing rule, not incidental to how the first few features happened to be built. - RAG ingestion pipeline. An uploaded document is extracted (
pdf-parse/mammoth/ plain text read depending on type), chunked (~1000 characters with ~200 character overlap so no sentence is orphaned at a chunk boundary), embedded withtext-embedding-3-small, and upserted into Pinecone with per-chunk metadata (docId,filename,chunkIndex,uploadedAt). Vector IDs are minted as<docId>#<chunkIndex>— a deliberate choice so that deleting a document can list-then-delete by ID prefix, since Pinecone serverless indexes don't reliably support metadata-filter deletion. See How the RAG pipeline works for the retrieval half. - Ticket auto-resolution is a state machine, not a flag.
Ticket.statusgains two internal-only values,newandprocessing, used only while the auto-resolve pipeline is deciding what to do with a ticket — they're never shown in the UI and can never be set manually (the ticket list always excludes them, and edits to them are rejected server-side). A ticket lands onresolved(AI answered) oropen(needs a human) once the attempt finishes, and from there on behaves exactly like a human-resolved or human-touched ticket — there's no separate "resolved by AI, hide forever" flag. An agent can still tell it apart from the reply thread itself, which carries an AI Assistant–authored reply. - Fail loud, never auto-provision. If the configured Pinecone index doesn't exist, or its dimension doesn't match the embedding model's output, ingestion (and auto-resolution's retrieval step) fails immediately with a clear error rather than the app silently creating or resizing infrastructure on your behalf — an auto-resolution failure specifically falls back to leaving the ticket
openfor a human, never stuck mid-pipeline. - Soft delete over hard delete. Deleting a user never removes the row — it sets
deletedAt, revokes every session, and frees the email address for reuse by overwriting it, so ticket history stays intact for anyone who worked a ticket in the past. - Knowledge-base storage is S3 with a local-disk fallback, not a hard dependency on either.
lib/knowledge-base/storage.tspicks its backend once, at module load, from whetherKNOWLEDGE_BASE_S3_BUCKETis set: S3 when it is, otherwise local disk (gitignored, the original and still-simplest option for running without AWS at all).KnowledgeDoc.pathis treated as an opaque handle (an S3 key or a local path) everywhere downstream — never read or constructed directly outside this module. - Outbound email is a direct AWS SES call, not LangChain. The LangChain-only rule above is scoped to AI/LLM access — SES is a transactional email provider, so
lib/email/send-email.tscalls@aws-sdk/client-sesv2directly, the same waylib/knowledge-base/pinecone.tscalls the Pinecone SDK directly. Sending is always handed to thesend-reply-emailpg-boss job rather than done inline, so neither an agent's reply submission nor the auto-resolve job ever blocks on SES. - Reply emails are templated with React Email, not hand-built HTML strings. Every outbound reply is wrapped in a branded template (
lib/email/templates/reply-email.tsx) before being handed to SES. A human agent's reply already has rich-textbodyHtmlfrom the client's editor; the AI auto-resolve job's reply doesn't (it's plain text from the model), so it's rendered into HTML through its own React Email component first, rather than string-concatenating<p>tags by hand. - Error tracking covers more than the request/response path. Sentry (
@sentry/nodeserver-side,@sentry/reactclient-side) is wired up beyond the automatic cases (uncaught exceptions, unhandled promise rejections, Express route errors, React render crashes via a top-levelSentry.ErrorBoundary). A lot of this codebase deliberately catches an error, logs it, and keeps going instead of letting it bubble up — a failed job-enqueue, auto-resolve's fallback-to-opencatch, a pg-boss worker that intentionally rethrows so pg-boss owns retry/backoff — and each of those sites callsSentry.captureException(err)explicitly, since Sentry's automatic integrations would never otherwise see them. Both DSNs are optional;Sentry.init()silently no-ops without one, so it's safe to leave configured in every environment including local dev.
Think of it like a mailroom with a small robot helper. Every new ticket gets a secret, invisible stamp — NEW — until the robot picks it up and stamps it PROCESSING while it thinks. Nobody sees a ticket while it wears either stamp. The robot doesn't just skim the team's rulebook (the knowledge base) — it turns the customer's question into a vector and searches for the rulebook pages that are the closest match, held as embeddings in Pinecone. That's the "retrieval" half of RAG. It then hands only those exact pages to a second robot that writes the reply, under one strict rule: answer only from what's written on those pages, never from outside knowledge or a guess. If those pages fully and confidently answer the question, it writes the reply, signs it "Customer Support," and stamps the ticket RESOLVED — done, no human needed. If the pages don't cover it, or anything goes wrong, it just stamps the ticket OPEN and hands it to a human agent, exactly as if it had never tried. (See How the RAG pipeline works below for the full mechanics.)
email arrives
|
v
NEW -------- hidden from agents
|
v
PROCESSING -------- still hidden, robot is thinking
|
v
turn the question into
a vector (OpenAI embedding)
|
v
find the closest-matching
rulebook pages (similarity
search against Pinecone)
|
v
hand only those exact pages to
the writer robot, with one rule:
answer from them, nothing else
|
+----------------------+
| |
v v
the pages fully the pages don't
answer it cover it
| |
v v
RESOLVED OPEN
(AI wrote a reply, (needs a human)
grounded in those
exact pages)
| |
+---- visible to agents now ----+
At the very same moment, a second, completely separate little robot reads that same new ticket to figure out what kind of question it is. It doesn't wait for the first robot and doesn't block anything — they both just quietly work on the ticket in the background. If a human agent already picked a category by hand before this robot finishes, the robot's guess is thrown away instead of overwriting the human's choice.
email arrives
|
v
category: none -------- shown as "Uncategorized"
|
v
robot reads the subject + message
|
v
robot picks ONE:
- General Question
- Technical Question
- Refund Request
|
v
did a human already pick one?
|
+------+------+
| |
yes no
| |
v v
keep the save the
human's pick robot's label
Same mailroom, but now picture an admin dropping a new policy PDF into the robot's rulebook. The robot doesn't just staple it in — it reads it, copies out all the small pieces, and files each piece somewhere it can find it again fast later.
admin uploads PDF
|
v
PROCESSING -------- saved to disk, not searchable yet
|
v
extract the text
|
v
split into overlapping
chunks
|
v
embed each chunk into
a vector (OpenAI)
|
v
store the vectors in
Pinecone, tagged with
this document's id
|
+------------------+
| |
v v
everything something
worked broke
| |
v v
READY FAILED
(chunks are now (error saved,
searchable by the admin can see
auto-resolve robot) why it failed)
Only once a document reaches READY can its chunks actually be found by the auto-resolution search above — a document stuck in PROCESSING or FAILED is invisible to it, the same way a NEW/PROCESSING ticket is invisible to agents.
Deskwise's knowledge base is a complete retrieval-augmented generation loop — an ingestion (write) half and a retrieval-and-generation (read) half — not just a document store bolted onto a chatbot.
1. Ingestion — turning a document into searchable vectors. When an admin uploads a policy document (POST /api/knowledge-docs), the file is saved and a KnowledgeDoc row is created as processing, then handed to a pg-boss job so the upload request returns immediately instead of blocking on extraction/embedding. The worker (jobs/ingest-document-job.ts → lib/knowledge-base/ingest-document.ts) extracts plain text, splits it into overlapping ~1000-character chunks, embeds each chunk with text-embedding-3-small, and upserts the vectors into Pinecone (@langchain/pinecone's PineconeStore) — each one tagged with its source document and chunk index. The KnowledgeDoc row flips to ready (or failed, with the error saved) once that finishes.
2. Retrieval + generation — answering a ticket from those vectors. When a new support ticket arrives (POST /api/tickets/inbound-email), it's enqueued onto the auto-resolve-ticket job without blocking the webhook response (jobs/auto-resolve-ticket-job.tsx). That job:
- Embeds the ticket's subject + body with the same embedding model and runs a similarity search against Pinecone (
lib/knowledge-base/search-knowledge-base.ts) to pull back the top-K most relevant chunks across every ingested document — the read-side counterpart to the ingestion pipeline above. - Hands those chunks to an LLM (
lib/tickets/auto-resolve-ticket.ts) with a strict instruction: answer only from the retrieved excerpts, never from outside knowledge, and say so honestly (canResolve: false) if they don't fully cover the question. This grounding step is what stops the model from confidently inventing a shipping or refund policy that doesn't exist. - If the model is confident the excerpts fully answer the ticket, it drafts a complete, ready-to-send email — greeting, grounded answer, "Customer Support" sign-off — which is posted as a reply from a synthetic AI Assistant account, and the ticket is marked
resolved. If not, or if anything in the pipeline fails (empty knowledge base, a Pinecone or OpenAI error), the ticket is simply leftopenfor a human, exactly as if auto-resolution had never been attempted.
This read path only ever runs once, at ticket creation — a follow-up email to an already-open ticket reuses that ticket instead of re-triggering resolution.
Deskwise's AWS footprint is deliberately small — a handful of services doing one job each, not a sprawling architecture:
| Service | Used For | Notes |
|---|---|---|
| EC2 | Hosts the running app — a single instance running postgres + app + web (Caddy) via Docker Compose |
Not ECS/Fargate or RDS — a deliberate cost/simplicity trade-off for this demo; see Roadmap |
| ECR | Stores the deskwise-server / deskwise-client Docker images CI builds on every push to master |
Tagged by both commit SHA and latest |
| IAM | A deploy role GitHub Actions assumes via OIDC, scoped to just ECR push + SSM send-command |
No long-lived AWS access keys stored as GitHub secrets |
| SSM (Systems Manager) | GitHub Actions runs the EC2 instance's /opt/deskwise/deploy.sh remotely via aws ssm send-command |
No SSH key to manage or rotate — just the SSM agent + the IAM role above |
| SES | Sends every outbound reply email (agent + AI auto-resolution), and receives inbound mail in production via a receipt rule on a live domain | Outbound is optional locally — unconfigured, sendEmail() logs and skips instead of failing (see Getting Started). Inbound receiving is always on in production |
| S3 | Two separate jobs: optional backend for knowledge-base documents, and the landing zone for raw inbound email before Lambda parses it | KB storage falls back to local disk when KNOWLEDGE_BASE_S3_BUCKET is unset; the inbound-mail bucket is a fixed part of production, not optional |
| Lambda | Parses the raw MIME email SES drops in S3 and POSTs it to the WEBHOOK_SECRET-protected /api/tickets/inbound-email endpoint |
Deployed directly, source isn't checked into this repo |
Same idea as the robots above, except the "ticket" is a code change and the "mailroom" is GitHub Actions. On every push to master, Actions builds and ships the change itself — nobody







0 comments
log in to comment.