← Personal projects Applied AI Proof of concept · 2026

One assistant, three roles, and a tool layer that says no

A finance assistant where employees, clients and the owner all talk to the same chatbot and each sees only what their role allows. A small model routes each message to one of three agents, and authorization lives in the tool code, never in the prompt.

  • TypeScript
  • Next.js
  • Vercel AI SDK
  • Supabase
  • pgvector
3
Specialist agents
3
Roles
pgvector, keyword fallback
Retrieval
Swappable — hosted or local
Model

Architecture

  1. 01Chat UIstreaming
  2. 02Auth contextsigned session → role
  3. 03Routermini-model, 3 intents
  4. 04Agentfinance · documents · actions
  5. 05Toolsrole-scoped, audited
  6. 06PostgresSupabase + pgvector

The problem

A small company’s finance questions come from three kinds of people. An employee wants to know why their payslip is lower this month. A client wants a copy of an overdue invoice. The owner wants to know the cash position and who to chase. One assistant should handle all three — and an employee must never be able to talk it into showing someone else’s salary.

That last sentence is the whole design problem. Building a chatbot that answers finance questions is straightforward. Building one that can’t be argued out of its permissions is not.

Authorization is not a prompt

The tempting approach is to put the rules in the system prompt: you are talking to an employee, only discuss their own records. The prompt in this project does say that, because it makes the model’s refusals read naturally. But nothing depends on it.

The session is resolved to a user and a role on the server before any model is called. The tools each agent receives are built for that role: an employee’s get_payslips tool is constructed with their employee ID already bound, and there is no argument the model can pass to ask for anyone else’s. The owner’s tool set contains company-wide queries; the employee’s doesn’t contain them at all.

So a successful prompt injection gets the model to want to fetch another person’s data, and then discover it has no tool that can.

Why route instead of one big agent

Each message first goes to a small, cheap model that classifies it into one of three intents, and the conversation is handed to the matching agent:

  • Structured finance — reads numbers from Postgres: payslips, invoices, balances, cash.
  • Document knowledge — retrieves policy text by vector search: expense rules, payment terms.
  • Actions — the only agent that writes: submit a claim, request an advance, queue a reminder.

Three reasons for the split, in the order they mattered:

  1. Reliability. Each agent gets a short prompt and only the three to six tools it needs. Tool selection is noticeably better with a short menu, especially on small or local models.
  2. Safety. The agent that can write is only invoked when the user asked for an action. A question about policy can’t accidentally submit a claim.
  3. Cost. Classification is about fifty tokens on a mini model.

If classification fails — some local models don’t support structured output — a keyword heuristic takes over, and the default is the read-only finance agent.

Retrieval with a fallback

Policy documents are embedded into pgvector and retrieved by similarity. Until the embedding step has been run, the same tool falls back to keyword search, so the application works on a fresh database and gets better when the vectors exist.

Uploaded files — a receipt, an invoice — are stored immediately, filed into the right folder by a tool call, and receipt images are read by the model to pre-fill an expense claim.

What I’d do differently

An evaluation set for the router. Three intents and a heuristic fallback is easy to test: a few hundred labelled messages would give routing accuracy, and show which phrasings land in the wrong agent.

Adversarial tests for the tool layer. The claim that permissions hold under prompt injection should be a test suite, not a paragraph — a set of hostile prompts per role, asserting on which rows the tools returned.

Row-level security in the database. The tools enforce scope in application code. Postgres can enforce the same rule underneath, so that a bug in one tool is not a data leak.

Contact

Want the longer version?

Happy to walk through any of this in detail — the parts that broke are usually the interesting bit.