Case study
Multitenant SaaS platform with AI features
A B2B platform where every customer's data, knowledge and AI usage has to stay separate and predictable. I owned the backend architecture and the cloud infrastructure.
- Role
- Lead backend & infrastructure engineer
- Period
- Ongoing engagement
- Stack
Highlights
- Separate databases per customer, provisioned automatically at signup
- AI responses streamed to the browser, with cancellation
- AI spend per customer metered, billed and cut by about a third
- Workers that scale with queue depth instead of a fixed size
- All infrastructure in code, deployed without long-lived cloud keys
The challenge
The product lets business customers work with AI over their own data. That raises three questions every multitenant AI product runs into sooner or later: how to guarantee one customer’s data never reaches another, how to keep responses fast when a model takes seconds to answer, and how to stop AI costs from growing faster than revenue.
Architecture
The platform is split into a few services with clear jobs: an API that owns customers, permissions and billing; a service that handles real-time conversations; a worker service that runs the AI workloads; and a web frontend. I led the backend and designed how the pieces fit together.
What I built
Isolation by design. Instead of separating customers with a column in shared tables, each customer gets their own databases, created automatically from templates when they sign up. A connection pooler keeps the number of database connections under control as customers grow, and a customer can be moved to dedicated capacity without code changes. Each customer’s AI knowledge is stored separately as well.
Real-time responses. Model output is streamed to the browser as it is generated, so users see an answer start within moments rather than waiting for the full reply, and they can stop a response midway.
Data ingestion. Customers bring their own documents and connect third-party tools through OAuth. Background jobs extract, split and index that content so the AI can use it, and keep it in sync when the source changes.
Cost and billing. Every AI call is metered per customer and fed into subscription billing. When one customer’s usage outgrew their plan, I broke the spend down by type of call and type of token, found that most of it was overhead rather than useful output, and cut their cost by about a third with a configuration change, with further savings identified.
Infrastructure. The whole cloud setup is defined in Terraform and runs on Kubernetes. Background workers scale with the length of their queue, deploys run from CI without long-lived credentials, and errors, traces and AI calls are all observable end to end.
Outcome
- Running in production for paying business customers
- Customer data separated at the database level, not by convention
- Predictable AI costs per customer, visible before the invoice arrives
- An environment that can be rebuilt from code, with no manual steps
Working on something similar?
I help agencies and product teams with this kind of work.
Book a 20-minute call