How it works
How on-premise AI works, step by step.
01
The hardware
A server you own. Sized to your number of users and documents before you buy anything. The exact hardware depends on how many people use it and which model you run.
02
The model
An open-weight model you choose. Models whose weights you can download and run yourself. In the lab demo an open-weight model of about 30 billion parameters ran on a single small desktop AI box (an NVIDIA DGX Spark). You aren’t locked to one provider, and you can switch models later.
03
Your documents
You choose which documents the assistant can use. It answers from those.
Access set per person. Your admin decides who can open each document library. People who aren’t members don’t see it, and the assistant won’t answer them from it. Sign-in supports single sign-on over SAML and two-factor authentication.
04
Offline
After installation it can run fully offline, with no internet connection at all.
05
Running it
An admin console you control. Your admin decides which models run and sees system status at a glance. An audit log, kept on your own server, records actions such as sign-ins, document uploads, admin changes and the tools the assistants use. It doesn’t record the questions themselves: chats are stored on your server, where authorised admins can search them. A built-in check shows whether the system is network-isolated. The admin console is in English.
Rollout
Rollout and ongoing care
Offered in tiers based on how many people will use it. We’ll go through what fits on a call.