AI SERVERS — ON-PREM TOKENS SERVER
The On-Prem Tokens Server is the Zanus AI Quantum hardware with the Zanus OS — without a specific industry tenant — working as your private GPT-class generative endpoint. The best open-weight models run on your own enterprise GPUs, served through an OpenAI-compatible API: anything that talks to GPT, Claude or Grok points at your server by changing one URL. Cursor, your apps, your agents, your Zanus tenants.
Price by RFQ — sized on GPU memory (models), RAM/context, and tokens per day

Your models, your GPUs, your tokens — behind your firewall.
OpenAI-compatible endpoint. Cursor, your apps, agents, scripts and tools keep working — they just stop sending your data (and your money) to the cloud.
Front Office and Zanus AI Cloud can run on YOUR local tokens instead of datacenter credits. Cloud convenience, on-prem economics.
No per-token meter. After payback, intelligence costs you cents of electricity. Models improve? New weights are a download, not a new bill.
A business that embeds AI in daily operations — front office, documents, agents — easily consumes hundreds of millions of tokens a month. At typical blended API rates that is thousands of dollars every month, forever, and it grows exactly with your success. Per-token pricing is the cloud’s per-seat tax in disguise.

Same award-winning machine as the Private AI tiers — Quantum-class GPUs, RAID 10 NVMe, mission-critical design with redundant power and network — configured as a pure generative endpoint. Final sizing depends on the models you want, the context you need, and your tokens per day: that’s why it’s RFQ.
| Sizing dimension | What it drives | Typical question we ask |
|---|---|---|
| GPU memory | Which model classes run (and how many in parallel) | Which models do you want on Day 1? |
| RAM / context | How long a document or conversation the AI can hold | Longest contracts, transcripts, codebases you process? |
| Tokens per day | Throughput and concurrency headroom | Interactions per month across all apps and tenants? |
The leading open-weight families — chosen and sized with you at configuration, installed and optimized by Zanus engineers. Swap or add models any time: new weights are a download, not a new machine.
For grounded business work — answering from YOUR knowledge, documents, quotes, agents — today’s open models are excellent, and the Zanus grounding + verification layer is what guarantees precision, not the brand of the model.
Yes. It is a standard OpenAI-compatible API on your network — any language, any framework, any tool that speaks that API. Point Cursor at it and code with your own tokens.
No — that is the difference by design. The Tokens Server is the Quantum + Zanus OS as a pure generative endpoint. If you want the complete working system with your industry tenant included, that is the Private On-Premises AI — which contains this endpoint too.
Yes — that is the most popular hybrid: keep the hosted Front Office / Back Office convenience, run them on YOUR local tokens. Cloud software, on-prem economics.
← Previous: Overview · Sharing with a colleague? 📄 Get the PDF · ✉️ Email this page ·