AI Servers

AI SERVERS — ON-PREM TOKENS SERVER

Stop renting intelligence.
Generate your own tokens.

The On-Prem Tokens Server is the Zanus AI Quantum hardware with the Zanus OS — without a specific industry tenant — working as your private GPT-class generative endpoint. The best open-weight models run on your own enterprise GPUs, served through an OpenAI-compatible API: anything that talks to GPT, Claude or Grok points at your server by changing one URL. Cursor, your apps, your agents, your Zanus tenants.

Price by RFQ — sized on GPU memory (models), RAM/context, and tokens per day

Zanus On-Prem Tokens Server — the Quantum as a generative endpoint

Your models, your GPUs, your tokens — behind your firewall.

What it does

One URL change

OpenAI-compatible endpoint. Cursor, your apps, agents, scripts and tools keep working — they just stop sending your data (and your money) to the cloud.

Feeds your Zanus tenant

Front Office and Zanus AI Cloud can run on YOUR local tokens instead of datacenter credits. Cloud convenience, on-prem economics.

Yours forever

No per-token meter. After payback, intelligence costs you cents of electricity. Models improve? New weights are a download, not a new bill.

The economics: rental vs ownership

A business that embeds AI in daily operations — front office, documents, agents — easily consumes hundreds of millions of tokens a month. At typical blended API rates that is thousands of dollars every month, forever, and it grows exactly with your success. Per-token pricing is the cloud’s per-seat tax in disguise.

  • One purchase — call for price, sized to your workload
  • Full-load electricity is roughly $1/hour (6 kW max) — and it only draws that while working
  • Typical payback: 12–24 months for AI-heavy operations
  • After payback: tokens at the cost of electricity
Token usage economics

Built on the Zanus AI Quantum

Same award-winning machine as the Private AI tiers — Quantum-class GPUs, RAID 10 NVMe, mission-critical design with redundant power and network — configured as a pure generative endpoint. Final sizing depends on the models you want, the context you need, and your tokens per day: that’s why it’s RFQ.

Sizing dimensionWhat it drivesTypical question we ask
GPU memoryWhich model classes run (and how many in parallel)Which models do you want on Day 1?
RAM / contextHow long a document or conversation the AI can holdLongest contracts, transcripts, codebases you process?
Tokens per dayThroughput and concurrency headroomInteractions per month across all apps and tenants?
🔇 Silent — office-friendly, no server room ⚡ Standard 50A circuit @ 115/220V — 6 kW max 📦 Delivered configured in ~3 weeks 🔒 Air-gap capable — your data never leaves

Questions, answered

Which models run on it?

The leading open-weight families — chosen and sized with you at configuration, installed and optimized by Zanus engineers. Swap or add models any time: new weights are a download, not a new machine.

Is an open model good enough vs GPT?

For grounded business work — answering from YOUR knowledge, documents, quotes, agents — today’s open models are excellent, and the Zanus grounding + verification layer is what guarantees precision, not the brand of the model.

Can my developers use it directly?

Yes. It is a standard OpenAI-compatible API on your network — any language, any framework, any tool that speaks that API. Point Cursor at it and code with your own tokens.

Does it include an industry tenant?

No — that is the difference by design. The Tokens Server is the Quantum + Zanus OS as a pure generative endpoint. If you want the complete working system with your industry tenant included, that is the Private On-Premises AI — which contains this endpoint too.

Can it power my Zanus cloud tenants?

Yes — that is the most popular hybrid: keep the hosted Front Office / Back Office convenience, run them on YOUR local tokens. Cloud software, on-prem economics.

Sized on your models, your context, your tokens per day.

Prefer a human? +1 (954) 736-3939

VIEW ALL NEXT

← Previous: Overview  ·  Sharing with a colleague? 📄 Get the PDF · ✉️ Email this page ·