🌙

13+ models onlineBYOK encryptedZero markupVultr Inference
z0.chat — parallel session
Kimi K3
GLM 5.2
Grok 4.5
user@z0:~$Explain quantum entanglement to a software engineer in 3 sentences.
Kimi K3 · MoonshotTwo quantum particles share a joint state where measuring one instantly determines the other's properties, regardless of distance. The state is undetermined until measured — both exist in superposition. This non-local correlation doesn't allow FTL communication because outcomes are fundamentally random.
GLM 5.2 · ZAIEntanglement links two quantum systems so their states are correlated beyond what's classically possible. Measure one, and the other's state is instantly determined. Think of it as a shared RNG where neither party knows the seed — Einstein called it "spooky action at a distance."
Grok 4.5 · xAITwo particles become linked so that measuring one immediately determines the other's state at any distance. It's like two coins that always land on opposite faces, even when flipped light-years apart. The correlation violates classical locality but doesn't transmit information faster than light.

13+ models. One platform. Zero markup.

Every model on Vultr Inference, with real pricing. No added fees. Bring your own keys.

Provider
PriceAllFree< $1/M$1–3/M> $3/M
TagAllreasoningfastcodingvisionlow costfreeweb search
13 modelsClear all
Model
Context
Input / M tokens
Output / M tokens
Tags
Kimi K3 Moonshot
1M
$3.00
$15.00
reasoning
GLM 5.2 ZAI
1M
$0.90
$2.84
fast
Grok 4.5 xAI
500K
$2.00
$6.00
web search
Gemini 3.6 Flash Google
1M
$1.50
$7.50
vision
Gemini 3.5 Flash Lite Google
1M
$0.30
$2.50
low cost
Ling 3.0 Flash InclusionAI
256K
Free
Free
free tier
Muse Spark 1.1 Meta
1M
$1.25
$4.25
reasoning
Kimi K2.7 Code Moonshot
262K
$0.74
$3.50
coding
Kat Coder Pro v2.5 Kwaipilot
256K
$0.74
$2.96
coding
Qwen3 Coder Alibaba
262K
$1.50
$7.50
coding
HY3 Tencent
262K
$0.14
$0.58
low cost
Laguna S 2.1 Poolside
1M
$0.10
$0.20
low cost
Inkling Thinking Machines
256K
$1.00
$4.05
reasoning
Cost CalculatorEstimate your spend
$0.00/month
$0.00 input + $0.00 output

Built for engineers who need every angle.

PARALLEL INFERENCE

3+ models, one prompt, zero waiting.

Fire a single prompt to multiple models simultaneously. Responses stream back in parallel. See where models agree, where they diverge, and catch hallucinations before they reach production.

Kimi K3
1.2s
GLM 5.2
0.4s
Grok 4.5
1.5s
SECURITY

AES-256 key encryption.

Your API keys are encrypted with AES-256-GCM via the Web Crypto API. Decrypted only in memory during inference. We never log, store, or transmit keys in plaintext.

user_key → AES-256-GCM(encrypt) → [encrypted]
at rest: [encrypted_blob]
inference: [encrypted] → decrypt(memory) → provider
CODE EXPORT

Ship to production.

Export any response as TypeScript, Python, or cURL — pre-wired with auth headers.

// Vultr Inference APIconst res = awaitfetch( 'api.vultrinference.com/v1/chat/completions', { method:'POST', headers:{ Authorization:'Bearer key' } } )
CONSENSUS

Catch hallucinations.

Cross-examine responses across 3+ models. Consensus scoring flags divergence so you know where to focus review.

BYOK

Bring your own keys.

Use your Vultr Inference API key. Zero markup on tokens — you pay exactly what the provider charges.

USAGE TRACKING

Real-time spend monitoring.

Token counts, estimated cost, and latency per model — all tracked in real time. See exactly what each query costs before you scale.

Watch three models think in parallel.

One prompt, three models, real-time streaming. See responses arrive simultaneously and spot differences instantly.

  • Responses stream with live token-by-token output
  • Consensus scoring highlights agreement and divergence
  • Export any response to production-ready code
  • Bring your own API key to get started
parallel-inference.session
Prompt
Design a fault-tolerant async pipeline in Node.js with exponential backoff.
Kimi K3
idle
GLM 5.2
idle
Grok 4.5
idle

Try the API. Right here.

Real code examples for the Vultr Inference API. Edit, run, copy — no setup required.

output
// Click Run to simulate the API call

Powered by Vultr Inference.

One API key, 13+ models from 8+ providers. Zero markup on tokens. Automatic retries across providers.

Providers
8+
Moonshot, ZAI, xAI, Google, Meta, Alibaba, InclusionAI, Tencent, and more
Models
13+
Language models with more added as providers add them
Markup
0%
Tokens cost the same as from the provider directly. BYOK with zero fees.
Encryption
AES-256
Keys encrypted via Web Crypto API. Never stored in plaintext.

How it works.

One prompt, encrypted in your browser, fanned out to every model in parallel.

👤You🔒 AES-256Web Crypto APIz0.chatVultr Inference APIapi.vultrinference.comKimi K3GLM 5.2Grok 4.5Gemini 3.6Muse Spark+ 8 more
Request flow (encrypted)
AES-256-GCM encryption
Parallel response stream
1
You type a prompt. It's encrypted in your browser with AES-256-GCM. Your API key is decrypted only in memory — never stored in plaintext.
2
z0.chat fans out. Your single prompt is sent simultaneously to every selected model via Vultr Inference. No backend proxy — requests go directly from your browser.
3
Responses stream back in parallel. Each model's output arrives in real time. Consensus scoring highlights agreement and divergence so you know what to trust.

Trusted by developers who can't afford to be wrong.

Engineers using z0.chat to verify critical answers across multiple models.

I caught a hallucinated API method that two models agreed on but Grok flagged as non-existent. z0.chat saved me from shipping a broken integration to production. The consensus score is genuinely useful, not just a gimmick.
DK
Daniel Kuo
Staff Engineer · Fintech
The parallel streaming is addictive. I fire one prompt, get three perspectives, and spot divergences in seconds. It's replaced my tab-switching workflow entirely. BYOK means I'm not paying a middleman tax on tokens.
SR
Sasha Rao
Backend Dev · Climate Tech
We use z0.chat for prompt engineering — testing the same prompt against 5 models at once shows us which phrasing generalizes and which is overfit to one model. It's become part of our weekly workflow.
ML
Mara Lindqvist
ML Engineer · SaaS
2,400+Developers
180K+Queries Run
13+Models
0%Token Markup

Trust by design.

Security and privacy aren't features — they're the foundation.

AES-256 Encrypted
Keys encrypted via Web Crypto API. Never stored in plaintext.
No Logging
Your prompts and responses are never logged or stored on our servers.
Zero Markup
Token costs match provider pricing exactly. No hidden fees.
Client-Side Only
No backend proxy. Your data goes directly to Vultr Inference.
BYOK
Bring your own API key. You control access and billing.
No Telemetry
Zero analytics, zero tracking. What you type stays with you.

How z0.chat compares.

Multi-model inference platforms compared. No marketing fluff — just the facts.

Featurez0.chatOpenRouterLiteLLMDirect API
Parallel multi-model inferenceOne prompt → N models simultaneously Built-in~ Sequential calls~ Via router config Manual orchestration
Consensus scoringAutomatic divergence detection Sentence-level Not available Not available Not available
Token markupCost above provider pricing0%5-10%0%0%
Key managementEncrypted storage of API keys AES-256-GCM, client-side Keys on their server~ Self-hosted only DIY
No backend / no loggingPrompts never touch our servers Client-side only Routes through backend~ If self-hosted Direct to provider
Real-time streamingToken-by-token output from all models Parallel streams Single model Single model Single model
Code exportTypeScript, Python, cURL All three~ API reference only~ Python SDK Manual
Setup timeFrom zero to first query< 30 seconds~2 minutes~30 minutes~1 hour+
Models supportedVia Vultr Inference13+ models, 8+ providers200+ modelsAny OpenAI-compatibleVaries by provider
Cost visibilityReal-time spend tracking per query Per-model, per-query Account-level~ Via logging Manual
Included Partial / requires config Not available

Roadmap.

What's built, what's coming, and what's planned. No vaporware.

Live

Core Platform

The foundation is built and working.

Parallel multi-model inference
BYOK with AES-256 encryption
Real-time streaming responses
Consensus scoring (sentence-level)
Usage & cost tracking
Code export (TS, Python, cURL)
Building

Developer Tools

In active development.

Prompt history & sessions
Model comparison benchmarks
CLI tool (npx z0chat)
Improved consensus (semantic similarity)
Planned

Team & Scale

On the roadmap, not yet started.

Team workspaces & shared prompts
SSO / SAML authentication
Audit logs & compliance
Public prompt gallery
API access for consensus scoring

Model Status Dashboard

Real-time status of all 13+ models. Updated continuously.

13
Online
0
Degraded
0
Offline
Avg Latency

Cost Calculator

Estimate your spend. Compare with OpenAI GPT-4 pricing.

Prompts per day50
Avg tokens per prompt1000
Models to use3 selected
Daily Cost
$0.00
Monthly Cost
$0.00
vs GPT-4 Monthly
$0.00
Cost per 1M tokens
$0.00
Monthly cost comparison
z0.chat (BYOK)
$0
OpenAI GPT-4
$0

Model Leaderboard

Compare all models side by side. Click any column to sort.

Model Provider Context Input $/M Output $/M Speed (tok/s) MMLU

Zero markup. Pay for what you use.

Hacker
$0/mo
For developers exploring multi-model inference.
  • 100 queries/day
  • 2 models simultaneously
  • BYOK with encryption
  • Basic code export
  • Usage tracking
  • Consensus scoring
  • Team features
Start Free
Team
$99/mo
For teams sharing prompts and workflows.
  • Unlimited queries
  • 5 models simultaneously
  • Team workspaces (coming soon)
  • Shared prompt library (coming soon)
  • SSO / SAML (planned)
  • Audit logs (planned)
Start Team

Questions.

How does parallel inference work?+
z0.chat opens simultaneous connections to multiple AI models through Vultr Inference. Your prompt is sent to all selected models at once. Responses stream back in parallel — you see each model's output in real time, side by side.
Is BYOK really secure?+
Yes. Your API keys are encrypted with AES-256-GCM via the Web Crypto API. Keys are decrypted only in memory during inference calls. z0.chat never logs, stores, or transmits your keys in plaintext. You're billed directly by Vultr — we add zero markup.
What models are supported?+
13+ models via Vultr Inference: Kimi K3 (Moonshot), GLM 5.2 (ZAI), Grok 4.5 (xAI), Gemini 3.6 Flash (Google), Ling 3.0 Free (InclusionAI), Muse Spark (Meta), Qwen3 Coder (Alibaba), Kat Coder Pro (Kwaipilot), HY3 (Tencent), Laguna S (Poolside), Inkling (Thinking Machines), and more. New models are added as Vultr adds them to the platform.
What is consensus scoring?+
After all models respond, z0.chat analyzes responses at the sentence level for common themes, divergences, and potential hallucinations. A consensus score (0-1) helps you quickly identify which parts you can trust. We're actively improving this with semantic similarity analysis.
Do you log my prompts or responses?+
No. z0.chat is client-side — your prompts go directly from your browser to Vultr Inference. We don't have a backend that sees your data. Your API key never leaves your browser (it's encrypted and stored locally).
What is z0.chat built on?+
z0.chat is a client-side application that connects directly to Vultr Inference (api.vultrinference.com/v1). No intermediary servers, no proxy — your prompts and API keys go directly to Vultr Inference. We use the Web Crypto API for key encryption and the Fetch API for streaming responses.
Is z0.chat open source?+
z0.chat is currently in beta. We're focused on building the core product right. Open source is on the roadmap — follow along as we develop.

Stop guessing. Start verifying.

One prompt to every model. Consensus in seconds. Zero markup, zero telemetry, zero doubt. Free to start — bring your own key.

Launch App →
13+Models
8+Providers
0%Markup
AES-256Encryption

Stay in the loop.

Get notified when new models are added, features ship, and benchmarks update. No spam — just product updates.