Build AI fast with Zai Node's AI compute + models

We design, build, and operate distributed AI infrastructure with optional custom models to power your AI journey.

brown and grey trees and rock formation painting

BUILT AND Trusted by engineering teams at

Large format datacentres

Reduce the need for more large format, problematic DCs

US big tech AI control

Be more in control of your AI, data and IP, not US big tech

Use what we already have

Use existing infra, unused power, buildings and networking

USE AI COMPUTE

Control your AI with your own AI compute

Zai Node's purpose built AI compute solutions keeps your inference reliable, fast and cost effective.

High performance NVIDIA & AMD

Most efficient, cost effective, reliable HPC in Aus

Any model (Kimi, Qwen, Llama, Mistral, custom)

24/7 expert engineering and networking support

request log

● live

POST /v1/chat · llama-4-70b

92ms · 200

POST /v1/images · flux-pro

418ms · 200

POST /v1/embeddings · bge-large

31ms · 200

POST /v1/chat · custom-ft-7b

67ms · 200

uptime 99.99%

1.2M req / min

deployments

3 regions

melb-me2-1 · dedicated

● healthy

bne-2 · dedicated

● healthy

111Col-14 · edge

● healthy

p50 latency 92ms

autoscaling on

HOST AI COMPUTE

Earn income, reduce costs by hosting AI compute

Zai Node's unique GPU hosting model allows building owners to host GPUs with existing, unused power.

Using headroom power

Outlays paid

100% heat reuse

Reduced heating cost

Low cost, sovereign AI compute

///

USE AI MODELS

Zai Node's model serving

Running open and custom models on optimised GPU clusters, Zai Node gets you fast, private responses at any scale.

Low latency

Optimised kernels and smart routing for sub-second responses

At scale

Scale from zero to many GPUs on demand

Latest models

Serve Kimi, Qwen, Mistral or your own fine-tuned weights

Low cost

Pay per token with no idle costs or hidden fees

api.zainode.ai

$ curl https://api.zainode.ai/v1/chat

-d ‘{ model: llama-4-70b, stream: true }’

✓ 200 OK · TTFT 84ms · 128 tok/s

Streaming tokens from the nearest GPU cluster. Autoscaled to 3 replicas for burst traffic.

melb-me2 · GTX600 x8

$0.0004 / 1K tokens

ZAI NODE'S FULL STACK

Distributed AI compute + models

Zai Node's HPC and model solutions lets you ship agents, products and work faster and cheaper. Your dedicated platform lets you scale compute, change models, prompt models, compare outputs, and tune parameters live.

///

AI SOLUTIONS

AI solutions for users, landlords, tenants & society

A complete AI platform combining model serving, GPU orchestration, monitoring, and distributed deployment.

1

Managed AI as a service

Inference that scales with low latency and high throughput mutli-models and flexible deployment options.

What does a host need to do?

Inference that scales with low latency and high throughput mutli-models and flexible deployment options.

2

Use Zai Node's AI cloud

Infrastructure and platform as a service so you can build and provision GPU clusters in minutes.

What does a host need to do?

Infrastructure and platform as a service so you can build and provision GPU clusters in minutes.

3

Host Zai Node's AI compute

State-of-the-art GPUs and HPC hosted at your premises, fully operated and maintained by Zai Node's expert technicians.

What does a host need to do?

State-of-the-art GPUs and HPC hosted at your premises, fully operated and maintained by Zai Node's expert technicians.

4

Zai Node's energy first approach

The most efficient AI compute, using headroom power, existing utilities, solar, EMS, with heat reuse.

What does a host need to do?

The most efficient AI compute, using headroom power, existing utilities, solar, EMS, with heat reuse.

///

Success Stories

Proven in production

Real teams, real workloads, real gains after moving inference onto Inferno.

4.2x

Faster time to first token

“For the first time in years, I wasn’t the one getting paged at 3am.”

Marcus Delgado

VP of Machine Learning

80%

Lower cost per million tokens

“I kept checking the dashboard waiting for something to break. It never did.”

Elena Whitfield

Leads Platform Engineering

“I kept checking the dashboard waiting for something to break. It never did.”

Elena Whitfield

Leads Platform Engineering

80%

Lower cost per million tokens

“For the first time in years, I wasn’t the one getting paged at 3am.”

Marcus Delgado

VP of Machine Learning

4.2x

Faster time to first token

Updates

Latest platform updates

A concise rollout log of new infrastructure, routing, and developer experience improvements shipping across Inferno.

2

Private Networking

Keep inference traffic inside your VPC with managed private endpoint support.

2

Cache Purge

Clear stale responses instantly and force fresh inference on the next request.

2

Prompt Cache

Reuse repeated responses automatically to lower latency and reduce token spend.

2

Smart Routing

Send requests to the fastest healthy region without changing client-side code.

2

Usage Guardrails

Apply per-key quotas, concurrency caps, and safety thresholds from one control layer.

2

Benchmark Runs

Compare latency, throughput, and cost across regions before promoting a model live.

2

Dedicated Deployments

Launch isolated model endpoints with reserved capacity and cleaner environment controls.

2

Autoscale Policies

Set traffic-aware replica rules that expand during spikes and settle after demand drops.

Showcase

Built on Inferno

Built on Inferno

Inferno turns raw models into production endpoints in minutes. Teams around the world are shipping AI products on Inferno.

Inferno turns raw models into production endpoints in minutes. Teams around the world are shipping AI products on Inferno.

Features

Deploy, Scale, Infer

Deploy, Scale, Infer

A unified platform for deploying models, scaling GPU capacity, and shipping inference faster with full observability.

A unified platform for deploying models, scaling GPU capacity, and shipping inference faster with full observability.

Multi-Region

Deploy endpoints across regions and route traffic to the fastest cluster automatically.

Multi-Region

Deploy endpoints across regions and route traffic to the fastest cluster automatically.

Any Framework

Serve models from vLLM, TensorRT, or your own custom inference stack.

Any Framework

Serve models from vLLM, TensorRT, or your own custom inference stack.

Cost Forecasting

See projected spend before you deploy, based on expected traffic and model size.

Cost Forecasting

See projected spend before you deploy, based on expected traffic and model size.

Output Monitoring

Catch latency spikes, error rates, and quality regressions before they reach users.

Output Monitoring

Catch latency spikes, error rates, and quality regressions before they reach users.

Batch Inference

Process large offline workloads efficiently without paying for idle GPU time.

Batch Inference

Process large offline workloads efficiently without paying for idle GPU time.

Fast Mode

Prioritize latency-sensitive requests while keeping throughput high under load.

Fast Mode

Prioritize latency-sensitive requests while keeping throughput high under load.

Metrics

Slow inference is costing you users

Teams shipping on Inferno see faster responses, fewer failures, and infrastructure that scales with them.

30M

Requests served daily

50K+

Active developers

12K

Models deployed

100%

Uptime SLA

100%

Uptime SLA

Models

All the latest models and tools

Switch between leading AI models and connect to the tools that power your workflow.

  • minimax
  • z.ai
  • stripe

Testimonials

Trusted by 300+ people

"The integrated sound and video generation have completely streamlined our creative output."

heashot image

Elena Rossi

Creative Director, Studio Aura

"The integrated sound and video generation have completely streamlined our creative output."

heashot image

Elena Rossi

Creative Director, Studio Aura

"Integration was incredibly fast, and the tools are surprisingly intuitive for the team. The real-time collaboration features are a standout."

heashot image

Casey Wright

Technical Director, Core Creative

"Integration was incredibly fast, and the tools are surprisingly intuitive for the team. The real-time collaboration features are a standout."

heashot image

Casey Wright

Technical Director, Core Creative

"This platform is a total game-changer for our creative workflow. It seamlessly combines image, video, and sound generation into one efficient tool."

heashot image

Sarah Chen

CTO, InnoCorp

"This platform is a total game-changer for our creative workflow. It seamlessly combines image, video, and sound generation into one efficient tool."

heashot image

Sarah Chen

CTO, InnoCorp

"The precision of the image editing tools is unmatched. We went from weeks of manual pipeline work to days."

heashot image

Julian Thorne

Head of Product, NexaStream

"The precision of the image editing tools is unmatched. We went from weeks of manual pipeline work to days."

heashot image

Julian Thorne

Head of Product, NexaStream

"The speed and precision of the image generation and editing tools are incredible. We went from weeks of manual content pipeline work to days."

heashot image

Marcus Rivera

Data lead, DataLabs

"The speed and precision of the image generation and editing tools are incredible. We went from weeks of manual content pipeline work to days."

heashot image

Marcus Rivera

Data lead, DataLabs

"Generating high-fidelity video and audio in one click is vital. We went from weeks of manual pipeline work to days."

heashot image

Taylor Brooks

VP of Design, Shift Digital

"Generating high-fidelity video and audio in one click is vital. We went from weeks of manual pipeline work to days."

heashot image

Taylor Brooks

VP of Design, Shift Digital

Pricing

Pricing plans

No hidden fees. No complicated calculations. Just clear, transparent pricing that grows with you.

Monthly

Yearly

-20%

Starter

For prototyping and side projects

Free

WHAT’S INCLUDED

Model-dependent token pricing

Open-model rates by input + output

Small and mid-size open-source models

Llama, Mistral, Qwen class

5 requests / second

Standard latency, shared GPU pool

Best for hobby projects

MVPs and early API testing

Community support

Discord plus 7-day usage logs

Growth

For prototyping and side projects

$

99

/mo

WHAT’S INCLUDED

Discounted model-based rates

Lower input + output token costs

Flagship + open-source models

GPT, Claude, Gemini-class access

50 requests / second

Prompt caching up to 90% off

Batch API

50% off non-real-time jobs

Email + chat support

24-hour SLA + analytics

POPULAR

Enterprise

For prototyping and side projects

Custom

WHAT’S INCLUDED

Custom token rates

Committed volume tiers

Dedicated GPU capacity

Guaranteed throughput

Custom rate limits

Multi-region deployment options

Model routing engine

Route by cost and complexity

99.9% uptime SLA

Slack, engineer, compliance

Monthly

Yearly

-20%

Starter

For prototyping and side projects

Free

WHAT’S INCLUDED

Model-dependent token pricing

Open-model rates by input + output

Small and mid-size open-source models

Llama, Mistral, Qwen class

5 requests / second

Standard latency, shared GPU pool

Best for hobby projects

MVPs and early API testing

Community support

Discord plus 7-day usage logs

Growth

For prototyping and side projects

$

99

/mo

WHAT’S INCLUDED

Discounted model-based rates

Lower input + output token costs

Flagship + open-source models

GPT, Claude, Gemini-class access

50 requests / second

Prompt caching up to 90% off

Batch API

50% off non-real-time jobs

Email + chat support

24-hour SLA + analytics

POPULAR

Enterprise

For prototyping and side projects

Custom

WHAT’S INCLUDED

Custom token rates

Committed volume tiers

Dedicated GPU capacity

Guaranteed throughput

Custom rate limits

Multi-region deployment options

Model routing engine

Route by cost and complexity

99.9% uptime SLA

Slack, engineer, compliance

Monthly

Yearly

-20%

Starter

For prototyping and side projects

Free

WHAT’S INCLUDED

Model-dependent token pricing

Open-model rates by input + output

Small and mid-size open-source models

Llama, Mistral, Qwen class

5 requests / second

Standard latency, shared GPU pool

Best for hobby projects

MVPs and early API testing

Community support

Discord plus 7-day usage logs

Growth

For prototyping and side projects

$

99

/mo

WHAT’S INCLUDED

Discounted model-based rates

Lower input + output token costs

Flagship + open-source models

GPT, Claude, Gemini-class access

50 requests / second

Prompt caching up to 90% off

Batch API

50% off non-real-time jobs

Email + chat support

24-hour SLA + analytics

POPULAR

Enterprise

For prototyping and side projects

Custom

WHAT’S INCLUDED

Custom token rates

Committed volume tiers

Dedicated GPU capacity

Guaranteed throughput

Custom rate limits

Multi-region deployment options

Model routing engine

Route by cost and complexity

99.9% uptime SLA

Slack, engineer, compliance

Platform Pricing

Pay for exactly what you run

Transparent, usage-based rates across inference, compute, and model shaping. Pick a product to see its pricing.

GPU Instances

CPU Instances

Storage

Fine-Tuning

Serverless Inference

Dedicated Inference

Serverless Inference

Most teams start with serverless inference and move to dedicated endpoints at scale. Price per 1M tokens.

Model

Input

Output

DeepSeek V4 Pro

$3.46 · $0.30 cached

$6.96

DeepSeek V3 0324

$1.00 · $0.50 cached

$3.00

DeepSeek V4 Flash

$0.28 · $0.06 cached

$0.56

Nemotron 3.5 Lightning

$0.10 · $0.40 cached

$0.06

Nemotron 3 Ultra 550B

$2.00 · $0.50 cached

$6.40

Nemotron 3 Nano 30B-A3B-FP8

$0.20 · $0.06 cached

$0.40

Kimi K2.7 Code

$0.95 · $0.19 cached

$4.00

Kimi K2.6

$1.40 · $0.70 cached

$7.00

GLM 5.1

$2.40 · $0.50 cached

$8.80

GLM-5.2

$2.80 · $0.52 cached

$8.80

Gemma 4 31B-it

$0.28 · $0.28 cached

$0.80

GPT-OSS 120B

$0.10 · $0.10 cached

$0.40

Llama 3.3 70B Instruct

$0.50 · $0.26 cached

$1.50

Qwen3 235B A22B Instruct 2507

$0.44 · $0.22 cached

$1.60

GPU Instances

CPU Instances

Storage

Fine-Tuning

Serverless Inference

Dedicated Inference

Serverless Inference

Most teams start with serverless inference and move to dedicated endpoints at scale. Price per 1M tokens.

Model

Input

Output

DeepSeek V4 Pro

$3.46 · $0.30 cached

$6.96

DeepSeek V3 0324

$1.00 · $0.50 cached

$3.00

DeepSeek V4 Flash

$0.28 · $0.06 cached

$0.56

Nemotron 3.5 Lightning

$0.10 · $0.40 cached

$0.06

Nemotron 3 Ultra 550B

$2.00 · $0.50 cached

$6.40

Nemotron 3 Nano 30B-A3B-FP8

$0.20 · $0.06 cached

$0.40

Kimi K2.7 Code

$0.95 · $0.19 cached

$4.00

Kimi K2.6

$1.40 · $0.70 cached

$7.00

GLM 5.1

$2.40 · $0.50 cached

$8.80

GLM-5.2

$2.80 · $0.52 cached

$8.80

Gemma 4 31B-it

$0.28 · $0.28 cached

$0.80

GPT-OSS 120B

$0.10 · $0.10 cached

$0.40

Llama 3.3 70B Instruct

$0.50 · $0.26 cached

$1.50

Qwen3 235B A22B Instruct 2507

$0.44 · $0.22 cached

$1.60

GPU Instances

CPU Instances

Storage

Fine-Tuning

Serverless Inference

Dedicated Inference

Serverless Inference

Most teams start with serverless inference and move to dedicated endpoints at scale. Price per 1M tokens.

Model

Input

Output

DeepSeek V4 Pro

$3.46 · $0.30 cached

$6.96

DeepSeek V3 0324

$1.00 · $0.50 cached

$3.00

DeepSeek V4 Flash

$0.28 · $0.06 cached

$0.56

Nemotron 3.5 Lightning

$0.10 · $0.40 cached

$0.06

Nemotron 3 Ultra 550B

$2.00 · $0.50 cached

$6.40

Nemotron 3 Nano 30B-A3B-FP8

$0.20 · $0.06 cached

$0.40

Kimi K2.7 Code

$0.95 · $0.19 cached

$4.00

Kimi K2.6

$1.40 · $0.70 cached

$7.00

GLM 5.1

$2.40 · $0.50 cached

$8.80

GLM-5.2

$2.80 · $0.52 cached

$8.80

Gemma 4 31B-it

$0.28 · $0.28 cached

$0.80

GPT-OSS 120B

$0.10 · $0.10 cached

$0.40

Llama 3.3 70B Instruct

$0.50 · $0.26 cached

$1.50

Qwen3 235B A22B Instruct 2507

$0.44 · $0.22 cached

$1.60

PRICE ESTIMATOR

Estimate it. Then get exactly that.

Drag the sliders to model your month — tokens, GPU hours, storage, and fine-tuning runs. The invoice on the right updates live, and it’s the same math your real bill uses.

Per-second metering, rounded down

Credits applied before you pay

Hard budget caps, zero overage shocks

INFERNO CLOUDJUNE 2026
Serverless tokens
$71.75
$1.75 / M tokens
41 M tokens
GPU hours
$185.38
$2.99 / GPU hrs
62 GPU hrs
Managed storage
$40.80
$0.12 / GB
340 GB
Fine-tuning runs
$48.00
$24.00 / runs
2 runs
Subtotal
$345.93
Starter credits$50.00
TOTAL DUE$295.93
Prices in USD. Metered per second, billed monthly.

Controls

Product security

Production System User Review Situational Awareness For Incidents Vulnerability Remediation Process Audit Logging Data Security Integrations MFA Process Product Architecture Role-Based Access Control SLA

Data security

Access Monitoring Data Backups Data Erasure Encryption-at-rest Encryption-in-transit Geographic Location of Data Physical Security

Network security

Firewall Spoofing Protection Virtual Private Cloud Web Application Firewall Wireless Security

App security

Conspicuous Link to Privacy Notice Secure system modification Approval of Changes

Endpoint security

Anti-Malware Disk Encryption Mobile Device Management

Corporate security

Asset Management Practices Email Protection Employee Training Incident Response Internal Assessments Internal SSO Penetration Testing

Risk management

Risk Framework Risk Assessments Supply Chain Risk Management Third-Party Dependence

Policies

Upon request

Controls

Product security

Production System User Review Situational Awareness For Incidents Vulnerability Remediation Process Audit Logging Data Security Integrations MFA Process Product Architecture Role-Based Access Control SLA

Data security

Access Monitoring Data Backups Data Erasure Encryption-at-rest Encryption-in-transit Geographic Location of Data Physical Security

Network security

Firewall Spoofing Protection Virtual Private Cloud Web Application Firewall Wireless Security

App security

Conspicuous Link to Privacy Notice Secure system modification Approval of Changes

Endpoint security

Anti-Malware Disk Encryption Mobile Device Management

Corporate security

Asset Management Practices Email Protection Employee Training Incident Response Internal Assessments Internal SSO Penetration Testing

Risk management

Risk Framework Risk Assessments Supply Chain Risk Management Third-Party Dependence

Policies

Upon request

Controls

Product security

Production System User Review Situational Awareness For Incidents Vulnerability Remediation Process Audit Logging Data Security Integrations MFA Process Product Architecture Role-Based Access Control SLA

Data security

Access Monitoring Data Backups Data Erasure Encryption-at-rest Encryption-in-transit Geographic Location of Data Physical Security

Network security

Firewall Spoofing Protection Virtual Private Cloud Web Application Firewall Wireless Security

App security

Conspicuous Link to Privacy Notice Secure system modification Approval of Changes

Endpoint security

Anti-Malware Disk Encryption Mobile Device Management

Corporate security

Asset Management Practices Email Protection Employee Training Incident Response Internal Assessments Internal SSO Penetration Testing

Risk management

Risk Framework Risk Assessments Supply Chain Risk Management Third-Party Dependence

Policies

Upon request

Stop creating the hard way

Join thousands using AI to generate stunning visuals and ideas in seconds.

Join thousands using AI to generate stunning visuals and ideas in seconds.

desktop app
desktop app git view
desktop app sidebar