Home/Pricing

Pricing

8 plans across token credits and GPU compute. Every price is quoted in USD and invoiced by MCLY TECHNOLOGY INC. Larger customers are billed by bank transfer or corporate card — place an order and our team sends the payment details.

01 / TOKEN PLANS

Token Plans

Prepaid token credits for every frontier model — one API key, one bill.

from $99.00
Starter Pack

For prototypes and side projects — get your first million tokens flowing today.

$99.00 / one-time
  • $110 in API credits (10% bonus included)
  • Access to the standard model pool
  • OpenAI-compatible endpoint
  • Credits valid for 12 months
  • Email support
Credits
$110
Bonus
+10%
Models
Standard pool
Validity
12 months
Growth Pack Most Popular

For production workloads that need premium models and priority routing.

$499.00 / one-time
  • $625 in API credits (25% bonus included)
  • Standard + premium model pools
  • Priority routing & automatic failover
  • Credits valid for 24 months
  • Email support with 24-hour response
Credits
$625
Bonus
+25%
Models
Standard + Premium
Validity
24 months
Scale Pack

For high-volume platforms running millions of requests per month.

$1,999.00 / one-time
  • $2,700 in API credits (35% bonus included)
  • All models including the flagship tier
  • Dedicated capacity & lowest-latency routing
  • 99.9% uptime SLA
  • Priority support with a named engineer
Credits
$2,700
Bonus
+35%
Models
All models
SLA
99.9% uptime
Enterprise Tokens Contact Sales

Custom token volume, private routing, and procurement-friendly terms.

Custom
  • Custom credit volume & volume pricing
  • Private deployment & VPC peering options
  • Custom data-retention policies
  • Dedicated technical account manager
  • MSA, DPA & security review support
Credits
Custom
Models
All models
Support
Dedicated TAM
Billing
Invoice / PO
02 / COMPUTE PLANS

Compute Plans

Reserved GPU instances for training, fine-tuning and self-hosted inference.

from $349.00
RTX 4090 Instance

A dedicated RTX 4090 for inference, LoRA fine-tuning, and development work.

$349.00 / per month
  • Dedicated NVIDIA RTX 4090 (24 GB)
  • 16 vCPU · 64 GB RAM · 1 TB NVMe
  • Root access — any CUDA stack
  • 99.9% uptime SLA
  • Cancel or resize with 30-day notice
GPU
NVIDIA RTX 4090
VRAM
24 GB GDDR6X
CPU
16 vCPU
Memory
64 GB
Storage
1 TB NVMe
A100 Node

Production-grade A100 capacity for fine-tuning and batch inference.

$899.00 / per month
  • Dedicated NVIDIA A100 80 GB (SXM)
  • 32 vCPU · 256 GB RAM · 2 TB NVMe
  • NVLink & high-bandwidth interconnect
  • 99.9% uptime SLA
  • Pre-configured PyTorch or bring your own image
GPU
NVIDIA A100 SXM
VRAM
80 GB HBM2e
CPU
32 vCPU
Memory
256 GB
Storage
2 TB NVMe
H100 Node Most Popular

Flagship Hopper performance for frontier training and low-latency serving.

$1,499.00 / per month
  • Dedicated NVIDIA H100 80 GB (SXM)
  • 48 vCPU · 512 GB RAM · 4 TB NVMe
  • NVLink · InfiniBand ready
  • 99.9% uptime SLA
  • Lowest-latency serving for large models
GPU
NVIDIA H100 SXM
VRAM
80 GB HBM3
CPU
48 vCPU
Memory
512 GB
Storage
4 TB NVMe
H100 8-GPU Cluster

An 8×H100 NVLink cluster for distributed training at scale.

$9,999.00 / per month
  • 8× NVIDIA H100 80 GB with NVLink
  • 640 GB of aggregate VRAM
  • 400 Gb/s InfiniBand fabric
  • Reserved capacity — runs are never preempted
  • Priority support & cluster engineering hours
GPU
8× NVIDIA H100 SXM
VRAM
640 GB total
Fabric
400 Gb/s InfiniBand
CPU
128 vCPU
Storage
16 TB NVMe
+ / PAYMENT

How payment works

Large accounts settle by bank transfer or corporate card: place the order, email us the order number, and billing returns wire details or a secure card link — with a proper invoice — within one business day.

Online card payment Coming soon
+ / VOLUME

Committing real budget?

Annual commitments, reserved GPU pools and private deployments are quoted individually, with per-model rate cards and negotiated latency targets written into the agreement.

Request a quote
Prices exclude applicable taxes. Where tax applies it is shown on the invoice. Credits are valid for the window stated on the plan — see the refund policy for the unused-credit rules.

Start routing in minutes

One endpoint. Every frontier model.

Create an account, pick a plan, and point your existing OpenAI client at our base URL. No SDK rewrite, no lock-in.