One inference API.
The whole network behind it.
Run open models through one simple, reliable API on managed serverless capacity — and optionally route workloads to community GPUs when cheaper is what matters. No queues, no provisioning.
The node executes. The cloud thinks.
# One endpoint, the whole network
POST https://api.spidernetwork.ai/v1/jobs
{
"model": "llama3.1:8b",
"security": { "tier": "COMMUNITY" // cheapest route
},
"input": { …your chat messages… }
}
# → community GPU if available, serverless if not
202 Accepted · $1 / $3 per 1M tokens“Your model. A global network. One low-latency API.”
Run leading open models or upload your own fine-tuned weights. Spider routes every request to nearby healthy capacity for fast inference anywhere.
Bring the model. We bring the network.
Upload a fine-tune once, expose it through a familiar API, and make it available across a distributed global GPU fleet.
“Llama, Mistral, Qwen, Gemma—or your own fine-tune.”
“Upload once. Deploy across a worldwide GPU network.”
Pricing comparisons depend on model, usage, context length, and subscription tier. Final Spider pricing has not been published.
One network. Two sides.
Spider gives AI teams one inference API, backed by managed serverless capacity and a growing network of community GPUs — and pays the people who supply that hardware.
Ship inference at scale
- One API, always-on serverless capacity
- Opt into community GPUs for cheaper workloads
- Pay only for the tokens you use
Spider
Orchestrator
auth · routing · scheduling · health
Earn from your gaming PC
- Install a lightweight app in minutes
- Get paid in real currency — no crypto
- Stay fully in control of your machine
Any model. Global, low-latency inference.
Choose a leading open model or upload your own fine-tune, then serve it near your users through one endpoint your team can adopt in an afternoon.
Leading open models
Run Llama, Mistral, Qwen, Gemma, and more through one OpenAI-compatible API.
Upload your fine-tune
Bring custom weights and make your fine-tuned model available across the global network.
Choose your capacity tier
Verified serverless capacity by default, or opt into community GPUs when cheaper workloads matter more than guarantees.
Elastic scale
Burst from one request to thousands. Capacity flexes with demand, no provisioning required.
Usage-based pricing
Pay per token you consume — no idle reservations, no minimums, no capacity planning.
Reliability built in
Continuous health checks and heartbeats route around slow or offline nodes before they hit you.
Lower prices by design—not by cutting capability.
Spider removes per-seat subscriptions, reserved-capacity commitments, and closed-platform markup. Choose a leading open model or upload your own fine-tune, then pay only when it runs.
The Spider price advantage
Pay only when your model runs.
$0
per-seat fees
$0
capacity reservations
Usage
based billing
More GPU owners competing for workloads means lower-cost compute without locking your team into another subscription.
Subscription + markup
- Pay for every team member
- A model catalog chosen by the provider
- Capacity and rate limits you cannot control
Compute you actually use
- Pay for inference usage, not seats
- Llama, Mistral, Qwen, Gemma, and more
- Upload and serve your own fine-tuned model
- Low-latency routing across a global network
Final pricing is not yet published. Actual comparisons will vary by model, token volume, context length, and provider tier.
From install to online in minutes.
The node stays simple. Everything hard — auth, registration, health, and coordination — happens in the cloud.
- 01
Create an account
Sign up on the web and open your Spider dashboard.
- 02
Install the agent
Download the lightweight Node Agent and log in.
- 03
Register your node
The agent detects your hardware and registers securely.
- 04
Stay connected
A persistent link streams heartbeats and metrics live.
- 05
Go live
Your node appears online, ready to earn.
The friendly face
of the network.
Every machine that joins Spider adds another strand to a resilient, worldwide inference cloud.

A thin node. A thinking cloud.
The Node Agent only knows who it is, whether it's connected, and what it's been assigned. Every decision lives in the Spider Orchestrator.
Node Agent
A lightweight execution agent. It authenticates, detects hardware, registers, holds a persistent connection, and reports in. No business logic.
- Authenticate
- Detect hardware
- Heartbeat
- Execute (soon)
Spider Orchestrator
Auth
Accounts, login, node tokens
Node Registry
Source of truth for every node
Scheduler
Decides which node runs what
Dispatcher
Assigns and delivers work
Metrics
Heartbeats, health, telemetry
Billing
Usage accounting and payouts
API
Dashboard and developer surface
The foundation we build on.
Thin node
The agent stays as small as possible — no billing, scheduling, or knowledge of other nodes.
Cloud intelligence
Scheduling, routing, health, reliability, and model placement all live in the cloud.
API first
Everything is exposed through clean APIs — the same ones future developer tools will use.
Modular services
Loosely coupled services keep responsibilities separate and independently evolvable.
Built for evolution
Simple today, with a clear upgrade path to distributed scheduling and global routing.
Traditional payouts
No blockchain, no crypto. Node operators are paid in traditional currency.