The distributed inference cloud

One inference API.
The whole network behind it.

Run open models through one simple, reliable API on managed serverless capacity — and optionally route workloads to community GPUs when cheaper is what matters. No queues, no provisioning.

The node executes. The cloud thinks.

spider — inference
# One endpoint, the whole network
POST https://api.spidernetwork.ai/v1/jobs

{
  "model": "llama3.1:8b",
  "security": { "tier": "COMMUNITY" // cheapest route
  },
  "input": { …your chat messages… }
}

# → community GPU if available, serverless if not
202 Accepted · $1 / $3 per 1M tokens
A lower-cost inference layer for AI builders
“Your model. A global network. One low-latency API.”

Run leading open models or upload your own fine-tuned weights. Spider routes every request to nearby healthy capacity for fast inference anywhere.

$

Bring the model. We bring the network.

Upload a fine-tune once, expose it through a familiar API, and make it available across a distributed global GPU fleet.

Llama, Mistral, Qwen, Gemma—or your own fine-tune.
Upload once. Deploy across a worldwide GPU network.

Pricing comparisons depend on model, usage, context length, and subscription tier. Final Spider pricing has not been published.

The two-sided network

One network. Two sides.

Spider gives AI teams one inference API, backed by managed serverless capacity and a growing network of community GPUs — and pays the people who supply that hardware.

DemandAI builders

Ship inference at scale

  • One API, always-on serverless capacity
  • Opt into community GPUs for cheaper workloads
  • Pay only for the tokens you use
Explore the API

Spider
Orchestrator

auth · routing · scheduling · health

SupplyNode owners

Earn from your gaming PC

  • Install a lightweight app in minutes
  • Get paid in real currency — no crypto
  • Stay fully in control of your machine
Start earning
For AI builders

Any model. Global, low-latency inference.

Choose a leading open model or upload your own fine-tune, then serve it near your users through one endpoint your team can adopt in an afternoon.

Read the docs

Leading open models

Run Llama, Mistral, Qwen, Gemma, and more through one OpenAI-compatible API.

Upload your fine-tune

Bring custom weights and make your fine-tuned model available across the global network.

Choose your capacity tier

Verified serverless capacity by default, or opt into community GPUs when cheaper workloads matter more than guarantees.

Elastic scale

Burst from one request to thousands. Capacity flexes with demand, no provisioning required.

Usage-based pricing

Pay per token you consume — no idle reservations, no minimums, no capacity planning.

Reliability built in

Continuous health checks and heartbeats route around slow or offline nodes before they hit you.

Lower-cost AI inference

Lower prices by design—not by cutting capability.

Spider removes per-seat subscriptions, reserved-capacity commitments, and closed-platform markup. Choose a leading open model or upload your own fine-tune, then pay only when it runs.

The Spider price advantage

Pay only when your model runs.

$0

per-seat fees

$0

capacity reservations

Usage

based billing

More GPU owners competing for workloads means lower-cost compute without locking your team into another subscription.

Closed AI platforms

Subscription + markup

  • Pay for every team member
  • A model catalog chosen by the provider
  • Capacity and rate limits you cannot control
Spider

Compute you actually use

  • Pay for inference usage, not seats
  • Llama, Mistral, Qwen, Gemma, and more
  • Upload and serve your own fine-tuned model
  • Low-latency routing across a global network

Final pricing is not yet published. Actual comparisons will vary by model, token volume, context length, and provider tier.

How it works

From install to online in minutes.

The node stays simple. Everything hard — auth, registration, health, and coordination — happens in the cloud.

  1. 01

    Create an account

    Sign up on the web and open your Spider dashboard.

  2. 02

    Install the agent

    Download the lightweight Node Agent and log in.

  3. 03

    Register your node

    The agent detects your hardware and registers securely.

  4. 04

    Stay connected

    A persistent link streams heartbeats and metrics live.

  5. 05

    Go live

    Your node appears online, ready to earn.

Meet the Spider mascot

The friendly face
of the network.

Every machine that joins Spider adds another strand to a resilient, worldwide inference cloud.

Spider, the friendly network mascot, descending on a web
Every node strengthens the web
Architecture

A thin node. A thinking cloud.

The Node Agent only knows who it is, whether it's connected, and what it's been assigned. Every decision lives in the Spider Orchestrator.

Data plane

Node Agent

A lightweight execution agent. It authenticates, detects hardware, registers, holds a persistent connection, and reports in. No business logic.

  • Authenticate
  • Detect hardware
  • Heartbeat
  • Execute (soon)
Control plane

Spider Orchestrator

Auth

Accounts, login, node tokens

Node Registry

Source of truth for every node

Scheduler

Decides which node runs what

Dispatcher

Assigns and delivers work

Metrics

Heartbeats, health, telemetry

Billing

Usage accounting and payouts

API

Dashboard and developer surface

Principles

The foundation we build on.

01

Thin node

The agent stays as small as possible — no billing, scheduling, or knowledge of other nodes.

02

Cloud intelligence

Scheduling, routing, health, reliability, and model placement all live in the cloud.

03

API first

Everything is exposed through clean APIs — the same ones future developer tools will use.

04

Modular services

Loosely coupled services keep responsibilities separate and independently evolvable.

05

Built for evolution

Simple today, with a clear upgrade path to distributed scheduling and global routing.

06

Traditional payouts

No blockchain, no crypto. Node operators are paid in traditional currency.