Architecture

A global GPU network built for creative AI.

Each generation request is routed to the optimal available node based on model requirements, VRAM availability, and geographic proximity. If a node drops mid-job, the scheduler reroutes automatically — the client receives output without retrying.

How it happens

Step-by-step: from request to result

1

You send a generation request via API or app

Include your model choice, prompt, resolution, and any ControlNet or LoRA parameters. Requests follow the OpenAI Images API schema — no new SDK required.

2

Sogni scheduler finds the best available node

Our scheduler evaluates latency, VRAM, queue depth, and model availability across all online nodes to select the optimal one for your job parameters. Geographic proximity to your endpoint is factored in.

3

Node runs the model, streams output

The selected node executes inference locally. For image generation, the result is returned as a URL-addressable object. For video, frames are streamed progressively.

4

Output delivered — typically under 2 seconds

Typical p50 latency for a 1024×1024 diffusion generation is under 2 seconds. p95 stays below 5 seconds under normal load. Results are cached at the CDN edge for fast retrieval.

5

Provider gets credited for the compute

The node that completed your job earns a credit allocation proportional to compute used. Credits accumulate and are redeemable monthly by the provider.

Technical specs

Network characteristics

Model support

50+ models

Stable Diffusion 1.5, 2.1, XL, SD 3.5, FLUX, Wan-Video, and more. Custom LoRA and ControlNet adapters supported.

API compatibility

OpenAI format

Fully OpenAI Images API v1 compatible. REST + JSON. SDK wrappers for Python, Node.js, and curl available.

Job queue

Priority routing

Standard, Priority, and Premium queues. Premium tier customers get first-pick node allocation and dedicated fallback nodes.

Node requirements

8GB VRAM min

Minimum 8GB VRAM GPU. Recommended 16GB+. Linux or Windows 10+, stable internet required.

Geographic distribution

APAC & Global

Nodes active across Southeast Asia, East Asia, Europe, and North America. APAC-optimized routing for Singapore origin requests.

Uptime SLA

99.2% observed

Network-level redundancy. Individual node failures are transparent to clients — jobs automatically reroute within milliseconds.

Provider economics

Earn for idle compute

Every completed generation job earns the contributing node a credit allocation. Credits accumulate in your provider dashboard and can be redeemed monthly via your preferred payment method.

  • Earnings scale with your GPU tier and availability hours
  • No lock-in — connect and disconnect your node anytime
  • Lightweight provider client with minimal system overhead
  • Monthly redemptions via supported payment methods
Become a provider

Provider earnings model

Per-job credits

Earned for every completed generation

Variable

Uptime bonus

Extra credits for high-availability periods

Bonus

Monthly redemption

Redeem via preferred payment method

Monthly