Architecture
A global GPU network built for creative AI.
Each generation request is routed to the optimal available node based on model requirements, VRAM availability, and geographic proximity. If a node drops mid-job, the scheduler reroutes automatically — the client receives output without retrying.
How it happens
Step-by-step: from request to result
You send a generation request via API or app
Include your model choice, prompt, resolution, and any ControlNet or LoRA parameters. Requests follow the OpenAI Images API schema — no new SDK required.
Sogni scheduler finds the best available node
Our scheduler evaluates latency, VRAM, queue depth, and model availability across all online nodes to select the optimal one for your job parameters. Geographic proximity to your endpoint is factored in.
Node runs the model, streams output
The selected node executes inference locally. For image generation, the result is returned as a URL-addressable object. For video, frames are streamed progressively.
Output delivered — typically under 2 seconds
Typical p50 latency for a 1024×1024 diffusion generation is under 2 seconds. p95 stays below 5 seconds under normal load. Results are cached at the CDN edge for fast retrieval.
Provider gets credited for the compute
The node that completed your job earns a credit allocation proportional to compute used. Credits accumulate and are redeemable monthly by the provider.
Technical specs
Network characteristics
Model support
50+ models
Stable Diffusion 1.5, 2.1, XL, SD 3.5, FLUX, Wan-Video, and more. Custom LoRA and ControlNet adapters supported.
API compatibility
OpenAI format
Fully OpenAI Images API v1 compatible. REST + JSON. SDK wrappers for Python, Node.js, and curl available.
Job queue
Priority routing
Standard, Priority, and Premium queues. Premium tier customers get first-pick node allocation and dedicated fallback nodes.
Node requirements
8GB VRAM min
Minimum 8GB VRAM GPU. Recommended 16GB+. Linux or Windows 10+, stable internet required.
Geographic distribution
APAC & Global
Nodes active across Southeast Asia, East Asia, Europe, and North America. APAC-optimized routing for Singapore origin requests.
Uptime SLA
99.2% observed
Network-level redundancy. Individual node failures are transparent to clients — jobs automatically reroute within milliseconds.
Provider economics
Earn for idle compute
Every completed generation job earns the contributing node a credit allocation. Credits accumulate in your provider dashboard and can be redeemed monthly via your preferred payment method.
- Earnings scale with your GPU tier and availability hours
- No lock-in — connect and disconnect your node anytime
- Lightweight provider client with minimal system overhead
- Monthly redemptions via supported payment methods
Provider earnings model
Per-job credits
Earned for every completed generation
Uptime bonus
Extra credits for high-availability periods
Monthly redemption
Redeem via preferred payment method