AI IaaS · In the WiLine suite

The infrastructure
AI actually needs.

Bare-metal GPU and CPU, Kubernetes, VPC, and storage - all delivered across 15 U.S. regions on WiLine's own backbone. No hypervisor tax, no egress fees, no sovereignty compromises.

Bare metal. Your own network underneath. Zero egress.

WILINE · IAAS CONTROL PLANE
15
U.S. regions
$0.61
GPU / hour
0
Egress fees
SOC 2
Type II · HIPAA
COMPUTE · STORAGE · NETWORK · 15/15 NOMINAL
WiLine Edge Cloud
AI IaaSInference Engine

Inference at the edge

Run 60+ open models on an OpenAI-compatible API - served from the US edge POP nearest your users. Single-digit-millisecond network overhead, data that never leaves the country.

POST /v1/chat/completionsDeepSeek-V3.1
DeepSeek-V3.1READY
Type a prompt and run it - tokens stream back live.
deepseek-v3.1 · us-west-2 · stream on
Open live playground ↗
No key connected - runs in simulated demo mode. Connect a WEC key for live inference.

Drop-in compatible with

vLLMSGLangOllamaLangChainLlamaIndexHugging FaceOpenAI-compatiblevLLMSGLangOllamaLangChainLlamaIndexHugging FaceOpenAI-compatible

01 · Deploy

Two ways to serve.

Start serverless and pay per token, or reserve dedicated capacity for steady, compliance-grade workloads.

Serverless

Pay per token, scale to zero.

Hit the API and go - no infrastructure to manage, no minimums. Ideal for spiky traffic, prototypes, and getting to production fast.

  • Per-token billing, no idle cost
  • Instant access to the full model catalog
  • Autoscaling throughput built in
  • $50 in starter credits
Start serverless

Dedicated endpoints

Reserved capacity, fixed latency.

Reserved H100 / MI300X capacity with single-tenant isolation - best for steady high throughput, compliance, and contractual SLAs.

  • Reserved H100 / MI300X capacity
  • 99.99% uptime SLA
  • Single-tenant isolation & fixed latency
  • Custom autoscaling & regional pinning
  • Bring fine-tuned checkpoints
Reserve capacity

02 · Catalog

The latest models, the day they drop.

Weekly open-source refreshes and Day-0 access to select frontier releases. One base URL, one-line model switching - no migrations, no re-tooling. Filter by modality, then open any model in the portal.

ModelDeveloperTypeContextParametersStatus
DeepSeekDeepSeek-V3.1DeepSeekReasoning128K671B MoEAvailable
AlibabaQwen3-235B-A22BAlibabaReasoning256K235B MoEAvailable
MetaLlama 4 MaverickMetaText + Vision1M400B MoEAvailable
Moonshot AIKimi K2Moonshot AIReasoning256K1T MoEDay 0
Mistral AIMistral Large 3Mistral AIChat128K123BAvailable
OpenAIgpt-oss-120bOpenAIReasoning128K120B MoEAvailable
GoogleGemma 3 27BGoogleText + Vision128K27BAvailable
AlibabaQwen3-VL 72BAlibabaVision128K72BAvailable
OpenAIWhisper Large v3OpenAISpeech → text-1.5BAvailable
AlibabaQwen3-Embedding-8BAlibabaEmbeddings32K8BAvailable
MetaLlama Guard 4MetaGuardrail128K12BAvailable

60+ models served · 6 modalities · 1 endpoint

03 · Multimodal

One API for every modality.

Text, images, video, speech and embeddings - served through the same endpoint, billed the same way, running on the same edge fleet. Hover a panel to expand it.

01 / 05
Why is inference faster at the edge?
REASONINGRequests reach a POP just miles from the user, so the network round-trip is ~3-8 ms instead of 40-80 ms. Tokens stream back the moment they're generated - keeping time-to-first-token under 200 ms even under load.

Text & reasoning

Chat, code, long-context reasoning and tool use through one OpenAI-compatible endpoint.

AI-generated image sample02 / 05FLUX.1.1 · 1024²

Image generation

FLUX and SDXL-class diffusion served on H100 for production-grade image pipelines.

03 / 05Wan 2.2 · 720p

Video generation

Wan-class text-to-video on reserved GPU clusters with frame-consistent output.

04 / 05
Whisper v3 · live

Speech

Real-time transcription and TTS with Whisper-class accuracy, streamed at the edge.

05 / 05Qwen3-Embedding · 4096-d

Embeddings

High-throughput embedding models for search, RAG and clustering - built for vector stores.

04 · Integrate

Compatible with your SDK.

If you've called OpenAI, you've called WiLine. Keep your existing client - swap the base URL and key. REST, streaming, function calling and structured outputs all work out of the box.

# Chat completion against the nearest US edge POP
curl https://inference.wiline.com/v1/chat/completions \
  -H "Authorization: Bearer $WEC_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v3.1",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

05 · Use cases

Everything chatbots and agents need.

One inference fabric - from conversational AI and retrieval to autonomous agents, vision and transcription - with the embeddings, function calling, structured outputs and guardrails to assemble them already on the platform.

Conversational AI

Low-latency chatbots and copilots that respond from the POP nearest your users.

RAG pipelines

Embeddings, retrieval and generation in one governed plane - no data leaves the US.

Autonomous agents

Native function calling and structured JSON for reliable multi-step agent loops.

Vision & OCR

Document understanding, image classification and multimodal extraction at scale.

Transcription

Batch and streaming speech-to-text for media, support and compliance workloads.

Classification

High-throughput scoring, moderation and routing with sub-10ms network overhead.

Building blocks · RAG & agents

Embedding models

State-of-the-art multilingual embeddings, served at high throughput for indexing and retrieval.

Function calling

Native tool-use schemas so agents can call your APIs with validated arguments.

Structured outputs

Constrain responses to JSON Schema for parser-free, production-safe pipelines.

Guardrails

Llama Guard and policy filters screen inputs and outputs before they reach users.

06 · Why WiLine

Built for inference, not borrowed from it.

Hyperscalers retrofit inference onto centralized regions. WiLine starts at the edge: compute next to your users, weights that stay in-country, and an API that costs what it says it does.

01 / 06

Inference at the edge

Tokens stream from the POP nearest your users - single-digit-ms network overhead, nationwide.

02 / 06

US data sovereignty

Every request, log and weight stays inside US borders. Zero-retention mode by default.

03 / 06

OpenAI-compatible

Point your existing SDK at our base URL. No migration, no re-tooling, no lock-in.

04 / 06

Autoscaling throughput

Burst to hundreds of millions of tokens per minute with speculative decoding built in.

05 / 06

No egress fees

Per-token billing with transparent rates. No surprise data-transfer charges, ever.

06 / 06

Enterprise security

SOC 2 Type II, HIPAA and ISO 27001. SSO, RBAC and dedicated isolation on request.

Serve your first token from the edge today.

OpenAI-compatible · $50 starter credits · no egress fees. Run 60+ models from 15 US metros, engineered for inference at the speed of light.

Get Connected with WiLine.

Dedicated internet, IP transit, and carrier ethernet - delivered from our hybrid fiber-wireless network.

*We'll only use your details to verify coverage and follow up. No obligations. No spam.

AI IaaS · The infrastructure plane

Build on infrastructure you can actually inspect.

Tell us the workloads. We'll design the stack — GPU, CPU, storage, Kubernetes, and VPC — across the 15 U.S. regions that fit your latency and compliance needs.