Inference at the edge
Run 60+ open models on an OpenAI-compatible API - served from the US edge POP nearest your users. Single-digit-millisecond network overhead, data that never leaves the country.
Drop-in compatible with
01 · Deploy
Two ways to serve.
Start serverless and pay per token, or reserve dedicated capacity for steady, compliance-grade workloads.
Serverless
Pay per token, scale to zero.
Hit the API and go - no infrastructure to manage, no minimums. Ideal for spiky traffic, prototypes, and getting to production fast.
- Per-token billing, no idle cost
- Instant access to the full model catalog
- Autoscaling throughput built in
- $50 in starter credits
Dedicated endpoints
Reserved capacity, fixed latency.
Reserved H100 / MI300X capacity with single-tenant isolation - best for steady high throughput, compliance, and contractual SLAs.
- Reserved H100 / MI300X capacity
- 99.99% uptime SLA
- Single-tenant isolation & fixed latency
- Custom autoscaling & regional pinning
- Bring fine-tuned checkpoints
02 · Catalog
The latest models, the day they drop.
Weekly open-source refreshes and Day-0 access to select frontier releases. One base URL, one-line model switching - no migrations, no re-tooling. Filter by modality, then open any model in the portal.
60+ models served · 6 modalities · 1 endpoint
03 · Multimodal
One API for every modality.
Text, images, video, speech and embeddings - served through the same endpoint, billed the same way, running on the same edge fleet. Hover a panel to expand it.
04 · Integrate
Compatible with your SDK.
If you've called OpenAI, you've called WiLine. Keep your existing client - swap the base URL and key. REST, streaming, function calling and structured outputs all work out of the box.
# Chat completion against the nearest US edge POP
curl https://inference.wiline.com/v1/chat/completions \
-H "Authorization: Bearer $WEC_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v3.1",
"messages": [{"role": "user", "content": "Hello"}]
}'05 · Use cases
Everything chatbots and agents need.
One inference fabric - from conversational AI and retrieval to autonomous agents, vision and transcription - with the embeddings, function calling, structured outputs and guardrails to assemble them already on the platform.
Conversational AI
Low-latency chatbots and copilots that respond from the POP nearest your users.
RAG pipelines
Embeddings, retrieval and generation in one governed plane - no data leaves the US.
Autonomous agents
Native function calling and structured JSON for reliable multi-step agent loops.
Vision & OCR
Document understanding, image classification and multimodal extraction at scale.
Transcription
Batch and streaming speech-to-text for media, support and compliance workloads.
Classification
High-throughput scoring, moderation and routing with sub-10ms network overhead.
Building blocks · RAG & agents
Embedding models
State-of-the-art multilingual embeddings, served at high throughput for indexing and retrieval.
Function calling
Native tool-use schemas so agents can call your APIs with validated arguments.
Structured outputs
Constrain responses to JSON Schema for parser-free, production-safe pipelines.
Guardrails
Llama Guard and policy filters screen inputs and outputs before they reach users.
06 · Why WiLine
Built for inference, not borrowed from it.
Hyperscalers retrofit inference onto centralized regions. WiLine starts at the edge: compute next to your users, weights that stay in-country, and an API that costs what it says it does.
Inference at the edge
Tokens stream from the POP nearest your users - single-digit-ms network overhead, nationwide.
US data sovereignty
Every request, log and weight stays inside US borders. Zero-retention mode by default.
OpenAI-compatible
Point your existing SDK at our base URL. No migration, no re-tooling, no lock-in.
Autoscaling throughput
Burst to hundreds of millions of tokens per minute with speculative decoding built in.
No egress fees
Per-token billing with transparent rates. No surprise data-transfer charges, ever.
Enterprise security
SOC 2 Type II, HIPAA and ISO 27001. SSO, RBAC and dedicated isolation on request.
Serve your first token from the edge today.
OpenAI-compatible · $50 starter credits · no egress fees. Run 60+ models from 15 US metros, engineered for inference at the speed of light.

