Run any model,
close to every user
Deploy open or custom models on GPUs that scale to zero. We send each request to the region nearest your user.
4.9 average
from 1,200+ reviews
Serving production traffic for teams at
Everything between your weights and your users
Push weights from Hugging Face or your own bucket. We build the container and give you an endpoint.
Wire up an endpoint in a minute
Connect a model, hardware, regions and a prompt. Kestrel turns the graph into a live API.
- Model weights
- Compute
- Routing
- System prompt
Support widget
Waiting
Slack bot
Waiting
Eval suite
Waiting
Usage alerts
Waiting
Model
llama-3.1-70b-instruct
fp8 · 128k context
Hardware
8 × H100 80GB
Autoscale 2–64 replicas
Regions
System prompt
Answer from the help center. Link the page you used.
Endpoint
support-agent- Model weights
- Compute
- Routing
- System prompt
Awaiting inputs · 1/4
Support widget
help.northwind.dev
Chat on every page
Slack bot
#support-escalations
Replies in threads
Eval suite
1,200 golden prompts
Runs on every deploy
Usage alerts
$2,000 monthly budget
Email and PagerDuty
Drag from a dot to the matching input. 1 of 4 connected.
Requests / sec
48,210
Regions live
7
Virginia, us-east, 38 milliseconds median latency.
One endpoint, seven regions
Deploy once. Kestrel serves your model from every region on the map.
Nearest-region routing
Each request goes to the closest healthy region. You change no DNS.
Automatic failover
If a region slows down, traffic moves to the next nearest one in seconds.
Health checks every 10 sec
We test every GPU node and drain bad hardware before it serves a token.
Change one URL.
Keep the code you have.
Kestrel speaks the OpenAI API, so the SDKs and tools you use today work without changes.
1import os2from openai import OpenAI34client = OpenAI(5 base_url="https://api.kestrel.dev/v1",6 api_key=os.environ["KESTREL_API_KEY"],7)89stream = client.chat.completions.create(10 model="llama-3.1-70b-instruct",11 messages=[{"role": "user", "content": "Where is my order?"}],12 stream=True,13)1415for chunk in stream:16 print(chunk.choices[0].delta.content, end="")
Where is my order?
- First token
- 84 ms
- Throughput
- 2,410 tok/s
- Cost
- $0.0004
Model catalog
Every open model,
ready to serve
Pick from more than a hundred tuned builds, or bring your own weights. Each one runs on the same API.
- Cold starts under 2 sec
- Private weights
- H100, H200 and L40S
- fp8 and int4 builds
- Pinned model versions
- Your own fine-tunes
llama-3.1-70b-instruct
General chat and tool use. Tuned for support teams and agents.
Jordan
Which tickets need a person today?
llama-3.1-70b-instruct
- Read 214 open tickets
- Checked order status in Shopify
These tickets need a person, as of 09:40:
Escalation queue
Pricing
Start free, pay for the tokens you serve, and reserve GPUs when your traffic is steady.
Hobby
For side projects and trying out models.
Free
No credit card needed
Includes:
- $5 in free credits each month
- Serverless endpoints
- 60 requests per minute
- Community support
Pro
Save 20%For teams that ship their first AI features.
$39
per workspace per month, billed yearly
Everything in Hobby, plus:
- $25 in monthly credits
- Every model in the catalog
- 600 requests per minute
- LoRA fine-tuning
- All 7 regions
Dedicated
Save 20%For steady traffic that needs fixed capacity.
$1.99
per H100 hour, 12-month term
Everything in Pro, plus:
- Reserved H100 and H200 GPUs
- No rate limits
- Autoscaling floors and caps
- Private model weights
- Priority support
$500 in creditsIncluded
Reserve GPUsEnterprise
For large teams with security and compliance needs.
Custom
Volume pricing
Everything in Dedicated, plus:
- Private regions and VPC peering
- SSO and audit logs
- 99.99% uptime SLA
- A named support engineer
- Invoicing and custom terms
| Features | HobbyFree | Pro$39 per workspace per month | Dedicated$1.99 per H100 hour | EnterpriseCustom |
|---|---|---|---|---|
| Usage | ||||
| Monthly credits | $5 | $25 | $500 | Custom |
| Requests per minute | 60 | 600 | No limit | No limit |
| Concurrent requests | 5 | 50 | 500 | Custom |
| Batch API | Not included | Included | Included | Included |
| Models | ||||
| Catalog models | Top 10 | All | All | All |
| LoRA fine-tuning | Not included | Included | Included | Included |
| Upload your own weights | Not included | Not included | Included | Included |
| Pinned model versions | Not included | Included | Included | Included |
| Infrastructure | ||||
| Regions | 3 | 7 | 7 | 7 + private |
| Reserved GPUs | Not included | Not included | Included | Included |
| Cold start | < 10 s | < 2 s | None | None |
| VPC peering | Not included | Not included | Not included | Included |
| Security and support | ||||
| SSO and audit logs | Not included | Not included | Not included | Included |
| SOC 2 Type II report | Not included | Included | Included | Included |
| Uptime SLA | Not included | 99.9% | 99.95% | 99.99% |
| Support | Community | Priority | Named engineer | |






