Report · llm-inference · updated 2026-09-24
LLM inference platforms
A timely slice of AI infrastructure: platforms that serve tokens and models to other software. Mix of neoclouds, GPU serverless, and the open-source serving engine many of them wrap. Only public sources are cited.
Placements
10 of 10 entities have a Jev placement. Axes: Proven Execution × Validated Direction.
| Entity | Status | Proven Execution | Validated Direction | uX | uY | Region | P(region) |
|---|---|---|---|---|---|---|---|
| Groq | scored | 0.475 | 0.432 | 0.351 | 0.379 | DIRECTED | OPERATIONAL:0.31 FORMING:0.05 DIRECTED:0.39 ANCHORED:0.25 |
| Fireworks AI | scored | 0.355 | 0.273 | 0.346 | 0.349 | OPERATIONAL | OPERATIONAL:0.76 ANCHORED:0.02 FORMING:0.16 DIRECTED:0.06 |
| Together AI | scored | 0.365 | 0.212 | 0.316 | 0.251 | OPERATIONAL | DIRECTED:0.06 ANCHORED:0.04 OPERATIONAL:0.79 FORMING:0.11 |
| Cerebras Inference | scored | 0.305 | 0.318 | 0.354 | 0.366 | OPERATIONAL | ANCHORED:0.05 DIRECTED:0.01 FORMING:0.01 OPERATIONAL:0.93 |
| Modal | scored | 0.465 | 0.225 | 0.325 | 0.265 | OPERATIONAL | ANCHORED:0.03 OPERATIONAL:0.59 DIRECTED:0.16 FORMING:0.22 |
| Replicate | scored | 0.378 | 0.213 | 0.405 | 0.338 | OPERATIONAL | ANCHORED:0.12 OPERATIONAL:0.78 DIRECTED:0.04 FORMING:0.06 |
| Baseten | scored | 0.253 | 0.240 | 0.349 | 0.341 | OPERATIONAL | DIRECTED:0.04 OPERATIONAL:0.76 FORMING:0.17 ANCHORED:0.03 |
| vLLM | scored | 0.388 | 0.230 | 0.389 | 0.345 | OPERATIONAL | OPERATIONAL:0.83 FORMING:0.12 DIRECTED:0.02 ANCHORED:0.03 |
| Hugging Face Inference | scored | 0.285 | 0.202 | 0.363 | 0.326 | OPERATIONAL | OPERATIONAL:0.65 ANCHORED:0.02 DIRECTED:0.04 FORMING:0.28 |
| Cloudflare Workers AI | scored | 0.230 | 0.247 | 0.345 | 0.353 | OPERATIONAL | ANCHORED:0.02 DIRECTED:0.10 FORMING:0.37 OPERATIONAL:0.51 |
Entities
- Groq
LPU-based inference cloud (GroqCloud) and silicon now in the NVIDIA stack.
2 evidence · scored · DIRECTED
- Fireworks AI
Training and inference platform for open and specialized models.
1 evidence · scored · OPERATIONAL
- Together AI
Open-model inference, fine-tuning, and GPU cloud.
2 evidence · scored · OPERATIONAL
- Cerebras Inference
Wafer-scale hosted inference with an OpenAI-compatible Chat Completions API.
1 evidence · scored · OPERATIONAL
- Modal
Programmable serverless GPU compute used for training and serving.
2 evidence · scored · OPERATIONAL
- Replicate
API for running and hosting public and custom models, often billed by runtime.
1 evidence · scored · OPERATIONAL
- Baseten
Model APIs plus dedicated deployments for custom serving.
1 evidence · scored · OPERATIONAL
- vLLM
Open-source high-throughput LLM serving engine used under many hosted stacks.
1 evidence · scored · OPERATIONAL
- Hugging Face Inference
Hosted inference and serverless endpoints on the Hugging Face Hub.
1 evidence · scored · OPERATIONAL
- Cloudflare Workers AI
Inference on Cloudflare’s edge network via Workers bindings.
1 evidence · scored · OPERATIONAL
Payment does not affect scores. Submit evidence or read the rubric.