Open frameworkJev QuadrantLive map

Report · llm-inference · updated 2026-09-24

LLM inference platforms

A timely slice of AI infrastructure: platforms that serve tokens and models to other software. Mix of neoclouds, GPU serverless, and the open-source serving engine many of them wrap. Only public sources are cited.

DIRECTEDANCHOREDFORMINGOPERATIONALGroqFireworks AITogether AICerebras InferenceModalReplicateBasetenvLLMHugging Face InferenceCloudflare Workers AIPROVEN EXECUTION →VALIDATED DIRECTION →

Placements

10 of 10 entities have a Jev placement. Axes: Proven Execution × Validated Direction.

Data table alternative to the LLM inference platforms quadrant chart
EntityStatusProven ExecutionValidated DirectionuXuYRegionP(region)
Groqscored0.4750.4320.3510.379DIRECTEDOPERATIONAL:0.31 FORMING:0.05 DIRECTED:0.39 ANCHORED:0.25
Fireworks AIscored0.3550.2730.3460.349OPERATIONALOPERATIONAL:0.76 ANCHORED:0.02 FORMING:0.16 DIRECTED:0.06
Together AIscored0.3650.2120.3160.251OPERATIONALDIRECTED:0.06 ANCHORED:0.04 OPERATIONAL:0.79 FORMING:0.11
Cerebras Inferencescored0.3050.3180.3540.366OPERATIONALANCHORED:0.05 DIRECTED:0.01 FORMING:0.01 OPERATIONAL:0.93
Modalscored0.4650.2250.3250.265OPERATIONALANCHORED:0.03 OPERATIONAL:0.59 DIRECTED:0.16 FORMING:0.22
Replicatescored0.3780.2130.4050.338OPERATIONALANCHORED:0.12 OPERATIONAL:0.78 DIRECTED:0.04 FORMING:0.06
Basetenscored0.2530.2400.3490.341OPERATIONALDIRECTED:0.04 OPERATIONAL:0.76 FORMING:0.17 ANCHORED:0.03
vLLMscored0.3880.2300.3890.345OPERATIONALOPERATIONAL:0.83 FORMING:0.12 DIRECTED:0.02 ANCHORED:0.03
Hugging Face Inferencescored0.2850.2020.3630.326OPERATIONALOPERATIONAL:0.65 ANCHORED:0.02 DIRECTED:0.04 FORMING:0.28
Cloudflare Workers AIscored0.2300.2470.3450.353OPERATIONALANCHORED:0.02 DIRECTED:0.10 FORMING:0.37 OPERATIONAL:0.51

Entities

  • Groq

    LPU-based inference cloud (GroqCloud) and silicon now in the NVIDIA stack.

    2 evidence · scored · DIRECTED

  • Fireworks AI

    Training and inference platform for open and specialized models.

    1 evidence · scored · OPERATIONAL

  • Together AI

    Open-model inference, fine-tuning, and GPU cloud.

    2 evidence · scored · OPERATIONAL

  • Cerebras Inference

    Wafer-scale hosted inference with an OpenAI-compatible Chat Completions API.

    1 evidence · scored · OPERATIONAL

  • Modal

    Programmable serverless GPU compute used for training and serving.

    2 evidence · scored · OPERATIONAL

  • Replicate

    API for running and hosting public and custom models, often billed by runtime.

    1 evidence · scored · OPERATIONAL

  • Baseten

    Model APIs plus dedicated deployments for custom serving.

    1 evidence · scored · OPERATIONAL

  • vLLM

    Open-source high-throughput LLM serving engine used under many hosted stacks.

    1 evidence · scored · OPERATIONAL

  • Hugging Face Inference

    Hosted inference and serverless endpoints on the Hugging Face Hub.

    1 evidence · scored · OPERATIONAL

  • Cloudflare Workers AI

    Inference on Cloudflare’s edge network via Workers bindings.

    1 evidence · scored · OPERATIONAL

Payment does not affect scores. Submit evidence or read the rubric.