Skip to content
cllm

// use cases

Where a no-OS inference kernel earns its keep.

cllm is deliberately the opposite end of the spectrum from a managed Python endpoint. These are the workloads where deleting the host OS is a feature — and we are honest about the ones where it is not.

// where cllm is the wrong tool

If you write Python and want a managed endpoint, need production GPU serving today, or run on macOS, Windows, or non-x86 Linux, cllm is not the right answer — and will not be. It targets the other end of the spectrum on purpose. See the vLLM comparison for that trade-off.