An OpenAI-compatible endpoint. Two inference engines, same weights, opposite memory strategies.
| id | engine | notes |
|---|---|---|
kimi-linear | WASTE | quantised experts streamed into a bounded RAM cache, CPU. Streaming supported. |
kimi-linear-airllm | AirLLM | bf16 experts streamed into VRAM, GPU. Minutes per reply; no streaming. |
GET /chat | browser chat console — bring your token |
GET /status | which models are up — no token needed |
GET /docs | interactive API reference (redoc, openapi.json) |
GET /v1/models | list what is currently serving |
POST /v1/chat/completions | chat, stream supported on kimi-linear |
curl https://waste.tinymachines.ai/v1/chat/completions \
-H "authorization: Bearer $TOKEN" \
-H "content-type: application/json" \
-d '{"model":"kimi-linear",
"messages":[{"role":"user","content":"Capital of France?"}]}'
A bearer token is required for /v1/. Generation is limited to 30
requests/minute per address; each engine runs one generation at a time and returns
503 rather than queueing. Over the limit is 429.