waste.tinymachines.ai

An OpenAI-compatible endpoint. Two inference engines, same weights, opposite memory strategies.

Models

idenginenotes
kimi-linearWASTE quantised experts streamed into a bounded RAM cache, CPU. Streaming supported.
kimi-linear-airllmAirLLM bf16 experts streamed into VRAM, GPU. Minutes per reply; no streaming.

Endpoints

GET /chatbrowser chat console — bring your token
GET /statuswhich models are up — no token needed
GET /docsinteractive API reference (redoc, openapi.json)
GET /v1/modelslist what is currently serving
POST /v1/chat/completionschat, stream supported on kimi-linear

Usage

curl https://waste.tinymachines.ai/v1/chat/completions \
  -H "authorization: Bearer $TOKEN" \
  -H "content-type: application/json" \
  -d '{"model":"kimi-linear",
       "messages":[{"role":"user","content":"Capital of France?"}]}'

A bearer token is required for /v1/. Generation is limited to 30 requests/minute per address; each engine runs one generation at a time and returns 503 rather than queueing. Over the limit is 429.