Home / APIs & tokens / Thinking Machines Lab

Thinking Machines Lab

Thinking Machines Lab

Offers for this service 0 offers

Every offer for this service at once. Tap a product type to narrow it down — the page address changes, so the link can be shared.

Nobody has listed anything for this service yet. We do not invent prices: a displayed price is a public offer, and it only ever comes from a live seller.

What you will be able to buy

Plans

Pay-as-you-go

  • You pay for the tokens you actually use — separate rates for prefill, sampling and training
  • Inkling and Inkling-Small in two context variants: 64K and 256K (the long-context variant is selected explicitly for LoRA runs)
  • More than 28 open-weights models: Qwen3.5-4B, Qwen3.5-9B, Qwen3.6-35B, Qwen3.5-397B, Kimi-K2.6, DeepSeek-V3.1, GPT-OSS, Nemotron Lightning-30B, Super-120B, Ultra-550B
  • LoRA training for models from 1B to 1T+ parameters, dense and MoE, text and vision; SFT, RL (GRPO, PPO), DPO and distillation
  • Sampling via SamplingClient and OpenAI- and Anthropic-compatible endpoints; serverless inference is still in beta and only for the Inkling models
  • 80% discount on cached prefill tokens; checkpoint storage is billed per gigabyte per month

Plan contents as published by the vendor; seller prices arrive at launch.

About the service

Access to Thinking Machines Lab models through their Tinker platform: the flagship Inkling (975B parameters, 41B active) and the lighter Inkling-Small (276B parameters, 12B active) — open-weights MoE models that natively handle text, images and audio and hold a context window of up to 1M tokens. On Tinker both come in two context variants, 64K and 256K, and thinking effort is adjustable, trading answer quality against token spend. Beyond Inkling the platform hosts other open-weights models: Qwen3.5-4B, Qwen3.5-9B, Qwen3.6-35B, Qwen3.5-397B, Kimi-K2.6, DeepSeek-V3.1, GPT-OSS and Nemotron from Lightning-30B to Ultra-550B. Tinker is not only inference but a training API: it manages the GPUs and recovery while you write your own training loop, supporting LoRA over models from 1B to 1T+ parameters, SFT, RL (GRPO, PPO), DPO and distillation. Billing is per token, charged separately for prefill, sampling and training.

The same from another vendor

More in this category APIs & tokens