About the service
Access to Thinking Machines Lab models through their Tinker platform: the flagship Inkling (975B parameters, 41B active) and the lighter Inkling-Small (276B parameters, 12B active) — open-weights MoE models that natively handle text, images and audio and hold a context window of up to 1M tokens. On Tinker both come in two context variants, 64K and 256K, and thinking effort is adjustable, trading answer quality against token spend. Beyond Inkling the platform hosts other open-weights models: Qwen3.5-4B, Qwen3.5-9B, Qwen3.6-35B, Qwen3.5-397B, Kimi-K2.6, DeepSeek-V3.1, GPT-OSS and Nemotron from Lightning-30B to Ultra-550B. Tinker is not only inference but a training API: it manages the GPUs and recovery while you write your own training loop, supporting LoRA over models from 1B to 1T+ parameters, SFT, RL (GRPO, PPO), DPO and distillation. Billing is per token, charged separately for prefill, sampling and training.