Inkling-Small: Open-weight multimodal model for enterprise AI deployment
venturebeat.com

Inkling-Small: Open-weight multimodal model for enterprise AI deployment

Tech News
2 min read

Published by AINave Editorial • Reviewed by Ramit

TL;DRThinking Machines releases Inkling-Small, a 276B-parameter open-weight multimodal model that approaches Inkling's performance with a smaller footprint and Apache 2.0 license.

Inkling-Small, a 276-billion-parameter open-weight multimodal reasoning model from Thinking Machines, delivers performance close to its larger predecessor Inkling while using roughly one quarter of the total parameters and a fraction of the active parameters per token. Released under Apache 2.0, the model is designed for enterprise self-hosting and custom deployment, with full weights on Hugging Face and fine-tuning via the Tinker API.

What happened

Inkling-Small is a sparse Mixture-of-Experts model with 276 billion total parameters and 12 billion active parameters per token. It uses a 42-layer decoder that routes each token to six of 256 specialized experts. The model accepts text, image, and audio inputs and supports a context window of up to one million tokens.

Thinking Machines released full BF16 and NVFP4 checkpoints on Hugging Face and added fine-tuning support through its Tinker API. At launch, the company offered a limited-time 50% discount on API pricing, bringing rates to $0.58 per million prefill tokens, $1.44 per million sampled tokens, and $1.73 per million training tokens. A 256K-context variant is also available at higher rates.

On benchmarks, Inkling-Small scores 80.2% on SWE-bench Verified compared to Inkling's 77.6%, and 64.7% on Terminal Bench 2.1 compared to 63.8%. It also edges

Sources

Latest Tech News