TL;DR
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
Create a free accountAs an affiliate, we earn on qualifying purchases.
Thinking Machines Lab released the full weights for Inkling, its first foundation model, on July 15 under the Apache 2.0 license before offering a closed API. The release signals that model ownership and deployment control may be more valuable to some buyers than leading every benchmark, though hardware demands and possible use restrictions require closer review.
Thinking Machines Lab, the 17-month-old company founded by former OpenAI technology chief Mira Murati, released the full weights for its first foundation model, Inkling, on July 15 under Apache 2.0 before launching a closed API. The open-first release gives developers more control over deployment and modification, even as the lab acknowledges that Inkling is not the strongest available model.
Inkling’s BF16 and NVFP4 checkpoints were published on Hugging Face with first-day support for Transformers, vLLM, SGLang and llama.cpp. Apache 2.0 generally allows users to download, modify and commercialize the weights, making the release materially different from models whose parameters remain accessible only through hosted services.
The flagship is a 975-billion-parameter Mixture-of-Experts model with 41 billion active parameters, a one-million-token context window and native support for text, images and audio as inputs. Thinking Machines says it was pretrained on 45 trillion tokens. The company also previewed Inkling-Small, with 276 billion total and 12 billion active parameters, but its full weights have not yet been released.
Vendor-published results place Inkling at 97.1% on AIME 2026 and 87.2% on GPQA Diamond. It trails named rivals on several coding and agent benchmarks, including SWE-bench Pro and Terminal-Bench 2.1. Some scores used a prerelease checkpoint, and independent replication has not been published.
Ownership Moves Ahead of Rankings
The order of release makes model ownership the main commercial signal. Organizations can inspect, fine-tune and host Inkling without depending solely on a vendor-controlled endpoint, offering a Western open-weight option for buyers concerned about service access, data handling or long-term dependence on one provider.
Inkling also offers a reasoning-effort setting from 0.2 to 0.99, allowing operators to trade output quality against tokens, latency and cost. The source material reports that the model matched Nemotron 3 Ultra on Terminal-Bench 2.1 while using about one-third of the tokens, but that comparison still needs outside testing.

Samsung 55-Inch Class U8000H Series Crystal UHD, 4K, Smart TV, 2026 Model
- Crystal Processor 4K: Enhances colors and sharpens details
- Endless Free Content: Access 2,700+ free streaming options
- Samsung TV Plus: 750+ free channels included
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Inside Inkling’s Open Release
Thinking Machines Lab was founded by Mira Murati and employs former OpenAI personnel who worked on ChatGPT. Its first model arrives as developers weigh hosted frontier systems against open-weight models that can be operated on private infrastructure.
The model routes each token through six of 256 experts, alongside two shared experts, across a 66-layer decoder. Its image and audio inputs are projected into a shared representation rather than handled through a separately attached language-model adapter. The release does not include Inkling’s training data or complete training pipeline, so it is more accurately described as open weight than fully open source.
“Inkling is not the strongest model available today, closed or open.”
— Thinking Machines Lab, in its Inkling announcement
Usage Rules and Scores Need Verification
A reported Model Acceptable Use Policy may apply to the original parameters and modified versions, including restrictions covering surveillance, deception and automated decisions affecting rights. That policy was not verified in the supplied analysis, leaving open how it interacts with the Apache 2.0 license for commercial and public-sector deployments.
Practical accessibility is another open issue. The source estimates that BF16 deployment needs at least two terabytes of aggregate VRAM, while NVFP4 still requires about 600 gigabytes. Inkling is open to download, but most local developers cannot run the flagship without costly infrastructure or heavy quantization.
Independent Tests and Smaller Weights
Developers will now test Inkling’s benchmark claims, reasoning-cost controls and multimodal performance on production workloads. Legal teams and regulated users are also likely to examine the model card and any separate use policy before deployment.
The next major milestone is the release of Inkling-Small’s full weights after testing. Its lower active-parameter count could make it more relevant to smaller operators, although its final hardware needs, license terms and real-world performance remain unconfirmed.
Key Questions
Are weights the key to AI thinking?
No single component explains AI reasoning. Weights store learned model parameters, while architecture, training, post-training and inference-time compute shape performance. Inkling’s news value lies in making those weights available for independent control.
Can anyone run Inkling locally?
Not on ordinary consumer hardware. Reported requirements reach at least two terabytes of VRAM for BF16 or about 600 gigabytes for NVFP4, placing the flagship beyond most workstations.
Is Inkling fully open source?
Inkling has published weights under Apache 2.0, but its training data and full pipeline were not released. It is best described as an open-weight model.
Does Inkling outperform other open models?
It leads or competes closely on some vendor-reported tests but trails GLM-5.2 and other models on several coding and agent tasks. The results are awaiting independent replication.
Source: Thorsten Meyer AI
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.