Meet Inkling: A Geordie’s Glimpse at the New Contender
If you’re searching for a top-tier open weights model, options outside Chinese stores are somewhat limited. Thinking Machines Lab believes it’s about to shift the landscape with Inkling, their newest innovation. Established in early 2025 by Mira Murati, a former high-ranking OpenAI official, their debut model is massive. With 975 billion parameters, it’s a tech powerhouse requiring over two terabytes of GPU memory—approximately eight of Nvidia’s B300 heavyweights or sixteen H200s. If that’s too demanding for your system, there’s a NVFP4 quantized alternative, needing only half the GPUs. It has now become the largest American open weights model, competing fiercely with Chinese titans like DeepSeek V4.
Stiff Competitions and Doubtful Metrics
Take their assertions with a grain of salt—gaming AI benchmarks is like stealing a pie from Greggs. Thinking Machines claims Inkling competes well with those massive Chinese models, although their data indicates it trails behind Anthropic’s Claude and OpenAI’s GPT. They portray Inkling as adaptable as a contortionist, perfect for developers, chatbots, and beyond. With a user-friendly Apache 2.0 license, users can adjust it to fit snugly. The Tinker platform even allows the model to compose its own scripts for refinement, adding a touch of Geordie toughness.
The Classic Components: Architecture and Training
Inkling features a 1 million-token context, akin to its short-term memory. This should aid in handling extensive codebases and complex queries. Thinking Machines acknowledges that the mixture of experts (MoE) framework was inspired by DeepSeek-V3, yet insists they developed Inkling from the ground up using Nvidia GB300 NVL72 systems and an impressive 45 trillion tokens of text, images, audio, and video. The model incorporates 256 routed exports along with two shared ones. Despite its large footprint, it’s as speedy as DeepSeek V4 on comparable hardware.
Thinking Tokens: Blessing or Bane?
Similar to modern LLMs, Inkling is a “reasoning model” trained via reinforcement learning to “process” requests. The developer asserts it’s been optimized to utilize thinking tokens effectively, rivaling Nvidia’s Nemotron 3 Ultra on Terminal Bench 2.1 while using only a third of the tokens. These thinking tokens could reduce AI hallucinations, but they come at a cost. The longer the model deliberates, the deeper into your wallet you may have to dig. Inkling is currently available on the Tinker platform, providing tools for customization and fine-tuning. They also plan to get it onto third-party API services as quickly as possible.
Hardware, Downloads, and Upcoming Releases
If experimenting in your own environment appeals to you, it’s downloadable from popular repositories like Hugging Face. It is compatible with a variety of engines, including vLLM and Llama.cpp at launch. Inkling is merely one of several models in the Thinking Machines lineup. Accompanying it is Inkling-Small, a speedy 276-billion-parameter MoE model for those who value low latency over sheer size and power. Named after the fictional supercomputer manufacturer from Jurassic Park, Thinking Machines is ensuring this model is flawless before unleashing its weights.
Conclusion: Big Innovations, Big Geordie Shenanigans
So, there you have it, Inkling in all its splendor. A huge contender from the rejuvenated team at Thinking Machines, headed by Mira Murati. If you’re enthusiastic about AI models that devour parameters like a Geordie at a buffet, you’re in luck. It may not be the most economical choice with those thinking tokens, but if you possess the GPUs and the funds, it could turn you into an AI legend right in your own living room. Go on, give it a shot at gadgetlad.co.uk.